Hook-enforced development discipline for Claude Code. Skills teach methodology; hooks enforce compliance mechanically (git gates, planning mode, step budgets…

Hook-enforced development discipline for Claude Code.
Skills teach methodology. Hooks enforce compliance. Store survives compaction.
Agent development plugins today rely on prose instructions the LLM reads and decides to follow. This works 80-90% of the time. It fails exactly when it matters most — under context pressure, after compaction, in long sessions, when the LLM rationalizes past the rules.
Proctor is the first development discipline plugin where critical rules are enforced by code, not compliance. The skills teach why. The hooks enforce when. The store guarantees what survives.
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1From inside Claude Code:
/plugin marketplace add nalyk/nalyk-skills
/plugin install proctor@nalyk-skills
Or from the terminal:
claude plugin marketplace add nalyk/nalyk-skills
claude plugin install proctor@nalyk-skills
claude plugin install proctor@nalyk-skills --scope project
This writes to .claude/settings.json which you commit to version control. Teammates get Proctor when they trust the project folder.
Add to your project's .claude/settings.json:
{
"extraKnownMarketplaces": {
"nalyk-skills": {
"source": {
"source": "github",
"repo": "nalyk/nalyk-skills"
}
}
}
}
git clone https://github.com/nalyk/nalyk-skills.git
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir ./nalyk-skills/plugins/proctor
Proctor's enforcement layer requires the function hooks runtime. Set the environment variable before launching:
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude
Or export it in your shell profile:
echo 'export CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1' >> ~/.bashrc
Without the flag, the skills still load and work as prose guidance (like any other skills plugin), but the hooks — git gates, branch protection, SDD dashboard, state persistence — do not fire.
| Gate | What it blocks | What it requires | |||
|---|---|---|---|---|---|
| Test evidence | git commit, git push | Fresh passing test run in this session — unless the project has no test suite, the change is prose-only, or you said proctor: no tests | |||
| Test freshness | git commit, git push | Test run within 5 minutes (configurable) | |||
| Test passing | git commit, git push | Last test run exit code 0 | |||
| Proven test status | git commit, git push | The last test run finished in the foreground and its own exit status reached the result — not hidden by a following `\ | , ;, \ | \ | or & (set -o pipefail / set -e` count), not sent to the background, not moved there by the Bash timeout, not interrupted |
| Branch protection | git commit/push/merge/rebase/reset on a protected branch — including a line that switches onto one first (git checkout main && git merge feat), and a push that writes one from elsewhere (git push origin HEAD:main, :main, --all, --mirror). The remote's default branch (origin/HEAD) is protected alongside the configured list | Feature branch or explicit human consent | |||
| SDD merge | git merge while the SDD run has tasks unmarked or a fix round open | Finish and mark the tasks, or proctor: sdd stop | |||
| Fix-round cap | An Agent dispatch past the cap — a round the ledger recorded, or one the dispatch's prompt names for the current task | Adjudicate with Ruling: lines, complete the task, or proctor: sdd stop | |||
| Step budget | Bash, Write, Edit or NotebookEdit once an SDD task hits 100% of its step budget | Complete the task, proctor: budget extend, or proctor: sdd stop | |||
| Time budget | The same four tools once an SDD task hits 100% of its time budget | Same three exits | |||
| Planning mode (Bash) | Shell writes to implementation files during a design phase | A design doc, or exit planning mode | |||
| Secret detection | git commit with staged credentials | No AWS/OpenAI/GitHub/GitLab/Slack tokens, private keys, or password-shaped assignments (quoted or not) among the added lines — the commit that removes a leaked key is not the one to block |
An SDD run works through the whole plan in one turn, so what happens during it is reported on the result of the tool call that caused it: task progress when a ledger line is written, budget warnings on the call that crosses 80% or reaches 100%, model advice on the Agent result. A note produced when a turn ends (the watchdog, context pressure, a done-check from the answer) cannot reach the model then — a turn's result carries nothing the model reads — so it arrives with the next prompt.
Quiet mode (proctor: quiet on) suppresses the nudges: the watchdog, model selection, context pressure, the destructive-command and diff-size warnings, and the step-aside and run-complete lines. Notes that carry the run's state — task progress, budgets, the fix-round cap, the done-check, the ruling aggregation — still arrive, and hard gates always enforce.
| Enforcer | When it fires | What it says | |
|---|---|---|---|
| Skill watchdog | 4+ turns without invoking a skill | "Check if brainstorming, TDD, debugging, or review applies" (once) | |
| Model selection | Agent spawn without explicit model during SDD | "Consider a cheaper model for mechanical tasks" | |
| Fix-round cap | Round N of 5 reached | "Decide on each open finding — skip debatable ones, resolve critical ones" | |
| Escalation | Round 4-5 (the last two of the cap) dispatched on the model the stuck implementer used | "Try a more capable model" | |
| Step budget | 80% / 100% of tool call limit per task | Warning at 80% (once), wrap-up at 100% (once per task), on the tool result | |
| Time budget | 80% / 100% of wall-clock limit per task | Warning at 80% (once), wrap-up at 100% (once per task) | |
| Context pressure | Turn 50, 70, 90 | "Progress preserved automatically — focus on current task" | |
| Ruling aggregation | The last task marked complete; the finishing skill loaded | Full list of rulings and deferred minors | |
| SDD done-check | The finishing skill loaded, or an answer claiming the finish | Each unmet condition: tasks unmarked, fix round open, tests missing/failing/stale/unproven | |
| Destructive command | rm -rf, chmod 777, `curl\ | bash`, etc. | Shows the actual command — "verify this is intentional" |
| Diff size | >500 lines staged (pre-commit) | "Consider splitting into smaller commits" | |
| Test hint | No tests run this session | Shows detected test command with actionable next step | |
| Phase indicator | Skill invocation changes lifecycle phase | Shows current phase in context | |
| Task advance | SDD task marked complete | "Next: Task N" with budget reset notification |
| Feature | What it does |
|---|---|
| SDD tracking | A run starts when subagent-driven-development or executing-plans loads (or is announced); the plan's ### Task N headings size it; the ledger (progress.md) drives it — Plan: <path> — <N> tasks, Task N: complete, Task N: fix round M — approach: …, Ruling: …, minor (deferred): …, Task N: added, read as they are written through Write, Edit or the shell, each counted once however often the ledger is rewritten |
| SDD state persistence | Task completion, fix rounds, rulings, failed approaches tracked in $.store |
| Live status | A [PROCTOR] block rides on every prompt as context the model reads: test verdict and age, planning mode, phase, SDD progress |
| Compaction recovery | [PROCTOR — SDD STATE] is one of the conversation's context blocks, which the engine re-reads at compaction — with its "DO NOT REDO" section and the last test run's failure output |
| SDD session recovery | Active SDD state detected and resumed on session restart |
| Failed approach tracking | Fix round descriptions captured and injected post-compaction to prevent retries |
| Task completion evidence | Completion evidence recorded per task for audit trail |
| Progress dashboard | Status bar above prompt during SDD |
| Test run tracking | Records every test execution for the git gate; a passing command is remembered for the project's next session |
| Branch tracking | Detects branch changes for protection enforcement |
| Agent counting | Tracks agents spawned for dashboard and diagnostics |
| Commit attribution | Appends task references and test evidence to commit messages |
| Phase lifecycle | Tracks idle → brainstorming → planning → implementing → reviewing → finishing |
| Quality metrics | Cross-session counters: commits, gate denials, gates passed, fix rounds, test runs, autonomy rate |
| Ruling persistence | Last 20 rulings preserved across sessions |
| Command | What it does |
|---|---|
proctor: status | Full status dashboard: phase, branch, quiet mode, SDD state, test evidence, quality metrics |
proctor: show trace | Last 25 structured trace events with timestamps |
proctor: allow <branch> | Grant consent for protected branch operations |
proctor: approve design | Exit planning mode |
proctor: no tests | Stand the test gate down for this session — for projects that genuinely have no suite |
proctor: budget extend | Grant the current SDD task one more full step and time budget |
proctor: quiet on | Suppress soft warnings (hard gates still enforce) |
proctor: quiet off | Re-enable all warnings |
proctor: check | Pre-flight gate status: test evidence, branch protection, planning mode, secret scan |
proctor: diagnose | Self-analysis: gate autonomy rate, fix round patterns, recommendations |
proctor: tasks N | Update SDD total task count (scope change) |
sdd done / proctor: sdd stop | Deactivate SDD session |
Proctor includes 14 development skills. Each is a methodology document the agent reads and follows. The hooks enforce the critical gates that skills alone cannot guarantee.
| Skill | Purpose |
|---|---|
using-proctor | Bootstrap — establishes skill discovery and hook awareness |
brainstorming | Turn ideas into designs (spike/bounded/architectural paths) |
test-driven-development | Red-green-refactor cycle. Git gate blocks commits without tests |
systematic-debugging | Root cause before fixes. Four-phase investigation |
verification-before-completion | Evidence before claims. Git gate enforces mechanically |
subagent-driven-development | Execute plans with fresh agents per task. Dashboard + state persistence |
executing-plans | Execute plans inline (cheaper). Same enforcement as SDD |
writing-plans | Create implementation plans from specs |
requesting-code-review | Dispatch reviewers with proper packages |
receiving-code-review | Evaluate feedback technically, not performatively |
finishing-a-development-branch | Verify → present options → execute → clean up |
using-git-worktrees | Workspace isolation. Branch protection hooks enforce |
dispatching-parallel-agents | Independent concurrent tasks |
writing-skills | TDD applied to skill creation |
Plugin options (via plugin.json userConfig):
| Option | Default | Description |
|---|---|---|
protectedBranches | ["main","master","production","release"] | Branches protected from destructive git ops |
executableDocPatterns | [] | Globs for prose-shaped files that are really behaviour, e.g. runbooks/**; a glob without / (*.runbook.md) matches at any depth |
testFreshnessMinutes | 5 | How many minutes before test evidence expires |
watchdogTurnThreshold | 4 | Turns without a skill before the watchdog fires |
fixRoundCap | 5 | Maximum fix-loop rounds in SDD |
stepBudgetPerTask | 100 | Tool call limit per SDD task (warn 80%, block 100%) |
timeBudgetPerTaskMinutes | 30 | Wall-clock limit per SDD task in minutes (warn 80%, block 100%) |
A project with no test suite could otherwise never commit, so the gate steps aside — visibly, with a line in the transcript — when:
package.json whose test script is missing or the npm init stub is not a marker, orproctor: no tests this session."The change" is whatever the operation actually sends: for a commit, the working tree; for a push, the commits the upstream does not have. A push is never excused by an unrelated edit sitting in the working tree, and when there is no upstream to compare against the gate enforces rather than guesses. A merge is never excused on prose grounds at all.
Every other gate keeps enforcing regardless: branch protection, secret detection and planning mode are untouched by this.
"Inert prose" is narrower than "a .md file". These stay behaviour and keep the gate enforcing:
| Kind | Examples |
|---|---|
| Agent and tool instructions | SKILL.md, CLAUDE.md, AGENTS.md, anything under .claude/, .cursor/, .github/ |
| Runbooks and playbooks | RUNBOOK.md, runbooks/, playbooks/ |
| Anything a test reads | tests/, spec/, fixtures/, testdata/, __snapshots__/, e2e/, golden/ |
| A plugin's own behaviour | commands/, agents/, skills/, prompts/, references/, templates/ |
| Prose a toolchain executes | every .md/.rst/.qmd in a repo holding book.toml, _quarto.yml, runme.yaml, mkdocs.yml, jupytext.toml or a Docusaurus config |
| Whatever you declare | executableDocPatterns |
Quiet mode is toggled at runtime via proctor: quiet on/off — it suppresses soft warnings while hard gates continue to enforce. This is session-scoped and does not persist across sessions.
proctor/
├── .claude-plugin/plugin.json # Plugin manifest
├── hooks/
│ ├── hooks.json # Module registration
│ └── proctor.tsx # All hook registrations
├── skills/ # 14 methodology skills
│ ├── using-proctor/
│ ├── brainstorming/
│ ├── test-driven-development/
│ ├── systematic-debugging/
│ ├── verification-before-completion/
│ ├── subagent-driven-development/
│ ├── executing-plans/
│ ├── writing-plans/
│ ├── requesting-code-review/
│ ├── receiving-code-review/
│ ├── finishing-a-development-branch/
│ ├── using-git-worktrees/
│ ├── dispatching-parallel-agents/
│ └── writing-skills/
├── README.md
└── package.json
Teaching layer (skills): Prose documents that explain methodology. Rationalization tables, core principles, checklists. The agent reads these and follows them. This is what existing plugins do.
Enforcement layer (hooks): TypeScript middleware that fires on engine events. Hard gates deny operations that violate rules. Soft enforcers inject reminders when patterns suggest non-compliance. This is what only function hooks can do.
Infrastructure layer (store): Persistent state that survives compaction. Task progress, test evidence, rulings, fix-round counts. Injected into context automatically so the agent never forgets where it is.
$.store is one namespace shared by every session on the machine, so session state, test evidence, the SDD run and the trace log are each filed under the project root — the repo git rev-parse --show-toplevel reports, not the cwd of the moment. A second session in another repo no longer overwrites the first one's branch consents, quiet mode or evidence, and cd-ing into a subdirectory does not make a session's own test run look foreign. Two sessions in the same repo still share one record: they share the branch and the consents that go with it, so the sharing is the accurate reading.
Every write goes through a single per-key queue. Counters are the one exception to "state is load-bearing": they are telemetry, they are written best-effort, and a counter that cannot be written can never stop a gate from firing.
The teaching layer handles nuance — when to use TDD, how to classify a task as spike vs architectural, what makes a good ruling. Hooks cannot express this.
The enforcement layer handles compliance — did you run tests before committing, are you on a protected branch, have you invoked a skill. Prose cannot guarantee this.
The infrastructure layer handles memory — what tasks are complete, what rulings were made, when was the last test run. Neither prose nor hooks alone can persist state across compaction.
Proctor is inspired by Superpowers and shares its philosophy. Source: nalyk/proctor.
The shared philosophy: TDD, systematic debugging, verification before completion, subagent-driven development with ledger-based recovery.
The differences:
| Aspect | Superpowers | Proctor |
|---|---|---|
| Platform | 9+ harnesses (universal) | Claude Code CLI only |
| Enforcement | Prose instructions | Function hooks (hard gates + soft nudges) |
| State | Ledger files (agent-managed) | $.store + ledger files (hook-managed + agent-managed) |
| Compaction recovery | Agent must read ledger | Hook injects state into context automatically |
| Verification | Prose rule | Git gate (mechanical denial) |
| Branch protection | Prose rule | Hook gate (mechanical denial) |
| Progress visibility | Post-hoc (diagnosing-superpowers) | Real-time dashboard |
| Skill invocation | Bootstrap re-reading | Hook watchdog |
| Fix-round tracking | Agent counting | Hook state machine |
| Ruling aggregation | Agent scanning ledger | Hook aggregation at session end |
Superpowers is the right choice for Codex, Cursor, Gemini CLI, and other harnesses where function hooks are not available. Proctor is for Claude Code CLI sessions where mechanical enforcement matters.
make test at the repo root runs four layers:
claude plugin validate — the manifest and hooks module as the engine's loader reads themtests/*.test.mjs — the pure helpers, a feature ratchet, and the hooks against a small fake enginetests/engine/*.test.ts — claude plugin test: every feature driven through the real engine, with the host (git, files, store, the tools' own results) answered beneath the plugin by tests/engine/world.ts. The fake engine could only ever agree with Proctor's own reading of the API; this layer is where five delivery channels that never reached the model were found.tests/skill-contract.test.mjs — the skills against the hooks: every ledger line a skill teaches parses as the hooks read it, every command taught is answered, every mechanism a skill claims existstests/proctor-configured/ (repo root) — the same module loaded with every userConfig option set away from its default, through the engine's real options pipelinetests/COVERAGE.md maps every feature to the tests that prove it.
An end-to-end check against a live model loads the working tree in place of the installed copy:
claude -p --plugin-dir plugins/proctor \
--settings '{"enabledPlugins":{"proctor@nalyk-skills":false}}' \
--output-format stream-json --verbose "..."
Proctor requires the function hooks runtime, which is behind CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 and has not officially shipped. The runtime exists in Claude Code ≥ 2.1.260. Build against it to learn the shape, not to run production on it until Anthropic ships the feature.
MIT
hooks/proctor.tsx 4008 lines1// Proctor — Hook-enforced development discipline for Claude Code
2// EARLY ACCESS: requires CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1
3//
4// Skills teach methodology. Hooks enforce compliance. Store survives compaction.
5//
6// Four layers:
7// Teaching (skills) → prose methodology the agent reads and follows
8// Enforcement (hooks) → TypeScript middleware that gates or nudges
9// Infrastructure ($) → persistent state that survives compaction
10// Observability (tracing) → structured event log for audit trail
11//
12// Run /plugin-types to regenerate the declarations for your build.
13
14import type { Register } from "claude-code";
15
16// ─────────────────────────────────────────────────────────────────────
17// State schemas
18// ─────────────────────────────────────────────────────────────────────
19
20interface TestEvidence {
21 command: string;
22 timestamp: number;
23 exitCode: number;
24 tailOutput: string;
25 /** The run reported success without proving the suite passed: its
26 * exit status was hidden by what followed it on the command line
27 * (`| tail`, `; echo`, `|| true`, `&`), or it had not finished — sent
28 * to the background, timed out into it, or interrupted. Recorded with
29 * exitCode -1, and `unproven` says which. */
30 masked?: boolean;
31 unproven?: string;
32 /** Project the run happened in. Evidence from another project must not
33 * unblock this one's commit gate. */
34 cwd: string | null;
35}
36
37interface Ruling {
38 task: number;
39 text: string;
40 costIfWrong: string;
41 phase: "preflight" | "fix-loop" | "final";
42}
43
44interface DeferredMinor {
45 task: number;
46 finding: string;
47}
48
49interface SDDState {
50 active: boolean;
51 /** Project the run belongs to. A run does not follow the store into
52 * another repo. */
53 cwd: string | null;
54 plan: string;
55 startedAt: number;
56 totalTasks: number;
57 currentTask: number;
58 completedTasks: number[];
59 currentFixRound: number;
60 totalFixRounds: number;
61 totalAgents: number;
62 lastImplementerModel: string | null;
63 rulings: Ruling[];
64 deferredMinors: DeferredMinor[];
65 toolCallsThisTask: number;
66 totalToolCalls: number;
67 taskStartedAt: number;
68 stepWarned80: boolean;
69 stepWarned100: boolean;
70 timeWarned80: boolean;
71 timeWarned100: boolean;
72 failedApproaches: string[];
73 completedEvidence: Record<number, string>;
74 /** Signals already absorbed, by key. The ledger is rewritten whole and
75 * repeated in the final answer, so every signal arrives more than once;
76 * a fix round or ruling must still be counted once. */
77 seen?: string[];
78 /** The last task completed and the run ended itself, but nothing has
79 * checked it against the finish conditions yet — the finishing skill,
80 * a finishing answer or a merge still has to. */
81 pendingDoneCheck?: boolean;
82}
83
84interface SessionState {
85 startedAt: number;
86 skillInvoked: boolean;
87 lastSkillName: string | null;
88 watchdogNudgeSent: boolean;
89 testCommand: string | null;
90 branch: string | null;
91 /** The remote's default branch (origin/HEAD), protected alongside the
92 * configured list: a repo whose trunk is `develop` or `trunk` is not
93 * left unguarded because it is not called main. */
94 defaultBranch?: string | null;
95 isWorktree: boolean;
96 protectedBranches: string[];
97 branchConsents: Record<string, boolean>;
98 turnsSinceSkill: number;
99 agentsSpawned: number;
100 turnCount: number;
101 planningMode: boolean;
102 planningSkill: string | null;
103 hasTestInfrastructure: boolean;
104 testsAcknowledgedAbsent: boolean;
105 executableDocs: boolean;
106 currentPhase:
107 | "idle"
108 | "brainstorming"
109 | "planning"
110 | "implementing"
111 | "reviewing"
112 | "finishing";
113 quietMode: boolean;
114 /** Notes for the model produced after its turn ended — the watchdog,
115 * task progress, budget warnings, the done-check. `turn.complete`
116 * cannot reach the model, so they wait here for the next prompt. */
117 pendingNotes?: string[];
118}
119
120interface SessionHistory {
121 lastTestCommand: string | null;
122 /** The last passing test command per project root. `lastTestCommand`
123 * was one value for the whole machine, so a command learned in one repo
124 * became another repo's required suite. */
125 learnedTestCommands: Record<string, string>;
126 projectPath: string | null;
127 skillUsage: Record<string, number>;
128 sessionsCount: number;
129 recentRulings: Array<{ text: string; ts: number }>;
130 qualityMetrics: {
131 totalCommits: number;
132 gateDenials: number;
133 gatesPassed: number;
134 fixRounds: number;
135 testsRun: number;
136 };
137}
138
139interface TraceEvent {
140 ts: number;
141 kind: string;
142 detail: string;
143}
144
145// ─────────────────────────────────────────────────────────────────────
146// Store helpers — $.store is JSON-backed KV under
147// ~/.claude/plugins/store/, persisting across sessions.
148// Module-level variables are session-scoped.
149// ─────────────────────────────────────────────────────────────────────
150
151// The per-project keys carry a version suffix: their shape changed from a
152// bare record to a book keyed by project root, and a stale record read as
153// a book would be read as an empty one anyway. History stays unversioned —
154// `loadHistory` migrates it field by field, and its counters are the one
155// thing worth carrying forward.
156const KEYS = {
157 session: "proctor:session:v3",
158 test: "proctor:test-evidence:v3",
159 sdd: "proctor:sdd-state:v3",
160 history: "proctor:history",
161 trace: "proctor:trace:v3",
162} as const;
163
164const DEFAULT_HISTORY: SessionHistory = {
165 lastTestCommand: null,
166 learnedTestCommands: {},
167 projectPath: null,
168 skillUsage: {},
169 sessionsCount: 0,
170 recentRulings: [],
171 qualityMetrics: {
172 totalCommits: 0,
173 gateDenials: 0,
174 gatesPassed: 0,
175 fixRounds: 0,
176 testsRun: 0,
177 },
178};
179
180const TRACE_CAP = 50;
181
182async function load<T>(
183 $: any,
184 key: string,
185 fallback: T,
186): Promise<T> {
187 const raw = await $.store.get(key);
188 if (!raw) return fallback;
189 try {
190 return JSON.parse(raw) as T;
191 } catch {
192 return fallback;
193 }
194}
195
196/** A stored counter that is usable as a number, whatever the store holds.
197 * `undefined++` wrote NaN, which JSON stores as null, which the next
198 * read had to cope with in turn. */
199function counter(value: unknown): number {
200 return typeof value === "number" && Number.isFinite(value) ? value : 0;
201}
202
203/**
204 * SessionHistory with every field the current code expects, merged over
205 * whatever the store holds. `load` returns a stored object verbatim, so a
206 * history written by an older version is missing fields added since — and
207 * `hist.qualityMetrics.x++` on it throws. A throwing hook is skipped
208 * entirely, which silently disabled every gate it contained, because only
209 * the DENIAL paths mutate those counters outside a try.
210 *
211 * Every write goes through `mutateHistory`, so this is the only door.
212 */
213async function loadHistory($: any): Promise<SessionHistory> {
214 const stored = await load<Partial<SessionHistory> | null>(
215 $,
216 KEYS.history,
217 null,
218 );
219 const metrics = (stored?.qualityMetrics ?? {}) as Partial<
220 SessionHistory["qualityMetrics"]
221 >;
222
223 return {
224 ...structuredClone(DEFAULT_HISTORY),
225 ...(stored ?? {}),
226 skillUsage: { ...(stored?.skillUsage ?? {}) },
227 learnedTestCommands: { ...(stored?.learnedTestCommands ?? {}) },
228 recentRulings: [...(stored?.recentRulings ?? [])],
229 qualityMetrics: {
230 totalCommits: counter(metrics.totalCommits),
231 gateDenials: counter(metrics.gateDenials),
232 gatesPassed: counter(metrics.gatesPassed),
233 fixRounds: counter(metrics.fixRounds),
234 testsRun: counter(metrics.testsRun),
235 },
236 };
237}
238
239// One in-flight write per key. Every hook did load-whole-object, mutate
240// one field, save-whole-object, so two hooks running for the same
241// assistant message (three parallel Edits, or a Bash and a Read) both read
242// the same snapshot and the second write erased the first. Lost that way:
243// step-budget increments, a branch change, and planningMode being set.
244//
245// A queue only helps if every writer uses it: `agent.spawn`, `skill.prompt`,
246// `turn.complete`, `trace` and the branch tracker each did their own raw
247// load+save straight past it, so the increments they raced with were lost
248// exactly as before. Nothing below calls `save` outside this chain.
249const writeQueues = new Map<string, Promise<unknown>>();
250
251/** Run `job` with no other write to `key` interleaved. */
252function enqueue<T>(key: string, job: () => Promise<T>): Promise<T> {
253 const queued = (writeQueues.get(key) ?? Promise.resolve()).then(job);
254 // Keep the chain alive even if one link rejects.
255 writeQueues.set(key, queued.catch(() => undefined));
256 return queued;
257}
258
259/**
260 * Read, mutate and write a key with no other mutation interleaved.
261 * `fn` may mutate its argument in place (return nothing) or return a
262 * replacement — `null` included. `fn(x) ?? x` read a returned null as
263 * "mutated in place", so every `put…(null)` was a no-op: session.start's
264 * reset of the test evidence kept the last session's run standing.
265 *
266 * The fallback is cloned before `fn` sees it: `load` returns the fallback
267 * itself when the key is empty, so a shared constant handed in here
268 * (DEFAULT_HISTORY was) gets mutated in place and every later reader
269 * inherits the mutation for the life of the module.
270 */
271async function mutate<T>(
272 $: any,
273 key: string,
274 fallback: T,
275 fn: (value: T) => T | void,
276): Promise<T> {
277 return enqueue(key, async () => {
278 const current = await load<T>($, key, structuredClone(fallback));
279 const out = fn(current);
280 const updated = (out === undefined ? current : out) as T;
281 await save($, key, updated);
282 return updated;
283 });
284}
285
286/**
287 * Counters are telemetry. A gate must never fail to fire because a number
288 * could not be written, so this swallows its own errors and reads through
289 * `loadHistory` rather than `load` — the two halves of the bug that let a
290 * stale stored history throw inside a denial path and skip the denial.
291 */
292async function mutateHistory(
293 $: any,
294 fn: (history: SessionHistory) => void,
295): Promise<void> {
296 try {
297 await enqueue(KEYS.history, async () => {
298 const history = await loadHistory($);
299 fn(history);
300 await save($, KEYS.history, history);
301 });
302 } catch {
303 // Best-effort by design — see above.
304 }
305}
306
307async function save($: any, key: string, value: unknown): Promise<void> {
308 await $.store.set(key, JSON.stringify(value));
309}
310
311// ─────────────────────────────────────────────────────────────────────
312// Project scoping — $.store is one global namespace shared by every
313// session on the machine. Session state, test evidence, the SDD run and
314// the trace log all belong to ONE project and ONE session; keeping them
315// under a bare key meant a second pane's session.start overwrote them,
316// taking this session's branch consents, quiet mode and test evidence
317// with it. Each is now a book keyed by project root.
318// ─────────────────────────────────────────────────────────────────────
319
320interface Shelf<T> {
321 at: number;
322 value: T;
323}
324type Book<T> = Record<string, Shelf<T>>;
325
326const BOOK_CAP = 12;
327
328// The repo root, not the cwd: a session that cds into a subdirectory is
329// still the same project, and evidence recorded before the cd must not
330// turn foreign. Re-read only when the cwd moves.
331let rootCache: { cwd: string | null; root: string } | null = null;
332
333async function projectRoot($: any): Promise<string> {
334 let cwd: string | null = null;
335 try {
336 cwd = (await $.session.cwd()) ?? null;
337 } catch {
338 cwd = null;
339 }
340
341 if (rootCache && rootCache.cwd === cwd) return rootCache.root;
342
343 let root: string | null = null;
344 try {
345 const res = await $.process.run(["git", "rev-parse", "--show-toplevel"]);
346 const top = (res.stdout ?? "").trim();
347 if (res.exitCode === 0 && top) root = top;
348 } catch {
349 // Not a repo, or git is unavailable — fall back to the cwd.
350 }
351
352 const resolved = root ?? cwd ?? "unknown";
353 rootCache = { cwd, root: resolved };
354 return resolved;
355}
356
357async function loadScoped<T>($: any, key: string, fallback: T): Promise<T> {
358 const book = await load<Book<T>>($, key, {});
359 const root = await projectRoot($);
360 const shelf = book?.[root];
361 return shelf && "value" in shelf ? shelf.value : fallback;
362}
363
364async function mutateScoped<T>(
365 $: any,
366 key: string,
367 fallback: T,
368 fn: (value: T) => T | void,
369): Promise<T> {
370 const root = await projectRoot($);
371 let result = fallback;
372
373 await mutate<Book<T>>($, key, {}, (book) => {
374 const shelf = book[root];
375 const current =
376 shelf && "value" in shelf ? shelf.value : structuredClone(fallback);
377 const out = fn(current);
378 result = (out === undefined ? current : out) as T;
379 book[root] = { at: Date.now(), value: result };
380
381 // Projects come and go; the book must not grow forever.
382 const roots = Object.keys(book);
383 if (roots.length > BOOK_CAP) {
384 roots
385 .sort((a, b) => (book[a]?.at ?? 0) - (book[b]?.at ?? 0))
386 .slice(0, roots.length - BOOK_CAP)
387 .forEach((stale) => {
388 delete book[stale];
389 });
390 }
391 });
392
393 return result;
394}
395
396async function trace(
397 $: any,
398 kind: string,
399 detail: string,
400): Promise<void> {
401 try {
402 await mutateScoped<TraceEvent[]>($, KEYS.trace, [], (log) => {
403 log.push({ ts: Date.now(), kind, detail });
404 if (log.length > TRACE_CAP) log.splice(0, log.length - TRACE_CAP);
405 });
406 } catch {
407 // Tracing is best-effort — never block the hook chain
408 }
409}
410
411// The three per-project records, always read and written through the book.
412const loadSession = ($: any) =>
413 loadScoped<SessionState | null>($, KEYS.session, null);
414
415const mutateSession = ($: any, fn: (session: SessionState) => void) =>
416 mutateScoped<SessionState | null>($, KEYS.session, null, (session) => {
417 if (session) fn(session);
418 });
419
420const putSession = ($: any, session: SessionState | null) =>
421 mutateScoped<SessionState | null>($, KEYS.session, null, () => session);
422
423/** Hold notes for the model until the next prompt carries them. */
424const queueNotes = ($: any, notes: string[]) =>
425 mutateSession($, (session) => {
426 session.pendingNotes = [...(session.pendingNotes ?? []), ...notes];
427 });
428
429/** Take the held notes, leaving none behind. */
430async function drainNotes($: any): Promise<string[]> {
431 let notes: string[] = [];
432 await mutateSession($, (session) => {
433 notes = session.pendingNotes ?? [];
434 session.pendingNotes = [];
435 });
436 return notes;
437}
438
439const loadSDD = ($: any) => loadScoped<SDDState | null>($, KEYS.sdd, null);
440
441const mutateSDD = ($: any, fn: (sdd: SDDState) => void) =>
442 mutateScoped<SDDState | null>($, KEYS.sdd, null, (sdd) => {
443 if (sdd) fn(sdd);
444 });
445
446const putSDD = ($: any, sdd: SDDState | null) =>
447 mutateScoped<SDDState | null>($, KEYS.sdd, null, () => sdd);
448
449const loadEvidence = ($: any) =>
450 loadScoped<TestEvidence | null>($, KEYS.test, null);
451
452const putEvidence = ($: any, evidence: TestEvidence | null) =>
453 mutateScoped<TestEvidence | null>($, KEYS.test, null, () => evidence);
454
455// ─────────────────────────────────────────────────────────────────────
456// Test command detection — language-agnostic file heuristics
457// ─────────────────────────────────────────────────────────────────────
458
459const TEST_PATTERNS: Array<{ file: string; command: string }> = [
460 { file: "package.json", command: "npm test" },
461 { file: "Cargo.toml", command: "cargo test" },
462 { file: "pyproject.toml", command: "pytest" },
463 { file: "setup.py", command: "pytest" },
464 { file: "go.mod", command: "go test ./..." },
465 { file: "Makefile", command: "make test" },
466 { file: "Gemfile", command: "bundle exec rspec" },
467 { file: "mix.exs", command: "mix test" },
468 { file: "build.gradle", command: "./gradlew test" },
469 { file: "build.gradle.kts", command: "./gradlew test" },
470 { file: "pom.xml", command: "mvn test" },
471 { file: "CMakeLists.txt", command: "ctest" },
472 { file: "Rakefile", command: "rake test" },
473 { file: "deno.json", command: "deno test" },
474 { file: "bun.lockb", command: "bun test" },
475 { file: "composer.json", command: "vendor/bin/phpunit" },
476 { file: "phpunit.xml", command: "vendor/bin/phpunit" },
477 { file: "phpunit.xml.dist", command: "vendor/bin/phpunit" },
478 { file: "Package.swift", command: "swift test" },
479 { file: "pubspec.yaml", command: "dart test" },
480 { file: "build.zig", command: "zig build test" },
481 { file: "project.clj", command: "lein test" },
482 { file: "build.sbt", command: "sbt test" },
483 { file: "stack.yaml", command: "stack test" },
484 { file: "cabal.project", command: "cabal test" },
485];
486
487// Runners anchored to the START of a command segment. The old pattern
488// matched a bare `pytest|jest|vitest|...` anywhere, so `pip install
489// pytest`, `cat jest.config.js` or `echo pytest` were all recorded as
490// passing test runs and satisfied the commit gate with nothing run.
491const TEST_RUN_RE = new RegExp(
492 "^(?:" +
493 [
494 String.raw`npm\s+(?:run\s+[\w:.-]*test[\w:.-]*|test|t)\b`,
495 String.raw`(?:yarn|pnpm|bun)\s+(?:run\s+[\w:.-]*test[\w:.-]*|test)\b`,
496 String.raw`npx\s+(?:jest|vitest|mocha|playwright|cypress|ava|tap)\b`,
497 String.raw`(?:jest|vitest|mocha|ava|tap|cypress|playwright)\b`,
498 String.raw`deno\s+test\b`,
499 String.raw`(?:pytest|py\.test)\b`,
500 String.raw`python[0-9.]*\s+-m\s+(?:pytest|unittest)\b`,
501 String.raw`(?:poetry|uv|pipenv|hatch|rye|pdm)\s+run\s+\S*(?:pytest|test)\S*\b`,
502 String.raw`tox\b`,
503 String.raw`cargo\s+(?:test|nextest\s+run)\b`,
504 String.raw`go\s+test\b`,
505 String.raw`(?:bundle\s+exec\s+)?rspec\b`,
506 String.raw`rake\s+[\w:]*test[\w:]*\b`,
507 String.raw`(?:elixir\s+-S\s+)?mix\s+test\b`,
508 String.raw`make\s+[\w-]*test[\w-]*\b`,
509 String.raw`(?:\./)?gradlew?\s+[\w:]*test\b`,
510 String.raw`gradle\w*\s+[\w:]*test\b`,
511 String.raw`mvn\s+(?:-\S+\s+)*test\b`,
512 String.raw`dotnet\s+test\b`,
513 String.raw`swift\s+test\b`,
514 String.raw`ctest\b`,
515 String.raw`(?:vendor/bin/)?phpunit\b`,
516 String.raw`(?:dart|flutter)\s+test\b`,
517 String.raw`zig\s+build\s+test\b`,
518 String.raw`(?:lein|sbt|stack|cabal)\s+test\b`,
519 String.raw`nimble\s+test\b`,
520 String.raw`nim\s+c\s+-r\b`,
521 String.raw`busted\b`,
522 String.raw`bazel\s+test\b`,
523 // Node's built-in runner, and Claude Code's own plugin test runner.
524 String.raw`node\s+(?:-[-\w]+(?:=\S+)?\s+)*--test\b`,
525 String.raw`claude\s+plugin\s+test\b`,
526 ].join("|") +
527 ")",
528);
529
530/**
531 * True when the command actually runs a test suite. Splits on shell
532 * separators and checks each segment's FIRST word, after stripping
533 * leading env assignments and wrappers, so `cd x && npm test` counts and
534 * `grep -rn vitest src/` does not.
535 */
536function isTestRun(skel: string): boolean {
537 return skel
538 .split(/(?:&&|\|\||[;|&\n])+/)
539 .map((seg) =>
540 seg
541 .trim()
542 .replace(/^\(+\s*/, "")
543 .replace(/^(?:[A-Za-z_][A-Za-z0-9_]*=\S*\s+)*/, "")
544 .replace(/^(?:time|command|exec|nice|stdbuf\s+\S+)\s+/, ""),
545 )
546 .some((seg) => TEST_RUN_RE.test(seg));
547}
548
549/**
550 * True when the test run's own exit status cannot reach the tool result.
551 * The result carries one status for the whole line, so in `npm test |
552 * tail`, `npm test; echo done`, `npm test || true` or `npm test &` it is
553 * the last command's, and a failing suite reads as a pass. `&&` after the
554 * test is fine: a failure stops the chain and the line fails with it.
555 */
556function testStatusMasked(skel: string): boolean {
557 // Redirections carry an `&` that is not a separator: 2>&1, &>, >&2, |&.
558 const s = skel.replace(/\d*>&\d*|&>>?|<&\d*/g, " ").trim();
559 const parts = s.split(/(&&|\|\||\|&?|[;&\n])/);
560
561 let test = -1;
562 for (let i = 0; i < parts.length; i += 2) {
563 if (isTestRun(parts[i])) test = i;
564 }
565 if (test < 0) return false;
566
567 const before = parts.slice(0, test).join("");
568 const pipefail = /\bset\s+-[A-Za-z]*o\s+pipefail\b/.test(before);
569 const errexit =
570 /\bset\s+-[A-Za-z]*e[A-Za-z]*\b/.test(before) ||
571 /\bset\s+-o\s+errexit\b/.test(before);
572
573 let piped = true; // still inside the test's own pipeline
574 for (let j = test + 1; j < parts.length; j += 2) {
575 const sep = parts[j];
576 const rest = parts.slice(j + 1).join("").trim();
577 if (sep === "&") return true;
578 if (sep === "||") return true;
579 if (sep === "|" || sep === "|&") {
580 if (piped && !pipefail) return true;
581 continue;
582 }
583 piped = false;
584 if ((sep === ";" || sep === "\n") && rest && !errexit) return true;
585 }
586 return false;
587}
588
589/**
590 * Why a Bash result that reads as success proves nothing about the suite,
591 * or null when it does prove it. The status is one for the whole line, and
592 * a command that has not finished reports none of its own: Bash answers at
593 * once for `run_in_background`, and moves a command that hits its timeout
594 * to the background and answers then — on a slow machine, the usual fate
595 * of a full suite.
596 */
597function unprovenBy(e: any, record: any, skel: string): string | null {
598 if (e?.run_in_background === true || record?.backgroundTaskId) {
599 return record?.timedOutAfterMs
600 ? "it hit the Bash timeout and was moved to the background, so it had not finished"
601 : "it was sent to the background, so it had not finished";
602 }
603 if (record?.interrupted === true) return "it was interrupted before it finished";
604 if (testStatusMasked(skel)) {
605 return "it was followed by a pipe, `;`, `||` or `&`, so the result reported that command's status, not the tests'";
606 }
607 return null;
608}
609
610/** How a status line names a test run: a hidden status is not a failure. */
611function testVerdict(t: TestEvidence): string {
612 if (t.masked) return "UNPROVEN (exit status hidden)";
613 return t.exitCode === 0 ? "passing" : `FAILING (exit ${t.exitCode})`;
614}
615
616// Global options sit between `git` and its subcommand: `git -c k=v commit`,
617// `git -C dir push`, `git --git-dir=... merge`. Matching `git\s+commit`
618// alone lets every one of those slip past the gates untouched.
619const GIT_OPTS = String.raw`(?:` +
620 [
621 String.raw`-[cC]\s+\S+\s+`,
622 String.raw`--(?:git-dir|work-tree|namespace|exec-path|config-env|attr-source|super-prefix)(?:=\S+|\s+\S+)\s+`,
623 String.raw`--(?:paginate|no-pager|bare|no-replace-objects|no-lazy-fetch|no-optional-locks|literal-pathspecs|glob-pathspecs|noglob-pathspecs|icase-pathspecs|no-advice)\s+`,
624 String.raw`-[pP]\s+`,
625 ].join("|") +
626 String.raw`)*`;
627
628// `\b` after the verb matches before a hyphen too, so `git commit-tree` and
629// `git checkout-index` — plumbing that commits nothing — tripped every gate
630// keyed off these. A verb is the verb only when nothing word-like follows.
631const GIT_VERB_END = String.raw`(?![\w-])`;
632
633const GIT_COMMIT_PUSH_RE = new RegExp(
634 String.raw`\bgit\s+` + GIT_OPTS + String.raw`(commit|push|merge)` + GIT_VERB_END,
635);
636
637const GIT_PUSH_RE = new RegExp(
638 String.raw`\bgit\s+` + GIT_OPTS + String.raw`push` + GIT_VERB_END,
639);
640
641const GIT_MERGE_RE = new RegExp(
642 String.raw`\bgit\s+` + GIT_OPTS + String.raw`merge` + GIT_VERB_END,
643);
644
645// Files that normally carry no behaviour: prose, assets and licences.
646// Extension alone is not enough to conclude that, so BEHAVIORAL_DOC_RE and
647// the executable-doc checks below can each take a file back out of this set.
648const PROSE_FILE_RE =
649 /(\.(md|markdown|mdx|txt|rst|adoc|asciidoc|org|tex|svg|png|jpe?g|gif|webp|ico|pdf|woff2?|ttf|otf)$|^(LICENSE|COPYING|NOTICE|AUTHORS|CONTRIBUTORS|CHANGELOG|CODEOWNERS)([.\-](?:md|markdown|txt|rst|adoc))?$)/i;
650
651// Prose-shaped files that are nothing of the sort: a runbook something
652// runs, instructions an agent reads as its prompt, a fixture or snapshot a
653// test compares against. Editing one changes behaviour, so the commit gate
654// keeps enforcing even though the extension says prose.
655const BEHAVIORAL_DOC_RE = new RegExp(
656 [
657 // Read by tools and agents as instructions, not by people as prose.
658 String.raw`(^|/)(SKILL|CLAUDE|AGENTS?|GEMINI|CURSOR|COPILOT[-_]INSTRUCTIONS|WARP|RUNBOOK|PLAYBOOK)\.[^/]*$`,
659 // A plugin's own behaviour lives in these folders as markdown.
660 String.raw`(^|/)(commands|agents|skills|prompts|references|templates)/`,
661 // A tool's own directory: its contents are configuration.
662 String.raw`(^|/)\.(claude|cursor|github|gitlab|gemini|aider|continue|devcontainer)/`,
663 // Anything a test can read: fixtures, snapshots, golden files.
664 String.raw`(^|/)(tests?|spec|specs|fixtures?|testdata|__tests__|__snapshots__|__fixtures__|e2e|integration|golden|snapshots?)/`,
665 // Runbooks and playbooks kept together by folder.
666 String.raw`(^|/)(runbooks?|playbooks?)/`,
667 ].join("|"),
668 "i",
669);
670
671// Prose formats a documentation toolchain executes or compiles rather than
672// merely renders. Only treated as behaviour when the repo actually has such
673// a toolchain — see EXECUTABLE_DOC_MARKERS.
674const EXECUTABLE_DOC_EXT_RE = /\.(md|markdown|mdx|rst|adoc|asciidoc|org|qmd|ipynb)$/i;
675
676// A repo holding one of these runs, tests or builds its prose, so a change
677// to that prose can break the build the same way code can.
678const EXECUTABLE_DOC_MARKERS = [
679 "book.toml",
680 "_quarto.yml",
681 "_quarto.yaml",
682 "quarto.yml",
683 "runme.yaml",
684 "runme.yml",
685 "jupytext.toml",
686 "mkdocs.yml",
687 "mkdocs.yaml",
688 "docusaurus.config.js",
689 "docusaurus.config.ts",
690];
691
692/**
693 * Paths git reports as changed, staged or not, plus untracked files.
694 * Returns null when git cannot be read, so callers can tell "nothing
695 * changed" apart from "could not tell".
696 */
697async function changedPaths($: any): Promise<string[] | null> {
698 try {
699 const res = await $.process.run(["git", "status", "--porcelain"]);
700 if (res.exitCode !== 0) return null;
701 return res.stdout
702 .split("\n")
703 .map((l: string) => l.slice(3).trim())
704 .filter(Boolean)
705 .map((p: string) => {
706 // Renames read "old -> new"; the destination is what matters.
707 const arrow = p.indexOf(" -> ");
708 return arrow === -1 ? p : p.slice(arrow + 4);
709 })
710 // git quotes any path needing it (spaces, non-ASCII under
711 // core.quotepath). Unquote so extension matching still works.
712 .map((p: string) =>
713 p.startsWith('"') && p.endsWith('"') ? p.slice(1, -1) : p,
714 );
715 } catch {
716 return null;
717 }
718}
719
720/**
721 * The paths a push would send: what HEAD has that its upstream does not.
722 * Returns null when there is no upstream to compare against, or git
723 * cannot be read — "could not tell", which the gate treats as "enforce".
724 */
725async function pushedPaths($: any): Promise<string[] | null> {
726 try {
727 const res = await $.process.run([
728 "git",
729 "diff",
730 "--name-only",
731 "@{u}..HEAD",
732 ]);
733 if (res.exitCode !== 0) return null;
734 return res.stdout
735 .split("\n")
736 .map((l: string) => l.trim())
737 .filter(Boolean);
738 } catch {
739 return null;
740 }
741}
742
743/**
744 * The added lines of a diff. `containsSecret` is documented as reading
745 * what a change ADDS, and was handed the whole diff — so the commit that
746 * removes a leaked key was the one that got blocked, with remediation
747 * text telling you to remove it.
748 */
749function addedLines(diff: string): string {
750 return diff
751 .split("\n")
752 .filter((l) => l.startsWith("+") && !l.startsWith("+++"))
753 .join("\n");
754}
755
756/**
757 * A shell command with its heredoc bodies and quoted literals blanked out,
758 * so the command regexes match commands actually being run rather than text
759 * that merely mentions one. Without this, a commit whose message says
760 * "make test" is recorded as passing test evidence, and editing a document
761 * that quotes a git subcommand is treated as running it.
762 */
763function commandSkeleton(cmd: string): string {
764 const text = unquotedSkeleton(cmd);
765 let out = "";
766 let at = 0;
767 for (const [start, end] of quotedSpans(text)) {
768 out += text.slice(at, start) + text.slice(start, start + 1).repeat(2);
769 at = end;
770 }
771 return out + text.slice(at);
772}
773
774/**
775 * The command with each heredoc's body removed, and nothing else.
776 *
777 * Two regexes used to do this. The first replaced a whole heredoc with a
778 * `<<HEREDOC` token at the end of its line — which the second then read
779 * as an unterminated heredoc and erased everything after it, so a
780 * `git commit` or `git push` on the lines after any heredoc passed every
781 * gate unseen. The first also ate the rest of the delimiter line, so
782 * `cat <<EOF > src/x.ts` lost its write target.
783 *
784 * Now: the delimiter line is kept (`<<` itself replaced), the body up to
785 * the terminator is dropped, and scanning resumes after the terminator. A
786 * `<<` inside quotes on its own line is a mention, not a heredoc. With no
787 * terminator, the rest of the command is body.
788 */
789function stripHeredocs(cmd: string): string {
790 const re = /<<(-?)[ \t]*(['"]?)([A-Za-z_][A-Za-z0-9_]*)\2([^\n]*)/g;
791 let out = "";
792 let at = 0;
793 for (const m of cmd.matchAll(re)) {
794 const start = m.index ?? 0;
795 if (start < at) continue; // inside a body already dropped
796
797 // Quoted on its own line — "a << b" — is text, not a heredoc.
798 const lineStart = cmd.lastIndexOf("\n", start - 1) + 1;
799 const before = cmd.slice(lineStart, start).replace(/\\./g, "");
800 const quotes = (q: string) => before.split(q).length - 1;
801 if (quotes("'") % 2 === 1 || quotes('"') % 2 === 1) continue;
802
803 const delimiter = m[3];
804 const lineEnd = start + m[0].length;
805 const bodyStart = cmd.indexOf("\n", lineEnd);
806 out += cmd.slice(at, start) + " HEREDOC " + m[4];
807 if (bodyStart === -1) {
808 at = lineEnd;
809 continue;
810 }
811 const terminator = new RegExp(String.raw`^[ \t]*${delimiter}[ \t]*$`, "m");
812 const rest = cmd.slice(bodyStart + 1);
813 const end = rest.search(terminator);
814 if (end === -1) {
815 at = cmd.length;
816 break;
817 }
818 const afterTerminator = rest.indexOf("\n", end);
819 out += "\n";
820 at = afterTerminator === -1 ? cmd.length : bodyStart + 1 + afterTerminator + 1;
821 }
822 return out + cmd.slice(at);
823}
824
825/**
826 * The same command with heredoc bodies removed and shell `-c` payloads
827 * unwrapped, but quoted literals left standing. Write-target detection
828 * needs this: blanking quotes first erased the path in `> 'src/x.ts'`
829 * along with the mention it was meant to erase.
830 */
831function unquotedSkeleton(cmd: string): string {
832 let out = cmd;
833
834 out = stripHeredocs(out);
835 // A shell's -c payload is a command, not a literal: unwrap it so what
836 // it runs is still seen. Only shells — `python3 -c "..."` stays opaque.
837 out = out.replace(
838 /\b(?:(?:ba|z|k|da)?sh|fish)\s+(?:-[a-zA-Z]+\s+)*-c\s+('(?:[^'\\]|\\.)*'|"(?:[^"\\]|\\.)*")/g,
839 (_m: string, q: string) => ` ${q.slice(1, -1)} `,
840 );
841
842 return out;
843}
844
845/**
846 * The character spans covered by quoted literals, scanned left to right.
847 *
848 * Two passes used to do this, every single-quoted span first and then
849 * every double-quoted one. A double-quoted string holding an apostrophe
850 * (`claude -p "don't commit"`) mis-paired: the apostrophe opened a span
851 * that ran on to the next one, and the quoted text between them stayed
852 * bare — so a command merely NAMED inside a quoted argument was read as a
853 * command being run, and the git gates fired on it.
854 */
855function quotedSpans(text: string): Array<[number, number]> {
856 const spans: Array<[number, number]> = [];
857 for (let i = 0; i < text.length; i++) {
858 const c = text[i];
859 if (c === "\\") {
860 i++;
861 continue;
862 }
863 if (c !== "'" && c !== '"') continue;
864 let j = i + 1;
865 for (; j < text.length; j++) {
866 if (c === '"' && text[j] === "\\") {
867 j++;
868 continue;
869 }
870 if (text[j] === c) break;
871 }
872 // An unclosed quote runs to the end of the command.
873 spans.push([i, Math.min(j + 1, text.length)]);
874 i = j;
875 }
876 return spans;
877}
878
879const GIT_COMMIT_RE = new RegExp(
880 String.raw`\bgit\s+` + GIT_OPTS + String.raw`commit` + GIT_VERB_END,
881);
882
883const GIT_BRANCH_SWITCH_RE = new RegExp(
884 String.raw`\bgit\s+` + GIT_OPTS + String.raw`(checkout|switch)` + GIT_VERB_END,
885);
886
887// Commands that stage as they go: `git add -A && git commit`, or
888// `git commit -am`. The hook runs before any of it, so at scan time the
889// index does not yet hold what is about to be committed.
890const GIT_STAGING_RE = new RegExp(
891 String.raw`\bgit\s+` +
892 GIT_OPTS +
893 String.raw`(?:add|stage)\b|\bgit\s+` +
894 GIT_OPTS +
895 String.raw`commit\b[^\n;&|]*(?:\s-[a-zA-Z]*a[a-zA-Z]*\b|\s--(?:all|include|only)\b)`,
896);
897
898// Untracked files read per scan, so a large working tree cannot stall a
899// commit while the gate reads it.
900/**
901 * A path glob as a RegExp, built in one pass. A chain of .replace() calls
902 * lets a later pass rewrite regex syntax an earlier one emitted: the `?`
903 * rule corrupted the non-capturing group the double-star rule had just
904 * produced, so a declared pattern silently matched nothing.
905 *
906 * A single star stays inside one segment, a double star followed by a
907 * slash spans any number of directories, and a trailing double star takes
908 * the rest of the path.
909 */
910function globToRegExp(glob: string): RegExp {
911 let out = "";
912
913 for (let i = 0; i < glob.length; i++) {
914 const c = glob[i];
915
916 if (c === "*") {
917 if (glob[i + 1] === "*") {
918 i++;
919 if (glob[i + 1] === "/") {
920 i++;
921 out += "(?:.*/)?";
922 } else {
923 out += ".*";
924 }
925 } else {
926 out += "[^/]*";
927 }
928 } else if (c === "?") {
929 out += "[^/]";
930 } else if (".+^${}()|[]\\/".includes(c)) {
931 out += "\\" + c;
932 } else {
933 out += c;
934 }
935 }
936
937 return new RegExp("^" + out + "$", "i");
938}
939
940/**
941 * A design document, which planning mode allows writing. Matched against
942 * the BASENAME: matching the whole path meant one ancestor directory
943 * named design/, plan/, spec/, proposal/ or rfc/ exempted every file in
944 * the repo and switched the planning hard-gate off wholesale.
945 */
946function isDesignDoc(path: string): boolean {
947 const base = path.split("/").pop() ?? path;
948
949 // A prose extension is a design doc outright.
950 if (/\.(md|markdown|mdx|rst|adoc|txt)$/i.test(base)) return true;
951
952 // Any other extension is code or data, whatever the name says —
953 // `Plan.tsx` and `api-spec.go` are implementation files.
954 if (/\.[A-Za-z0-9]+$/.test(base)) return false;
955
956 // Extensionless, so judge by name: DESIGN, rfc-0001, SPEC.
957 return /\b(design|plan|spec|proposal|rfc)\b/i.test(base);
958}
959
960/**
961 * The branch HEAD is on right now. The gate used to trust the branch
962 * cached at session start, which a `git checkout main && git merge feat`
963 * in one command line has not updated yet — and which was never updated
964 * at all while the post-hook that maintained it was unreachable (S25).
965 * Falls back to the cached value when git cannot be read.
966 */
967async function currentBranch($: any, cached: string | null): Promise<string | null> {
968 try {
969 const res = await $.process.run(["git", "branch", "--show-current"]);
970 const live = (res.stdout ?? "").trim();
971 return live || cached;
972 } catch {
973 return cached;
974 }
975}
976
977/**
978 * A tool.call result reports failure through `isError` and carries its
979 * output as text. It has no exitCode, stdout or stderr — those belong to
980 * `$.process.run`, which is a different shape. Reading them here recorded
981 * every failing test run as a pass (`exitCode ?? 0`), left the failure
982 * output empty, and made the `result.exitCode === 0` trackers unreachable,
983 * so a branch switch was never noticed and commits were never counted.
984 *
985 * A numeric exitCode is still honoured if a future build supplies one.
986 */
987function toolFailed(result: any): boolean {
988 if (typeof result?.exitCode === "number") return result.exitCode !== 0;
989 return result?.isError === true;
990}
991
992function toolOutput(result: any): string {
993 if (typeof result?.stdout === "string" || typeof result?.stderr === "string") {
994 return `${result.stdout ?? ""}${result.stderr ?? ""}`;
995 }
996 return String(result?.text ?? result?.result ?? "");
997}
998
999// Shell forms that write a file: a redirection, an in-place edit, a tee.
1000// A path may be quoted, so the alternatives accept a quoted literal —
1001// `> 'src/x.ts'` is a write to src/x.ts, and matching against a skeleton
1002// with the quotes blanked out saw no path at all.
1003const WRITE_PATH = String.raw`'[^']+'|"[^"]+"|[\w./~@+=-]+`;
1004
1005const SHELL_WRITE_RE = new RegExp(
1006 String.raw`>>?\s*(?!&)(` +
1007 WRITE_PATH +
1008 String.raw`)|\b(?:sed|perl|ruby)\s+(?:-\S+\s+)*-i\S*\s+(?:-\S+\s+)*(?:'[^']*'|"[^"]*"|\S+)\s+(` +
1009 WRITE_PATH +
1010 String.raw`)|\btee\s+(?:-\S+\s+)*(` +
1011 WRITE_PATH +
1012 String.raw`)`,
1013 "g",
1014);
1015
1016// A redirection to one of these writes nothing a planning gate cares
1017// about; `ls > /dev/null` was being denied as an implementation write.
1018const DEV_SINK_RE = /^\/dev\/(?:null|stdout|stderr|tty|fd\/\d+)$/;
1019
1020/**
1021 * Every file the command writes.
1022 *
1023 * Three bugs lived in the single-match version this replaces: `tee` is the
1024 * third capture group and only the first two were read; a quoted path was
1025 * erased before the match; and `String.match` without /g returned one
1026 * target, so `echo a > notes.md && echo b > src/x.ts` was judged entirely
1027 * by notes.md. A write named inside a quoted string is still only a
1028 * mention — the operator itself has to be outside the quotes.
1029 */
1030function shellWriteTargets(cmd: string): string[] {
1031 const text = unquotedSkeleton(cmd);
1032 const spans = quotedSpans(text);
1033 const mention = (at: number) => spans.some(([a, b]) => at > a && at < b);
1034
1035 const targets: string[] = [];
1036 for (const m of text.matchAll(SHELL_WRITE_RE)) {
1037 if (m.index !== undefined && mention(m.index)) continue;
1038 const path = (m[1] ?? m[2] ?? m[3] ?? "")
1039 .replace(/^['"]|['"]$/g, "")
1040 .trim();
1041 if (!path || DEV_SINK_RE.test(path)) continue;
1042 targets.push(path);
1043 }
1044 return targets;
1045}
1046
1047/**
1048 * The untracked files a command line would stage, as pathspecs: "all" for
1049 * `git add -A|.|-u`, a list for named paths, and null when the line stages
1050 * no untracked file at all (`git commit -a` takes tracked changes only).
1051 *
1052 * Scanning every untracked file instead denied a commit over a key in a
1053 * stray log the line never touched.
1054 */
1055function stagedPathspecs(skel: string): string[] | "all" | null {
1056 const add = new RegExp(
1057 String.raw`\bgit\s+` + GIT_OPTS + String.raw`(?:add|stage)` + GIT_VERB_END + String.raw`([^;&|\n]*)`,
1058 "g",
1059 );
1060 const specs: string[] = [];
1061 let staging = false;
1062 for (const m of skel.matchAll(add)) {
1063 staging = true;
1064 for (const arg of (m[1] ?? "").trim().split(/\s+/).filter(Boolean)) {
1065 if (/^(?:-A|--all|-u|--update|--|\.)$/.test(arg)) return "all";
1066 if (arg.startsWith("-")) continue;
1067 specs.push(arg.replace(/^['"]|['"]$/g, ""));
1068 }
1069 }
1070 if (staging) return specs.length > 0 ? specs : "all";
1071
1072 // No `git add`: a commit's own pathspecs, or none for `-a`.
1073 const commit = new RegExp(
1074 String.raw`\bgit\s+` + GIT_OPTS + String.raw`commit` + GIT_VERB_END + String.raw`([^;&|\n]*)`,
1075 ).exec(skel);
1076 if (!commit) return null;
1077 const args = (commit[1] ?? "").trim().split(/\s+/).filter(Boolean);
1078 const paths: string[] = [];
1079 let all = false;
1080 for (let i = 0; i < args.length; i++) {
1081 const a = args[i];
1082 if (/^-[a-zA-Z]*a[a-zA-Z]*$/.test(a) || a === "--all") all = true;
1083 if (/^(?:-m|--message|-c|-C|--reuse-message|--reedit-message|--author|--date|--file|-F)$/.test(a)) {
1084 i++;
1085 continue;
1086 }
1087 if (a.startsWith("-")) continue;
1088 paths.push(a.replace(/^['"]|['"]$/g, ""));
1089 }
1090 if (paths.length > 0) return paths;
1091 return all ? null : "all";
1092}
1093
1094const MAX_UNTRACKED_SCAN = 50;
1095
1096const GIT_DESTRUCTIVE_RE = new RegExp(
1097 String.raw`\bgit\s+` +
1098 GIT_OPTS +
1099 String.raw`(?:(commit|push|merge|rebase|force-push)\b|reset\s+--hard\b|(?:checkout|restore)\s+(?:--\s+)?[.*]|checkout\s+--\s)`,
1100);
1101
1102/**
1103 * The branch a destructive git command in this line will run on, when the
1104 * line switches branch first: `git checkout main && git merge feat`. The
1105 * gate runs before the line does, when git still reports the branch it
1106 * started on — so without this, one line walked onto a protected branch
1107 * and merged there unchecked. The last switch before the destructive
1108 * command wins; a path checkout (`checkout main -- file`) is not a switch.
1109 */
1110function lineSwitchTarget(skel: string): string | null {
1111 const at = skel.search(GIT_DESTRUCTIVE_RE);
1112 if (at <= 0) return null;
1113 const switchRe = new RegExp(
1114 String.raw`\bgit\s+` + GIT_OPTS + String.raw`(?:checkout|switch)` + GIT_VERB_END + String.raw`([^;&|\n]*)`,
1115 "g",
1116 );
1117 let target: string | null = null;
1118 for (const m of skel.slice(0, at).matchAll(switchRe)) {
1119 const args = (m[1] ?? "").trim().split(/\s+/).filter(Boolean);
1120 if (args.includes("--")) continue;
1121 const named = args.findIndex((a) => /^-[bBcC]$/.test(a) || a === "--orphan");
1122 const name = named >= 0 ? args[named + 1] : args.find((a) => !a.startsWith("-"));
1123 if (name) target = name;
1124 }
1125 return target;
1126}
1127
1128// Destructive non-git bash commands — soft warning
1129const DESTRUCTIVE_BASH_RE =
1130 /(?:rm\s+(?:-[^\s]*r[^\s]*\s|--recursive\s)|chmod\s+(?:-R\s+)?777\s|curl\s[^|]*\|\s*(?:sudo\s+)?(?:bash|sh|zsh)|wget\s[^|]*\|\s*(?:sudo\s+)?(?:bash|sh|zsh)|dd\s+if=|mkfs\.|>\s*\/dev\/sd)/;
1131
1132// Secret/credential patterns — hard gate on git commit
1133// Values that are obviously not a real credential, so an unquoted
1134// assignment carrying one is not treated as a leak.
1135const SECRET_PLACEHOLDER_RE =
1136 /^(?:[*x.]{3,}|<[^>]*>|\$\{?[A-Za-z_]|\{\{|%[A-Za-z_]|your[-_]?|changeme|example|placeholder|redacted|dummy|sample|test|fake|none|null|true|false)/i;
1137
1138// A value that reads a secret rather than being one: `process.env.X`,
1139// `os.environ["X"]`, `getenv(...)`, `config.token`. Any dotted identifier
1140// or call is code, not a credential.
1141const SECRET_REFERENCE_RE = /^[A-Za-z_$][\w$]*(?:\.[\w$]+|\[|\()/;
1142
1143const SECRET_PATTERNS: RegExp[] = [
1144 // AWS long-lived, temporary (ASIA) and the other documented prefixes.
1145 /(?:AKIA|ASIA|ABIA|ACCA)[0-9A-Z]{16}/,
1146 // OpenAI legacy and project keys. `+` is in the lead-in class because
1147 // this runs against diff output, where every added line starts with one.
1148 /(?:^|[\s'"=:+])sk-[a-zA-Z0-9]{20,}/m,
1149 /(?:^|[\s'"=:+])sk-proj-[a-zA-Z0-9_-]{20,}/m,
1150 /gh[pousr]_[a-zA-Z0-9]{36}/,
1151 /github_pat_[a-zA-Z0-9_]{22,}/,
1152 /glpat-[a-zA-Z0-9\-_]{20,}/,
1153 /xox[abprs]-[a-zA-Z0-9-]{10,}/,
1154 /-----BEGIN\s+(?:RSA\s+|EC\s+|DSA\s+|OPENSSH\s+|ENCRYPTED\s+|PGP\s+)?PRIVATE\s+KEY/,
1155 // Quoted assignments.
1156 /(?:password|passwd|pwd)\s*[:=]\s*['"][^'"]{8,}['"]/i,
1157 /(?:api[_-]?key|apikey|secret[_-]?key|auth[_-]?token)\s*[:=]\s*['"][^'"]{12,}['"]/i,
1158];
1159
1160// Unquoted assignments — the shape a committed `.env` has, which the
1161// quoted patterns above could never match. Checked separately so the
1162// value can be tested against SECRET_PLACEHOLDER_RE first.
1163const BARE_SECRET_ASSIGN_RE =
1164 /(?:^|[\s+])(?:[A-Za-z_][A-Za-z0-9_]*_)?(?:password|passwd|pwd|secret|api[_-]?key|apikey|access[_-]?key|auth[_-]?token|token|credential)s?\s*[:=]\s*([^\s'"#]{8,})/gim;
1165
1166/**
1167 * True when the added lines of a diff carry something credential-shaped.
1168 * Kept as a function so the unquoted-assignment case can discount
1169 * placeholders without that logic living inside a regex.
1170 */
1171function containsSecret(text: string): RegExp | null {
1172 for (const pattern of SECRET_PATTERNS) {
1173 if (pattern.test(text)) return pattern;
1174 }
1175
1176 BARE_SECRET_ASSIGN_RE.lastIndex = 0;
1177 for (const m of text.matchAll(BARE_SECRET_ASSIGN_RE)) {
1178 const value = m[1];
1179 if (SECRET_PLACEHOLDER_RE.test(value)) continue;
1180 if (SECRET_REFERENCE_RE.test(value)) continue;
1181 return BARE_SECRET_ASSIGN_RE;
1182 }
1183
1184 return null;
1185}
1186
1187// Skill → lifecycle phase mapping
1188const SKILL_PHASE_MAP: Record<string, SessionState["currentPhase"]> = {
1189 brainstorming: "brainstorming",
1190 "proctor:brainstorming": "brainstorming",
1191 "writing-plans": "planning",
1192 "proctor:writing-plans": "planning",
1193 "test-driven-development": "implementing",
1194 "proctor:test-driven-development": "implementing",
1195 "subagent-driven-development": "implementing",
1196 "proctor:subagent-driven-development": "implementing",
1197 "executing-plans": "implementing",
1198 "proctor:executing-plans": "implementing",
1199 "systematic-debugging": "implementing",
1200 "proctor:systematic-debugging": "implementing",