SLOPSHOPPER

Proctor

Hook-enforced development discipline for Claude Code. Skills teach methodology; hooks enforce compliance mechanically (git gates, planning mode, step budgets…

newbandguardpromptmodelprocess
★ 2v3.0.0MITupdated 2026-09-22nalyk/nalyk-skills/plugins/proctor
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · proctor
› fix the failing auth test and add an audit log call ● proctor: Proctor active (tests: npm test | branch: feat/auth-refresh | protected: main,master,production,release,auth-refresh) ● proctor: Proctor: ✗ tests failing (exit 1) — git commit blocked until fixed. ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(rm -rf build && git push --force origin main) ⎿ Denied by proctor: Proctor gate: tests are failing (exit 1). Last output: src/auth.test.ts: ✓ refreshes ex ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

Proctor

Hook-enforced development discipline for Claude Code.

Skills teach methodology. Hooks enforce compliance. Store survives compaction.

What Problem This Solves

Agent development plugins today rely on prose instructions the LLM reads and decides to follow. This works 80-90% of the time. It fails exactly when it matters most — under context pressure, after compaction, in long sessions, when the LLM rationalizes past the rules.

Proctor is the first development discipline plugin where critical rules are enforced by code, not compliance. The skills teach why. The hooks enforce when. The store guarantees what survives.

Requirements

  • Claude Code CLI ≥ 2.1.260
  • Function hooks enabled: CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1
  • A git repository (for git gates and branch protection)

Installation

Option 1: Marketplace (recommended)

From inside Claude Code:

/plugin marketplace add nalyk/nalyk-skills
/plugin install proctor@nalyk-skills

Or from the terminal:

claude plugin marketplace add nalyk/nalyk-skills
claude plugin install proctor@nalyk-skills

Option 2: Project-scoped (shared with your team)

claude plugin install proctor@nalyk-skills --scope project

This writes to .claude/settings.json which you commit to version control. Teammates get Proctor when they trust the project folder.

Option 3: Team settings (auto-install for all team members)

Add to your project's .claude/settings.json:

{
  "extraKnownMarketplaces": {
    "nalyk-skills": {
      "source": {
        "source": "github",
        "repo": "nalyk/nalyk-skills"
      }
    }
  }
}

Option 4: Local development

git clone https://github.com/nalyk/nalyk-skills.git
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir ./nalyk-skills/plugins/proctor

Enable function hooks

Proctor's enforcement layer requires the function hooks runtime. Set the environment variable before launching:

CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude

Or export it in your shell profile:

echo 'export CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1' >> ~/.bashrc

Without the flag, the skills still load and work as prose guidance (like any other skills plugin), but the hooks — git gates, branch protection, SDD dashboard, state persistence — do not fire.

What the Hooks Enforce

Hard Gates (deny — the agent cannot proceed)

GateWhat it blocksWhat it requires
Test evidencegit commit, git pushFresh passing test run in this session — unless the project has no test suite, the change is prose-only, or you said proctor: no tests
Test freshnessgit commit, git pushTest run within 5 minutes (configurable)
Test passinggit commit, git pushLast test run exit code 0
Proven test statusgit commit, git pushThe last test run finished in the foreground and its own exit status reached the result — not hidden by a following `\, ;, \\ or & (set -o pipefail / set -e` count), not sent to the background, not moved there by the Bash timeout, not interrupted
Branch protectiongit commit/push/merge/rebase/reset on a protected branch — including a line that switches onto one first (git checkout main && git merge feat), and a push that writes one from elsewhere (git push origin HEAD:main, :main, --all, --mirror). The remote's default branch (origin/HEAD) is protected alongside the configured listFeature branch or explicit human consent
SDD mergegit merge while the SDD run has tasks unmarked or a fix round openFinish and mark the tasks, or proctor: sdd stop
Fix-round capAn Agent dispatch past the cap — a round the ledger recorded, or one the dispatch's prompt names for the current taskAdjudicate with Ruling: lines, complete the task, or proctor: sdd stop
Step budgetBash, Write, Edit or NotebookEdit once an SDD task hits 100% of its step budgetComplete the task, proctor: budget extend, or proctor: sdd stop
Time budgetThe same four tools once an SDD task hits 100% of its time budgetSame three exits
Planning mode (Bash)Shell writes to implementation files during a design phaseA design doc, or exit planning mode
Secret detectiongit commit with staged credentialsNo AWS/OpenAI/GitHub/GitLab/Slack tokens, private keys, or password-shaped assignments (quoted or not) among the added lines — the commit that removes a leaked key is not the one to block

Soft Enforcers (context injection — the agent is reminded)

An SDD run works through the whole plan in one turn, so what happens during it is reported on the result of the tool call that caused it: task progress when a ledger line is written, budget warnings on the call that crosses 80% or reaches 100%, model advice on the Agent result. A note produced when a turn ends (the watchdog, context pressure, a done-check from the answer) cannot reach the model then — a turn's result carries nothing the model reads — so it arrives with the next prompt.

Quiet mode (proctor: quiet on) suppresses the nudges: the watchdog, model selection, context pressure, the destructive-command and diff-size warnings, and the step-aside and run-complete lines. Notes that carry the run's state — task progress, budgets, the fix-round cap, the done-check, the ruling aggregation — still arrive, and hard gates always enforce.

EnforcerWhen it firesWhat it says
Skill watchdog4+ turns without invoking a skill"Check if brainstorming, TDD, debugging, or review applies" (once)
Model selectionAgent spawn without explicit model during SDD"Consider a cheaper model for mechanical tasks"
Fix-round capRound N of 5 reached"Decide on each open finding — skip debatable ones, resolve critical ones"
EscalationRound 4-5 (the last two of the cap) dispatched on the model the stuck implementer used"Try a more capable model"
Step budget80% / 100% of tool call limit per taskWarning at 80% (once), wrap-up at 100% (once per task), on the tool result
Time budget80% / 100% of wall-clock limit per taskWarning at 80% (once), wrap-up at 100% (once per task)
Context pressureTurn 50, 70, 90"Progress preserved automatically — focus on current task"
Ruling aggregationThe last task marked complete; the finishing skill loadedFull list of rulings and deferred minors
SDD done-checkThe finishing skill loaded, or an answer claiming the finishEach unmet condition: tasks unmarked, fix round open, tests missing/failing/stale/unproven
Destructive commandrm -rf, chmod 777, `curl\bash`, etc.Shows the actual command — "verify this is intentional"
Diff size>500 lines staged (pre-commit)"Consider splitting into smaller commits"
Test hintNo tests run this sessionShows detected test command with actionable next step
Phase indicatorSkill invocation changes lifecycle phaseShows current phase in context
Task advanceSDD task marked complete"Next: Task N" with budget reset notification

Infrastructure (invisible — the agent doesn't manage these)

FeatureWhat it does
SDD trackingA run starts when subagent-driven-development or executing-plans loads (or is announced); the plan's ### Task N headings size it; the ledger (progress.md) drives it — Plan: <path> — <N> tasks, Task N: complete, Task N: fix round M — approach: …, Ruling: …, minor (deferred): …, Task N: added, read as they are written through Write, Edit or the shell, each counted once however often the ledger is rewritten
SDD state persistenceTask completion, fix rounds, rulings, failed approaches tracked in $.store
Live statusA [PROCTOR] block rides on every prompt as context the model reads: test verdict and age, planning mode, phase, SDD progress
Compaction recovery[PROCTOR — SDD STATE] is one of the conversation's context blocks, which the engine re-reads at compaction — with its "DO NOT REDO" section and the last test run's failure output
SDD session recoveryActive SDD state detected and resumed on session restart
Failed approach trackingFix round descriptions captured and injected post-compaction to prevent retries
Task completion evidenceCompletion evidence recorded per task for audit trail
Progress dashboardStatus bar above prompt during SDD
Test run trackingRecords every test execution for the git gate; a passing command is remembered for the project's next session
Branch trackingDetects branch changes for protection enforcement
Agent countingTracks agents spawned for dashboard and diagnostics
Commit attributionAppends task references and test evidence to commit messages
Phase lifecycleTracks idle → brainstorming → planning → implementing → reviewing → finishing
Quality metricsCross-session counters: commits, gate denials, gates passed, fix rounds, test runs, autonomy rate
Ruling persistenceLast 20 rulings preserved across sessions

Operator Commands

CommandWhat it does
proctor: statusFull status dashboard: phase, branch, quiet mode, SDD state, test evidence, quality metrics
proctor: show traceLast 25 structured trace events with timestamps
proctor: allow <branch>Grant consent for protected branch operations
proctor: approve designExit planning mode
proctor: no testsStand the test gate down for this session — for projects that genuinely have no suite
proctor: budget extendGrant the current SDD task one more full step and time budget
proctor: quiet onSuppress soft warnings (hard gates still enforce)
proctor: quiet offRe-enable all warnings
proctor: checkPre-flight gate status: test evidence, branch protection, planning mode, secret scan
proctor: diagnoseSelf-analysis: gate autonomy rate, fix round patterns, recommendations
proctor: tasks NUpdate SDD total task count (scope change)
sdd done / proctor: sdd stopDeactivate SDD session

Skills

Proctor includes 14 development skills. Each is a methodology document the agent reads and follows. The hooks enforce the critical gates that skills alone cannot guarantee.

SkillPurpose
using-proctorBootstrap — establishes skill discovery and hook awareness
brainstormingTurn ideas into designs (spike/bounded/architectural paths)
test-driven-developmentRed-green-refactor cycle. Git gate blocks commits without tests
systematic-debuggingRoot cause before fixes. Four-phase investigation
verification-before-completionEvidence before claims. Git gate enforces mechanically
subagent-driven-developmentExecute plans with fresh agents per task. Dashboard + state persistence
executing-plansExecute plans inline (cheaper). Same enforcement as SDD
writing-plansCreate implementation plans from specs
requesting-code-reviewDispatch reviewers with proper packages
receiving-code-reviewEvaluate feedback technically, not performatively
finishing-a-development-branchVerify → present options → execute → clean up
using-git-worktreesWorkspace isolation. Branch protection hooks enforce
dispatching-parallel-agentsIndependent concurrent tasks
writing-skillsTDD applied to skill creation

Configuration

Plugin options (via plugin.json userConfig):

OptionDefaultDescription
protectedBranches["main","master","production","release"]Branches protected from destructive git ops
executableDocPatterns[]Globs for prose-shaped files that are really behaviour, e.g. runbooks/**; a glob without / (*.runbook.md) matches at any depth
testFreshnessMinutes5How many minutes before test evidence expires
watchdogTurnThreshold4Turns without a skill before the watchdog fires
fixRoundCap5Maximum fix-loop rounds in SDD
stepBudgetPerTask100Tool call limit per SDD task (warn 80%, block 100%)
timeBudgetPerTaskMinutes30Wall-clock limit per SDD task in minutes (warn 80%, block 100%)

When the test gate stands down

A project with no test suite could otherwise never commit, so the gate steps aside — visibly, with a line in the transcript — when:

  • no test marker file exists anywhere in the repo (a docs, notes or config repo: there is no suite to run) — a package.json whose test script is missing or the npm init stub is not a marker, or
  • the change is inert prose only, or
  • you said proctor: no tests this session.

"The change" is whatever the operation actually sends: for a commit, the working tree; for a push, the commits the upstream does not have. A push is never excused by an unrelated edit sitting in the working tree, and when there is no upstream to compare against the gate enforces rather than guesses. A merge is never excused on prose grounds at all.

Every other gate keeps enforcing regardless: branch protection, secret detection and planning mode are untouched by this.

"Inert prose" is narrower than "a .md file". These stay behaviour and keep the gate enforcing:

KindExamples
Agent and tool instructionsSKILL.md, CLAUDE.md, AGENTS.md, anything under .claude/, .cursor/, .github/
Runbooks and playbooksRUNBOOK.md, runbooks/, playbooks/
Anything a test readstests/, spec/, fixtures/, testdata/, __snapshots__/, e2e/, golden/
A plugin's own behaviourcommands/, agents/, skills/, prompts/, references/, templates/
Prose a toolchain executesevery .md/.rst/.qmd in a repo holding book.toml, _quarto.yml, runme.yaml, mkdocs.yml, jupytext.toml or a Docusaurus config
Whatever you declareexecutableDocPatterns

Quiet mode is toggled at runtime via proctor: quiet on/off — it suppresses soft warnings while hard gates continue to enforce. This is session-scoped and does not persist across sessions.

Architecture

proctor/
├── .claude-plugin/plugin.json     # Plugin manifest
├── hooks/
│   ├── hooks.json                 # Module registration
│   └── proctor.tsx                # All hook registrations
├── skills/                        # 14 methodology skills
│   ├── using-proctor/
│   ├── brainstorming/
│   ├── test-driven-development/
│   ├── systematic-debugging/
│   ├── verification-before-completion/
│   ├── subagent-driven-development/
│   ├── executing-plans/
│   ├── writing-plans/
│   ├── requesting-code-review/
│   ├── receiving-code-review/
│   ├── finishing-a-development-branch/
│   ├── using-git-worktrees/
│   ├── dispatching-parallel-agents/
│   └── writing-skills/
├── README.md
└── package.json

The Three Layers

Teaching layer (skills): Prose documents that explain methodology. Rationalization tables, core principles, checklists. The agent reads these and follows them. This is what existing plugins do.

Enforcement layer (hooks): TypeScript middleware that fires on engine events. Hard gates deny operations that violate rules. Soft enforcers inject reminders when patterns suggest non-compliance. This is what only function hooks can do.

Infrastructure layer (store): Persistent state that survives compaction. Task progress, test evidence, rulings, fix-round counts. Injected into context automatically so the agent never forgets where it is.

$.store is one namespace shared by every session on the machine, so session state, test evidence, the SDD run and the trace log are each filed under the project root — the repo git rev-parse --show-toplevel reports, not the cwd of the moment. A second session in another repo no longer overwrites the first one's branch consents, quiet mode or evidence, and cd-ing into a subdirectory does not make a session's own test run look foreign. Two sessions in the same repo still share one record: they share the branch and the consents that go with it, so the sharing is the accurate reading.

Every write goes through a single per-key queue. Counters are the one exception to "state is load-bearing": they are telemetry, they are written best-effort, and a counter that cannot be written can never stop a gate from firing.

Why Each Layer Exists

The teaching layer handles nuance — when to use TDD, how to classify a task as spike vs architectural, what makes a good ruling. Hooks cannot express this.

The enforcement layer handles compliance — did you run tests before committing, are you on a protected branch, have you invoked a skill. Prose cannot guarantee this.

The infrastructure layer handles memory — what tasks are complete, what rulings were made, when was the last test run. Neither prose nor hooks alone can persist state across compaction.

How It Differs From Superpowers

Proctor is inspired by Superpowers and shares its philosophy. Source: nalyk/proctor.

The shared philosophy: TDD, systematic debugging, verification before completion, subagent-driven development with ledger-based recovery.

The differences:

AspectSuperpowersProctor
Platform9+ harnesses (universal)Claude Code CLI only
EnforcementProse instructionsFunction hooks (hard gates + soft nudges)
StateLedger files (agent-managed)$.store + ledger files (hook-managed + agent-managed)
Compaction recoveryAgent must read ledgerHook injects state into context automatically
VerificationProse ruleGit gate (mechanical denial)
Branch protectionProse ruleHook gate (mechanical denial)
Progress visibilityPost-hoc (diagnosing-superpowers)Real-time dashboard
Skill invocationBootstrap re-readingHook watchdog
Fix-round trackingAgent countingHook state machine
Ruling aggregationAgent scanning ledgerHook aggregation at session end

Superpowers is the right choice for Codex, Cursor, Gemini CLI, and other harnesses where function hooks are not available. Proctor is for Claude Code CLI sessions where mechanical enforcement matters.

Testing

make test at the repo root runs four layers:

  • claude plugin validate — the manifest and hooks module as the engine's loader reads them
  • tests/*.test.mjs — the pure helpers, a feature ratchet, and the hooks against a small fake engine
  • tests/engine/*.test.ts — claude plugin test: every feature driven through the real engine, with the host (git, files, store, the tools' own results) answered beneath the plugin by tests/engine/world.ts. The fake engine could only ever agree with Proctor's own reading of the API; this layer is where five delivery channels that never reached the model were found.
  • tests/skill-contract.test.mjs — the skills against the hooks: every ledger line a skill teaches parses as the hooks read it, every command taught is answered, every mechanism a skill claims exists
  • tests/proctor-configured/ (repo root) — the same module loaded with every userConfig option set away from its default, through the engine's real options pipeline

tests/COVERAGE.md maps every feature to the tests that prove it.

An end-to-end check against a live model loads the working tree in place of the installed copy:

claude -p --plugin-dir plugins/proctor \
  --settings '{"enabledPlugins":{"proctor@nalyk-skills":false}}' \
  --output-format stream-json --verbose "..."

Status

Proctor requires the function hooks runtime, which is behind CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 and has not officially shipped. The runtime exists in Claude Code ≥ 2.1.260. Build against it to learn the shape, not to run production on it until Anthropic ships the feature.

License

MIT

Source 1 files
hooks/proctor.tsx 4008 lines
1// Proctor — Hook-enforced development discipline for Claude Code
2// EARLY ACCESS: requires CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1
3//
4// Skills teach methodology. Hooks enforce compliance. Store survives compaction.
5//
6// Four layers:
7//   Teaching       (skills)  → prose methodology the agent reads and follows
8//   Enforcement    (hooks)   → TypeScript middleware that gates or nudges
9//   Infrastructure ($)       → persistent state that survives compaction
10//   Observability  (tracing) → structured event log for audit trail
11//
12// Run /plugin-types to regenerate the declarations for your build.
13
14import type { Register } from "claude-code";
15
16// ─────────────────────────────────────────────────────────────────────
17//  State schemas
18// ─────────────────────────────────────────────────────────────────────
19
20interface TestEvidence {
21  command: string;
22  timestamp: number;
23  exitCode: number;
24  tailOutput: string;
25  /** The run reported success without proving the suite passed: its
26   *  exit status was hidden by what followed it on the command line
27   *  (`| tail`, `; echo`, `|| true`, `&`), or it had not finished — sent
28   *  to the background, timed out into it, or interrupted. Recorded with
29   *  exitCode -1, and `unproven` says which. */
30  masked?: boolean;
31  unproven?: string;
32  /** Project the run happened in. Evidence from another project must not
33   *  unblock this one's commit gate. */
34  cwd: string | null;
35}
36
37interface Ruling {
38  task: number;
39  text: string;
40  costIfWrong: string;
41  phase: "preflight" | "fix-loop" | "final";
42}
43
44interface DeferredMinor {
45  task: number;
46  finding: string;
47}
48
49interface SDDState {
50  active: boolean;
51  /** Project the run belongs to. A run does not follow the store into
52   *  another repo. */
53  cwd: string | null;
54  plan: string;
55  startedAt: number;
56  totalTasks: number;
57  currentTask: number;
58  completedTasks: number[];
59  currentFixRound: number;
60  totalFixRounds: number;
61  totalAgents: number;
62  lastImplementerModel: string | null;
63  rulings: Ruling[];
64  deferredMinors: DeferredMinor[];
65  toolCallsThisTask: number;
66  totalToolCalls: number;
67  taskStartedAt: number;
68  stepWarned80: boolean;
69  stepWarned100: boolean;
70  timeWarned80: boolean;
71  timeWarned100: boolean;
72  failedApproaches: string[];
73  completedEvidence: Record<number, string>;
74  /** Signals already absorbed, by key. The ledger is rewritten whole and
75   *  repeated in the final answer, so every signal arrives more than once;
76   *  a fix round or ruling must still be counted once. */
77  seen?: string[];
78  /** The last task completed and the run ended itself, but nothing has
79   *  checked it against the finish conditions yet — the finishing skill,
80   *  a finishing answer or a merge still has to. */
81  pendingDoneCheck?: boolean;
82}
83
84interface SessionState {
85  startedAt: number;
86  skillInvoked: boolean;
87  lastSkillName: string | null;
88  watchdogNudgeSent: boolean;
89  testCommand: string | null;
90  branch: string | null;
91  /** The remote's default branch (origin/HEAD), protected alongside the
92   *  configured list: a repo whose trunk is `develop` or `trunk` is not
93   *  left unguarded because it is not called main. */
94  defaultBranch?: string | null;
95  isWorktree: boolean;
96  protectedBranches: string[];
97  branchConsents: Record<string, boolean>;
98  turnsSinceSkill: number;
99  agentsSpawned: number;
100  turnCount: number;
101  planningMode: boolean;
102  planningSkill: string | null;
103  hasTestInfrastructure: boolean;
104  testsAcknowledgedAbsent: boolean;
105  executableDocs: boolean;
106  currentPhase:
107    | "idle"
108    | "brainstorming"
109    | "planning"
110    | "implementing"
111    | "reviewing"
112    | "finishing";
113  quietMode: boolean;
114  /** Notes for the model produced after its turn ended — the watchdog,
115   *  task progress, budget warnings, the done-check. `turn.complete`
116   *  cannot reach the model, so they wait here for the next prompt. */
117  pendingNotes?: string[];
118}
119
120interface SessionHistory {
121  lastTestCommand: string | null;
122  /** The last passing test command per project root. `lastTestCommand`
123   *  was one value for the whole machine, so a command learned in one repo
124   *  became another repo's required suite. */
125  learnedTestCommands: Record<string, string>;
126  projectPath: string | null;
127  skillUsage: Record<string, number>;
128  sessionsCount: number;
129  recentRulings: Array<{ text: string; ts: number }>;
130  qualityMetrics: {
131    totalCommits: number;
132    gateDenials: number;
133    gatesPassed: number;
134    fixRounds: number;
135    testsRun: number;
136  };
137}
138
139interface TraceEvent {
140  ts: number;
141  kind: string;
142  detail: string;
143}
144
145// ─────────────────────────────────────────────────────────────────────
146//  Store helpers — $.store is JSON-backed KV under
147//  ~/.claude/plugins/store/, persisting across sessions.
148//  Module-level variables are session-scoped.
149// ─────────────────────────────────────────────────────────────────────
150
151// The per-project keys carry a version suffix: their shape changed from a
152// bare record to a book keyed by project root, and a stale record read as
153// a book would be read as an empty one anyway. History stays unversioned —
154// `loadHistory` migrates it field by field, and its counters are the one
155// thing worth carrying forward.
156const KEYS = {
157  session: "proctor:session:v3",
158  test: "proctor:test-evidence:v3",
159  sdd: "proctor:sdd-state:v3",
160  history: "proctor:history",
161  trace: "proctor:trace:v3",
162} as const;
163
164const DEFAULT_HISTORY: SessionHistory = {
165  lastTestCommand: null,
166  learnedTestCommands: {},
167  projectPath: null,
168  skillUsage: {},
169  sessionsCount: 0,
170  recentRulings: [],
171  qualityMetrics: {
172    totalCommits: 0,
173    gateDenials: 0,
174    gatesPassed: 0,
175    fixRounds: 0,
176    testsRun: 0,
177  },
178};
179
180const TRACE_CAP = 50;
181
182async function load<T>(
183  $: any,
184  key: string,
185  fallback: T,
186): Promise<T> {
187  const raw = await $.store.get(key);
188  if (!raw) return fallback;
189  try {
190    return JSON.parse(raw) as T;
191  } catch {
192    return fallback;
193  }
194}
195
196/** A stored counter that is usable as a number, whatever the store holds.
197 *  `undefined++` wrote NaN, which JSON stores as null, which the next
198 *  read had to cope with in turn. */
199function counter(value: unknown): number {
200  return typeof value === "number" && Number.isFinite(value) ? value : 0;
201}
202
203/**
204 * SessionHistory with every field the current code expects, merged over
205 * whatever the store holds. `load` returns a stored object verbatim, so a
206 * history written by an older version is missing fields added since — and
207 * `hist.qualityMetrics.x++` on it throws. A throwing hook is skipped
208 * entirely, which silently disabled every gate it contained, because only
209 * the DENIAL paths mutate those counters outside a try.
210 *
211 * Every write goes through `mutateHistory`, so this is the only door.
212 */
213async function loadHistory($: any): Promise<SessionHistory> {
214  const stored = await load<Partial<SessionHistory> | null>(
215    $,
216    KEYS.history,
217    null,
218  );
219  const metrics = (stored?.qualityMetrics ?? {}) as Partial<
220    SessionHistory["qualityMetrics"]
221  >;
222
223  return {
224    ...structuredClone(DEFAULT_HISTORY),
225    ...(stored ?? {}),
226    skillUsage: { ...(stored?.skillUsage ?? {}) },
227    learnedTestCommands: { ...(stored?.learnedTestCommands ?? {}) },
228    recentRulings: [...(stored?.recentRulings ?? [])],
229    qualityMetrics: {
230      totalCommits: counter(metrics.totalCommits),
231      gateDenials: counter(metrics.gateDenials),
232      gatesPassed: counter(metrics.gatesPassed),
233      fixRounds: counter(metrics.fixRounds),
234      testsRun: counter(metrics.testsRun),
235    },
236  };
237}
238
239// One in-flight write per key. Every hook did load-whole-object, mutate
240// one field, save-whole-object, so two hooks running for the same
241// assistant message (three parallel Edits, or a Bash and a Read) both read
242// the same snapshot and the second write erased the first. Lost that way:
243// step-budget increments, a branch change, and planningMode being set.
244//
245// A queue only helps if every writer uses it: `agent.spawn`, `skill.prompt`,
246// `turn.complete`, `trace` and the branch tracker each did their own raw
247// load+save straight past it, so the increments they raced with were lost
248// exactly as before. Nothing below calls `save` outside this chain.
249const writeQueues = new Map<string, Promise<unknown>>();
250
251/** Run `job` with no other write to `key` interleaved. */
252function enqueue<T>(key: string, job: () => Promise<T>): Promise<T> {
253  const queued = (writeQueues.get(key) ?? Promise.resolve()).then(job);
254  // Keep the chain alive even if one link rejects.
255  writeQueues.set(key, queued.catch(() => undefined));
256  return queued;
257}
258
259/**
260 * Read, mutate and write a key with no other mutation interleaved.
261 * `fn` may mutate its argument in place (return nothing) or return a
262 * replacement — `null` included. `fn(x) ?? x` read a returned null as
263 * "mutated in place", so every `put…(null)` was a no-op: session.start's
264 * reset of the test evidence kept the last session's run standing.
265 *
266 * The fallback is cloned before `fn` sees it: `load` returns the fallback
267 * itself when the key is empty, so a shared constant handed in here
268 * (DEFAULT_HISTORY was) gets mutated in place and every later reader
269 * inherits the mutation for the life of the module.
270 */
271async function mutate<T>(
272  $: any,
273  key: string,
274  fallback: T,
275  fn: (value: T) => T | void,
276): Promise<T> {
277  return enqueue(key, async () => {
278    const current = await load<T>($, key, structuredClone(fallback));
279    const out = fn(current);
280    const updated = (out === undefined ? current : out) as T;
281    await save($, key, updated);
282    return updated;
283  });
284}
285
286/**
287 * Counters are telemetry. A gate must never fail to fire because a number
288 * could not be written, so this swallows its own errors and reads through
289 * `loadHistory` rather than `load` — the two halves of the bug that let a
290 * stale stored history throw inside a denial path and skip the denial.
291 */
292async function mutateHistory(
293  $: any,
294  fn: (history: SessionHistory) => void,
295): Promise<void> {
296  try {
297    await enqueue(KEYS.history, async () => {
298      const history = await loadHistory($);
299      fn(history);
300      await save($, KEYS.history, history);
301    });
302  } catch {
303    // Best-effort by design — see above.
304  }
305}
306
307async function save($: any, key: string, value: unknown): Promise<void> {
308  await $.store.set(key, JSON.stringify(value));
309}
310
311// ─────────────────────────────────────────────────────────────────────
312//  Project scoping — $.store is one global namespace shared by every
313//  session on the machine. Session state, test evidence, the SDD run and
314//  the trace log all belong to ONE project and ONE session; keeping them
315//  under a bare key meant a second pane's session.start overwrote them,
316//  taking this session's branch consents, quiet mode and test evidence
317//  with it. Each is now a book keyed by project root.
318// ─────────────────────────────────────────────────────────────────────
319
320interface Shelf<T> {
321  at: number;
322  value: T;
323}
324type Book<T> = Record<string, Shelf<T>>;
325
326const BOOK_CAP = 12;
327
328// The repo root, not the cwd: a session that cds into a subdirectory is
329// still the same project, and evidence recorded before the cd must not
330// turn foreign. Re-read only when the cwd moves.
331let rootCache: { cwd: string | null; root: string } | null = null;
332
333async function projectRoot($: any): Promise<string> {
334  let cwd: string | null = null;
335  try {
336    cwd = (await $.session.cwd()) ?? null;
337  } catch {
338    cwd = null;
339  }
340
341  if (rootCache && rootCache.cwd === cwd) return rootCache.root;
342
343  let root: string | null = null;
344  try {
345    const res = await $.process.run(["git", "rev-parse", "--show-toplevel"]);
346    const top = (res.stdout ?? "").trim();
347    if (res.exitCode === 0 && top) root = top;
348  } catch {
349    // Not a repo, or git is unavailable — fall back to the cwd.
350  }
351
352  const resolved = root ?? cwd ?? "unknown";
353  rootCache = { cwd, root: resolved };
354  return resolved;
355}
356
357async function loadScoped<T>($: any, key: string, fallback: T): Promise<T> {
358  const book = await load<Book<T>>($, key, {});
359  const root = await projectRoot($);
360  const shelf = book?.[root];
361  return shelf && "value" in shelf ? shelf.value : fallback;
362}
363
364async function mutateScoped<T>(
365  $: any,
366  key: string,
367  fallback: T,
368  fn: (value: T) => T | void,
369): Promise<T> {
370  const root = await projectRoot($);
371  let result = fallback;
372
373  await mutate<Book<T>>($, key, {}, (book) => {
374    const shelf = book[root];
375    const current =
376      shelf && "value" in shelf ? shelf.value : structuredClone(fallback);
377    const out = fn(current);
378    result = (out === undefined ? current : out) as T;
379    book[root] = { at: Date.now(), value: result };
380
381    // Projects come and go; the book must not grow forever.
382    const roots = Object.keys(book);
383    if (roots.length > BOOK_CAP) {
384      roots
385        .sort((a, b) => (book[a]?.at ?? 0) - (book[b]?.at ?? 0))
386        .slice(0, roots.length - BOOK_CAP)
387        .forEach((stale) => {
388          delete book[stale];
389        });
390    }
391  });
392
393  return result;
394}
395
396async function trace(
397  $: any,
398  kind: string,
399  detail: string,
400): Promise<void> {
401  try {
402    await mutateScoped<TraceEvent[]>($, KEYS.trace, [], (log) => {
403      log.push({ ts: Date.now(), kind, detail });
404      if (log.length > TRACE_CAP) log.splice(0, log.length - TRACE_CAP);
405    });
406  } catch {
407    // Tracing is best-effort — never block the hook chain
408  }
409}
410
411// The three per-project records, always read and written through the book.
412const loadSession = ($: any) =>
413  loadScoped<SessionState | null>($, KEYS.session, null);
414
415const mutateSession = ($: any, fn: (session: SessionState) => void) =>
416  mutateScoped<SessionState | null>($, KEYS.session, null, (session) => {
417    if (session) fn(session);
418  });
419
420const putSession = ($: any, session: SessionState | null) =>
421  mutateScoped<SessionState | null>($, KEYS.session, null, () => session);
422
423/** Hold notes for the model until the next prompt carries them. */
424const queueNotes = ($: any, notes: string[]) =>
425  mutateSession($, (session) => {
426    session.pendingNotes = [...(session.pendingNotes ?? []), ...notes];
427  });
428
429/** Take the held notes, leaving none behind. */
430async function drainNotes($: any): Promise<string[]> {
431  let notes: string[] = [];
432  await mutateSession($, (session) => {
433    notes = session.pendingNotes ?? [];
434    session.pendingNotes = [];
435  });
436  return notes;
437}
438
439const loadSDD = ($: any) => loadScoped<SDDState | null>($, KEYS.sdd, null);
440
441const mutateSDD = ($: any, fn: (sdd: SDDState) => void) =>
442  mutateScoped<SDDState | null>($, KEYS.sdd, null, (sdd) => {
443    if (sdd) fn(sdd);
444  });
445
446const putSDD = ($: any, sdd: SDDState | null) =>
447  mutateScoped<SDDState | null>($, KEYS.sdd, null, () => sdd);
448
449const loadEvidence = ($: any) =>
450  loadScoped<TestEvidence | null>($, KEYS.test, null);
451
452const putEvidence = ($: any, evidence: TestEvidence | null) =>
453  mutateScoped<TestEvidence | null>($, KEYS.test, null, () => evidence);
454
455// ─────────────────────────────────────────────────────────────────────
456//  Test command detection — language-agnostic file heuristics
457// ─────────────────────────────────────────────────────────────────────
458
459const TEST_PATTERNS: Array<{ file: string; command: string }> = [
460  { file: "package.json", command: "npm test" },
461  { file: "Cargo.toml", command: "cargo test" },
462  { file: "pyproject.toml", command: "pytest" },
463  { file: "setup.py", command: "pytest" },
464  { file: "go.mod", command: "go test ./..." },
465  { file: "Makefile", command: "make test" },
466  { file: "Gemfile", command: "bundle exec rspec" },
467  { file: "mix.exs", command: "mix test" },
468  { file: "build.gradle", command: "./gradlew test" },
469  { file: "build.gradle.kts", command: "./gradlew test" },
470  { file: "pom.xml", command: "mvn test" },
471  { file: "CMakeLists.txt", command: "ctest" },
472  { file: "Rakefile", command: "rake test" },
473  { file: "deno.json", command: "deno test" },
474  { file: "bun.lockb", command: "bun test" },
475  { file: "composer.json", command: "vendor/bin/phpunit" },
476  { file: "phpunit.xml", command: "vendor/bin/phpunit" },
477  { file: "phpunit.xml.dist", command: "vendor/bin/phpunit" },
478  { file: "Package.swift", command: "swift test" },
479  { file: "pubspec.yaml", command: "dart test" },
480  { file: "build.zig", command: "zig build test" },
481  { file: "project.clj", command: "lein test" },
482  { file: "build.sbt", command: "sbt test" },
483  { file: "stack.yaml", command: "stack test" },
484  { file: "cabal.project", command: "cabal test" },
485];
486
487// Runners anchored to the START of a command segment. The old pattern
488// matched a bare `pytest|jest|vitest|...` anywhere, so `pip install
489// pytest`, `cat jest.config.js` or `echo pytest` were all recorded as
490// passing test runs and satisfied the commit gate with nothing run.
491const TEST_RUN_RE = new RegExp(
492  "^(?:" +
493    [
494      String.raw`npm\s+(?:run\s+[\w:.-]*test[\w:.-]*|test|t)\b`,
495      String.raw`(?:yarn|pnpm|bun)\s+(?:run\s+[\w:.-]*test[\w:.-]*|test)\b`,
496      String.raw`npx\s+(?:jest|vitest|mocha|playwright|cypress|ava|tap)\b`,
497      String.raw`(?:jest|vitest|mocha|ava|tap|cypress|playwright)\b`,
498      String.raw`deno\s+test\b`,
499      String.raw`(?:pytest|py\.test)\b`,
500      String.raw`python[0-9.]*\s+-m\s+(?:pytest|unittest)\b`,
501      String.raw`(?:poetry|uv|pipenv|hatch|rye|pdm)\s+run\s+\S*(?:pytest|test)\S*\b`,
502      String.raw`tox\b`,
503      String.raw`cargo\s+(?:test|nextest\s+run)\b`,
504      String.raw`go\s+test\b`,
505      String.raw`(?:bundle\s+exec\s+)?rspec\b`,
506      String.raw`rake\s+[\w:]*test[\w:]*\b`,
507      String.raw`(?:elixir\s+-S\s+)?mix\s+test\b`,
508      String.raw`make\s+[\w-]*test[\w-]*\b`,
509      String.raw`(?:\./)?gradlew?\s+[\w:]*test\b`,
510      String.raw`gradle\w*\s+[\w:]*test\b`,
511      String.raw`mvn\s+(?:-\S+\s+)*test\b`,
512      String.raw`dotnet\s+test\b`,
513      String.raw`swift\s+test\b`,
514      String.raw`ctest\b`,
515      String.raw`(?:vendor/bin/)?phpunit\b`,
516      String.raw`(?:dart|flutter)\s+test\b`,
517      String.raw`zig\s+build\s+test\b`,
518      String.raw`(?:lein|sbt|stack|cabal)\s+test\b`,
519      String.raw`nimble\s+test\b`,
520      String.raw`nim\s+c\s+-r\b`,
521      String.raw`busted\b`,
522      String.raw`bazel\s+test\b`,
523      // Node's built-in runner, and Claude Code's own plugin test runner.
524      String.raw`node\s+(?:-[-\w]+(?:=\S+)?\s+)*--test\b`,
525      String.raw`claude\s+plugin\s+test\b`,
526    ].join("|") +
527    ")",
528);
529
530/**
531 * True when the command actually runs a test suite. Splits on shell
532 * separators and checks each segment's FIRST word, after stripping
533 * leading env assignments and wrappers, so `cd x && npm test` counts and
534 * `grep -rn vitest src/` does not.
535 */
536function isTestRun(skel: string): boolean {
537  return skel
538    .split(/(?:&&|\|\||[;|&\n])+/)
539    .map((seg) =>
540      seg
541        .trim()
542        .replace(/^\(+\s*/, "")
543        .replace(/^(?:[A-Za-z_][A-Za-z0-9_]*=\S*\s+)*/, "")
544        .replace(/^(?:time|command|exec|nice|stdbuf\s+\S+)\s+/, ""),
545    )
546    .some((seg) => TEST_RUN_RE.test(seg));
547}
548
549/**
550 * True when the test run's own exit status cannot reach the tool result.
551 * The result carries one status for the whole line, so in `npm test |
552 * tail`, `npm test; echo done`, `npm test || true` or `npm test &` it is
553 * the last command's, and a failing suite reads as a pass. `&&` after the
554 * test is fine: a failure stops the chain and the line fails with it.
555 */
556function testStatusMasked(skel: string): boolean {
557  // Redirections carry an `&` that is not a separator: 2>&1, &>, >&2, |&.
558  const s = skel.replace(/\d*>&\d*|&>>?|<&\d*/g, " ").trim();
559  const parts = s.split(/(&&|\|\||\|&?|[;&\n])/);
560
561  let test = -1;
562  for (let i = 0; i < parts.length; i += 2) {
563    if (isTestRun(parts[i])) test = i;
564  }
565  if (test < 0) return false;
566
567  const before = parts.slice(0, test).join("");
568  const pipefail = /\bset\s+-[A-Za-z]*o\s+pipefail\b/.test(before);
569  const errexit =
570    /\bset\s+-[A-Za-z]*e[A-Za-z]*\b/.test(before) ||
571    /\bset\s+-o\s+errexit\b/.test(before);
572
573  let piped = true; // still inside the test's own pipeline
574  for (let j = test + 1; j < parts.length; j += 2) {
575    const sep = parts[j];
576    const rest = parts.slice(j + 1).join("").trim();
577    if (sep === "&") return true;
578    if (sep === "||") return true;
579    if (sep === "|" || sep === "|&") {
580      if (piped && !pipefail) return true;
581      continue;
582    }
583    piped = false;
584    if ((sep === ";" || sep === "\n") && rest && !errexit) return true;
585  }
586  return false;
587}
588
589/**
590 * Why a Bash result that reads as success proves nothing about the suite,
591 * or null when it does prove it. The status is one for the whole line, and
592 * a command that has not finished reports none of its own: Bash answers at
593 * once for `run_in_background`, and moves a command that hits its timeout
594 * to the background and answers then — on a slow machine, the usual fate
595 * of a full suite.
596 */
597function unprovenBy(e: any, record: any, skel: string): string | null {
598  if (e?.run_in_background === true || record?.backgroundTaskId) {
599    return record?.timedOutAfterMs
600      ? "it hit the Bash timeout and was moved to the background, so it had not finished"
601      : "it was sent to the background, so it had not finished";
602  }
603  if (record?.interrupted === true) return "it was interrupted before it finished";
604  if (testStatusMasked(skel)) {
605    return "it was followed by a pipe, `;`, `||` or `&`, so the result reported that command's status, not the tests'";
606  }
607  return null;
608}
609
610/** How a status line names a test run: a hidden status is not a failure. */
611function testVerdict(t: TestEvidence): string {
612  if (t.masked) return "UNPROVEN (exit status hidden)";
613  return t.exitCode === 0 ? "passing" : `FAILING (exit ${t.exitCode})`;
614}
615
616// Global options sit between `git` and its subcommand: `git -c k=v commit`,
617// `git -C dir push`, `git --git-dir=... merge`. Matching `git\s+commit`
618// alone lets every one of those slip past the gates untouched.
619const GIT_OPTS = String.raw`(?:` +
620  [
621    String.raw`-[cC]\s+\S+\s+`,
622    String.raw`--(?:git-dir|work-tree|namespace|exec-path|config-env|attr-source|super-prefix)(?:=\S+|\s+\S+)\s+`,
623    String.raw`--(?:paginate|no-pager|bare|no-replace-objects|no-lazy-fetch|no-optional-locks|literal-pathspecs|glob-pathspecs|noglob-pathspecs|icase-pathspecs|no-advice)\s+`,
624    String.raw`-[pP]\s+`,
625  ].join("|") +
626  String.raw`)*`;
627
628// `\b` after the verb matches before a hyphen too, so `git commit-tree` and
629// `git checkout-index` — plumbing that commits nothing — tripped every gate
630// keyed off these. A verb is the verb only when nothing word-like follows.
631const GIT_VERB_END = String.raw`(?![\w-])`;
632
633const GIT_COMMIT_PUSH_RE = new RegExp(
634  String.raw`\bgit\s+` + GIT_OPTS + String.raw`(commit|push|merge)` + GIT_VERB_END,
635);
636
637const GIT_PUSH_RE = new RegExp(
638  String.raw`\bgit\s+` + GIT_OPTS + String.raw`push` + GIT_VERB_END,
639);
640
641const GIT_MERGE_RE = new RegExp(
642  String.raw`\bgit\s+` + GIT_OPTS + String.raw`merge` + GIT_VERB_END,
643);
644
645// Files that normally carry no behaviour: prose, assets and licences.
646// Extension alone is not enough to conclude that, so BEHAVIORAL_DOC_RE and
647// the executable-doc checks below can each take a file back out of this set.
648const PROSE_FILE_RE =
649  /(\.(md|markdown|mdx|txt|rst|adoc|asciidoc|org|tex|svg|png|jpe?g|gif|webp|ico|pdf|woff2?|ttf|otf)$|^(LICENSE|COPYING|NOTICE|AUTHORS|CONTRIBUTORS|CHANGELOG|CODEOWNERS)([.\-](?:md|markdown|txt|rst|adoc))?$)/i;
650
651// Prose-shaped files that are nothing of the sort: a runbook something
652// runs, instructions an agent reads as its prompt, a fixture or snapshot a
653// test compares against. Editing one changes behaviour, so the commit gate
654// keeps enforcing even though the extension says prose.
655const BEHAVIORAL_DOC_RE = new RegExp(
656  [
657    // Read by tools and agents as instructions, not by people as prose.
658    String.raw`(^|/)(SKILL|CLAUDE|AGENTS?|GEMINI|CURSOR|COPILOT[-_]INSTRUCTIONS|WARP|RUNBOOK|PLAYBOOK)\.[^/]*$`,
659    // A plugin's own behaviour lives in these folders as markdown.
660    String.raw`(^|/)(commands|agents|skills|prompts|references|templates)/`,
661    // A tool's own directory: its contents are configuration.
662    String.raw`(^|/)\.(claude|cursor|github|gitlab|gemini|aider|continue|devcontainer)/`,
663    // Anything a test can read: fixtures, snapshots, golden files.
664    String.raw`(^|/)(tests?|spec|specs|fixtures?|testdata|__tests__|__snapshots__|__fixtures__|e2e|integration|golden|snapshots?)/`,
665    // Runbooks and playbooks kept together by folder.
666    String.raw`(^|/)(runbooks?|playbooks?)/`,
667  ].join("|"),
668  "i",
669);
670
671// Prose formats a documentation toolchain executes or compiles rather than
672// merely renders. Only treated as behaviour when the repo actually has such
673// a toolchain — see EXECUTABLE_DOC_MARKERS.
674const EXECUTABLE_DOC_EXT_RE = /\.(md|markdown|mdx|rst|adoc|asciidoc|org|qmd|ipynb)$/i;
675
676// A repo holding one of these runs, tests or builds its prose, so a change
677// to that prose can break the build the same way code can.
678const EXECUTABLE_DOC_MARKERS = [
679  "book.toml",
680  "_quarto.yml",
681  "_quarto.yaml",
682  "quarto.yml",
683  "runme.yaml",
684  "runme.yml",
685  "jupytext.toml",
686  "mkdocs.yml",
687  "mkdocs.yaml",
688  "docusaurus.config.js",
689  "docusaurus.config.ts",
690];
691
692/**
693 * Paths git reports as changed, staged or not, plus untracked files.
694 * Returns null when git cannot be read, so callers can tell "nothing
695 * changed" apart from "could not tell".
696 */
697async function changedPaths($: any): Promise<string[] | null> {
698  try {
699    const res = await $.process.run(["git", "status", "--porcelain"]);
700    if (res.exitCode !== 0) return null;
701    return res.stdout
702      .split("\n")
703      .map((l: string) => l.slice(3).trim())
704      .filter(Boolean)
705      .map((p: string) => {
706        // Renames read "old -> new"; the destination is what matters.
707        const arrow = p.indexOf(" -> ");
708        return arrow === -1 ? p : p.slice(arrow + 4);
709      })
710      // git quotes any path needing it (spaces, non-ASCII under
711      // core.quotepath). Unquote so extension matching still works.
712      .map((p: string) =>
713        p.startsWith('"') && p.endsWith('"') ? p.slice(1, -1) : p,
714      );
715  } catch {
716    return null;
717  }
718}
719
720/**
721 * The paths a push would send: what HEAD has that its upstream does not.
722 * Returns null when there is no upstream to compare against, or git
723 * cannot be read — "could not tell", which the gate treats as "enforce".
724 */
725async function pushedPaths($: any): Promise<string[] | null> {
726  try {
727    const res = await $.process.run([
728      "git",
729      "diff",
730      "--name-only",
731      "@{u}..HEAD",
732    ]);
733    if (res.exitCode !== 0) return null;
734    return res.stdout
735      .split("\n")
736      .map((l: string) => l.trim())
737      .filter(Boolean);
738  } catch {
739    return null;
740  }
741}
742
743/**
744 * The added lines of a diff. `containsSecret` is documented as reading
745 * what a change ADDS, and was handed the whole diff — so the commit that
746 * removes a leaked key was the one that got blocked, with remediation
747 * text telling you to remove it.
748 */
749function addedLines(diff: string): string {
750  return diff
751    .split("\n")
752    .filter((l) => l.startsWith("+") && !l.startsWith("+++"))
753    .join("\n");
754}
755
756/**
757 * A shell command with its heredoc bodies and quoted literals blanked out,
758 * so the command regexes match commands actually being run rather than text
759 * that merely mentions one. Without this, a commit whose message says
760 * "make test" is recorded as passing test evidence, and editing a document
761 * that quotes a git subcommand is treated as running it.
762 */
763function commandSkeleton(cmd: string): string {
764  const text = unquotedSkeleton(cmd);
765  let out = "";
766  let at = 0;
767  for (const [start, end] of quotedSpans(text)) {
768    out += text.slice(at, start) + text.slice(start, start + 1).repeat(2);
769    at = end;
770  }
771  return out + text.slice(at);
772}
773
774/**
775 * The command with each heredoc's body removed, and nothing else.
776 *
777 * Two regexes used to do this. The first replaced a whole heredoc with a
778 * `<<HEREDOC` token at the end of its line — which the second then read
779 * as an unterminated heredoc and erased everything after it, so a
780 * `git commit` or `git push` on the lines after any heredoc passed every
781 * gate unseen. The first also ate the rest of the delimiter line, so
782 * `cat <<EOF > src/x.ts` lost its write target.
783 *
784 * Now: the delimiter line is kept (`<<` itself replaced), the body up to
785 * the terminator is dropped, and scanning resumes after the terminator. A
786 * `<<` inside quotes on its own line is a mention, not a heredoc. With no
787 * terminator, the rest of the command is body.
788 */
789function stripHeredocs(cmd: string): string {
790  const re = /<<(-?)[ \t]*(['"]?)([A-Za-z_][A-Za-z0-9_]*)\2([^\n]*)/g;
791  let out = "";
792  let at = 0;
793  for (const m of cmd.matchAll(re)) {
794    const start = m.index ?? 0;
795    if (start < at) continue; // inside a body already dropped
796
797    // Quoted on its own line — "a << b" — is text, not a heredoc.
798    const lineStart = cmd.lastIndexOf("\n", start - 1) + 1;
799    const before = cmd.slice(lineStart, start).replace(/\\./g, "");
800    const quotes = (q: string) => before.split(q).length - 1;
801    if (quotes("'") % 2 === 1 || quotes('"') % 2 === 1) continue;
802
803    const delimiter = m[3];
804    const lineEnd = start + m[0].length;
805    const bodyStart = cmd.indexOf("\n", lineEnd);
806    out += cmd.slice(at, start) + " HEREDOC " + m[4];
807    if (bodyStart === -1) {
808      at = lineEnd;
809      continue;
810    }
811    const terminator = new RegExp(String.raw`^[ \t]*${delimiter}[ \t]*$`, "m");
812    const rest = cmd.slice(bodyStart + 1);
813    const end = rest.search(terminator);
814    if (end === -1) {
815      at = cmd.length;
816      break;
817    }
818    const afterTerminator = rest.indexOf("\n", end);
819    out += "\n";
820    at = afterTerminator === -1 ? cmd.length : bodyStart + 1 + afterTerminator + 1;
821  }
822  return out + cmd.slice(at);
823}
824
825/**
826 * The same command with heredoc bodies removed and shell `-c` payloads
827 * unwrapped, but quoted literals left standing. Write-target detection
828 * needs this: blanking quotes first erased the path in `> 'src/x.ts'`
829 * along with the mention it was meant to erase.
830 */
831function unquotedSkeleton(cmd: string): string {
832  let out = cmd;
833
834  out = stripHeredocs(out);
835  // A shell's -c payload is a command, not a literal: unwrap it so what
836  // it runs is still seen. Only shells — `python3 -c "..."` stays opaque.
837  out = out.replace(
838    /\b(?:(?:ba|z|k|da)?sh|fish)\s+(?:-[a-zA-Z]+\s+)*-c\s+('(?:[^'\\]|\\.)*'|"(?:[^"\\]|\\.)*")/g,
839    (_m: string, q: string) => ` ${q.slice(1, -1)} `,
840  );
841
842  return out;
843}
844
845/**
846 * The character spans covered by quoted literals, scanned left to right.
847 *
848 * Two passes used to do this, every single-quoted span first and then
849 * every double-quoted one. A double-quoted string holding an apostrophe
850 * (`claude -p "don't commit"`) mis-paired: the apostrophe opened a span
851 * that ran on to the next one, and the quoted text between them stayed
852 * bare — so a command merely NAMED inside a quoted argument was read as a
853 * command being run, and the git gates fired on it.
854 */
855function quotedSpans(text: string): Array<[number, number]> {
856  const spans: Array<[number, number]> = [];
857  for (let i = 0; i < text.length; i++) {
858    const c = text[i];
859    if (c === "\\") {
860      i++;
861      continue;
862    }
863    if (c !== "'" && c !== '"') continue;
864    let j = i + 1;
865    for (; j < text.length; j++) {
866      if (c === '"' && text[j] === "\\") {
867        j++;
868        continue;
869      }
870      if (text[j] === c) break;
871    }
872    // An unclosed quote runs to the end of the command.
873    spans.push([i, Math.min(j + 1, text.length)]);
874    i = j;
875  }
876  return spans;
877}
878
879const GIT_COMMIT_RE = new RegExp(
880  String.raw`\bgit\s+` + GIT_OPTS + String.raw`commit` + GIT_VERB_END,
881);
882
883const GIT_BRANCH_SWITCH_RE = new RegExp(
884  String.raw`\bgit\s+` + GIT_OPTS + String.raw`(checkout|switch)` + GIT_VERB_END,
885);
886
887// Commands that stage as they go: `git add -A && git commit`, or
888// `git commit -am`. The hook runs before any of it, so at scan time the
889// index does not yet hold what is about to be committed.
890const GIT_STAGING_RE = new RegExp(
891  String.raw`\bgit\s+` +
892    GIT_OPTS +
893    String.raw`(?:add|stage)\b|\bgit\s+` +
894    GIT_OPTS +
895    String.raw`commit\b[^\n;&|]*(?:\s-[a-zA-Z]*a[a-zA-Z]*\b|\s--(?:all|include|only)\b)`,
896);
897
898// Untracked files read per scan, so a large working tree cannot stall a
899// commit while the gate reads it.
900/**
901 * A path glob as a RegExp, built in one pass. A chain of .replace() calls
902 * lets a later pass rewrite regex syntax an earlier one emitted: the `?`
903 * rule corrupted the non-capturing group the double-star rule had just
904 * produced, so a declared pattern silently matched nothing.
905 *
906 * A single star stays inside one segment, a double star followed by a
907 * slash spans any number of directories, and a trailing double star takes
908 * the rest of the path.
909 */
910function globToRegExp(glob: string): RegExp {
911  let out = "";
912
913  for (let i = 0; i < glob.length; i++) {
914    const c = glob[i];
915
916    if (c === "*") {
917      if (glob[i + 1] === "*") {
918        i++;
919        if (glob[i + 1] === "/") {
920          i++;
921          out += "(?:.*/)?";
922        } else {
923          out += ".*";
924        }
925      } else {
926        out += "[^/]*";
927      }
928    } else if (c === "?") {
929      out += "[^/]";
930    } else if (".+^${}()|[]\\/".includes(c)) {
931      out += "\\" + c;
932    } else {
933      out += c;
934    }
935  }
936
937  return new RegExp("^" + out + "$", "i");
938}
939
940/**
941 * A design document, which planning mode allows writing. Matched against
942 * the BASENAME: matching the whole path meant one ancestor directory
943 * named design/, plan/, spec/, proposal/ or rfc/ exempted every file in
944 * the repo and switched the planning hard-gate off wholesale.
945 */
946function isDesignDoc(path: string): boolean {
947  const base = path.split("/").pop() ?? path;
948
949  // A prose extension is a design doc outright.
950  if (/\.(md|markdown|mdx|rst|adoc|txt)$/i.test(base)) return true;
951
952  // Any other extension is code or data, whatever the name says —
953  // `Plan.tsx` and `api-spec.go` are implementation files.
954  if (/\.[A-Za-z0-9]+$/.test(base)) return false;
955
956  // Extensionless, so judge by name: DESIGN, rfc-0001, SPEC.
957  return /\b(design|plan|spec|proposal|rfc)\b/i.test(base);
958}
959
960/**
961 * The branch HEAD is on right now. The gate used to trust the branch
962 * cached at session start, which a `git checkout main && git merge feat`
963 * in one command line has not updated yet — and which was never updated
964 * at all while the post-hook that maintained it was unreachable (S25).
965 * Falls back to the cached value when git cannot be read.
966 */
967async function currentBranch($: any, cached: string | null): Promise<string | null> {
968  try {
969    const res = await $.process.run(["git", "branch", "--show-current"]);
970    const live = (res.stdout ?? "").trim();
971    return live || cached;
972  } catch {
973    return cached;
974  }
975}
976
977/**
978 * A tool.call result reports failure through `isError` and carries its
979 * output as text. It has no exitCode, stdout or stderr — those belong to
980 * `$.process.run`, which is a different shape. Reading them here recorded
981 * every failing test run as a pass (`exitCode ?? 0`), left the failure
982 * output empty, and made the `result.exitCode === 0` trackers unreachable,
983 * so a branch switch was never noticed and commits were never counted.
984 *
985 * A numeric exitCode is still honoured if a future build supplies one.
986 */
987function toolFailed(result: any): boolean {
988  if (typeof result?.exitCode === "number") return result.exitCode !== 0;
989  return result?.isError === true;
990}
991
992function toolOutput(result: any): string {
993  if (typeof result?.stdout === "string" || typeof result?.stderr === "string") {
994    return `${result.stdout ?? ""}${result.stderr ?? ""}`;
995  }
996  return String(result?.text ?? result?.result ?? "");
997}
998
999// Shell forms that write a file: a redirection, an in-place edit, a tee.
1000// A path may be quoted, so the alternatives accept a quoted literal —
1001// `> 'src/x.ts'` is a write to src/x.ts, and matching against a skeleton
1002// with the quotes blanked out saw no path at all.
1003const WRITE_PATH = String.raw`'[^']+'|"[^"]+"|[\w./~@+=-]+`;
1004
1005const SHELL_WRITE_RE = new RegExp(
1006  String.raw`>>?\s*(?!&)(` +
1007    WRITE_PATH +
1008    String.raw`)|\b(?:sed|perl|ruby)\s+(?:-\S+\s+)*-i\S*\s+(?:-\S+\s+)*(?:'[^']*'|"[^"]*"|\S+)\s+(` +
1009    WRITE_PATH +
1010    String.raw`)|\btee\s+(?:-\S+\s+)*(` +
1011    WRITE_PATH +
1012    String.raw`)`,
1013  "g",
1014);
1015
1016// A redirection to one of these writes nothing a planning gate cares
1017// about; `ls > /dev/null` was being denied as an implementation write.
1018const DEV_SINK_RE = /^\/dev\/(?:null|stdout|stderr|tty|fd\/\d+)$/;
1019
1020/**
1021 * Every file the command writes.
1022 *
1023 * Three bugs lived in the single-match version this replaces: `tee` is the
1024 * third capture group and only the first two were read; a quoted path was
1025 * erased before the match; and `String.match` without /g returned one
1026 * target, so `echo a > notes.md && echo b > src/x.ts` was judged entirely
1027 * by notes.md. A write named inside a quoted string is still only a
1028 * mention — the operator itself has to be outside the quotes.
1029 */
1030function shellWriteTargets(cmd: string): string[] {
1031  const text = unquotedSkeleton(cmd);
1032  const spans = quotedSpans(text);
1033  const mention = (at: number) => spans.some(([a, b]) => at > a && at < b);
1034
1035  const targets: string[] = [];
1036  for (const m of text.matchAll(SHELL_WRITE_RE)) {
1037    if (m.index !== undefined && mention(m.index)) continue;
1038    const path = (m[1] ?? m[2] ?? m[3] ?? "")
1039      .replace(/^['"]|['"]$/g, "")
1040      .trim();
1041    if (!path || DEV_SINK_RE.test(path)) continue;
1042    targets.push(path);
1043  }
1044  return targets;
1045}
1046
1047/**
1048 * The untracked files a command line would stage, as pathspecs: "all" for
1049 * `git add -A|.|-u`, a list for named paths, and null when the line stages
1050 * no untracked file at all (`git commit -a` takes tracked changes only).
1051 *
1052 * Scanning every untracked file instead denied a commit over a key in a
1053 * stray log the line never touched.
1054 */
1055function stagedPathspecs(skel: string): string[] | "all" | null {
1056  const add = new RegExp(
1057    String.raw`\bgit\s+` + GIT_OPTS + String.raw`(?:add|stage)` + GIT_VERB_END + String.raw`([^;&|\n]*)`,
1058    "g",
1059  );
1060  const specs: string[] = [];
1061  let staging = false;
1062  for (const m of skel.matchAll(add)) {
1063    staging = true;
1064    for (const arg of (m[1] ?? "").trim().split(/\s+/).filter(Boolean)) {
1065      if (/^(?:-A|--all|-u|--update|--|\.)$/.test(arg)) return "all";
1066      if (arg.startsWith("-")) continue;
1067      specs.push(arg.replace(/^['"]|['"]$/g, ""));
1068    }
1069  }
1070  if (staging) return specs.length > 0 ? specs : "all";
1071
1072  // No `git add`: a commit's own pathspecs, or none for `-a`.
1073  const commit = new RegExp(
1074    String.raw`\bgit\s+` + GIT_OPTS + String.raw`commit` + GIT_VERB_END + String.raw`([^;&|\n]*)`,
1075  ).exec(skel);
1076  if (!commit) return null;
1077  const args = (commit[1] ?? "").trim().split(/\s+/).filter(Boolean);
1078  const paths: string[] = [];
1079  let all = false;
1080  for (let i = 0; i < args.length; i++) {
1081    const a = args[i];
1082    if (/^-[a-zA-Z]*a[a-zA-Z]*$/.test(a) || a === "--all") all = true;
1083    if (/^(?:-m|--message|-c|-C|--reuse-message|--reedit-message|--author|--date|--file|-F)$/.test(a)) {
1084      i++;
1085      continue;
1086    }
1087    if (a.startsWith("-")) continue;
1088    paths.push(a.replace(/^['"]|['"]$/g, ""));
1089  }
1090  if (paths.length > 0) return paths;
1091  return all ? null : "all";
1092}
1093
1094const MAX_UNTRACKED_SCAN = 50;
1095
1096const GIT_DESTRUCTIVE_RE = new RegExp(
1097  String.raw`\bgit\s+` +
1098    GIT_OPTS +
1099    String.raw`(?:(commit|push|merge|rebase|force-push)\b|reset\s+--hard\b|(?:checkout|restore)\s+(?:--\s+)?[.*]|checkout\s+--\s)`,
1100);
1101
1102/**
1103 * The branch a destructive git command in this line will run on, when the
1104 * line switches branch first: `git checkout main && git merge feat`. The
1105 * gate runs before the line does, when git still reports the branch it
1106 * started on — so without this, one line walked onto a protected branch
1107 * and merged there unchecked. The last switch before the destructive
1108 * command wins; a path checkout (`checkout main -- file`) is not a switch.
1109 */
1110function lineSwitchTarget(skel: string): string | null {
1111  const at = skel.search(GIT_DESTRUCTIVE_RE);
1112  if (at <= 0) return null;
1113  const switchRe = new RegExp(
1114    String.raw`\bgit\s+` + GIT_OPTS + String.raw`(?:checkout|switch)` + GIT_VERB_END + String.raw`([^;&|\n]*)`,
1115    "g",
1116  );
1117  let target: string | null = null;
1118  for (const m of skel.slice(0, at).matchAll(switchRe)) {
1119    const args = (m[1] ?? "").trim().split(/\s+/).filter(Boolean);
1120    if (args.includes("--")) continue;
1121    const named = args.findIndex((a) => /^-[bBcC]$/.test(a) || a === "--orphan");
1122    const name = named >= 0 ? args[named + 1] : args.find((a) => !a.startsWith("-"));
1123    if (name) target = name;
1124  }
1125  return target;
1126}
1127
1128// Destructive non-git bash commands — soft warning
1129const DESTRUCTIVE_BASH_RE =
1130  /(?:rm\s+(?:-[^\s]*r[^\s]*\s|--recursive\s)|chmod\s+(?:-R\s+)?777\s|curl\s[^|]*\|\s*(?:sudo\s+)?(?:bash|sh|zsh)|wget\s[^|]*\|\s*(?:sudo\s+)?(?:bash|sh|zsh)|dd\s+if=|mkfs\.|>\s*\/dev\/sd)/;
1131
1132// Secret/credential patterns — hard gate on git commit
1133// Values that are obviously not a real credential, so an unquoted
1134// assignment carrying one is not treated as a leak.
1135const SECRET_PLACEHOLDER_RE =
1136  /^(?:[*x.]{3,}|<[^>]*>|\$\{?[A-Za-z_]|\{\{|%[A-Za-z_]|your[-_]?|changeme|example|placeholder|redacted|dummy|sample|test|fake|none|null|true|false)/i;
1137
1138// A value that reads a secret rather than being one: `process.env.X`,
1139// `os.environ["X"]`, `getenv(...)`, `config.token`. Any dotted identifier
1140// or call is code, not a credential.
1141const SECRET_REFERENCE_RE = /^[A-Za-z_$][\w$]*(?:\.[\w$]+|\[|\()/;
1142
1143const SECRET_PATTERNS: RegExp[] = [
1144  // AWS long-lived, temporary (ASIA) and the other documented prefixes.
1145  /(?:AKIA|ASIA|ABIA|ACCA)[0-9A-Z]{16}/,
1146  // OpenAI legacy and project keys. `+` is in the lead-in class because
1147  // this runs against diff output, where every added line starts with one.
1148  /(?:^|[\s'"=:+])sk-[a-zA-Z0-9]{20,}/m,
1149  /(?:^|[\s'"=:+])sk-proj-[a-zA-Z0-9_-]{20,}/m,
1150  /gh[pousr]_[a-zA-Z0-9]{36}/,
1151  /github_pat_[a-zA-Z0-9_]{22,}/,
1152  /glpat-[a-zA-Z0-9\-_]{20,}/,
1153  /xox[abprs]-[a-zA-Z0-9-]{10,}/,
1154  /-----BEGIN\s+(?:RSA\s+|EC\s+|DSA\s+|OPENSSH\s+|ENCRYPTED\s+|PGP\s+)?PRIVATE\s+KEY/,
1155  // Quoted assignments.
1156  /(?:password|passwd|pwd)\s*[:=]\s*['"][^'"]{8,}['"]/i,
1157  /(?:api[_-]?key|apikey|secret[_-]?key|auth[_-]?token)\s*[:=]\s*['"][^'"]{12,}['"]/i,
1158];
1159
1160// Unquoted assignments — the shape a committed `.env` has, which the
1161// quoted patterns above could never match. Checked separately so the
1162// value can be tested against SECRET_PLACEHOLDER_RE first.
1163const BARE_SECRET_ASSIGN_RE =
1164  /(?:^|[\s+])(?:[A-Za-z_][A-Za-z0-9_]*_)?(?:password|passwd|pwd|secret|api[_-]?key|apikey|access[_-]?key|auth[_-]?token|token|credential)s?\s*[:=]\s*([^\s'"#]{8,})/gim;
1165
1166/**
1167 * True when the added lines of a diff carry something credential-shaped.
1168 * Kept as a function so the unquoted-assignment case can discount
1169 * placeholders without that logic living inside a regex.
1170 */
1171function containsSecret(text: string): RegExp | null {
1172  for (const pattern of SECRET_PATTERNS) {
1173    if (pattern.test(text)) return pattern;
1174  }
1175
1176  BARE_SECRET_ASSIGN_RE.lastIndex = 0;
1177  for (const m of text.matchAll(BARE_SECRET_ASSIGN_RE)) {
1178    const value = m[1];
1179    if (SECRET_PLACEHOLDER_RE.test(value)) continue;
1180    if (SECRET_REFERENCE_RE.test(value)) continue;
1181    return BARE_SECRET_ASSIGN_RE;
1182  }
1183
1184  return null;
1185}
1186
1187// Skill → lifecycle phase mapping
1188const SKILL_PHASE_MAP: Record<string, SessionState["currentPhase"]> = {
1189  brainstorming: "brainstorming",
1190  "proctor:brainstorming": "brainstorming",
1191  "writing-plans": "planning",
1192  "proctor:writing-plans": "planning",
1193  "test-driven-development": "implementing",
1194  "proctor:test-driven-development": "implementing",
1195  "subagent-driven-development": "implementing",
1196  "proctor:subagent-driven-development": "implementing",
1197  "executing-plans": "implementing",
1198  "proctor:executing-plans": "implementing",
1199  "systematic-debugging": "implementing",
1200  "proctor:systematic-debugging": "implementing",