SLOPSHOPPER

workbench-dev-team

The Index-driven development pipeline: Inspector Lestrade (triage), Dr. Watson (dev work), and Sherlock Holmes (code review). A launchd job called Dispatch…

newpaneguardcommandtoastprompt
v0.52.0no licenseupdated 2026-10-09mike-bronner/workbench-dev-team
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · workbench-dev-team
│ ┃ Dev-team runs ✕ › fix the failing auth test and add an audit log call │ ┃ No dispatched runs in the logs yet. │ ⏺ Read(src/auth.ts) │ ⎿ Read 6 lines │ ⏺ Update(src/auth.ts) │ ⎿ Added 2 lines, removed 1 line │ ⏺ Edit(/work/app/src/auth.ts) │ ⎿ Denied by workbench-dev-team: 🛑 Blocked: a call the guar │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /dev-team-runs │ ⎿ workbench-dev-team: Dev-team runs pane opened. │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · Dev-team runs
No dispatched runs in the logs yet.
Pane · The Index board
Fetching the board from The Index.
README

workbench-dev-team

A Claude Code plugin that runs a three-agent development pipeline against a GitHub project board. Items flow from triage → development → review without you touching them. Part of the claude-workbench marketplace.

What it does

You work out of a GitHub project board. New issues land with no acceptance criteria. Well-refined ones sit in Ready. Work-in-progress has open draft PRs. PRs waiting on review pile up.

This plugin runs a team of agents on a local 20-minute clock that move items through the pipeline for you — and records what the reviews teach as it goes, so the team stops repeating itself:

  • Inspector Lestrade (claude-opus-5-5[1m] at medium effort, Sonnet lenses) — triage. Reads items in the Inbox lane, writes acceptance criteria as a managed follow-up comment on the issue (the description is left untouched). Before scoring, the draft AC is checked by four blind lens sub-agents (malicious-compliance, testability, completeness, edge-case) that each try to find a gap Holmes and Watson would otherwise hit downstream — a much more expensive place to catch it, since Lestrade only ever sees an issue once while Watson and Holmes cycle back on it every bounce. Real gaps get folded in with one bounded tightening pass, never a re-verification loop; a lightweight comment marks when the AC changed this way. Scores WSJF, moves them to Backlog for your review. The WSJF write also lands two GitHub-native issue attributes The Index derives server-side: the issue Type (PBI) and an issue-level Priority (Urgent/High/Medium/Low) mapped from the WSJF — org-repo-only and best-effort. Also runs blocker + consolidation sweeps: after a repo gets fresh triage work, he re-reads all of its open issues and (1) marks blocked-by dependencies (native GitHub issue dependencies, additive only) so blocked items stay out of Dr. Watson's queue, and (2) consolidates follow-ups — folds expand-from comments into the issue they target and merges unmistakable near-duplicate follow-ups into the earliest anchor (native duplicate-close, high bar, ambiguous clusters flagged not closed) so the backlog stops sprawling.
  • Dr. Watson (claude-opus-5-5[1m] at medium effort, $10/run cap) — development. Has two modes: Direct mode, the default (invocable as a sub-agent from Claude Code or Cowork for ad-hoc dev work — no The Index calls, just runs the /develop skill in a sub-agent context, and hands its work back as an uncommitted working tree for the dispatching session to commit) and The Index mode, entered only on an explicit Item ID: <n> token (Dispatch-driven, picks the top Ready/In Progress item, clones the repo, writes code and tests against AC, opens a PR, moves to In Review). Ambiguous prose resolves to Direct mode, never to the board. Both modes follow the /develop skill for the actual coding.
  • Sherlock Holmes (claude-opus-5-5[1m] parent at medium effort, Sonnet lenses, $10/run cap) — code review. Has two modes, the same way Watson does: Local mode, the default (invocable as a sub-agent on a six-slot prose brief — reviews the uncommitted working tree in the given workdir, tracked changes and untracked files, against the brief's Acceptance: list as the rubric it never amends, and returns the verdict as prose) and The Index mode, entered only on an explicit Item ID: <n> token (Dispatch-driven, reviews the board item's PR and posts an App-signed verdict). Ambiguous prose resolves to Local mode, never to the board. Local mode makes no The Index call and no GitHub write at all — no review, no comment, no issue — so it is safe on any repo, and it closes the loop on Watson's Direct mode, whose output is exactly an uncommitted tree with no review path otherwise. It replaces the board-coupled steps rather than skipping them: the brief is the rubric, the workdir is the evidence room (never cloned, never written to — a harness-level guard enforces that, see Review guard below), and the repo's own suite is run locally in place of reading CI, because the toolchain objection that keeps Index mode off local test runs does not hold when the agent is already in the repo's directory on your machine. It records a vault note per local verdict but deliberately never touches the top-lessons.md digest, since local reviews are a separate population whose counts would skew a frequency ranking. Everything in between — the lens fan-out, adversarial verification, the memory pass, the finding-routing matrix — runs unchanged. The rest of this entry describes The Index mode: reviews open PRs, approves or requests changes. Escalates to you after 3 change rounds — but your input resets that count: comment, review, or weigh in on the PR and the window restarts from your last word, so an escalated PR you've decided on gets a fresh review instead of bouncing straight back. Reviews fan out across blind, read-only lens sub-agents (AC conformance, correctness, security, test honesty), with every blocker adversarially verified before it lands — a 3-agent red-team/blue-team/auditor pipeline handles security-lens findings every round and hard defects on the PR's first review, and a single skeptic handles the rest, soft observations included (the fullest check lands where a wrong verdict costs a round, not on a naming nit). After verification, Holmes (and only Holmes — no sub-agent touches the vault) checks each surviving finding against the memory vault for relevant context — a documented decision that reframes it, a past incident that reinforces it — always re-verified against the current tree before it's trusted, never used to waive a real defect or mark an AC item met. Only the parent writes, so there's still exactly one App-signed verdict. Falls back to a single inline pass when the fan-out is unavailable. Findings route by the coherent unit of work — what the issue is really about — then coupling and locality: a finding that belongs to the unit blocks and is fixed in this PR, even in untouched code the diff never caused, because a half-delivered unit is itself the defect. On top of that, anything actionable in the code a PR touched blocks (request changes, however minor), a hard correctness/security/test defect blocks wherever it lives, and — coupling beating locality — untouched code the diff made stale, inconsistent, or wrong blocks too. The self-test: block and fix here if EITHER the diff caused it OR it belongs to the coherent unit — a follow-up only when both are false. The precise routing rule lives in Holmes's canonical review contract (agents/holmes.md, §4e/§5). The non-blocking follow-up tier is then gated by materiality, default-deny: it holds only findings unrelated to the unit, and most of those — one-off cosmetics (naming, small duplication, style) — are noted in the verdict, not tracked. Only an unrelated latent hazard (security/data-integrity/correctness not live enough to block) or systemic/substantial debt (a schedulable chunk with its own testable "done") earns one tracked issue, tagged Tracked under:, capped at one new anchor per PR — the materiality bar that stops the follow-up flood. When Holmes does track, he expands the earliest related open issue in place (a comment Lestrade folds into its acceptance criteria) rather than opening a near-duplicate, and opens a new anchor via create_issue (App-signed as Holmes, board-added and PBI-typed) only when nothing related exists. A class of sites violating one invariant (a containment guard, a null-check, a helper every caller owes) is swept whole: if the class belongs to the unit it's folded into the PR; if it's an unrelated anti-pattern agents will replicate it becomes one umbrella issue with a checkbox per site — either way closing the class at once rather than minting a fresh single-site issue every review, the treadmill that otherwise turns one finding into an endless #A → #B → #C chain. A finding gets the same disposition regardless of verdict, so a clean PR never generates more tracked work than a messy one: on a change request, unit-belonging findings are blockers Watson folds into the same PR, unrelated cosmetics are optional (fix if cheap, else skip), and the unrelated hazard/debt tier is tracked exactly as on approval.

A fourth component — Dispatch — is a local launchd job that polls the board every 20 minutes and fires the right agent for each pending item. It is a shell script with no model in it, so a tick with nothing to do costs no tokens. Dispatch is the only thing that's scheduled; the three agents run as dispatched subprocesses.

The feedback loop

Lestrade, Watson, and Holmes used to have no memory of their own reviews: nothing recorded what Holmes rejected, why, or how it got fixed, so the same lessons got re-taught every review (test-honesty is ~38% of all rejections, fail-open ~11%, doc-drift ~11%). There's no separate harvesting agent for this — Holmes is the only one who holds both halves of a rejection (what he flagged, and whether the next push actually fixed it), so he records it himself, live, at re-review: one atomic vault note per bounce or AC-dispute event, categorized against a fixed taxonomy, plus an incrementally-refreshed dev-team/top-lessons.md digest (recurring categories, frequency-ranked, each with the concrete prevention rule). Watson reads that digest and searches the vault for anything task-specific before coding; Lestrade reads it before writing acceptance criteria, so the pipeline gets smarter instead of repeating itself on both sides — what gets built and what gets asked for. Your own corrections take a second channel, because no review rejection records them: they live under feedback/ in the vault, and all three agents are required to search it: Watson before building (/develop requires it in every lane, Direct mode included), Holmes before judging, in both Local and Index mode, and Lestrade before writing acceptance criteria. Your corrections bind every stage, and a stage that never reads them never catches a violation of one.

A review note never replaces an existing one. Holmes builds each learnings note's path himself, and the memory MCP's write replaces any note already at that path. A read cannot prove a path is free, because the server answers Document not found both for a missing file and for a note it cannot parse. So each name ends with the time to the second and a token from openssl rand -hex 3, and he writes only after a read of that path answers Document not found. A write that still answers created: false replaced a note, and his report names the path so the old note can be restored from the vault's git history. agents/lint-vault-note-paths.sh fails when a note write by a built path, in any Markdown file, loses its unique name or its read, or stands outside a code fence.

Install

/plugin marketplace add mike-bronner/claude-workbench
/plugin install workbench-dev-team@claude-workbench

That installs the agents, the Dispatch scripts, and the bundled skills (see below). Nothing is scheduled yet.

Bundled skills

The plugin ships three skills for general use, plus one skill the agents write by:

  • develop and git-commit — universal development standards. They register themselves globally via session-warmup.md, which workbench-core picks up at session start and injects into ~/.claude/CLAUDE.md. They apply to every Claude Code / Cowork session, not just dev-team agents. Both are also packageable as .skill files for Claude Chat (Mac app) where plugins aren't supported but skills are. Require workbench-core 0.2.0+ for the session-warmup discovery mechanism — install it first if you don't already have it (Claude Code does not enforce plugin install order).
  • orchestrate — runs the team as background sub-agents from any interactive session (see below). A session-warmup.md hint makes every session aware the team is available for delegation.

One more skill exists for the agents rather than for you. comms-style is how Lestrade, Watson, and Holmes write every piece of prose that isn't code — ticket comments, PR bodies, review verdicts — modeled on ASD-STE100 (Simplified Technical English).

Each agent's long, situational procedure is not a skill. It lives in references/<agent>/, and the agent reads it by path: each agent prompt in agents/ is a thin router that keeps the always-relevant rules inline and points at its reference file for the detail it only needs at one moment — Holmes's review phases and sub-agent prompt skeletons plus his Local-mode path (references/holmes/), Watson's Index-mode pipeline (references/watson/), Lestrade's acceptance-criteria lenses and Sweep mode (references/lestrade/). Keeping them out of skills/ keeps their descriptions out of every session's skill listing, and the router pattern keeps the per-dispatch prompt small without putting any rule out of reach.

Plugin configuration lives in a slash command (/workbench-dev-team:setup), not a skill — see the Setup section below.

develop

Universal dev workflow + standards: orient before writing (the repo's CLAUDE.md, AGENTS.md, CONTRIBUTING.md, .ai/ rules, and review wiki), plan before coding (including a required feedback/ vault search), atomic commits, every change gets a test, no committed secrets, lint before pushing. Triggers whenever code is being implemented, fixed, refactored, or tested — manual or agent-driven.

It names three lanes (foreground, sub-agent, scheduled pipeline) and ends each step the way that lane can: a sub-agent finishes with an uncommitted tree and a report, never a commit or a PR. Includes a decision protocol that presents three options in one table (Option, Pros, Cons, and a Grade against every acceptance criterion), then a one- or two-sentence recommendation, for a meaningful fork the human has not already decided — implementation approach, library choice, scope decisions, a public interface. A fork they already decided is not asked again. In the foreground the human decides through AskUserQuestion, with each option's grade and any warning in its description; a sub-agent picks and records the assumption below the blocking-uncertainty bar and stops with the options above it. Trivial choices (mechanical translation, following existing repo conventions, naming, one-line obvious fixes) are exempt.

Also defines the commit approval lanes (see Commit approval below): a sub-agent commits, merges, and pushes nothing, a foreground commit waits for a "Commit it" pick in AskUserQuestion, once the human says their review is done, Index-mode development never asks about committing or pushing, with Claude Code's permission prompt on every git commit and git push as the backstop, and a push that forces or deletes remote refs is refused outright.

Used by Watson internally in both operating modes. Also invocable directly in any plugin-aware Claude session.

git-commit

Generates commit messages using Conventional Commits + Gitmoji format. Triggers whenever a commit message is being composed — manual, scripted, or agent-driven (including Watson's PRs).

Format example:

feat: ✨ Add email validation endpoint.

Fixes: #789

Full type and emoji references at skills/git-commit/references/.

orchestrate

Turns the current session into the team's orchestrator: dispatches Lestrade, Watson, and Holmes as background sub-agents (Agent tool, run_in_background), passing each agent's model from the shared config; maintains a roster table of who's working on what; relays verdicts and decision forks back to you; follows up on running agents via SendMessage. The main conversation stays lean — sub-agents do the heavy work in their own contexts and return summaries.

Watson supports brief-driven Direct mode for ad-hoc dev work with no board item, and Holmes a brief-driven Local mode that reviews the resulting uncommitted tree. Only Lestrade is Index-coupled and needs a board item ID.

The skill also fixes who gets dispatched, and what they're told. Routing covers all three specialists: development goes to Watson, triage to Lestrade, review to Holmes, and the shape of the request says which. Anything ending in a changed file is Watson's, never general-purpose — a specialist loads develop and discovers the repo's conventions and test framework for itself, a generic agent doesn't. Read-only dispatches (Explore, Plan, general-purpose) stay legitimate and are called out as such.

Every handoff is written to a fixed six-slot brief — Workdir: / Goal: / Context: / Constraints: / Acceptance: / Done when: — read-only research dispatches included, which is what makes classifying a prompt as code work unnecessary in the first place. Workdir: is the absolute path, plus the branch or worktree when the human settled one (Workdir: /Users/mike/Developer/foo (branch: fix/retry-backoff)); a bare path stays valid and means there was no workspace decision to record. Goal: is bounded at one or two sentences, because it's the only slot tight enough to check a result against; Context: is prose and deliberately unbounded, since that's where the reasoning goes so Goal: doesn't have to carry it; Constraints: is bullets that each state their own reason, because a constraint without its reason gets obeyed literally and defeated in spirit. Acceptance: is the numbered list of criteria the work is graded against: the receiver grades its forks and its report against it, Holmes reviews against it, and it is where the criteria from /workbench-core:intake cross the handoff. For a Watson or Holmes Item ID: <n> run, the item's acceptance criteria from triage are that list. No length limit is stated anywhere — length was only ever a proxy for prescriptiveness, and the must-omit list (shell commands, numbered steps, named test paths, framework choices) attacks that directly. Constraints: may read "none"; Context: may not, and carries at least one sentence on why the task exists. The brief binds the receiver too, and in two ways: every agent in agents/ refuses a brief missing a slot and names what's missing, and an agent handed a complete-but-unusable brief stops and sends its questions back to the orchestrator instead of guessing — the bar there is blocking uncertainty only, so anything short of it proceeds with the assumption stated. The two fixed dispatch tokens (Item ID: <n>, Repo sweep: <owner/repo>) are exempt from both, since they aren't briefs. A third exemption is the orchestrator boundary itself: the template governs a dispatch that leaves an orchestrator for a specialist, and a specialist's own fan-out to internal workers is outside it. The measurement behind the rule already excluded that traffic — of 622 dispatches, the 422 sent from sessions that were themselves agent runs were counted as correct behaviour — the parent holds every fact those workers need, so Context: has nothing to recover, and their prompts are written against measured cost rather than to a template. It's stated as a boundary and not as a list of agents, so a specialist that grows a fan-out later inherits it with no edit. A companion PreToolUse gate in workbench-core checks one thing — that the slots are present — and refuses a handoff that drops one; judging whether the prose inside them is any good is the receiving agent's job, not a shell script's.

The orchestrator asks before a branch or a worktree is created. Not a prohibition — branches and worktrees are the wanted outcome of most dev work, and a PR is a normal finish line. What changed is who decides. Ahead of any dispatch whose work ends in a commit, the skill reads the target tree and asks in three cases: on main/master/trunk it proposes a branch name; on a feature branch already carrying unrelated work it names what's there and proposes a branch off the base; inside a worktree it confirms that worktree is the one meant for this task. All three end in a question to you, never in a refusal, and your answer is recorded in the brief's Workdir: slot so the sub-agent works where you said. Two Watsons on one repo still get separate worktrees, asked for the same way: you create them, and the orchestrator then dispatches one Watson per worktree. It never creates a worktree and never passes the Agent tool's worktree isolation, which workbench-core's provisioning guard denies. The reasoning, including why the answer widened Workdir: instead of adding a slot, sits in skills/orchestrate/references/brief-rationale.md, which loads only when a rule is challenged rather than on every orchestration.

Independent units dispatch together. Work in different repos or in files that do not overlap, read-only research, and reviews of a finished tree all go out at once, in one message. Local work (interactive sessions, Watson's Direct mode, Holmes's Local mode, research) has no count cap on sub-agents. The Index pipeline stays bounded on purpose: scheduled Dispatch starts one Watson per tick, and every Index-mode cap stays.

The skill also routes GitHub actions to the right executor. Two rules: (1) agent work products (formal reviews, AC, status moves) only ever go through The Index, signed as the dispatched agent — never gh; (2) your own actions (comments you dictate, merges you order) go through gh under your identity, on any repo. Whether a repo is Index-governed is answered by check_repo_access — a server-side tool that checks The Index GitHub App's installation list (spec in THE_INDEX_HANDOFF_ROUTING.md; until it ships, the skill degrades to a list_items scan and says so). Merges are never delegated to agents and only happen on your explicit request.

Commit approval

A sub-agent does not commit, merge, or push, and never asks to. In the foreground session, the agent commits only after a "Commit it" pick in AskUserQuestion, once the human says their review is done. That is the approval. A typed "commit it" in chat does not count. The orchestrator asks the commit question alone, and recommends "Not yet" until your review is done. Claude Code's permission prompt on each git commit, git push, and gh pr merge is the mechanical backstop, because a prompt that appears mid-flow gets answered without a review. After a commit you approved, the agent attempts the push itself and lets the prompt ask you. The scheduled pipeline commits and pushes unattended, inside its own clone and the scratch roots, and never merges.

No agent disguises a command to get past a gate. Watson, Holmes, and Lestrade each carry a ## When a gate or guard refuses you section: never reword, split, encode, or rebuild a command to get past a gate or guard, and report the refusal instead. It exists because Holmes once got past the installed commit gate by building the words "commit" and "push" from pieces at run time. agents/lint-gate-refusal.sh pins the section in every agent file.

Every agent's scratch lands in a folder of its own, and the dev-team mod deletes it. Each agent file carries a ## Scratch folders section: make scratch with a bare mktemp -d. The mod points a dev-team agent's bare mktemp at a folder made for that run under a scratch root, the session scratchpad or ~/Developer/scratchpad when the harness names none, and deletes the folder at the end of a turn of that agent that finds no child it spawned still live (hooks/mods/scratch.ts). A child's end never deletes its parent's folder, since an ended agent may resume and read it again. A folder kept that way for a run that never resumes goes when the session ends: the mod records every folder it makes and has not deleted, and at `sess

Source 12 files
hooks/register.ts 641 lines
1// workbench-dev-team's hooks module, beside the command hooks in hooks.json.
2// It builds on workbench-core's $.workbench noun (dependencies in plugin.json).
3//
4//   agent.spawn    one dispatch of an agent, in order:
5//                  1. the dispatch gate: a main-session dispatch whose brief
6//                     lacks a slot is refused, as hooks/agent-dispatch-gate.sh
7//                     in workbench-core refuses it, and so is one whose brief
8//                     cannot be checked
9//                  2. the helper rule: a spawn Holmes makes, in any mode, runs
10//                     on holmes-lens whatever type it named, and a fork or a
11//                     teammate of his is refused
12//                  3. the workspace check: a gated Watson Direct-mode brief
13//                     whose Workdir: is a bare path, on a repo on main, master
14//                     or trunk, is refused with an ask-for-a-branch reason
15//                  4. routing: a dispatch of watson, holmes or lestrade runs the
16//                     mode agent its token picks (bin/compose-agents.sh builds
17//                     them), so the public names keep working
18//                  5. model and effort from the plugin's /config rows (its
19//                     options), and for holmes-local, holmes-index and
20//                     lestrade-item the config line (fanout, lensModel) added
21//                     to the prompt
22//                  Each dev-team agent's type is kept by its agentId.
23//   prompt.submit  the config line, as context, for the top-level loop of a
24//                  `claude -p --agent` run of those three modes: the run
25//                  bin/dispatch-agent.sh starts raises no agent.spawn
26//   turn.step      applies the effort step 5 recorded, to that sub-agent's
27//                  loop, and measures each dev-team request's working context
28//                  against the budget, notifying the human once per run
29//   turn.complete  deletes the run's scratch folder unless a child it spawned is
30//                  still live, and resets its budget notice
31//   session.end    deletes every scratch folder the mod still records, within
32//                  the end's short time budget: a run whose last turn ended
33//                  with a live child, and was never resumed, left one
34//   session.start  registers /dev-team-runs and /dev-team-board, and restarts
35//                  the panes' refresh timer when a reload finds a pane open
36//   command.run    those two commands open their pane (mods/panes.tsx)
37//   ui.render      draws the two panes, each read-only:
38//                  - runs: the newest dispatch logs (mods/runs.ts), read again
39//                    every 15 s while the pane is open
40//                  - board: The Index's three lanes from bin/dispatch-tick.sh
41//                    --board (mods/board.ts), fetched when the pane opens and
42//                    then once per dispatch cadence at most, and the items the
43//                    breaker escalated, from the log folder
44//   tool.call      on Bash, Edit, Write and NotebookEdit, before the call:
45//                  1. scratch: a dev-team agent's bare mktemp is pointed at a
46//                     folder of its own under a scratch root (mods/scratch.ts)
47//                  2. the commit guard (mods/commit-guard.ts), on every Bash
48//                     line in every lane: a forced or deleting push, a merge by
49//                     an agent or the pipeline, a sub-agent's commit or push,
50//                     and a commit, push or merge the ask rules cannot see
51//                  3. the commit subject check (mods/commit-subject.ts) on
52//                     every commit the guard lets through
53//                  4. the review guard (mods/review-guard.ts): a Holmes
54//                     reviewer, or anything he spawned, writes only in scratch
55//                  The guards read the line through $.workbench.parseShell,
56//                  take the lane from $.workbench.callerLane and isUnattended,
57//                  and refuse the call when they throw.
58//                  On Agent, from the main session (mods/dispatch.ts): a Watson
59//                  `Item ID: <n>` call runs bin/dispatch-agent.sh in its place
60//                  and answers with the dispatcher's first line, and a dev-team
61//                  call that did not ask for the foreground runs in the
62//                  background. After the call ran: the gate's advisory hint on
63//                  a complete brief that dictates method.
64//
65// The logic is pure and lives in mods/. Every hook that touches `$` lives in
66// this file, because the engine follows `$` into no imported function, and a
67// plugin registers each event once.
68//
69// The scheduled path never reaches agent.spawn: it starts the mode type
70// directly (`claude -p --agent workbench-dev-team:watson-index`), and
71// bin/dispatch-agent.sh passes the rows' model and effort as flags. The
72// other hooks run in that process too.
73//
74// The /config rows reach this module as its options. Claude Code reloads the
75// module when one changes, so `register` runs again with the new values.
76
77import { atom, read, update } from 'claude-code'
78import type { EngineInterface, Register, Timer, TurnUsage } from 'claude-code'
79
80import type { RunsView } from '../types'
81
82import { BOARD_TIMEOUT_MS, EMPTY_BOARD, afterFetch, boardArgv, boardOutcome, cadenceMsOf, escalatedOf, isDue } from './mods/board'
83
84import { budgetNotice, contextOf, isOverBudget } from './mods/budget'
85import { commitVerdict } from './mods/commit-guard'
86import type { ShellParse } from './mods/commit-guard'
87import { subjectVerdict } from './mods/commit-subject'
88import { DISPATCH_CONTEXT, dispatchDeny, dispatcherArgv, dispatchOutcome, dispatchResult, indexDispatchOf, isForcedBackground } from './mods/dispatch'
89import type { Disk } from './mods/review-guard'
90import { isReviewerType, judgeWrites, refusalOf, reviewBash, reviewEdit } from './mods/review-guard'
91import { BOARD_PANE, RUNS_PANE, boardTree, runsTree } from './mods/panes'
92import type { LogEnd } from './mods/runs'
93import { LOG_DIR, RUNS_REFRESH_MS, endsOf, livePairsOf, markersOf, newestRuns, rowsOf, scanArgv } from './mods/runs'
94import { hasBareMktemp, hasLiveChild, isDeletable, pointMktemp, prefixOf, readsAsPointed, rootOf, sweepOf } from './mods/scratch'
95import type { CallerLane } from './mods/spawn'
96import {
97  CONFIG_MODES,
98  DEFAULT_BRANCHES,
99  branchDeny,
100  branchUnreadDeny,
101  configLineOf,
102  configTextOf,
103  denyOf,
104  familyOf,
105  hintOf,
106  isDevTeamType,
107  isGated,
108  isHolmesMode,
109  knobsOf,
110  lensSpawnOf,
111  modeTypeOf,
112  uncheckedDeny,
113  withConfigLine,
114  workdirOf,
115} from './mods/spawn'
116
117// Whether the dispatch gate judges a call made in the loop `agentId` names.
118// The lane rejects while it is unknown; a gate then reads it as the main
119// session, so the brief is still checked. orchestratorIsOn already fails toward
120// on, and its .catch does the same.
121async function gated($: EngineInterface, agentId: string | undefined): Promise<boolean> {
122  const lane = await $.workbench.callerLane(agentId === undefined ? {} : { agentId }).catch(() => 'main' as const)
123  const isOn = await $.workbench.orchestratorIsOn().catch(() => true)
124  return isGated(lane, isOn)
125}
126
127// ── Whose run a loop is ──────────────────────────────────────────────────────
128
129// A dev-team run, as the per-run state keys it: its agent type, and its agentId,
130// or `run` for the top-level loop of a `claude -p --agent` run. Undefined for
131// the main session and for every agent outside the dev-team.
132type Run = { type: string; key: string }
133
134async function runOf($: EngineInterface, agentId: string | undefined): Promise<Run | undefined> {
135  if (agentId !== undefined) {
136    const { value: type } = await $.state.get({ plugin: 'workbench-dev-team', key: 'agentType', id: agentId })
137    return isDevTeamType(type) && type !== undefined ? { type, key: agentId } : undefined
138  }
139  if ((await $.workbench.callerLane({})) !== 'top-level-agent') return undefined
140  const type = await $.env.get('CLAUDE_CODE_AGENT')
141  return isDevTeamType(type) && type !== undefined ? { type, key: 'run' } : undefined
142}
143
144// The type of the loop a spawn happens in: the top-level `--agent` type, or
145// none, for a main-loop spawn; the type the mod kept, or the one
146// $.agent.list() gives, for a sub-agent's. Undefined when it cannot be told.
147async function spawnerTypeOf($: EngineInterface, parentAgentId: string | undefined): Promise<string | undefined> {
148  if (parentAgentId === undefined) return $.env.get('CLAUDE_CODE_AGENT').catch(() => undefined)
149  const { value: kept } = await $.state.get({ plugin: 'workbench-dev-team', key: 'agentType', id: parentAgentId })
150  if (kept !== undefined) return kept
151  const agents = await $.agent.list().catch(() => undefined)
152  return agents?.find(agent => agent.id === parentAgentId)?.type
153}
154
155// ── Scratch folders ──────────────────────────────────────────────────────────
156
157// The run's scratch folder: the one it has, or a new one under the first
158// scratch root, made with mktemp. Undefined when none can be made.
159async function scratchFolderOf($: EngineInterface, run: Run): Promise<string | undefined> {
160  const { value: kept } = await $.state.get({ plugin: 'workbench-dev-team', key: 'scratch', id: run.key })
161  if (kept && (await $.fs.exists(kept).catch(() => false))) return kept
162  const roots = await $.workbench.scratchRoots()
163  const root = rootOf(roots)
164  if (root === undefined) return undefined
165  const made = await $.process.run(['mktemp', '-d', `${root}/${prefixOf(run.type)}.XXXXXX`])
166  const folder = made.stdout.trim()
167  if (made.exitCode !== 0 || !isDeletable(folder, roots)) return undefined
168  await $.state.set({ plugin: 'workbench-dev-team', key: 'scratch', id: run.key }, folder)
169  await recordFolder($, [folder], 'add')
170  return folder
171}
172
173const FOLDERS = { plugin: 'workbench-dev-team', key: 'scratchFolders' } as const
174
175// Adds folders to the record, or takes them out. Two runs can write at once,
176// so the write lands only on the version it read, and is tried again when
177// another write came first.
178async function recordFolder($: EngineInterface, folders: readonly string[], change: 'add' | 'remove'): Promise<void> {
179  for (let attempt = 0; attempt < 8; attempt++) {
180    const { value = [], version } = await $.state.get(FOLDERS)
181    const rest = value.filter(kept => !folders.includes(kept))
182    if ((await $.state.set(FOLDERS, change === 'add' ? [...rest, ...folders] : rest, { ifVersion: version })).isSet) return
183  }
184}
185
186// The end of the session: every folder still recorded is deleted, in one rm,
187// when it lies under a scratch root. Each was kept because its run had not
188// ended, or a child of its run was live at the run's last turn, and nothing is
189// live once the session ends.
190// The rm is held to what the end's one short budget leaves, and skipped when
191// too little is left. A folder outside every root is never deleted. Only what
192// an rm that exited 0 deleted leaves the record: a folder the rm skipped, or
193// one an rm that failed or ran out of time may have left, stays recorded.
194async function endSession($: EngineInterface, budget: { remainingMs: number }): Promise<void> {
195  const { value: folders = [] } = await $.state.get(FOLDERS)
196  if (folders.length === 0) return
197  const sweep = sweepOf(folders, await $.workbench.scratchRoots(), budget.remainingMs)
198  if (sweep === undefined) return
199  const removed = await $.process.run(['rm', '-rf', '--', ...sweep.doomed], { timeoutMs: sweep.timeoutMs }).catch(() => undefined)
200  if (removed?.exitCode === 0) await recordFolder($, sweep.doomed, 'remove')
201}
202
203// The Bash line with a dev-team agent's bare mktemp pointed at its folder, or
204// the line as written when it has none, or when anything here fails: a mktemp
205// left in $TMPDIR is how the line ran before the mod.
206async function withScratch($: EngineInterface, agentId: string | undefined, line: string): Promise<string> {
207  if (!line.includes('mktemp')) return line
208  try {
209    const before = await $.workbench.parseShell(line)
210    if (!hasBareMktemp(before)) return line
211    const run = await runOf($, agentId)
212    if (run === undefined) return line
213    const folder = await scratchFolderOf($, run)
214    if (folder === undefined) return line
215    const pointed = pointMktemp(line, folder)
216    if (pointed === undefined) return line
217    return readsAsPointed(before, await $.workbench.parseShell(pointed), folder) ? pointed : line
218  } catch {
219    return line
220  }
221}
222
223// The end of a turn: the run's scratch folder deleted, unless a child the run
224// spawned is still live, and its budget notice reset. A child's end never
225// deletes its parent's folder: an ended agent may resume, as Holmes does when
226// his background helpers report, and read it again. A folder kept for a live
227// child goes at the end of the next turn that finds none. When the agent list
228// cannot be read, the folder stays.
229async function endRun($: EngineInterface, agentId: string | undefined): Promise<void> {
230  const key = agentId ?? ((await $.workbench.callerLane({})) === 'top-level-agent' ? 'run' : undefined)
231  if (key === undefined) return
232  const { value: folder } = await $.state.get({ plugin: 'workbench-dev-team', key: 'scratch', id: key })
233  const agents = folder ? await $.agent.list().catch(() => undefined) : undefined
234  if (folder && agents !== undefined && !hasLiveChild(agents, agentId)) {
235    if (isDeletable(folder, await $.workbench.scratchRoots())) await $.process.run(['rm', '-rf', '--', folder])
236    await $.state.set({ plugin: 'workbench-dev-team', key: 'scratch', id: key }, '')
237    await recordFolder($, [folder], 'remove')
238  }
239  const { value: told } = await $.state.get({ plugin: 'workbench-dev-team', key: 'overBudget', id: key })
240  if (told) await $.state.set({ plugin: 'workbench-dev-team', key: 'overBudget', id: key }, false)
241}
242
243// ── The context budget ───────────────────────────────────────────────────────
244
245async function measure($: EngineInterface, agentId: string | undefined, usage: TurnUsage | null): Promise<void> {
246  const tokens = contextOf(usage)
247  if (!isOverBudget(tokens) || tokens === undefined) return
248  const run = await runOf($, agentId)
249  if (run === undefined) return
250  const { value: told } = await $.state.get({ plugin: 'workbench-dev-team', key: 'overBudget', id: run.key })
251  if (told) return
252  await $.state.set({ plugin: 'workbench-dev-team', key: 'overBudget', id: run.key }, true)
253  const notice = budgetNotice(run.type, run.key, tokens)
254  $.ui.toast(notice, { timeoutMs: 15000 })
255  $.ui.log(notice, { to: 'transcript' })
256}
257
258// ── The Index dispatcher ─────────────────────────────────────────────────────
259
260// The Agent call's answer for a Watson Index-mode dispatch: the dispatcher's
261// first line, or a refusal. It never falls through to a spawn, which would
262// start a run without its pipeline flag.
263async function runDispatcher($: EngineInterface, itemId: string, prompt: string) {
264  const home = await $.env.get('HOME').catch(() => undefined)
265  if (!home) return { deny: dispatchDeny('HOME is not set, so the dispatcher cannot be found') }
266  const run = await $.process.run(dispatcherArgv(home, itemId), { timeoutMs: 120_000 }).catch(() => undefined)
267  if (run === undefined) return { deny: dispatchDeny('the dispatcher could not run') }
268  const outcome = dispatchOutcome(run)
269  if ('deny' in outcome) return outcome
270  return { result: dispatchResult(outcome.line, prompt), context: [DISPATCH_CONTEXT] }
271}
272
273// ── The guards ───────────────────────────────────────────────────────────────
274
275const GUARDED: ReadonlySet<string> = new Set(['Bash', 'Edit', 'Write', 'NotebookEdit'])
276
277const GUARD_FAILED =
278  '🛑 Blocked: a call the guards could not judge.\n\nworkbench-dev-team. The commit guard or the review guard failed while reading this call, so it is refused rather than let an unread commit, push, merge, or write through. Report this to the human as a guard defect. Do not try another spelling.'
279
280// The input fields the guards read. The engine types them per tool; these are
281// read as unknown, so a field of the wrong type is judged as absent.
282type GuardInput = { tool: string; agentId?: string; command?: unknown; file_path?: unknown; notebook_path?: unknown }
283
284// Whether a call comes from a Holmes reviewer, or from anything a reviewer
285// spawned: true, false, or undefined when it cannot be told. A top-level
286// `--agent` run names its type in CLAUDE_CODE_AGENT, and every loop in it is
287// held. A sub-agent's type, and its parents', come from $.agent.list().
288async function isReviewer($: EngineInterface, agentId: string | undefined, lane: CallerLane | undefined): Promise<boolean | undefined> {
289  if (isReviewerType(await $.env.get('CLAUDE_CODE_AGENT').catch(() => undefined))) return true
290  if (lane === 'main' || lane === 'top-level-agent') return false
291  if (lane === undefined || agentId === undefined) return undefined
292  const agents = await $.agent.list().catch(() => undefined)
293  if (agents === undefined) return undefined
294  const seen = new Set<string>()
295  for (let id: string | undefined = agentId; id !== undefined && !seen.has(id); ) {
296    seen.add(id)
297    const agent = agents.find(a => a.id === id)
298    if (agent === undefined) return undefined
299    if (isReviewerType(agent.type)) return true
300    id = agent.parentId
301  }
302  return false
303}
304
305// The scratch roots: core's, and $TMPDIR, each a physical directory.
306async function scratchRootsOf($: EngineInterface): Promise<string[]> {
307  const roots = [...(await $.workbench.scratchRoots().catch(() => []))]
308  const tmp = await $.env.get('TMPDIR').catch(() => undefined)
309  const real = tmp ? (await $.fs.stat(tmp, { resolve: true }).catch(() => undefined))?.realPath : undefined
310  if (real !== undefined && (await $.fs.stat(real).catch(() => undefined))?.kind === 'dir') roots.push(real)
311  return roots
312}
313
314// The review guard's refusal for a reviewer's call, or undefined.
315async function reviewDeny($: EngineInterface, e: GuardInput, parse: ShellParse | undefined): Promise<string | undefined> {
316  const cwd = await $.session.cwd().catch(() => undefined)
317  const context = { cwd, home: await $.env.get('HOME').catch(() => undefined) }
318  let verdict
319  if (e.tool === 'Bash') {
320    if (parse === undefined) return undefined
321    verdict = reviewBash(parse, context)
322  } else {
323    const raw = e.tool === 'NotebookEdit' ? e.notebook_path : e.file_path
324    verdict = reviewEdit(e.tool, typeof raw === 'string' ? raw : '', context)
325  }
326  const roots = await scratchRootsOf($)
327  if ('finding' in verdict) return refusalOf(verdict.finding, roots)
328  const disk: Disk = {
329    stat: path => $.fs.stat(path, { resolve: true }).then(stat => ({ realPath: stat.realPath }), () => undefined),
330    names: dir => $.fs.list(dir).then(entries => entries.map(entry => entry.name), () => undefined),
331  }
332  const finding = await judgeWrites(verdict.writes, roots, disk)
333  return finding === undefined ? undefined : refusalOf(finding, roots)
334}
335
336// The guards on one call: the commit guard and the subject check on a Bash line
337// in every lane, then the review guard when a reviewer makes the call, or when
338// who makes it cannot be told and the call writes outside scratch.
339async function guard($: EngineInterface, e: GuardInput): Promise<string | undefined> {
340  const lane = await $.workbench.callerLane(e.agentId === undefined ? {} : { agentId: e.agentId }).catch(() => undefined)
341  let parse: ShellParse | undefined
342  if (e.tool === 'Bash') {
343    const line = typeof e.command === 'string' ? e.command : ''
344    parse = await $.workbench.parseShell(line)
345    const isUnattended = await $.workbench.isUnattended().catch(() => true)
346    const refused = commitVerdict(parse, line, { lane, isUnattended }) ?? subjectVerdict(parse)
347    if (refused !== undefined) return refused.deny
348  }
349  const reviewer = await isReviewer($, e.agentId, lane)
350  if (reviewer === false) return undefined
351  const deny = await reviewDeny($, e, parse)
352  if (deny === undefined || reviewer) return deny
353  return `${deny}\n\nThe guard could not tell which agent made this call, so it held the call to the reviewer's rule. Report this to the human as a guard defect.`
354}
355
356// ── The panes ────────────────────────────────────────────────────────────────
357
358// Both panes only read. The runs pane lists the log folder and runs awk and
359// ps. The board pane runs the tick's board mode, which uses the cached token
360// only: it never mints, never reads the Keychain, and writes nothing.
361
362const RUNS_COMMAND = 'dev-team-runs'
363const BOARD_COMMAND = 'dev-team-board'
364const PANE_TITLE = { [RUNS_PANE]: 'Dev-team runs', [BOARD_PANE]: 'The Index board' } as const
365
366const runsAtom = atom({ plugin: 'workbench-dev-team', key: 'runs' } as const, { rows: [] } as RunsView)
367const boardAtom = atom({ plugin: 'workbench-dev-team', key: 'board' } as const, EMPTY_BOARD)
368
369// What the panes keep between refreshes in one load of the module: each read
370// log's end by name, with the size and mtime it was read at, whether a refresh
371// of each pane is under way, and the refresh timer.
372type PaneMemory = { ends: Map<string, { stamp: string; end: LogEnd }>; isRunsBusy: boolean; isBoardBusy: boolean; timer?: Timer }
373
374async function logDirOf($: EngineInterface): Promise<string | undefined> {
375  const home = await $.env.get('HOME').catch(() => undefined)
376  return home ? `${home}/${LOG_DIR}` : undefined
377}
378
379const stampOf = (run: { mtimeMs: number; size: number }): string => `${run.mtimeMs}:${run.size}`
380const NO_END: LogEnd = { tail: [], refusals: 0 }
381
382// The runs view from the newest logs. Only a log whose size or mtime changed
383// since its last read is read again.
384async function readRuns($: EngineInterface, memory: PaneMemory): Promise<RunsView> {
385  const dir = await logDirOf($)
386  if (dir === undefined) return { rows: [], error: 'HOME is not set, so the dispatch logs cannot be found.' }
387  const entries = await $.fs.list(dir).catch(() => undefined)
388  if (entries === undefined) return { rows: [], error: `There is no dispatch log folder at ${dir}.` }
389  const runs = newestRuns(entries)
390  if (runs.length === 0) return { rows: [] }
391  const changed = runs.filter(run => memory.ends.get(run.name)?.stamp !== stampOf(run))
392  if (changed.length > 0) {
393    const scan = await $.process.run(scanArgv($.plugin.root, changed.map(run => `${dir}/${run.name}`)), { timeoutMs: 20_000 }).catch(() => undefined)
394    // The scan skips a log it cannot open, so a non-zero exit with output is
395    // still the ends of the logs it read. No output and a failure is no read.
396    if (scan === undefined || (scan.exitCode !== 0 && scan.stdout === '')) return { rows: [], error: 'The dispatch logs could not be read.' }
397    const scanned = endsOf(scan.stdout)
398    for (const run of changed) {
399      const end = scanned.get(`${dir}/${run.name}`)
400      if (end === undefined) memory.ends.delete(run.name)
401      else memory.ends.set(run.name, { stamp: stampOf(run), end })
402    }
403  }
404  // A changed log the scan could not open was deleted since the listing: its
405  // run leaves the pane, and the others stay.
406  const read = runs.filter(run => memory.ends.get(run.name)?.stamp === stampOf(run))
407  for (const name of memory.ends.keys()) if (!read.some(run => run.name === name)) memory.ends.delete(name)
408  const ps = await $.process.run(['ps', '-axo', 'command='], { timeoutMs: 10_000 }).catch(() => undefined)
409  if (ps === undefined || ps.exitCode !== 0) return { rows: [], error: 'The process list could not be read, so which runs are live is unknown.' }
410  const ends = new Map(read.map(run => [`${dir}/${run.name}`, memory.ends.get(run.name)?.end ?? NO_END]))
411  return { rows: rowsOf(dir, read, ends, livePairsOf(ps.stdout), markersOf(entries)) }
412}
413
414async function refreshRuns($: EngineInterface, memory: PaneMemory): Promise<void> {
415  if (memory.isRunsBusy) return
416  memory.isRunsBusy = true
417  try {
418    const view = await readRuns($, memory)
419    await update($, runsAtom, () => view)
420  } finally {
421    memory.isRunsBusy = false
422  }
423}
424
425// The escalated items, read from the log folder on every refresh, and the
426// board from The Index when the cadence allows. The attempt is recorded before
427// the client runs, so no later refresh, and no reload, starts a second one
428// inside the cadence.
429async function refreshBoard($: EngineInterface, memory: PaneMemory, cadenceMs: number): Promise<void> {
430  const dir = await logDirOf($)
431  const entries = dir === undefined ? [] : await $.fs.list(dir).catch(() => [])
432  const escalated = escalatedOf(markersOf(entries))
433  await update($, boardAtom, view => ({ ...view, escalated }))
434  if (memory.isBoardBusy) return
435  const now = await $.clock.now()
436  if (!isDue(await read($, boardAtom), now, cadenceMs)) return
437  memory.isBoardBusy = true
438  try {
439    await update($, boardAtom, view => ({ ...view, attemptedAt: now }))
440    const run = await $.process.run(boardArgv($.plugin.root), { timeoutMs: BOARD_TIMEOUT_MS }).catch(() => undefined)
441    const outcome = boardOutcome(run)
442    await update($, boardAtom, view => afterFetch(view, now, outcome))
443  } finally {
444    memory.isBoardBusy = false
445  }
446}
447
448// One refresh of every open pane. With none open, the timer stops.
449async function refreshPanes($: EngineInterface, memory: PaneMemory, cadenceMs: number): Promise<void> {
450  const open = new Set((await $.ui.panes()).map(pane => pane.id))
451  if (!open.has(RUNS_PANE) && !open.has(BOARD_PANE)) {
452    memory.timer?.cancel()
453    memory.timer = undefined
454    return
455  }
456  await Promise.all([open.has(RUNS_PANE) ? refreshRuns($, memory) : undefined, open.has(BOARD_PANE) ? refreshBoard($, memory, cadenceMs) : undefined])
457}
458
459function startRefresh($: EngineInterface, memory: PaneMemory, cadenceMs: number): void {
460  memory.timer ??= $.clock.every(RUNS_REFRESH_MS, () => void refreshPanes($, memory, cadenceMs).catch(() => undefined))
461}
462
463export const register: Register = (on, options) => {
464  // The rows, read once: the options are fixed for this activation.
465  const config = configTextOf(options)
466
467  // The panes. Neither one writes: see refreshRuns and refreshBoard.
468  const cadenceMs = cadenceMsOf(options.dispatchCadenceMinutes)
469  const memory: PaneMemory = { ends: new Map(), isRunsBusy: false, isBoardBusy: false }
470
471  on('session.start', async ($, e, next) => {
472    await $.command.register({ name: RUNS_COMMAND, description: 'Show live and recent dev-team runs from the dispatch logs', immediate: true }).catch(() => undefined)
473    await $.command.register({ name: BOARD_COMMAND, description: "Show The Index board's lanes, fetched once per dispatch cadence at most", immediate: true }).catch(() => undefined)
474    const open = await $.ui.panes().catch(() => [])
475    if (open.some(pane => pane.id === RUNS_PANE || pane.id === BOARD_PANE)) startRefresh($, memory, cadenceMs)
476    return next(e)
477  })
478
479  on('command.run', async ($, e, next) => {
480    const pane = e.command === RUNS_COMMAND ? RUNS_PANE : e.command === BOARD_COMMAND ? BOARD_PANE : undefined
481    if (pane === undefined) return next(e)
482    await $.ui.open({ id: pane, title: PANE_TITLE[pane] })
483    startRefresh($, memory, cadenceMs)
484    void refreshPanes($, memory, cadenceMs).catch(() => undefined)
485    return { text: `${PANE_TITLE[pane]} pane opened.` }
486  }).catch(($, e, next) => (e.command === RUNS_COMMAND || e.command === BOARD_COMMAND ? { text: 'The dev-team pane could not open. Try the command again.' } : next(e)))
487
488  on('ui.render', { component: 'Pane' }, async ($, e, next) => {
489    if (e.requestId === RUNS_PANE) return runsTree($.ui.resolve(e), await read($, runsAtom), e.props.scroll.bodyRows)
490    if (e.requestId === BOARD_PANE) return boardTree($.ui.resolve(e), await read($, boardAtom), cadenceMs)
491    return next(e)
492  })
493
494  // The gate fails closed: a dispatch it judges whose brief cannot be checked
495  // is refused, and so is one whose gate throws before next is called.
496  // workbench-core is a declared dependency, so a check that rejects is a
497  // fault, and /orchestrator off is the way past it. The helper rule and the
498  // workspace check are refusals too, and each says how it fails below. What
499  // follows them (the config, the routing) is not a guard: when it fails, the
500  // dispatch goes on unchanged, unrouted and on the agent file's own model,
501  // which is how the public type runs without this module.
502  on('agent.spawn', async ($, e, next) => {
503    // Only the model's own Agent calls are judged, the ones the bash gate sees.
504    // A plugin's $.agent.spawn is the plugin's implementation, not a handoff.
505    const isModel = next.origin.plugin === 'engine'
506    const isGatedCall = isModel && (await gated($, e.parentAgentId))
507    if (isGatedCall) {
508      const check = await $.workbench.briefCheck(e.prompt).catch(() => undefined)
509      if (check === undefined) return { deny: uncheckedDeny() }
510      if (!check.isComplete) {
511        const slots = await $.workbench.briefSlots().catch(() => [])
512        return { deny: denyOf(check.missing, slots) }
513      }
514    }
515
516    // The helper rule. A spawner the mod cannot name is let through: the
517    // review guard still holds every write by anything Holmes spawned.
518    let input = e
519    if (isModel && isHolmesMode(await spawnerTypeOf($, e.parentAgentId))) {
520      const lens = lensSpawnOf(e)
521      if ('deny' in lens) return lens
522      input = { ...e, subagentType: lens.subagentType }
523    }
524
525    // The workspace check. git answering with no branch (not a repository, or
526    // a detached HEAD) passes; git that cannot run at all refuses, since the
527    // answer is then unknown, and recording a branch in Workdir: passes.
528    if (isGatedCall && familyOf(input.subagentType) === 'watson' && modeTypeOf('watson', input.prompt).endsWith(':watson-direct')) {
529      const workdir = workdirOf(input.prompt)
530      if (workdir !== undefined && !workdir.isRecorded && workdir.path.startsWith('/')) {
531        const head = await $.process
532          .run(['git', '-C', workdir.path, 'symbolic-ref', '--quiet', '--short', 'HEAD'], { timeoutMs: 10_000 })
533          .catch(() => undefined)
534        if (head === undefined) return { deny: branchUnreadDeny(workdir.path) }
535        const branch = head.stdout.trim()
536        if (head.exitCode === 0 && DEFAULT_BRANCHES.includes(branch)) return { deny: branchDeny(workdir.path, branch) }
537      }
538    }
539
540    // Nothing in this block calls next, so a failure here passes the dispatch
541    // on unchanged, and next is called once either way.
542    let routed: typeof e | undefined
543    let knobs: ReturnType<typeof knobsOf> = {}
544    try {
545      const family = familyOf(input.subagentType)
546      if (family !== undefined) {
547        knobs = knobsOf(config, family)
548        const subagentType = modeTypeOf(family, input.prompt)
549        const configFamily = CONFIG_MODES[subagentType]
550        const prompt = configFamily === undefined ? input.prompt : withConfigLine(input.prompt, configLineOf(config, configFamily))
551        // A model the caller named is the caller's choice, and stands.
552        const model = input.model ?? knobs.model
553        routed = { ...input, subagentType, prompt, ...(model === undefined ? {} : { model }) }
554      }
555    } catch {
556      routed = undefined
557    }
558    const spawned = routed ?? input
559    const result = await next(spawned)
560    if (result.agentId !== undefined) {
561      if (routed !== undefined && knobs.effort !== undefined) {
562        await $.state.set({ plugin: 'workbench-dev-team', key: 'effort', id: result.agentId }, knobs.effort)
563      }
564      if (isDevTeamType(spawned.subagentType)) {
565        await $.state.set({ plugin: 'workbench-dev-team', key: 'agentType', id: result.agentId }, spawned.subagentType)
566      }
567    }
568    return result
569  }).catch(($, e, next) => (next.called ? next(e) : { deny: uncheckedDeny() }))
570
571  // The config line for a top-level run of a mode that reads it. Nothing here
572  // calls next, so a failure leaves the prompt as it came.
573  on('prompt.submit', async ($, e, next) => {
574    let line: string | undefined
575    try {
576      if ((await $.workbench.callerLane({})) === 'top-level-agent') {
577        const family = CONFIG_MODES[(await $.env.get('CLAUDE_CODE_AGENT')) ?? '']
578        if (family !== undefined) line = configLineOf(config, family)
579      }
580    } catch {
581      line = undefined
582    }
583    return next(line === undefined ? e : { ...e, context: [...(e.context ?? []), line] })
584  })
585
586  on('turn.step', async function* ($, e, next) {
587    let input = e
588    if (e.agentId !== undefined) {
589      const { value: effort } = await $.state.get({ plugin: 'workbench-dev-team', key: 'effort', id: e.agentId })
590      if (effort !== undefined) input = { ...e, effort }
591    }
592    const result = yield* next(input)
593    await measure($, e.agentId, result.usage).catch(() => undefined)
594    return result
595  })
596
597  on('turn.complete', async ($, e, next) => {
598    const result = await next(e)
599    await endRun($, e.agentId).catch(() => undefined)
600    return result
601  })
602
603  // The sweep runs before the engine's own end step, and its failure leaves
604  // the folders where they are rather than holding the exit up.
605  on('session.end', async ($, e, next) => {
606    await endSession($, next.budget).catch(() => undefined)
607    return next(e)
608  })
609
610  // One tool.call hook, since a plugin registers each event once. A guard that
611  // throws refuses the call (fail closed), a scratch rewrite that fails leaves
612  // the line as written, and the hint's failure leaves the call as it ran.
613  on('tool.call', async ($, e, next) => {
614    if (e.tool === 'Bash' || e.tool === 'Edit' || e.tool === 'Write' || e.tool === 'NotebookEdit') {
615      const call = e.tool === 'Bash' && typeof e.command === 'string' ? { ...e, command: await withScratch($, e.agentId, e.command) } : e
616      const deny = await guard($, call)
617      return deny === undefined ? next(call) : { deny }
618    }
619    if (e.tool !== 'Agent') return next(e)
620    const isModel = next.origin.plugin === 'engine'
621    let call = e
622    if (isModel) {
623      // A set agentId is a sub-agent's call, whatever callerLane says, so a
624      // sub-agent never reaches the dispatcher. Only a call with no agentId
625      // asks the lane, and a lookup that fails there reads as the main session.
626      const isMain = e.agentId === undefined && (await $.workbench.callerLane({}).catch(() => 'main' as const)) === 'main'
627      if (isMain) {
628        const itemId = indexDispatchOf(e.subagent_type, e.prompt)
629        if (itemId !== undefined) return runDispatcher($, itemId, e.prompt)
630        if (isForcedBackground(e.subagent_type, e.run_in_background)) call = { ...e, run_in_background: true }
631      }
632    }
633    const result = await next(call)
634    if (result.deny !== undefined || result.isError || !isModel) return result
635    if (typeof e.prompt !== 'string' || !(await gated($, e.agentId))) return result
636    const check = await $.workbench.briefCheck(e.prompt).catch(() => undefined)
637    const hint = check === undefined ? undefined : hintOf(check, e.prompt)
638    return hint === undefined ? result : { ...result, context: [...(result.context ?? []), hint] }
639  }).catch(($, e, next) => (next.called || !GUARDED.has(e.tool) ? next(e) : { deny: GUARD_FAILED }))
640}
641
hooks/mods/board.ts 103 lines
1// The board pane's logic, as pure functions. hooks/register.ts holds the hooks;
2// hooks/mods/panes.tsx draws the lanes.
3//
4// The pane reaches The Index only through `bin/dispatch-tick.sh --board`, run
5// from this plugin's own folder. That mode uses the tick's cached token only
6// (it never mints, never reads the Keychain, and writes nothing), calls the
7// three list tools and nothing else, and prints the lanes as JSON. So no token
8// or secret ever reaches this module. The pane runs
9// it when it opens and then once per dispatch cadence at most, counting every
10// attempt, a failed one included.
11
12// The view types live in the type contract, types/index.d.ts.
13import type { Board, BoardItem, BoardLane, BoardView } from '../../types'
14
15export const LANES = ['unrefined', 'review', 'development'] as const
16
17export const EMPTY_BOARD: BoardView = { attemptedAt: 0, escalated: [] }
18
19// The dispatch cadence's default, as in .claude-plugin/plugin.json.
20export const DEFAULT_CADENCE_MINUTES = 20
21
22// How long a board fetch may take: three list calls of up to 120 s each in the
23// tick's own client.
24export const BOARD_TIMEOUT_MS = 400_000
25
26// The cadence in milliseconds, from the /config row. A value that is not a
27// number of at least one minute reads as the default, so a bad row never
28// makes the pane call The Index more often.
29export function cadenceMsOf(minutes: unknown): number {
30  const value = typeof minutes === 'number' && Number.isFinite(minutes) && minutes >= 1 ? minutes : DEFAULT_CADENCE_MINUTES
31  return value * 60_000
32}
33
34// Whether the pane may fetch the board now: never fetched, or the last attempt
35// started a whole cadence ago.
36export const isDue = (view: BoardView, now: number, cadenceMs: number): boolean => view.attemptedAt === 0 || now - view.attemptedAt >= cadenceMs
37
38export const boardArgv = (root: string): string[] => ['bash', `${root}/bin/dispatch-tick.sh`, '--board']
39
40const isObject = (value: unknown): value is Record<string, unknown> => typeof value === 'object' && value !== null && !Array.isArray(value)
41const text = (value: unknown): string | null => (typeof value === 'string' ? value : null)
42
43function itemOf(value: unknown): BoardItem | undefined {
44  if (!isObject(value) || typeof value.id !== 'number') return undefined
45  return {
46    id: value.id,
47    number: typeof value.number === 'number' ? value.number : null,
48    isPr: value.isPr === true,
49    repo: text(value.repo),
50    title: text(value.title),
51    claimedAt: text(value.claimedAt),
52  }
53}
54
55function laneOf(value: unknown): BoardLane | undefined {
56  if (!isObject(value)) return undefined
57  if (typeof value.error === 'string') return { error: value.error.slice(0, 300) }
58  if (typeof value.limit !== 'number' || !Array.isArray(value.items)) return undefined
59  const items = value.items.map(itemOf)
60  return items.every(item => item !== undefined) ? { limit: value.limit, items: items as BoardItem[] } : undefined
61}
62
63// The board the client printed, or undefined when its output is anything else.
64// Only the named fields are kept, each of its own type.
65export function boardOf(stdout: string): Board | undefined {
66  let parsed: unknown
67  try {
68    parsed = JSON.parse(stdout)
69  } catch {
70    return undefined
71  }
72  if (!isObject(parsed) || !isObject(parsed.lanes)) return undefined
73  const lanes = parsed.lanes
74  const [unrefined, review, development] = LANES.map(lane => laneOf(lanes[lane]))
75  return unrefined && review && development ? { unrefined, review, development } : undefined
76}
77
78// What one run of the client means: the board, or why there is none. The
79// client prints its reason as the last line on stderr, and never a token.
80export function boardOutcome(run: { exitCode: number; stdout: string; stderr: string } | undefined): { board: Board } | { error: string } {
81  if (run === undefined) return { error: 'the board client could not run' }
82  const board = run.exitCode === 0 ? boardOf(run.stdout) : undefined
83  if (board !== undefined) return { board }
84  const why = run.stderr.split('\n').findLast(line => line.trim() !== '')?.trim()
85  if (run.exitCode === 0) return { error: 'the board client printed something that is not a board' }
86  return { error: (why ?? `the board client exited ${run.exitCode}`).slice(0, 300) }
87}
88
89// The view after a fetch that started at `startedAt`. A failed fetch keeps the
90// last board, so the pane still shows it, with its age.
91export function afterFetch(view: BoardView, startedAt: number, outcome: { board: Board } | { error: string }): BoardView {
92  if ('board' in outcome) return { attemptedAt: startedAt, board: outcome.board, boardAt: startedAt, escalated: view.escalated }
93  const { board, boardAt } = view
94  return { attemptedAt: startedAt, ...(board ? { board, boardAt } : {}), error: outcome.error, escalated: view.escalated }
95}
96
97// "watson item 42" for each escalated pair "watson-42".
98export const escalatedOf = (markers: readonly string[]): string[] =>
99  markers.map(pair => {
100    const dash = pair.lastIndexOf('-')
101    return `${pair.slice(0, dash)} item ${pair.slice(dash + 1)}`
102  })
103
hooks/mods/budget.ts 25 lines
1// Each dev-team agent's working-context budget, as pure functions.
2// hooks/register.ts measures every model request of a dev-team agent's loop in
3// turn.step, and notifies the human once when a request passes the budget. It
4// stops nothing: the agents' rule is to finish the work and say what made it
5// expensive, never to buy the budget with the work.
6
7import type { TurnUsage } from 'claude-code'
8
9// The budget every dev-team agent's prose states: about 250k tokens of working
10// context.
11export const BUDGET_TOKENS = 250_000
12
13// A request's working context: every input token it was answered over, cached
14// or not. Undefined when the request carried no usage.
15export const contextOf = (usage: TurnUsage | null | undefined): number | undefined =>
16  usage == null ? undefined : usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
17
18export const isOverBudget = (tokens: number | undefined): boolean => tokens !== undefined && tokens > BUDGET_TOKENS
19
20// The notice, one line: which agent, how far past, and that nothing stopped.
21export function budgetNotice(type: string, id: string, tokens: number): string {
22  const k = (n: number) => `${Math.round(n / 1000)}k`
23  return `📏 workbench-dev-team: ${type.split(':').pop()} (${id}) passed its ${k(BUDGET_TOKENS)}-token working-context budget, at ${k(tokens)} in its latest request. Nothing was stopped.`
24}
25
hooks/mods/commit-guard.ts 520 lines
1// The commit guard, as pure functions over workbench-core's reading of a Bash
2// line ($.workbench.parseShell). hooks/register.ts calls them on every Bash
3// call, in every lane. It refuses what the permissions.ask rules cannot cover,
4// and is silent on everything else.
5//
6// The approval is a "Commit it" pick in `AskUserQuestion`, once the human says
7// their review is done. Claude Code's own prompt is the mechanical backstop.
8// /workbench-dev-team:setup installs the ask rules git commit *, git push *,
9// git * commit *, git * push *, git * commit, git * push, gh * pr merge *,
10// gh * pr merge, and gh api *pulls/*/merge*, beside workbench-core's own
11// gh pr merge:*. This guard does not ask and does not approve. It catches honest
12// mistakes. It is not a security boundary: a script file, an interpreter, or a
13// shell alias gets past it, and past the ask rules too. Review is the real gate.
14//
15// ACCEPTED LIMITS (Mike, 2026-10-07: the guards stay mistake-catchers, and
16// these gaps are accepted rather than chased). The guard reads the words of a
17// line, so a change it cannot see in them is out of scope: a value-only change
18// such as PATH=, HOME= or XDG_CONFIG_HOME= before git; git -C into a
19// repository whose own config is hostile; an arithmetic assignment such as
20// $((X=1)); a bash alias defined on the same line; a git alias in a config
21// file; a script file or a file passed to make; a command that code builds at
22// run time, such as python joining "git" and "push"; and a GraphQL mutation
23// held in a variable or a file (`-f query="$Q"`, `--input file`).
24//
25// It refuses, in this order:
26//   1. A push that forces or deletes, in every lane: --force, --force-with-lease,
27//      --mirror, --delete, --prune (and their prefixes), -f or -d in a short
28//      cluster, a refspec that starts with + or :, or a -c that makes a remote
29//      mirror or names its push refspec.
30//   2. A pull request merge from a sub-agent, from a top-level `--agent` run
31//      such as the pipeline, or from any turn nobody attends: gh pr merge in any
32//      spelling, or gh api on pulls/<n>/merge. Holmes and Watson never merge.
33//   3. A commit or push from a sub-agent. A sub-agent hands its work back
34//      uncommitted. The scheduled pipeline commits from its top-level
35//      `claude -p --agent` loop, which is not a sub-agent.
36//   4. A commit, push, or merge the ask rules cannot see, in every lane: behind
37//      a wrapper the harness does not strip before it matches (env, sudo,
38//      xargs, caffeinate, ...), a NAME=value assignment, a `bash -c` or `eval`
39//      script, a substitution, or a heredoc fed to a shell;
40//      a git or gh named by a path or in another case; an escaped word; a
41//      subcommand a `-c alias.<name>` supplies; or a subcommand, or a global
42//      option word before it, built at run time from a substitution or a
43//      variable (`git "$(echo commit)"`, `git $x`, `gh pr $(echo merge)`). The
44//      guard cannot tell which subcommand that is, so it counts as any of
45//      them. An option's value (`-C "$D"`, `--git-dir=$D`, `-R"$R"`) decides
46//      no subcommand. A gh api endpoint built at run time counts as a merge
47//      when the method is PUT or built, or a built word stands where a flag
48//      could (the REST merge takes a PUT), and a gh api line whose fields
49//      name the GraphQL mergePullRequest or enablePullRequestAutoMerge
50//      mutation counts as a merge whatever its endpoint and method.
51//   5. In a sub-agent, or in a run nobody attends (an attended top-level
52//      `--agent` run is neither): a line whose command name the
53//      reader cannot place (a wrapper option it cannot read) or that is built
54//      at run time from a substitution or a variable (`"$CMD" x`,
55//      `g$(echo it) push`), whatever the line names. Mike accepted the cost on
56//      2026-10-07: a rare `$CMD …` line refused there. A plain `"$NAME/…"`
57//      before a literal path names its program, so it is not refused. Since
58//      workbench-core 4e83554, core's guards refuse each of these lines first,
59//      in every lane. This rule stays as dev-team's own check in the lanes no
60//      human watches. A substitution in front of a literal path (`$(…)x/y`)
61//      is refused here through the reader's own `expansion` unknown, since
62//      workbench-core 47a5e27. So is any line with any unknown at all
63//      (HIDES_NAME): each can leave a command with no name to read, and the
64//      cost is a rare line refused, such as one whose $'…' holds an escape the
65//      reader does not decode (a Unicode or control escape).
66//
67// THE FOURTH RULE IS KEPT. workbench-core's commit approval gate reads past
68// every hidden form above, but only for a commit or push in the main loop of an
69// attended session. It stands aside for a sub-agent, a top-level `--agent` run,
70// a session nobody sits at, and a turn a schedule opened, and it never looks at
71// a merge. In those places a hidden form draws no prompt from the ask rules and
72// no question from core, so this rule is the only thing that stops it.
73//
74// LANES. The lane is $.workbench.callerLane's: `main`, `sub-agent`, or
75// `top-level-agent` (a `claude -p --agent` run, the pipeline among them). A
76// lane the noun cannot give is read as a sub-agent, the side that refuses. The
77// merge rule also reads $.workbench.isUnattended, and an unknown answer there is
78// read as unattended.
79//
80// READING. Each rule reads the statements parseShell returns, never the text of
81// the line, so `git commit -m "load env"`, `grep "git push" notes.md`, and a
82// heredoc data file that quotes `git push origin` are not refused. Text is
83// read in three places only, each where a command may be hidden from the
84// statements:
85//   - A line parseShell could not read whole (any unknown), or with a wrapper
86//     it could not place, that names commit, push, or merge anywhere: rule 4
87//     in every lane, and rule 1 when the text reads as a forced push.
88//   - Code a program runs from its own words or heredoc: the arguments of
89//     anything that is not a plain reader (python -c, awk system(), sed's e
90//     command, a runner the shell reader does not list, `rg --pre`), a git -c
91//     value, `git grep -O`, and a `-c alias.<name>=!…` shell alias. A mention of
92//     git commit, push, or gh pr merge there counts as that command behind a
93//     wrapper: rules 1 to 4. A mention that is only a quote is refused too,
94//     on purpose: a `node -e` that counts the string 'git push' reads the same
95//     as one that runs it through child_process, and the guard cannot show a
96//     script cannot run git. Node reaches child_process with no fixed word in
97//     the text (`require(['child', 'process'].join('_'))`), so no word list
98//     proves its absence. This is the cost of a mistake-catcher that does not
99//     read code (workbench-core report, 2026-10-09).
100//   - Text a plain reader handles (echo, cat, a heredoc data file) on a line
101//     that also runs a program that could read it back, such as
102//     `echo git push > x.sh; bash x.sh`: rules 1 to 3, not rule 4.
103// In the second and third, a whole line of a comment (# or //) is skipped, so
104// a script whose comment says "git commit or push" is not refused. The first
105// reads the raw line, comments included.
106
107import type { EngineInterface } from 'claude-code'
108
109type Workbench = EngineInterface['workbench']
110export type ShellParse = Awaited<ReturnType<Workbench['parseShell']>>
111export type Statement = ShellParse['statements'][number]
112export type CallerLane = Awaited<ReturnType<Workbench['callerLane']>>
113
114// Who runs the line. lane is undefined when callerLane could not answer.
115export type Caller = { lane: CallerLane | undefined; isUnattended: boolean }
116
117type Op = 'commit' | 'push' | 'merge'
118
119// The guard's verdict: the refusal the model reads, or undefined to pass.
120export type Refusal = { deny: string }
121
122// What a line runs, as the rules weigh it.
123type Found = {
124  // Commits, pushes and merges the rules 1 to 3 weigh.
125  ops: Set<Op>
126  // A push that forces or deletes.
127  isForce: boolean
128  // An op the ask rules cannot see (rule 4).
129  isHidden: boolean
130}
131
132// The text match of the bash guard before this one (tests/oracle/), kept for
133// the text the statements cannot show. A character that cannot be part of a
134// word, then git and commit or push, or gh and pr merge or the API merge
135// endpoint, within one command.
136const W = '[^A-Za-z0-9_.-]'
137const SEG = '[^;&|\\n]'
138const GIT_OP = `git${W}(?:${SEG}*[^A-Za-z0-9_-])?(commit|push)(?:$|[^A-Za-z0-9_-])`
139const MERGE_OP = `gh${W}(?:${SEG}*[^A-Za-z0-9_-])?(?:pr\\s+merge|api${W}${SEG}*pulls/[^\\s/]+/merge)(?:$|[^A-Za-z0-9_/-])`
140const MENTION = new RegExp(`(?:^|${W})(?:${GIT_OP}|${MERGE_OP})`, 'im')
141const FORCE_TEXT = new RegExp(
142  `(?:^|${W})git${W}(?:${SEG}*[^A-Za-z0-9_-])?push${SEG}*\\s(?:--(?:forc|m|de|pru)|-[A-Za-z0-9]*[fd]|[+:]\\S)`,
143  'im',
144)
145// Any of the three words, for a line the reader could not read whole.
146const LOOSE = /commit|push|merge/i
147
148// Programs that never run code from their own words: what they are handed is
149// data. Text in their words is read only beside a program that could run it.
150const READERS: ReadonlySet<string> = new Set([
151  'echo', 'printf', 'cat', 'head', 'tail', 'wc', 'ls', 'grep', 'egrep', 'fgrep', 'rg', 'jq', 'cd', 'pwd', 'true',
152  'false', 'test', '[', '[[', '((', ':', 'printenv', 'basename', 'dirname', 'realpath', 'readlink', 'stat', 'file',
153  'which', 'type', 'command', 'diff', 'cmp', 'comm', 'cut', 'tr', 'sort', 'uniq', 'column', 'nl', 'tee', 'mkdir',
154  'touch', 'rm', 'rmdir', 'mv', 'cp', 'ln', 'chmod', 'date', 'sleep', 'export', 'local', 'declare', 'readonly',
155  'unset', 'set',
156])
157
158// The wrappers Claude Code strips before it matches an ask rule, so a commit
159// behind one still prompts.
160const ASK_STRIPPED: ReadonlySet<string> = new Set(['timeout', 'time', 'nice', 'nohup', 'stdbuf', 'command', 'builtin', 'noglob'])
161
162// A command name the shell builds at run time: a brace (`{git,}`), a glob, or
163// zsh's `=name` path expansion. Which program runs is not written.
164const BUILT_NAME = /[{}*?[\]]|^=/
165
166// Whole-line comments, dropped before text is matched.
167const withoutComments = (text: string): string =>
168  text
169    .split('\n')
170    .filter(line => !/^\s*(?:#|\/\/)/.test(line))
171    .join('\n')
172
173const opsIn = (text: string): Op[] => {
174  const ops: Op[] = []
175  const match = new RegExp(MENTION.source, 'gim')
176  for (const m of text.matchAll(match)) {
177    // GIT_OP captures commit or push. MERGE_OP captures nothing.
178    ops.push(m[1] === undefined ? 'merge' : (m[1].toLowerCase() as Op))
179  }
180  return ops
181}
182
183// A git option word that sets configuration: `-c key=value` or `-ckey=value`.
184function configOf(args: readonly string[], end: number): { key: string; value: string }[] {
185  const config: { key: string; value: string }[] = []
186  for (let i = 0; i < end; i++) {
187    const arg = args[i] ?? ''
188    const pair = arg === '-c' ? args[++i] : arg.startsWith('-c') ? arg.slice(2) : undefined
189    if (pair === undefined) continue
190    const at = pair.indexOf('=')
191    config.push({ key: (at < 0 ? pair : pair.slice(0, at)).toLowerCase(), value: at < 0 ? '' : pair.slice(at + 1) })
192  }
193  return config
194}
195
196// A word the shell builds when the line runs: the reader's `$_` stands where a
197// substitution stood, and a `$` left in a word is an expansion.
198const isBuilt = (word: string): boolean => word.includes('$')
199
200// git's and gh's global options whose value is the next word. A value built at
201// run time cannot change which subcommand runs; the option word itself can.
202const GLOBAL_VALUES: ReadonlySet<string> = new Set([
203  '-C', '-c', '--git-dir', '--work-tree', '--namespace', '--super-prefix', '--config-env', '--exec-path', '-R', '--repo',
204])
205
206// A global option word that carries its own value, `--git-dir=$D`, `-C$D` or
207// `-R"$R"`: the value decides no subcommand.
208const carriesValue = (arg: string): boolean =>
209  [...GLOBAL_VALUES].some(o => (o.startsWith('--') ? arg.startsWith(`${o}=`) : arg.startsWith(o) && arg.length > o.length))
210
211// Whether a word that decides which subcommand runs is built at run time: an
212// option word before the subcommand (never a value of one, separate, after
213// `=` or attached), or the subcommand itself.
214function decidesBuilt(s: Statement): boolean {
215  const at = s.subcommandAt
216  const end = at < 0 ? s.args.length : Math.min(s.args.length, at + 1)
217  for (let i = 0; i < end; i++) {
218    const arg = s.args[i] ?? ''
219    if (i < at && GLOBAL_VALUES.has(arg)) {
220      i++
221      continue
222    }
223    if (i < at && carriesValue(arg)) continue
224    if (isBuilt(arg)) return true
225  }
226  return false
227}
228
229// gh api's options that take the next word as their value, so the endpoint is
230// the first word that is neither an option nor such a value.
231const GH_API_VALUES: ReadonlySet<string> = new Set([
232  '-f', '--raw-field', '-F', '--field', '-H', '--header', '-X', '--method', '--input', '-q', '--jq', '-t', '--template',
233  '--cache', '-p', '--preview', '--hostname',
234])
235
236// A gh api call's words as gh reads them (pflag, which takes no cut of a long
237// flag): every method value, and the words in endpoint position. A method is
238// `-X PUT`, `-XPUT`, `-X=PUT`, `--method PUT` or `--method=PUT`, and a short
239// cluster that ends in X (`-iX PUT`) takes the next word, as pflag gives it.
240// Every value is kept, since a later one wins and the guard reads them all.
241// gh api takes one endpoint, so a second word in that position is a flag the
242// shell built (`gh api $M repos/$X` with M=-XPUT).
243function apiWordsOf(rest: readonly string[]): { methods: string[]; positionals: string[] } {
244  const methods: string[] = []
245  const positionals: string[] = []
246  for (let i = 0; i < rest.length; i++) {
247    const arg = rest[i] ?? ''
248    if (arg === '--') {
249      positionals.push(...rest.slice(i + 1))
250      break
251    }
252    const cluster = /^-[A-Za-z]*X(.*)$/.exec(arg)
253    if (arg === '--method') methods.push(rest[++i] ?? '')
254    else if (arg.startsWith('--method=')) methods.push(arg.slice('--method='.length))
255    else if (cluster && !arg.startsWith('--')) methods.push(cluster[1] === '' ? (rest[++i] ?? '') : (cluster[1] ?? '').replace(/^=/, ''))
256    else if (GH_API_VALUES.has(arg)) i++
257    else if (!arg.startsWith('-')) positionals.push(arg)
258  }
259  return { methods, positionals }
260}
261
262// Whether a gh api call could merge a pull request through an endpoint built
263// at run time. The REST merge endpoint takes a PUT. gh sends GET, or POST with
264// fields, unless the method says otherwise. So with a PUT, a method built at
265// run time, or a built word where a flag could stand, any built endpoint
266// counts as a merge, since the guard cannot know where it points (fail
267// closed). A GraphQL merge goes over POST and is read from the fields
268// (MERGE_MUTATION), and a literal pulls/<n>/merge endpoint is matched as
269// written, whatever the method.
270function builtMerge(rest: readonly string[]): boolean {
271  const { methods, positionals } = apiWordsOf(rest)
272  if (!positionals.some(isBuilt)) return false
273  const isBuiltFlag = positionals.length > 1
274  return isBuiltFlag || methods.some(m => isBuilt(m) || m.toUpperCase() === 'PUT')
275}
276
277// The GraphQL mutations that merge a pull request, or set it to merge.
278const MERGE_MUTATION = /mergePullRequest|enablePullRequestAutoMerge/i
279
280// Whether a push's own words force or delete.
281const isForceArg = (arg: string): boolean => /^--(?:forc|m|de|pru)/i.test(arg) || /^-[A-Za-z0-9]*[fd]/.test(arg) || /^[+:]\S/.test(arg)
282
283// What one git or gh statement runs, read from its words.
284type Read = { ops: Op[]; isForce: boolean; isHidden: boolean; code: string[]; isUnread: boolean }
285
286function readGit(s: Statement): Read {
287  const read: Read = { ops: [], isForce: false, isHidden: false, code: [], isUnread: false }
288  const args = s.args
289  if (s.name === 'git-commit' || s.name === 'git-push') {
290    read.ops.push(s.name === 'git-commit' ? 'commit' : 'push')
291    read.isForce = s.name === 'git-push' && args.some(isForceArg)
292    return read
293  }
294  const at = s.subcommandAt
295  const config = configOf(args, at < 0 ? args.length : at)
296  read.code.push(...config.map(c => c.value))
297  // A subcommand built at run time, or configuration whose key is, could be
298  // commit or push: read as unread, every op and hidden.
299  if (decidesBuilt(s) || config.some(c => isBuilt(c.key))) read.isUnread = true
300  if (config.some(c => /^remote\..*\.(?:mirror|push)$/.test(c.key))) read.isForce = true
301  if (args.some(arg => arg.startsWith('--config-env'))) read.isUnread = true
302  if (at < 0) {
303    // Words xargs or parallel add could be the subcommand.
304    if (s.wrappers.includes('xargs')) read.isUnread = true
305    return read
306  }
307  let sub = (args[at] ?? '').toLowerCase()
308  let rest = args.slice(at + 1)
309  const alias = config.find(c => c.key === `alias.${sub}`)
310  if (alias !== undefined) {
311    read.isHidden = true
312    if (isBuilt(alias.value)) read.isUnread = true
313    if (alias.value.trimStart().startsWith('!')) {
314      // A shell alias: its text is a script the reader was never handed.
315      read.code.push(alias.value)
316      return read
317    }
318    const words = alias.value.trim().split(/\s+/)
319    sub = (words[0] ?? '').toLowerCase()
320    rest = [...words.slice(1), ...rest]
321  }
322  // git grep -O runs its value as a program on the matched files.
323  if (sub === 'grep' && rest.some(arg => arg.startsWith('-O') || arg.startsWith('--open-files'))) {
324    read.code.push(...rest.map(arg => arg.replace(/^(?:-O|--open-files[a-z-]*=?)/, ' ')))
325  }
326  if (sub === 'commit' || sub === 'push') read.ops.push(sub)
327  if (sub === 'push' && rest.some(isForceArg)) read.isForce = true
328  return read
329}
330
331function readGh(s: Statement): Read {
332  const read: Read = { ops: [], isForce: false, isHidden: false, code: [], isUnread: false }
333  const at = s.subcommandAt
334  if (decidesBuilt(s)) read.isUnread = true
335  if (at < 0) return read
336  const sub = (s.args[at] ?? '').toLowerCase()
337  const rest = s.args.slice(at + 1)
338  // `gh pr $(echo merge)`: the word after pr is the pr subcommand.
339  if (sub === 'pr' && isBuilt(rest[0] ?? '')) read.isUnread = true
340  if (sub === 'pr' && rest.some(arg => arg.toLowerCase() === 'merge')) read.ops.push('merge')
341  const isLiteralMerge = sub === 'api' && rest.some(arg => /pulls\/[^\s/]+\/merge(?![A-Za-z0-9_/-])/i.test(arg))
342  // A GraphQL merge, named in a field or a heredoc fed to --input, over any
343  // endpoint and method. The ask rules match no such line.
344  if (sub === 'api' && [...rest, ...s.heredocs.map(h => h.body)].some(word => MERGE_MUTATION.test(word))) {
345    read.ops.push('merge')
346    read.isHidden = true
347  }
348  if (isLiteralMerge) read.ops.push('merge')
349  // A built endpoint the ask rule's pulls/*/merge pattern cannot see.
350  if (sub === 'api' && !isLiteralMerge && builtMerge(rest)) {
351    read.ops.push('merge')
352    read.isHidden = true
353  }
354  return read
355}
356
357// Whether the ask rules would see this git or gh statement as written: a plain
358// name in command position, with nothing in front of it but a wrapper the
359// harness strips.
360function isPlain(s: Statement): boolean {
361  const named = s.words[s.nameAt] ?? ''
362  const escapedAt = new Set(s.escaped)
363  const lastRead = s.nameAt + 1 + Math.max(s.subcommandAt, 0)
364  for (let i = s.nameAt; i <= lastRead; i++) if (escapedAt.has(i)) return false
365  return s.assignments.length === 0 && s.wrappers.every(w => ASK_STRIPPED.has(w)) && s.source === 'line' && s.depth === 0 && named === s.name
366}
367
368// Everything the guard can say about a line: the ops it runs and how they are
369// written, from the statements and from the text they cannot show.
370export function findOps(parse: ShellParse, line: string): Found {
371  const found: Found = { ops: new Set(), isForce: false, isHidden: false }
372  const add = (ops: Iterable<Op>, isHidden: boolean) => {
373    for (const op of ops) {
374      found.ops.add(op)
375      if (isHidden) found.isHidden = true
376    }
377  }
378  const all: Op[] = ['commit', 'push', 'merge']
379
380  // A line the reader could not read whole may hide any of them.
381  const isUnread = parse.unknowns.length > 0 || parse.statements.some(s => !s.isPlaced)
382  if (isUnread && LOOSE.test(line)) {
383    add(all.filter(op => new RegExp(op, 'i').test(line)), true)
384    if (FORCE_TEXT.test(line.replace(/\\\n/g, '').replace(/["'\\]/g, ''))) found.isForce = true
385  }
386
387  // Text a reader handles, and whether anything on the line could run it.
388  const readerText: string[] = []
389  let hasRunner = false
390  for (const s of parse.statements) {
391    const heredocs = s.heredocs.filter(h => !h.feedsShell).map(h => h.body)
392    const code = (texts: readonly string[]) => {
393      const text = withoutComments(texts.join('\n'))
394      add(opsIn(text), true)
395      if (FORCE_TEXT.test(text)) found.isForce = true
396    }
397    if (s.nameAt < 0) {
398      readerText.push(...s.assignments, ...heredocs)
399      continue
400    }
401    if (s.name === 'git' || s.name === 'git-commit' || s.name === 'git-push' || s.name === 'gh') {
402      const read = s.name === 'gh' ? readGh(s) : readGit(s)
403      add(read.ops, read.isHidden || !isPlain(s))
404      if (read.isForce) found.isForce = true
405      if (read.isUnread) add(all, true)
406      if (read.code.length > 0) {
407        hasRunner = true
408        code(read.code)
409      }
410      readerText.push(...s.assignments, ...heredocs)
411      continue
412    }
413    if (BUILT_NAME.test(s.name) || s.escaped.includes(s.nameAt)) {
414      // The program is not written, so it may be git or gh itself.
415      hasRunner = true
416      code([s.words.slice(s.nameAt).join(' '), ...heredocs])
417      if (LOOSE.test(s.words.join(' '))) add(all.filter(op => new RegExp(op, 'i').test(s.words.join(' '))), true)
418      continue
419    }
420    const isReader = READERS.has(s.name) && !(s.name === 'rg' && s.args.some(arg => arg.startsWith('--pre')))
421    if (isReader) {
422      readerText.push(...s.args, ...s.assignments, ...heredocs)
423    } else {
424      hasRunner = true
425      code([...s.args, ...heredocs])
426      readerText.push(...s.assignments)
427    }
428  }
429  if (hasRunner) {
430    const text = withoutComments(readerText.join('\n'))
431    add(opsIn(text), false)
432    if (FORCE_TEXT.test(text)) found.isForce = true
433  }
434  return found
435}
436
437const READS =
438  'A command that only reads or quotes these words, such as git log --grep, a grep for them, or a heredoc of notes, is not refused here. The code of a program such as node -e or python3 -c is refused when it names them, even in a quote, because the guard cannot tell a quote there from a call. Read or count that text with grep or the Read tool. If any other read was refused, report it as a guard defect, and use the Read tool for the file meanwhile.'
439
440const refusal = (line: string, why: string): Refusal => ({ deny: `🛑 Blocked: ${line}\n\nCommit guard (workbench-dev-team). ${why}` })
441
442export const FORCE_REFUSAL = refusal(
443  'a push that forces or deletes.',
444  'Pushes that force or delete remote refs are refused outright, and no approval changes that. Push without the force or delete, or ask the human to do it.',
445)
446
447export const MERGE_REFUSAL = refusal(
448  'a sub-agent or the pipeline does not merge a pull request.',
449  `Merging is the human's own step, after review. Report that the pull request is ready to merge, and stop. Do not look for another route to a merge. ${READS}`,
450)
451
452export const SUBAGENT_REFUSAL = refusal(
453  'a sub-agent does not commit or push.',
454  `Leave the working tree uncommitted. Report the diff and a proposed commit message to the session that dispatched you. Do not ask to commit. Do not look for another route to a commit or push. ${READS}`,
455)
456
457export const PLAIN_REFUSAL = refusal(
458  'run the commit, push, or merge as a plain line, so you are asked.',
459  'The permission rules that prompt the human match a plain git or gh line, and they miss one behind a wrapper such as env or sudo, a leading NAME=value, bash -c, eval, a substitution, a path, an escape, or an alias, and one inside a program or a line the guard cannot read. Run it as git commit …, git push …, or gh pr merge …, with git -C <dir> for a directory and git -c <key>=<value> for configuration. Drop a variable prefix such as HUSKY=0. If the line only quotes these words in the code of a program such as node -e, the guard cannot tell that from a call either. Read or count that text with grep or the Read tool.',
460)
461
462export const BUILT_REFUSAL = refusal(
463  'a command name the guard cannot read, from a sub-agent or an unattended run.',
464  'The command this line runs is named by a variable or a substitution, or the guard cannot read the whole line: a wrapper option it cannot place, a quote, substitution, heredoc, array or case it cannot close, an escape it cannot decode, a script piped into a shell, or scripts nested past four levels. So the guard cannot tell whether it commits, pushes, or merges. Name the program in plain words, such as python3 script.py rather than "$PY" script.py, and write the line so every quote, substitution, heredoc, array and case closes plainly. Do not look for another spelling.',
465)
466
467export const FAILED_REFUSAL = refusal(
468  'a call the commit guard could not judge.',
469  'The guard failed while reading this command, so it refuses it rather than let a commit, push, or merge through unread. Report this to the human as a guard defect. Do not try another spelling.',
470)
471
472// Whether any statement's command name is unplaced or built at run time, or
473// may be misread, as parseShell reads it. Its `expansion` unknown covers a `$`
474// or a backtick anywhere in a command word, except a plain `"$NAME/…"` or
475// `"${NAME}/…"` in double quotes before a literal path, whose last path part
476// names the program. Since workbench-core 47a5e27 that includes a
477// substitution before a path, quoted or not (`$(…)x/y`, `$((…))x/y`), whose
478// output only the run decides.
479//
480// HIDES_NAME holds every unknown the reader can set, because each can leave a
481// command the shell runs with no name to read:
482//   quote         the reader and the shell can disagree on where a quote
483//                 ends: zsh closes the quote in `echo $[ "1 ]; $x$y` and
484//                 runs `$x$y`, which the reader reads as quoted text
485//   substitution  likewise for a `$(( … ))` or `$( … )` the reader reads as
486//                 unclosed
487//   heredoc       a delimiter the reader does not decode (`<<$'EOF'`), so it
488//                 reads the lines after the shell's terminator as body
489//   escape        a $'…' escape the reader does not decode (a Unicode or
490//                 control escape), which may spell any name, git included
491//   wrapper       a wrapper option the reader cannot place
492//   expansion     a name built from a variable or a substitution
493//   stdin         a script piped or fed into a shell (`curl … | sh`)
494//   depth         scripts nested past four levels, which are not read
495//   compound      an array or a case still open where the script ends, where
496//                 the reader cannot tell where it was meant to end
497// The record is keyed by core's unknown type, so tsc fails when core adds an
498// unknown this list does not hold. tests/shipped-lines.mjs reads the list, so
499// every shipped shell block is held to it too.
500const EVERY_UNKNOWN: Record<ShellParse['unknowns'][number], true> = {
501  quote: true, substitution: true, heredoc: true, escape: true, wrapper: true, expansion: true, stdin: true, depth: true, compound: true,
502}
503export const HIDES_NAME = Object.keys(EVERY_UNKNOWN) as readonly ShellParse['unknowns'][number][]
504const hasBuiltName = (parse: ShellParse): boolean =>
505  parse.unknowns.some(u => HIDES_NAME.includes(u)) || parse.statements.some(s => !s.isPlaced)
506
507// The verdict on one Bash line, in the order of the header's rules.
508export function commitVerdict(parse: ShellParse, line: string, caller: Caller): Refusal | undefined {
509  const found = findOps(parse, line)
510  // A sub-agent (or a lane callerLane cannot give), or a run nobody attends.
511  const isUnwatched = caller.lane === 'sub-agent' || caller.lane === undefined || caller.isUnattended
512  if (found.isForce) return FORCE_REFUSAL
513  if (found.ops.has('merge') && (caller.lane !== 'main' || caller.isUnattended)) return MERGE_REFUSAL
514  const writes = found.ops.has('commit') || found.ops.has('push')
515  if (writes && caller.lane !== 'main' && caller.lane !== 'top-level-agent') return SUBAGENT_REFUSAL
516  if (found.ops.size > 0 && found.isHidden) return PLAIN_REFUSAL
517  if (isUnwatched && hasBuiltName(parse)) return BUILT_REFUSAL
518  return undefined
519}
520
hooks/mods/commit-subject.ts 126 lines
1// The commit subject check, as a pure function over workbench-core's reading of
2// a Bash line. hooks/register.ts runs it after the commit guard passes a line,
3// so it judges the commits the guard lets through: the main loop's and the
4// pipeline's. A sub-agent's commit is already refused.
5//
6// It holds a `git commit -m` subject to the git-commit skill's format:
7//
8//   <type>: <gitmoji> <description>.
9//
10// A type from the skill's references/conventional-commits.md, an optional !
11// for a breaking change, no scope, one space, one gitmoji from
12// references/gitmoji.md, one space, and a description that ends with a period.
13//
14// What it judges: the first line of the first -m (or --message) value of each
15// git commit in the line, read from the statements, never from the raw text. A
16// commit with no -m, or with -F, --file, -C, -c, --reuse-message,
17// --reedit-message, --fixup or --squash, is not judged: its message is not on
18// the line. Neither is a message the line builds at run time (a `$` in it),
19// except one form: `-m "$(cat <<'EOF' … EOF)"`, whose subject is the first line
20// of the heredoc's body.
21
22import type { ShellParse, Statement } from './commit-guard'
23
24export const TYPES: readonly string[] = ['feat', 'fix', 'docs', 'style', 'refactor', 'perf', 'test', 'build', 'ci', 'chore']
25
26// The gitmoji, as references/gitmoji.md lists them.
27// tests/test-commit-subject-lists.sh holds this list to that file, and TYPES to
28// references/conventional-commits.md, both ways. A subject may write each
29// gitmoji with or without its trailing U+FE0F, as keyboards differ.
30export const GITMOJI: readonly string[] = [
31  '🎨', '⚡️', '🔥', '🐛', '🚑️', '✨', '📝', '🚀', '💄', '🎉', '✅', '🔒️', '🔐', '🔖', '🚨', '🚧', '💚', '⬇️', '⬆️', '📌',
32  '👷', '📈', '♻️', '➕', '➖', '🔧', '🔨', '🌐', '✏️', '💩', '⏪️', '🔀', '📦️', '👽️', '🚚', '📄', '💥', '🍱', '♿️', '💡',
33  '🍻', '💬', '🗃️', '🔊', '🔇', '👥', '🚸', '🏗️', '📱', '🤡', '🥚', '🙈', '📸', '⚗️', '🔍️', '🏷️', '🌱', '🚩', '🥅', '💫',
34  '🗑️', '🛂', '🩹', '🧐', '⚰️', '🧪', '👔', '🩺', '🧱', '🧑‍💻', '💸', '🧵', '🦺', '✈️', '🦖',
35]
36
37const bare = (emoji: string): string => emoji.replace(/\uFE0F/g, '')
38const KNOWN: ReadonlySet<string> = new Set(GITMOJI.map(bare))
39
40// Options of git commit whose message is not on the line.
41const NOT_ON_LINE = /^(-F|--file|-C|-c|--reuse-message|--reedit-message|--fixup|--squash)(=|$)/
42// Short options of git commit that take a value, attached or as the next word.
43const SHORT_VALUED = 'mFCct'
44// Short options of git commit whose value is optional and only ever attached
45// (`-unormal`, `-S<keyid>`): the rest of the word is the value, never another
46// option and never the message.
47const SHORT_ATTACHED = 'uS'
48
49type Message = { text: string } | 'not-judged'
50
51// The message of one git commit statement, or not-judged.
52function messageOf(statement: Statement, parse: ShellParse): Message | undefined {
53  const args = statement.args
54  let message: string | undefined
55  for (let i = statement.subcommandAt + 1; i < args.length; i++) {
56    const arg = args[i] ?? ''
57    if (arg === '--') break
58    if (NOT_ON_LINE.test(arg)) return 'not-judged'
59    let value: string | undefined
60    if (arg === '--message' || arg === '--message=') value = args[++i]
61    else if (arg.startsWith('--message=')) value = arg.slice('--message='.length)
62    else if (/^-[^-]/.test(arg)) {
63      for (let c = 1; c < arg.length; c++) {
64        const flag = arg[c] ?? ''
65        if (SHORT_ATTACHED.includes(flag)) break
66        if (!SHORT_VALUED.includes(flag)) continue
67        if (flag === 'F' || flag === 'C' || flag === 'c') return 'not-judged'
68        const attached = arg.slice(c + 1)
69        const taken = attached === '' ? args[++i] : attached
70        if (flag === 'm') value = taken
71        break
72      }
73    }
74    if (value !== undefined && message === undefined) message = value
75  }
76  if (message === undefined) return undefined
77  if (message === '$_') {
78    const bodies = parse.statements.filter(s => s.source === 'substitution')
79    const only = bodies.length === 1 ? bodies[0] : undefined
80    if (only?.name === 'cat' && only.args.length === 0 && only.heredocs.length === 1) return { text: only.heredocs[0]?.body ?? '' }
81    return 'not-judged'
82  }
83  return message.includes('$') ? 'not-judged' : { text: message }
84}
85
86// What is wrong with a subject, or undefined when it keeps the format.
87export function subjectFault(subject: string): string | undefined {
88  const head = /^([A-Za-z]+)(\([^)]*\))?(!)?:/.exec(subject)
89  if (head === null) return 'it does not start with a type and a colon'
90  if (head[2] !== undefined) return `it has a scope, ${head[2]}`
91  if (!TYPES.includes(head[1] ?? '')) return `"${head[1]}" is not a type`
92  const rest = subject.slice(head[0].length)
93  if (!rest.startsWith(' ') || rest.startsWith('  ')) return 'the colon is not followed by exactly one space'
94  const [emoji = '', ...words] = rest.slice(1).split(' ')
95  if (!KNOWN.has(bare(emoji))) return `"${emoji}" is not a gitmoji from the skill's list`
96  const description = words.join(' ')
97  if (!/[^.\s]/.test(description) || description.startsWith(' ')) return 'the gitmoji is not followed by one space and a description'
98  if (!description.endsWith('.')) return 'the description does not end with a period'
99  return undefined
100}
101
102// The refusal for a line whose commit subject breaks the format, or undefined.
103export function subjectVerdict(parse: ShellParse): { deny: string } | undefined {
104  for (const statement of parse.statements) {
105    if (statement.name !== 'git' || statement.subcommandAt < 0 || statement.args[statement.subcommandAt] !== 'commit') continue
106    const message = messageOf(statement, parse)
107    if (message === undefined || message === 'not-judged') continue
108    const subject = message.text.split('\n').find(line => line.trim() !== '')?.replace(/\s+$/, '') ?? ''
109    const fault = subjectFault(subject)
110    if (fault !== undefined) return { deny: subjectDeny(subject, fault) }
111  }
112  return undefined
113}
114
115export function subjectDeny(subject: string, fault: string): string {
116  return [
117    `🛑 Blocked: a commit subject that breaks the git-commit format, because ${fault}.`,
118    '',
119    `Commit guard (workbench-dev-team). The subject was: ${JSON.stringify(subject)}. ` +
120      'The expected shape is "<type>: <gitmoji> <Description>.", for example "feat: ✨ Add email validation endpoint.". ' +
121      `The type is one of ${TYPES.join(', ')}, with ! after it for a breaking change, and never a scope. ` +
122      'The gitmoji comes from the git-commit skill\'s references/gitmoji.md, and the description ends with a period. ' +
123      'Fix the message with the /workbench-dev-team:git-commit skill, and commit again.',
124  ].join('\n')
125}
126
hooks/mods/dispatch.ts 91 lines
1// What a main-session Agent call to a dev-team agent turns into, before it
2// reaches agent.spawn, as pure functions. hooks/register.ts holds the hook.
3//
4//   1. A Watson Index-mode run goes through the dispatcher. Its prompt is the
5//      whole token `Item ID: <n>`, and only bin/dispatch-agent.sh can give the
6//      run its pipeline flag, so the hook runs the installed dispatcher in place
7//      of the Agent call and answers the call with the dispatcher's first line:
8//      dispatched, SKIP, ESCALATE or REPRIEVE. agent.spawn cannot answer a call
9//      with text (its result is a started agent or a refusal), so this lives on
10//      the Agent tool call itself, which a hook may answer with its own record.
11//   2. A dev-team dispatch runs in the background unless the call asked for the
12//      foreground: `run_in_background: false` is the caller's explicit ask, and
13//      anything else is set to true.
14
15import type { ToolResultOf } from 'claude-code'
16
17import { familyOf, itemIdOf } from './spawn'
18
19// Where /workbench-dev-team:setup installs the dispatcher, under $HOME.
20export const DISPATCHER = '.claude-workbench/bin/dispatch-agent.sh'
21
22// The item id a Watson Index-mode dispatch names, or undefined when the call
23// is anything else: another agent, a brief, or a bare id (the dispatch gate
24// refuses a bare id from the main session, so it never reaches the dispatcher).
25export function indexDispatchOf(subagentType: unknown, prompt: unknown): string | undefined {
26  if (typeof subagentType !== 'string' || typeof prompt !== 'string') return undefined
27  if (familyOf(subagentType) !== 'watson') return undefined
28  return itemIdOf(prompt)
29}
30
31// The command the hook runs: bash and the installed script, no shell between.
32export const dispatcherArgv = (home: string, itemId: string): string[] => ['bash', `${home}/${DISPATCHER}`, 'watson', itemId]
33
34const firstLine = (text: string): string | undefined => text.split('\n').find(line => line.trim() !== '')?.trim()
35
36// What the dispatcher's run means for the Agent call: its first line, or a
37// refusal when it exited non-zero or printed nothing. The dispatcher exits 0
38// for every verdict, SKIP and ESCALATE included, so a non-zero exit is a bad
39// argument or a run folder it could not make.
40export function dispatchOutcome(run: { exitCode: number; stdout: string; stderr: string }): { line: string } | { deny: string } {
41  const line = firstLine(run.stdout)
42  if (run.exitCode === 0 && line !== undefined) return { line }
43  const why = firstLine(run.stderr) ?? line ?? `exit ${run.exitCode}, no output`
44  return { deny: dispatchDeny(`the dispatcher refused the run: ${why}`) }
45}
46
47export function dispatchDeny(why: string): string {
48  return [
49    `🛑 Blocked: a Watson Index-mode dispatch, because ${why}.`,
50    '',
51    'Dispatch gate (workbench-dev-team). A Watson `Item ID: <n>` dispatch runs through ~/.claude-workbench/bin/dispatch-agent.sh, which the dev-team mod runs in place of the Agent call. ' +
52      'Nothing was spawned and no board item moved. Relay this to the human, and do not retry through the Agent tool.',
53  ].join('\n')
54}
55
56// What the model reads beside the line: what each first line means.
57export const DISPATCH_CONTEXT =
58  'The dev-team mod ran ~/.claude-workbench/bin/dispatch-agent.sh in place of this Agent call, because only the dispatcher gives a Watson Index-mode run its pipeline flag. ' +
59  'No sub-agent started here, so no completion notification will come. The result is the dispatcher\'s first line. ' +
60  '"dispatched watson pid=... log=..." means the run started: track it from that log, and look there for "Permission denied:" lines, one per refused call. ' +
61  '"REPRIEVE" means a human re-activated an escalated item, and the run started with a raised budget. ' +
62  '"SKIP" (a run on that item is still alive) and "ESCALATE" (the breaker judges the item wedged) mean nothing was spawned: relay the line to the human rather than retrying.'
63
64// The Agent tool's own record for a call the hook answered: completed, with the
65// line as its text. Core checks it against the tool's output schema.
66export function dispatchResult(line: string, prompt: string): ToolResultOf<'Agent'> {
67  return {
68    status: 'completed',
69    agentId: 'dispatch-agent.sh',
70    content: [{ type: 'text', text: line }],
71    totalToolUseCount: 0,
72    totalDurationMs: 0,
73    totalTokens: 0,
74    prompt,
75    usage: {
76      input_tokens: 0,
77      output_tokens: 0,
78      cache_creation_input_tokens: null,
79      cache_read_input_tokens: null,
80      server_tool_use: null,
81      service_tier: null,
82      cache_creation: null,
83    },
84  }
85}
86
87// Whether a dev-team dispatch must be put in the background: every one whose
88// call did not set run_in_background to false.
89export const isForcedBackground = (subagentType: unknown, runInBackground: unknown): boolean =>
90  typeof subagentType === 'string' && familyOf(subagentType) !== undefined && runInBackground !== false && runInBackground !== true
91
hooks/mods/review-guard.ts 891 lines
1// The review guard, as pure functions over workbench-core's reading of a Bash
2// line ($.workbench.parseShell). hooks/register.ts calls them for Bash, Edit,
3// Write and NotebookEdit, resolves the paths they return on the disk, and
4// refuses the call of a Holmes reviewer that would write outside the scratch
5// roots.
6//
7// Holmes reviews code he must not change. In Local mode the tree he reads is the
8// human's live working directory, and the uncommitted change in it is the ONLY
9// copy of the work, so a reviewer that writes to it destroys the thing it was
10// sent to read. Prose in an agent prompt is advisory: on the mode's first real
11// exercise a lens sub-agent ran `chmod` against that directory with the rule in
12// its own prompt. This guard is not advisory.
13//
14// WHO IS HELD. A call from Holmes (`holmes`, or his mode agents `holmes-local`
15// and `holmes-index`) or his helper (`holmes-lens`), with or without the
16// `workbench-dev-team:` prefix, and from any sub-agent one of them spawned, at
17// any depth. A top-level `--agent` run of one of them (the scheduled pipeline)
18// is held on its main loop and in every sub-agent. The main loop of any other
19// session, and every other agent, keeps its tools.
20//
21// THE SCRATCH ROOTS are the places a reviewer may write: workbench-core's
22// $.workbench.scratchRoots (the session scratchpad, ~/Developer/scratchpad and
23// ~/.claude/plans) and $TMPDIR, where a bare `mktemp -d` lands. A write lands
24// in a root only when it is strictly beneath one, resolved both through every
25// symlink and with its final name left unresolved, so no root can be removed
26// and no link can lead out of one. With no root found, every write is refused.
27//
28// WHAT COUNTS AS A WRITE. Reads and test runs pass, and so does a write inside
29// a scratch root. The line is drawn at commands whose purpose is to change a
30// file's content, location, existence, or metadata:
31//   1. git, inverted: GIT_READ_ONLY lists the verbs that only read, and every
32//      other verb is refused. branch, tag, remote, config, stash and worktree
33//      pass in their listing forms only, symbolic-ref with one argument, and
34//      reflog in every form but expire, delete and drop. A git -c,
35//      --config-env or --exec-path=<dir> is refused, since configuration can
36//      name a program (core.pager, diff.external, an alias). The environment is
37//      configuration too (GIT_EXTERNAL_DIFF, GIT_SSH_COMMAND, GIT_CONFIG_COUNT
38//      each name a program), so a reviewer's git is refused when a GIT_*
39//      variable is set in front of it, and when ANY other statement on the
40//      line can change the environment git inherits: an assignment of any
41//      name, the export family, anything that sets a variable by name
42//      (printf -v, read, mapfile, getopts, a for loop, ${x:=y}), set, source,
43//      the dot command, eval, let and arithmetic. The rule is the category,
44//      not a list of GIT_ spellings, because each review round found a new
45//      spelling. It over-refuses a line that sets an unrelated variable before
46//      git, a cost Mike's fail-closed rule accepts for reviewers. A plain
47//      `FOO=1 git status`, a non-GIT_ name on git itself, still runs.
48//      An option of a reading verb that names or runs a program is refused
49//      (PROGRAM_OPTIONS: grep -O, ls-remote --upload-pack, archive --exec and
50//      --remote, --ext-diff, and --textconv, git grep's included). `git
51//      archive` and git's --output write the file they name, judged by path.
52//   2. The file writers in MUTATING, judged by the paths they write:
53//      cp's destination (never its sources), mv's every operand, rsync's
54//      destination, the directory tar and unzip extract into, sort -o, uniq's
55//      second operand, and so on. A writer fed by xargs or parallel is refused,
56//      since its paths arrive at run time.
57//   3. In-place editing: sed -i, perl -i and ruby -i, judged by the files they
58//      edit, so an edit on a scratch copy passes. A sed script with a w or e
59//      command, an awk program that prints to a file, pipes, or calls
60//      system(), and find with -delete, -exec, -ok or -fprint are refused.
61//   4. Formatters asked to write: --write, --fix, --in-place on any other
62//      command (git, the writers, find and an in-place edit meet the rules
63//      above), a formatter's short write flag, and a formatter that writes by
64//      default and was not put in check mode. Refused wherever they point.
65//   5. Redirects bash really performs (>, >>, >|, &>, <>, a >& to a file),
66//      judged by target. A > inside quotes, an awk or sed program, a heredoc
67//      body, arithmetic or [[ ]] is not a redirect, and the reader says so.
68//
69// FAIL CLOSED. A line parseShell cannot read whole (any unknown), a wrapper
70// it cannot place (isPlaced false), a command name built at run time, and a
71// program the guard does not know with a writer's name among its words
72// (`parallel chmod …`, a runner off every list) are refused: a false refusal
73// costs a reviewer one step, while a bypass costs a write to the tree under
74// review (vault: decisions/2026-10-05-holmes-guard-fail-closed-on-unknown-
75// wrappers.md). A cd, pushd, popd or chdir, or a wrapper option that changes
76// directory, makes every later relative path unresolvable.
77//
78// WHAT IT DOES NOT COVER. This guard catches honest mistakes, and is not a
79// security boundary (Mike, 2026-10-07: these gaps are accepted rather than
80// chased). It reads the words of a line, so a change it cannot see in them is
81// out of scope: a value-only change such as PATH=, HOME= or XDG_CONFIG_HOME=
82// before git; git -C into a repository whose own config is hostile; an
83// arithmetic assignment such as $((X=1)); a bash alias defined on the same
84// line; code inside an interpreter (`python -c`, `node -e`); a script file;
85// a package script (`npm run format`); and a zsh builtin outside its tables
86// that sets a variable by name (zformat -f or -a, zregexparse, and module
87// builtins such as sysread, zstat -A and zselect -a). Nor does it catch a
88// program off every list that writes on its own (`gh pr checkout`). The
89// reviewer prompts' no-write rule and the human's review of the diff are the
90// backstop there.
91
92import type { EngineInterface } from 'claude-code'
93
94type Workbench = EngineInterface['workbench']
95export type ShellParse = Awaited<ReturnType<Workbench['parseShell']>>
96export type Statement = ShellParse['statements'][number]
97
98// A refused call: the short action the person reads, and the reason the model
99// reads after it.
100export type Finding = { action: string; reason: string }
101
102// A path a call writes: as written, and absolute when it can be known.
103export type Write = { raw: string; path: string | undefined; writer: string }
104
105export type Verdict = { finding: Finding } | { writes: Write[] }
106
107// The agent types the guard holds. `u` folds `ſ` into `s`, as Python's
108// re.IGNORECASE did for the bash guard.
109const REVIEWER = /(^|[:/])holmes(-lens|-local|-index)?$/iu
110
111export const isReviewerType = (type: string | undefined): boolean => type !== undefined && REVIEWER.test(type.trim())
112
113// git verbs that only read. Every other verb is refused.
114const GIT_READ_ONLY: ReadonlySet<string> = new Set([
115  'annotate', 'blame', 'cat-file', 'check-attr', 'check-ignore', 'count-objects', 'describe', 'diff', 'diff-index',
116  'diff-tree', 'for-each-ref', 'grep', 'log', 'ls-files', 'ls-tree', 'merge-base', 'name-rev', 'rev-list',
117  'rev-parse', 'shortlog', 'show', 'show-ref', 'status', 'var', 'verify-commit', 'verify-tag',
118  'whatchanged', 'ls-remote', 'archive',
119])
120// Read-only only in some forms, judged in gitReads: symbolic-ref with one
121// argument, and reflog other than expire, delete and drop.
122
123const BRANCH_READ_FLAGS = new Set([
124  '--show-current', '-a', '--all', '-r', '--remotes', '-v', '-vv', '--verbose', '-l', '--list', '--no-color',
125  '--contains', '--no-contains', '--merged', '--no-merged', '--points-at',
126])
127const TAG_READ_FLAGS = new Set(['-l', '--list', '-n', '--no-color', '--contains', '--no-contains', '--merged', '--no-merged', '--points-at'])
128const LIST_VALUE_PREFIXES = ['--format=', '--sort=', '--contains=', '--no-contains=', '--merged=', '--no-merged=', '--points-at=']
129const LIST_VALUE_FLAGS = new Set(LIST_VALUE_PREFIXES.map(prefix => prefix.slice(0, -1)))
130const CONFIG_READ_FLAGS = new Set(['--get', '--get-all', '--get-regexp', '--get-urlmatch', '--list', '-l', '--get-color', '--get-colorbool'])
131const CONFIG_WRITE_FLAGS = new Set(['--add', '--unset', '--unset-all', '--replace-all', '--rename-section', '--remove-section', '--edit', '-e'])
132
133// Commands whose purpose is to change a file's content, location, existence,
134// or metadata, each judged by the paths writeTargets says it writes.
135const MUTATING: ReadonlySet<string> = new Set([
136  'chmod', 'chown', 'chgrp', 'rm', 'rmdir', 'unlink', 'mv', 'shred', 'truncate', 'touch', 'tee', 'cp', 'ln',
137  'install', 'dd', 'patch', 'mkdir', 'rsync', 'tar', 'unzip', 'sort', 'uniq',
138])
139
140// Options that take the next word as their value, on both GNU and BSD.
141const VALUE_OPTIONS: Readonly<Record<string, ReadonlySet<string>>> = {
142  install: new Set(['-m', '-o', '-g']),
143  truncate: new Set(['-s', '-r']),
144  touch: new Set(['-r', '-t', '-d']),
145  mkdir: new Set(['-m']),
146  rsync: new Set(['-e', '-f', '-T']),
147  uniq: new Set(['-f', '-s', '-w', '--skip-fields', '--skip-chars', '--check-chars']),
148  chmod: new Set(['--reference']),
149}
150
151const UNZIP_READ_FLAGS = new Set('lptvcZh')
152
153// The three in-place editors: the letters that mean in place, and the letters
154// that take the rest of their word (or the next word) as a value.
155const IN_PLACE_EDITORS: Readonly<Record<string, { inPlace: string; value: string }>> = {
156  sed: { inPlace: 'iI', value: 'ef' },
157  perl: { inPlace: 'i', value: 'eEFImMx' },
158  ruby: { inPlace: 'i', value: 'eCEFIrx' },
159}
160// The value letters whose value is the program, so no file operand is a script.
161const CODE_LETTERS: Readonly<Record<string, string>> = { sed: 'ef', perl: 'eE', ruby: 'e' }
162
163const IN_PLACE_FLAGS = new Set(['--write', '--fix', '--in-place'])
164
165const FORMATTER_WRITE_FLAGS: Readonly<Record<string, ReadonlySet<string>>> = {
166  gofmt: new Set(['-w']), goimports: new Set(['-w']), gofumpt: new Set(['-w']), shfmt: new Set(['-w']),
167  prettier: new Set(['-w']), 'clang-format': new Set(['-i']), autopep8: new Set(['-i']), yapf: new Set(['-i']),
168  'swift-format': new Set(['-i']),
169  rubocop: new Set(['-a', '-A', '-x', '--autocorrect', '--autocorrect-all', '--auto-correct', '--auto-correct-all', '--safe-auto-correct', '--fix-layout']),
170}
171const FORMATTER_CLUSTER_FLAGS: Readonly<Record<string, { write: string; value: string }>> = {
172  yapf: { write: 'i', value: 'le' },
173  autopep8: { write: 'i', value: 'jp' },
174  prettier: { write: 'w', value: '' },
175  rubocop: { write: 'aAx', value: 'corfCs' },
176}
177const STDIN_NAME_OPTIONS = new Set(['--stdin-filename', '--filename'])
178const FORMATTERS_WRITING_BY_DEFAULT: Readonly<Record<string, ReadonlySet<string>>> = {
179  black: new Set(['--check', '--diff', '-c', '--code']),
180  rustfmt: new Set(['--check', '--emit=stdout', '--print-config']),
181  'cargo fmt': new Set(['--check']),
182  'go fmt': new Set(['-n']),
183  'ruff format': new Set(['--check', '--diff']),
184  isort: new Set(['--check-only', '--check', '-c', '--diff', '--show-config', '--show-files', '--stdout', '-d']),
185  'terraform fmt': new Set(['-check', '-write=false']),
186  'mix format': new Set(['--check-formatted', '--dry-run']),
187  'dotnet format': new Set(['--verify-no-changes']),
188  'deno fmt': new Set(['--check']),
189  'zig fmt': new Set(['--check']),
190  stylua: new Set(['--check']),
191  pint: new Set(['--test']),
192  'php-cs-fixer fix': new Set(['--dry-run']),
193}
194const RUSTFMT_CONFIG_FILE_KINDS = new Set(['default', 'minimal'])
195const OFF_VALUES = new Set(['false', 'f', '0', 'no', 'off'])
196const FORMATTER_INFO_FLAGS = new Set(['-h', '--help', '-V', '--version'])
197
198// Wrappers and runners the shell reader does not place, stepped through here
199// so the program behind them is judged. Each runner lists the subcommands that
200// run another program. The value options take the next word.
201const PASS_THROUGH: ReadonlySet<string> = new Set(['npx', 'bunx', 'pnpx', 'noglob'])
202const RUNNERS: Readonly<Record<string, ReadonlySet<string>>> = {
203  bundle: new Set(['exec']), uv: new Set(['run']), poetry: new Set(['run']), pipx: new Set(['run']),
204  pnpm: new Set(['exec', 'dlx']), npm: new Set(['exec', 'x']), yarn: new Set(['exec', 'dlx', 'run']), composer: new Set(['exec']),
205}
206const YARN_OWN_WRITER_NAMES = new Set(['install', 'unlink', 'patch'])
207const RUNNER_VALUE_OPTIONS: Readonly<Record<string, readonly string[]>> = {
208  npx: ['-p', '--package'],
209  uv: ['-w', '--with', '--with-editable', '--with-requirements', '-p', '--python', '--package', '--extra', '--group', '--env-file', '--index', '-i', '--index-url', '--directory', '--project'],
210  poetry: ['-C', '--directory', '-P', '--project'],
211  pipx: ['--spec', '--python', '--pip-args', '--index-url'],
212  pnpm: ['-C', '--dir', '-F', '--filter', '--package'],
213  npm: ['-p', '--package', '--prefix', '-w', '--workspace'],
214  yarn: ['--cwd', '-p', '--package'],
215  composer: ['-d', '--working-dir'],
216}
217const RUNNER_CHDIR_OPTIONS: Readonly<Record<string, readonly string[]>> = {
218  uv: ['--directory'], poetry: ['-C', '--directory'], pnpm: ['-C', '--dir'], npm: ['-w', '--workspace', '--prefix'],
219  yarn: ['--cwd'], composer: ['-d', '--working-dir'],
220}
221// A runner option whose value is a command line, which the guard cannot read.
222const RUNNER_SCRIPT_OPTIONS: Readonly<Record<string, readonly string[]>> = { npm: ['-c', '--call'], npx: ['-c', '--call'] }
223
224// Builtins that change the directory later statements run in.
225const DIRECTORY_CHANGES = new Set(['cd', 'pushd', 'popd', 'chdir'])
226
227// Programs that never run a command named among their words, so a writer's
228// name there is data: `grep -rn chmod .`, `man rm`, `command -v tee`.
229const KNOWN: ReadonlySet<string> = new Set([
230  'echo', 'printf', 'cat', 'head', 'tail', 'wc', 'ls', 'grep', 'egrep', 'fgrep', 'rg', 'ag', 'ack', 'jq', 'yq', 'man',
231  'which', 'type', 'whereis', 'whatis', 'apropos', 'command', 'file', 'stat', 'diff', 'cmp', 'comm', 'cut', 'tr',
232  'column', 'nl', 'less', 'more', 'test', '[', '[[', '((', 'printenv', 'basename', 'dirname', 'realpath', 'readlink',
233  'true', 'false', ':', 'cd', 'pushd', 'popd', 'pwd', 'git', 'gh', 'date', 'sleep', 'help', 'hash', 'alias',
234  'export', 'local', 'declare', 'readonly', 'unset', 'set', 'trap', 'tldr', 'mktemp',
235  // A case pattern and a loop's word list are words, never commands.
236  'case', 'for', 'select',
237])
238
239// Names that run or write when they stand among another program's words.
240const WRITER_NAMES: ReadonlySet<string> = new Set([
241  ...MUTATING, ...Object.keys(IN_PLACE_EDITORS), ...Object.keys(FORMATTER_WRITE_FLAGS), 'black', 'rustfmt', 'cargo',
242  'go', 'ruff', 'isort', 'terraform', 'mix', 'dotnet', 'deno', 'zig', 'stylua', 'pint', 'php-cs-fixer', 'git', 'find',
243  'awk', 'gawk', 'mawk', 'nawk', 'bash', 'sh', 'zsh', 'dash', 'ksh', 'fish', 'csh', 'tcsh', 'eval', 'exec', 'source',
244  'env', 'sudo', 'doas', 'nohup', 'nice', 'ionice', 'stdbuf', 'timeout', 'gtimeout', 'xargs', 'parallel',
245  'caffeinate', 'chronic', 'unbuffer', 'flock', 'setsid', 'watch', 'time', 'builtin', ...PASS_THROUGH,
246  ...Object.keys(RUNNERS),
247])
248
249const AWK = new Set(['awk', 'gawk', 'mawk', 'nawk'])
250
251// A command name the shell builds at run time: a brace, a glob, or zsh's
252// `=name` path expansion.
253const BUILT_NAME = /[{}*?[\]]|^=/
254
255// A GIT_* assignment in front of git.
256const isGitEnv = (word: string): boolean => /^GIT_[A-Za-z0-9_]*\+?=/.test(word)
257
258// Commands that can change the environment later commands on the line
259// inherit: they set, export, or import variables by name, or run text that
260// could. printf counts only with -v.
261const ENV_CHANGERS: ReadonlySet<string> = new Set([
262  'export', 'declare', 'typeset', 'local', 'readonly', 'read', 'mapfile', 'readarray', 'getopts', 'for', 'select',
263  'set', 'source', '.', 'eval', 'let', '((', 'unset',
264  // zsh, Mike's shell: its options (allexport among them), its declarations,
265  // and the builtins that set a variable by name.
266  'setopt', 'unsetopt', 'emulate', 'integer', 'float', 'vared', 'zparseopts', 'getln',
267])
268
269// Builtins that set a variable by name only with one of these options:
270// printf -v and print -v, zsh's strftime -s, and zstyle's lookup forms.
271const SETS_BY_OPTION: Readonly<Record<string, RegExp>> = {
272  printf: /^-[A-Za-z]*v/,
273  print: /^-[A-Za-z]*v/,
274  strftime: /^-[A-Za-z]*s/,
275  zstyle: /^-/,
276}
277
278// A parameter expansion that assigns: ${x:=y} or ${x=y}.
279const ASSIGNING_EXPANSION = /\$\{[^}]*?:?=/
280
281// Whether a statement can change the environment of the others on its line.
282const changesEnv = (st: Statement): boolean =>
283  st.assignments.length > 0 ||
284  (ENV_CHANGERS.has(st.name) && st.nameAt >= 0) ||
285  (SETS_BY_OPTION[st.name] !== undefined && st.args.some(a => SETS_BY_OPTION[st.name]?.test(a) === true)) ||
286  st.words.some(w => ASSIGNING_EXPANSION.test(w))
287
288// A path holding any of these is shell the guard does not expand.
289const UNRESOLVABLE = /[$`*?[\]{}()'"\\<>\uE000]/
290
291// Redirect targets that write no file.
292const DEVICE = /^\/dev\/(?:null|stdout|stderr|tty|fd\/[0-9]+)$/
293
294const basename = (word: string): string => (word.split('/').pop() ?? word).toLowerCase()
295
296// The absolute path a command names, or undefined when it cannot be known.
297export function resolvePath(raw: string, cwd: string | undefined, home: string | undefined): string | undefined {
298  let target = raw
299  if (target === '~' || target.startsWith('~/')) {
300    if (!home) return undefined
301    target = home + target.slice(1)
302  }
303  if (target === '' || target.startsWith('~') || UNRESOLVABLE.test(target)) return undefined
304  if (!target.startsWith('/')) {
305    if (!cwd) return undefined
306    target = `${cwd.replace(/\/$/, '')}/${target}`
307  }
308  return target
309}
310
311// Options of a reading verb that run a program, name one, or run one the
312// configuration names, by verb: each refused as `git -c` is. A long option is
313// matched by any cut git accepts down to the length given, so `--open-files=`
314// is `--open-files-in-pager`. git log's `-O<file>` is an order file and reads.
315const PROGRAM_OPTIONS: Readonly<Record<string, readonly [string, number][]>> = {
316  // grep alone runs textconv only when asked.
317  grep: [['-O', 2], ['--open-files-in-pager', 6], ['--ext-grep', 7], ['--textconv', 7]],
318  'ls-remote': [['--upload-pack', 4], ['--exec', 5]],
319  archive: [['--exec', 5], ['--remote', 8]],
320  diff: [['--ext-diff', 6], ['--textconv', 7]],
321  log: [['--ext-diff', 6], ['--textconv', 7]],
322  show: [['--ext-diff', 6], ['--textconv', 7]],
323  whatchanged: [['--ext-diff', 6], ['--textconv', 7]],
324  'diff-tree': [['--ext-diff', 6], ['--textconv', 7]],
325  'diff-index': [['--ext-diff', 6], ['--textconv', 7]],
326  blame: [['--textconv', 7]],
327  annotate: [['--textconv', 7]],
328  'cat-file': [['--textconv', 7], ['--filters', 5]],
329}
330
331// The option of `rest` that runs a program, or undefined. `--` ends options.
332function programOption(verb: string, rest: readonly string[]): string | undefined {
333  for (const arg of rest) {
334    if (arg === '--') return undefined
335    const key = arg.split('=')[0] ?? arg
336    for (const [option, shortest] of PROGRAM_OPTIONS[verb] ?? []) {
337      // A short option counts anywhere in a single-dash cluster (`-nOcat`).
338      const isShort = !option.startsWith('--') && arg.startsWith('-') && !arg.startsWith('--') && arg.includes(option.slice(1))
339      if (option.startsWith('--') ? key.length >= shortest && option.startsWith(key) : isShort) return option
340    }
341  }
342  return undefined
343}
344
345// git's own verdict on `git <verb> <rest>`: true when it only reads.
346function gitReads(verb: string, rest: readonly string[]): boolean {
347  if (programOption(verb, rest) !== undefined) return false
348  if (GIT_READ_ONLY.has(verb)) return true
349  const flags = rest.filter(a => a.startsWith('-'))
350  let positionals = rest.filter(a => !a.startsWith('-'))
351  // symbolic-ref reads a ref with one argument, and sets or deletes it with
352  // two or with -d.
353  if (verb === 'symbolic-ref') {
354    return positionals.length === 1 && flags.every(f => ['-q', '--quiet', '--short', '--recurse', '--no-recurse'].includes(f))
355  }
356  // reflog shows by default; expire, delete and drop destroy the recovery trail.
357  if (verb === 'reflog') return !['expire', 'delete', 'drop'].includes(positionals[0] ?? '')
358  const isListing = flags.includes('-l') || flags.includes('--list')
359  if (verb === 'branch' || verb === 'tag') {
360    positionals = rest.filter((a, i) => !a.startsWith('-') && !(i > 0 && LIST_VALUE_FLAGS.has(rest[i - 1] ?? '')))
361    const allowed = verb === 'branch' ? BRANCH_READ_FLAGS : TAG_READ_FLAGS
362    const flagsOk = flags.every(f => allowed.has(f) || LIST_VALUE_PREFIXES.some(p => f.startsWith(p)) || (verb === 'tag' && /^-n\d+$/.test(f)))
363    return flagsOk && (positionals.length === 0 || isListing)
364  }
365  const first = positionals[0]
366  if (verb === 'remote') {
367    return flags.every(f => ['-v', '--verbose', '--all', '--push', '-n'].includes(f)) && (first === undefined || first === 'get-url' || first === 'show')
368  }
369  if (verb === 'config') {
370    if (flags.some(f => CONFIG_WRITE_FLAGS.has(f.split('=')[0] ?? f))) return false
371    return flags.some(f => CONFIG_READ_FLAGS.has(f)) || first === 'get' || first === 'list'
372  }
373  if (verb === 'stash') return first === 'list' || first === 'show'
374  if (verb === 'worktree') return first === 'list'
375  return false
376}
377
378// The files git writes beside its output: `--output <f>`, `-o <f>` for archive.
379function gitOutputs(verb: string, rest: readonly string[]): string[] {
380  const out: string[] = []
381  rest.forEach((arg, i) => {
382    if (arg.startsWith('--output=')) out.push(arg.slice('--output='.length))
383    else if (arg === '--output' || (verb === 'archive' && arg === '-o')) out.push(rest[i + 1] ?? '')
384    else if (verb === 'archive' && /^-o./.test(arg)) out.push(arg.slice(2))
385  })
386  return out
387}
388
389// The operands of a writer: the words that are not options or option values.
390function operandsOf(name: string, args: readonly string[]): string[] {
391  const operands: string[] = []
392  let isDone = false
393  for (let i = 0; i < args.length; i++) {
394    const arg = args[i] ?? ''
395    if (!isDone && arg === '--') isDone = true
396    else if (isDone || !arg.startsWith('-') || arg === '-') operands.push(arg)
397    else if (VALUE_OPTIONS[name]?.has(arg)) i++
398  }
399  return operands
400}
401
402function tarTargets(args: readonly string[], cwd: string): string[] {
403  const longModes: Record<string, string> = { '--extract': 'x', '--get': 'x', '--create': 'c', '--append': 'r', '--update': 'u', '--catenate': 'A', '--concatenate': 'A', '--delete': 'r' }
404  const modes = new Set<string>()
405  const archives: string[] = []
406  const dirs: string[] = []
407  let toStdout = false
408  for (let i = 0; i < args.length; i++) {
409    const arg = args[i] ?? ''
410    if (arg.startsWith('--')) {
411      const at = arg.indexOf('=')
412      const key = at < 0 ? arg : arg.slice(0, at)
413      for (const m of longModes[key] ?? '') modes.add(m)
414      toStdout ||= key === '--to-stdout'
415      if (key === '--file' || key === '--directory') {
416        const value = at < 0 ? (args[++i] ?? '') : arg.slice(at + 1)
417        ;(key === '--file' ? archives : dirs).push(value)
418      }
419    } else if ((arg.startsWith('-') && arg !== '-') || i === 0) {
420      const letters = arg.replace(/^-+/, '')
421      for (let j = 0; j < letters.length; j++) {
422        const char = letters[j] ?? ''
423        if (char === 'f' || char === 'C') {
424          let value = arg.startsWith('-') ? letters.slice(j + 1) : ''
425          if (!value) value = args[++i] ?? ''
426          ;(char === 'f' ? archives : dirs).push(value)
427          if (arg.startsWith('-')) break
428        } else {
429          if ('xcruA'.includes(char)) modes.add(char)
430          toStdout ||= char === 'O'
431        }
432      }
433    }
434  }
435  const targets: string[] = []
436  if ([...modes].some(m => 'cruA'.includes(m))) targets.push(...archives.filter(a => a !== '-'))
437  if (modes.has('x') && !toStdout) targets.push(...(dirs.length > 0 ? dirs : [cwd]))
438  return targets
439}
440
441function unzipTargets(args: readonly string[], cwd: string): string[] {
442  let letters = ''
443  let dest: string | undefined
444  const operands: string[] = []
445  for (let i = 0; i < args.length; i++) {
446    const arg = args[i] ?? ''
447    if (arg.startsWith('-') && !arg.startsWith('--') && arg.length > 1) {
448      for (let j = 1; j < arg.length; j++) {
449        if (arg[j] === 'd') {
450          dest = arg.slice(j + 1) || (args[++i] ?? '')
451          break
452        }
453        letters += arg[j]
454      }
455    } else if (!arg.startsWith('-')) operands.push(arg)
456  }
457  if (operands.length === 0 || [...letters].some(l => UNZIP_READ_FLAGS.has(l))) return []
458  return [dest ?? cwd]
459}
460
461function sortTargets(args: readonly string[]): string[] {
462  const targets: string[] = []
463  args.forEach((arg, i) => {
464    const next = args[i + 1] ?? ''
465    if (arg.startsWith('--output')) targets.push(arg.includes('=') ? arg.slice(arg.indexOf('=') + 1) : next)
466    else if (/^-[A-Za-z]*o$/.test(arg)) targets.push(next)
467    else if (/^-[A-Za-z]*o.+/.test(arg) && !arg.startsWith('--')) targets.push(arg.slice(arg.indexOf('o') + 1))
468  })
469  return targets
470}
471
472// The paths a MUTATING command writes, as written. An extra path can only
473// refuse more. `cwd` is '' when it is unknown, which no path resolves under.
474function writeTargets(name: string, args: readonly string[], cwd: string): string[] {
475  const operands = operandsOf(name, args)
476  if (name === 'dd') return args.filter(a => a.startsWith('of=')).map(a => a.slice(3))
477  if (name === 'cp' || name === 'ln' || name === 'install') {
478    for (let i = 0; i < args.length; i++) {
479      const arg = args[i] ?? ''
480      if (arg.startsWith('--target-directory=')) return [arg.slice(arg.indexOf('=') + 1)]
481      if (arg === '--target-directory' || /^-[A-Za-z]*t$/.test(arg)) return [args[i + 1] ?? '']
482      if (/^-[A-Za-z]*t.+/.test(arg) && !arg.startsWith('--')) return [arg.slice(arg.indexOf('t') + 1)]
483    }
484    if (name === 'install' && args.includes('-d')) return operands
485    if (name === 'ln' && operands.length === 1) return [cwd ? `${cwd}/${basename(operands[0] ?? '')}` : '']
486    // The destination is the last operand; a source is only read.
487    return operands.slice(-1)
488  }
489  if (name === 'chmod' || name === 'chown' || name === 'chgrp') {
490    return args.some(a => a.startsWith('--reference')) ? operands : operands.slice(1)
491  }
492  if (name === 'rsync') {
493    if (args.includes('--remove-source-files')) return operands
494    return operands.length > 1 ? operands.slice(-1) : []
495  }
496  if (name === 'tar') return tarTargets(args, cwd)
497  if (name === 'unzip') return unzipTargets(args, cwd)
498  if (name === 'sort') return sortTargets(args)
499  if (name === 'uniq') return operands.slice(1)
500  if (name === 'patch') {
501    let directory = cwd
502    args.forEach((arg, i) => {
503      if (arg === '-d' || arg === '--directory') directory = args[i + 1] ?? ''
504      else if (arg.startsWith('--directory=')) directory = arg.slice('--directory='.length)
505      else if (arg.startsWith('-d') && arg.length > 2) directory = arg.slice(2)
506    })
507    return [...operands, directory]
508  }
509  return operands
510}
511
512// The files an in-place sed, perl or ruby edits, or undefined when it edits
513// none in place.
514function inPlaceTargets(name: string, args: readonly string[]): string[] | undefined {
515  const editor = IN_PLACE_EDITORS[name]
516  if (editor === undefined) return undefined
517  let isInPlace = false
518  let hasCode = false
519  const operands: string[] = []
520  for (let i = 0; i < args.length; i++) {
521    const arg = args[i] ?? ''
522    if (arg === '--') {
523      operands.push(...args.slice(i + 1))
524      break
525    }
526    if (arg.startsWith('--')) {
527      if (arg.startsWith('--in-place')) isInPlace = true
528      if (name === 'sed' && (arg === '--expression' || arg === '--file')) {
529        hasCode = true
530        i++
531      } else if (name === 'sed' && (arg.startsWith('--expression=') || arg.startsWith('--file='))) hasCode = true
532      continue
533    }
534    if (!arg.startsWith('-') || arg === '-') {
535      operands.push(arg)
536      continue
537    }
538    for (let j = 1; j < arg.length; j++) {
539      const char = arg[j] ?? ''
540      if (editor.inPlace.includes(char)) {
541        isInPlace = true
542        // BSD sed takes the backup suffix as the next word: `sed -i '' …`.
543        if (name === 'sed' && j === arg.length - 1 && (args[i + 1] === '' || /^\.[A-Za-z0-9_~-]*$/.test(args[i + 1] ?? ''))) i++
544        // GNU's attached suffix (`-i.bak`) is the rest of the word.
545        if (name === 'sed' || name === 'perl' || name === 'ruby') {
546          if (j < arg.length - 1 && !/^[A-Za-z]/.test(arg.slice(j + 1))) break
547        }
548        continue
549      }
550      if (editor.value.includes(char)) {
551        if (CODE_LETTERS[name]?.includes(char)) hasCode = true
552        if (j === arg.length - 1) i++
553        break
554      }
555    }
556  }
557  if (!isInPlace) return undefined
558  // With no -e or -f, the first operand is the program, not a file.
559  return hasCode ? operands : operands.slice(1)
560}
561
562// A sed program that writes a file (w, W, s///w) or runs a command (e).
563const SED_WRITES = /(?:^|[;\n{}\d$/])\s*[wW]\s*\S|(?:^|[;\n{}])\s*e(?:\s|$|;)|\/[gpiImM0-9]*e[gpiImM0-9]*\s*(?:$|[;}\n])/
564// An awk program that prints to a file or a pipe, reads a command, or runs one.
565const AWK_WRITES = /\bsystem\s*\(|\|\s*(?:&\s*)?getline|\bprintf?\b[^;}\n]*(?:>|\|)|"\s*\|\s*getline/
566
567function sedPrograms(args: readonly string[]): string[] {
568  const programs: string[] = []
569  let hasFlag = false
570  for (let i = 0; i < args.length; i++) {
571    const arg = args[i] ?? ''
572    if (arg === '-e' || arg === '--expression') {
573      hasFlag = true
574      programs.push(args[++i] ?? '')
575    } else if (arg.startsWith('--expression=')) {
576      hasFlag = true
577      programs.push(arg.slice('--expression='.length))
578    } else if (/^-[A-Za-z]*e./.test(arg) && !arg.startsWith('--')) {
579      hasFlag = true
580      programs.push(arg.slice(arg.indexOf('e') + 1))
581    } else if (arg === '-f' || arg === '--file' || arg.startsWith('--file=')) {
582      hasFlag = true
583    }
584  }
585  if (!hasFlag) {
586    const first = args.find(a => !a.startsWith('-'))
587    if (first !== undefined) programs.push(first)
588  }
589  return programs
590}
591
592function formatterWrites(name: string, args: readonly string[]): string | undefined {
593  if (/^python[0-9.]*$/.test(name) && args[0] === '-m' && args.length > 1) {
594    name = args[1] ?? ''
595    args = args.slice(2)
596  }
597  for (const arg of args) {
598    let key = arg.split('=')[0] ?? arg
599    if (key.length === 3 && key.startsWith('--')) key = key.slice(1)
600    if (FORMATTER_WRITE_FLAGS[name]?.has(key)) return name
601  }
602  const cluster = FORMATTER_CLUSTER_FLAGS[name]
603  if (cluster !== undefined) {
604    for (const arg of args) {
605      if (arg.startsWith('--') || !arg.startsWith('-')) continue
606      for (const char of arg.slice(1)) {
607        if (cluster.write.includes(char)) return name
608        if (cluster.value.includes(char)) break
609      }
610    }
611  }
612  const sub = args.find(a => !a.startsWith('-') && !a.startsWith('+')) ?? ''
613  const key = name in FORMATTERS_WRITING_BY_DEFAULT ? name : `${name} ${sub}`
614  const checks = FORMATTERS_WRITING_BY_DEFAULT[key]
615  if (checks === undefined) return undefined
616  const words = new Set<string>(args)
617  args.forEach((a, i) => {
618    if (i + 1 < args.length) words.add(`${a}=${args[i + 1]}`)
619    const at = a.indexOf('=')
620    if (at >= 0 && !OFF_VALUES.has(a.slice(at + 1).toLowerCase())) words.add(a.slice(0, at))
621  })
622  if (key === 'rustfmt' && words.has('--print-config')) {
623    const operands = args.filter(a => !a.startsWith('-'))
624    const joined = args.filter(a => a.startsWith('--print-config=')).map(a => a.slice('--print-config='.length))
625    const kind = joined[0] ?? operands.shift() ?? ''
626    return operands.length > 0 && RUSTFMT_CONFIG_FILE_KINDS.has(kind) ? key : undefined
627  }
628  if ([...words].some(w => checks.has(w) || FORMATTER_INFO_FLAGS.has(w))) return undefined
629  const operands = args.filter((a, i) => !a.startsWith('-') && !a.startsWith('+') && (i === 0 || !STDIN_NAME_OPTIONS.has(args[i - 1] ?? '')))
630  if (key.includes(' ')) operands.splice(operands.indexOf(sub), 1)
631  if (operands.length === 0 && (args.includes('-') || key === 'rustfmt')) return undefined
632  return key
633}
634
635// The program behind the runners and wrappers the shell reader does not place:
636// its name, its words, whether it changes directory, and a finding when its
637// command cannot be read at all.
638function stepThrough(name: string, args: readonly string[]): { name: string; args: readonly string[]; chdir: boolean; finding?: Finding } {
639  let chdir = false
640  for (let guard = 0; guard < 16; guard++) {
641    const runner = RUNNERS[name]
642    if (!PASS_THROUGH.has(name) && runner === undefined) break
643    const values = RUNNER_VALUE_OPTIONS[name] ?? []
644    const matches = (options: readonly string[], key: string) => options.some(o => key === o || (o.startsWith('--') && key.length > 3 && o.startsWith(key)))
645    // Steps over the options from `i`, before and after a run subcommand.
646    let i = 0
647    let script: string | undefined
648    const skipOptions = () => {
649      while (i < args.length && (args[i] ?? '').startsWith('-')) {
650        const arg = args[i] ?? ''
651        const key = arg.split('=')[0] ?? arg
652        if (matches(RUNNER_SCRIPT_OPTIONS[name] ?? [], key)) script = key
653        if (matches(RUNNER_CHDIR_OPTIONS[name] ?? [], key)) chdir = true
654        i += matches(values, key) && !arg.includes('=') ? 2 : 1
655      }
656    }
657    skipOptions()
658    if (runner !== undefined) {
659      const sub = args[i]
660      if (sub !== undefined && runner.has(sub)) {
661        i++
662        skipOptions()
663      } else if (sub === undefined || !(name === 'yarn' && !YARN_OWN_WRITER_NAMES.has(sub))) {
664        return { name, args, chdir }
665      }
666    }
667    if (script !== undefined) {
668      return { name, args, chdir, finding: { action: `\`${name} ${script}\``, reason: `\`${name} ${script}\` hands over a command line the guard cannot read.` } }
669    }
670    if (i >= args.length) return { name, args, chdir }
671    name = basename(args[i] ?? '')
672    args = args.slice(i + 1)
673  }
674  return { name, args, chdir }
675}
676
677const unreadable = (what: string): Finding => ({
678  action: 'a command the guard cannot read',
679  reason: `${what}, so the guard cannot tell whether it writes. Run each command plainly, with literal paths.`,
680})
681
682// Options of env and sudo that run the command in another directory.
683const CHDIR_WORD = /^(?:-[A-Za-z]*[CD]|--ch)/
684
685export type Context = {
686  // The directory the line starts in; undefined when it is not known.
687  cwd: string | undefined
688  home: string | undefined
689  // Prose in a fenced block, read one line at a time by a lint: what the
690  // reader could not read is passed over, a word that is not a command name
691  // (an unknown program, a name built at run time) is not judged, and neither
692  // are redirects. The rules a known command meets still apply.
693  isDocs?: boolean
694}
695
696// The verdict on one Bash line: a finding, or the paths it writes.
697export function reviewBash(parse: ShellParse, context: Context): Verdict {
698  const isDocs = context.isDocs === true
699  if (!isDocs && parse.unknowns.length > 0) {
700    return { finding: unreadable(`The line holds shell the reader cannot read (${parse.unknowns.join(', ')})`) }
701  }
702  const writes: Write[] = []
703  let cwd = context.cwd
704  const write = (writer: string, raw: string) => writes.push({ writer, raw, path: resolvePath(raw, cwd, context.home) })
705  // Another statement on the line that can change what a git inherits.
706  const envChangers = parse.statements.filter(changesEnv)
707  for (const s of parse.statements) {
708    if (!s.isPlaced && !isDocs) return { finding: unreadable(`A wrapper option in front of \`${s.words[s.nameAt] ?? ''}\` cannot be placed`) }
709    if (!isDocs) {
710      for (const r of s.redirects) {
711        if (!r.isReal) continue
712        const isFile = ['>', '>>', '>|', '&>', '&>>', '<>'].includes(r.op) || (r.op === '>&' && !/^(?:[0-9]+-?|-)$/.test(r.target))
713        if (!isFile || DEVICE.test(r.target)) continue
714        // zsh reads `>!` and `>>!` as one operator that clobbers the next
715        // word, where bash reads a file named `!`: both are judged.
716        if (r.target === '!') writes.push({ writer: 'a redirect', raw: r.target, path: undefined })
717        else write('a redirect', r.target)
718        if (r.target.startsWith('!') && r.target.length > 1) write('a redirect', r.target.slice(1))
719      }
720    }
721    if (s.nameAt < 0) continue
722    const pre = s.words.slice(0, s.nameAt)
723    if ((s.wrappers.includes('env') || s.wrappers.includes('sudo')) && pre.some(w => CHDIR_WORD.test(w))) cwd = undefined
724    if (!isDocs && BUILT_NAME.test(s.name) && !['[', '[['].includes(s.name)) {
725      return { finding: unreadable(`The command name \`${s.words[s.nameAt] ?? ''}\` is built when the line runs`) }
726    }
727    const stepped = stepThrough(s.name, s.args)
728    if (stepped.finding !== undefined) return { finding: stepped.finding }
729    if (stepped.chdir) cwd = undefined
730    const { name, args } = stepped
731    const isFed = s.wrappers.includes('xargs') || s.wrappers.includes('parallel')
732    if (DIRECTORY_CHANGES.has(name)) {
733      cwd = undefined
734      continue
735    }
736    if (name === 'git') {
737      const at = s.subcommandAt
738      const globals = at < 0 ? s.args : s.args.slice(0, at)
739      if (globals.some(a => a === '-c' || /^-c./.test(a) || a.startsWith('--config-env') || a.startsWith('--exec-path='))) {
740        return { finding: { action: '`git -c`', reason: 'git is run with configuration on its command line, which can name a program git runs (core.pager, diff.external, an alias). Run git without -c.' } }
741      }
742      // The environment is configuration too: GIT_EXTERNAL_DIFF, GIT_SSH_COMMAND
743      // and GIT_CONFIG_COUNT/KEY/VALUE each name a program git runs.
744      if (s.assignments.some(isGitEnv) || s.words.some(w => ASSIGNING_EXPANSION.test(w)) || envChangers.some(st => st !== s)) {
745        return {
746          finding: {
747            action: '`git` beside a change to its environment',
748            reason:
749              'git is run with a GIT_ variable in front of it, or beside a command that can change the environment it inherits (an assignment, export, read, a for loop, set, source, eval). A GIT_ variable can name a program git runs (GIT_EXTERNAL_DIFF, GIT_SSH_COMMAND, GIT_CONFIG_COUNT). Run git on its own line, with no variable set before it.',
750          },
751        }
752      }
753      if (at >= 0) {
754        const verb = (s.args[at] ?? '').toLowerCase()
755        const rest = s.args.slice(at + 1)
756        if (!gitReads(verb, rest)) {
757          return {
758            finding: {
759              action: `\`git ${verb}\``,
760              reason: `\`git ${verb}\` is not one of git's read-only forms, so it is refused wherever it points. \`git restore\`, \`checkout\`, \`switch\`, \`reset\`, \`clean\`, and \`stash\` all discard exactly the uncommitted change a Local-mode review is sent to read.`,
761            },
762          }
763        }
764        for (const out of gitOutputs(verb, rest)) write(`git ${verb}`, out)
765      }
766      continue
767    }
768    if (MUTATING.has(name)) {
769      if (isFed) return { finding: { action: `\`${name}\``, reason: `\`${name}\` changes files, and its paths arrive from xargs or parallel at run time, where the guard cannot read them.` } }
770      for (const raw of writeTargets(name, args, cwd ?? '')) write(name, raw)
771      continue
772    }
773    if (name === 'find') {
774      const writer = args.find(a => ['-delete', '-exec', '-execdir', '-ok', '-okdir', '-fprint', '-fprint0', '-fprintf', '-fls'].includes(a))
775      if (writer !== undefined) return { finding: { action: `\`find ${writer}\``, reason: `\`find ${writer}\` deletes, writes, or runs a command against the files it matches.` } }
776      continue
777    }
778    const inPlace = inPlaceTargets(name, args)
779    if (inPlace !== undefined) {
780      if (isFed || inPlace.length === 0) return { finding: { action: `\`${name} -i\``, reason: `\`${name} -i\` rewrites files in place, and the guard cannot tell which.` } }
781      for (const raw of inPlace) write(`${name} -i`, raw)
782    }
783    if (name === 'sed' && sedPrograms(args).some(p => SED_WRITES.test(p))) {
784      return { finding: { action: '`sed` writing or running', reason: 'the sed program writes a file (w) or runs a command (e).' } }
785    }
786    if (AWK.has(name) && args.some(a => AWK_WRITES.test(a))) {
787      return { finding: { action: `\`${name}\` writing or running`, reason: `the ${name} program prints to a file or a pipe, or runs a command.` } }
788    }
789    if (inPlace === undefined && args.some(a => IN_PLACE_FLAGS.has(a.split('=')[0] ?? a))) {
790      return { finding: { action: `\`${name}\` with a rewrite flag`, reason: `\`${name}\` is being run with a rewrite flag. Run formatters and linters in check mode only.` } }
791    }
792    const formatter = formatterWrites(name, args)
793    if (formatter !== undefined) {
794      return { finding: { action: `\`${formatter}\` rewriting files`, reason: `\`${formatter}\` rewrites files unless it runs in check mode. Run formatters and linters in check mode only, such as \`--check\`.` } }
795    }
796    if (!isDocs && !KNOWN.has(name)) {
797      // Each word of an argument counts, after an `=` too: `--split-string=rm x`.
798      const behind = args.flatMap(a => a.split(/[\s=]+/)).find(word => WRITER_NAMES.has(basename(word)))
799      if (behind !== undefined) {
800        return { finding: unreadable(`\`${name}\` is a program the guard does not know, and \`${behind}\` stands among its words, where it may run`) }
801      }
802    }
803  }
804  return { writes }
805}
806
807// The verdict on an Edit, Write or NotebookEdit of `raw`.
808export function reviewEdit(tool: string, raw: string, context: Context): Verdict {
809  return { writes: [{ writer: tool, raw, path: resolvePath(raw, context.cwd, context.home) }] }
810}
811
812// The refusal the model reads: the action, the reason, and the way to work.
813export function refusalOf(finding: Finding, roots: readonly string[]): string {
814  const where = roots.map(root => `\`${root}\``).join(', ') || 'none could be found on this host'
815  return (
816    `🛑 Blocked: ${finding.action}. A Holmes reviewer writes only in scratch.\n\n` +
817    `Review guard (workbench-dev-team). This command is refused, because ${finding.reason}\n\n` +
818    'You are running as a Holmes reviewer, and a reviewer writes nothing outside ' +
819    `the scratch roots. The scratch roots here are: ${where}. The code under review ` +
820    "is read, never changed: in Local mode it is the human's live working " +
821    'directory, and the uncommitted change in it is the only copy of the work.\n\n' +
822    'Read it instead. `git status`, `git diff HEAD`, `git ls-files --others ' +
823    '--exclude-standard`, and `git show` are all allowed, and so is the test suite. ' +
824    'A probe that needs a changed tree runs on a copy in your own folder, made with ' +
825    '`mktemp -d` under the session scratchpad or ~/Developer/scratchpad, and deleted ' +
826    'before you report with `rm -rf` and its literal path as its own command. ' +
827    'If a failure looks pre-existing, say so in your findings — never ' +
828    'isolate it by changing the tree.\n\n' +
829    'There is no flag to clear and no path around this. If you believe you are not ' +
830    'a reviewer, report that to the session that dispatched you and stop.'
831  )
832}
833
834// The disk, as the guard reads it. stat answers where a path lands with every
835// symlink resolved (realPath undefined when it leads nowhere), or undefined
836// when nothing is there. names lists a directory's entries, links included.
837export type Disk = {
838  stat: (path: string) => Promise<{ realPath: string | undefined } | undefined>
839  names: (dir: string) => Promise<string[] | undefined>
840}
841
842// Where an absolute path lands with every symlink resolved: the deepest part
843// that exists, resolved, then the rest as written. Undefined when a part that
844// does not exist is `.` or `..`, or exists only as a link that leads nowhere,
845// since either could land anywhere.
846export async function landing(path: string, disk: Disk): Promise<string | undefined> {
847  const parts = path.split('/').filter(part => part !== '')
848  for (let i = parts.length; i >= 0; i--) {
849    const head = `/${parts.slice(0, i).join('/')}`
850    const stat = await disk.stat(head)
851    if (stat === undefined) continue
852    if (stat.realPath === undefined) return undefined
853    const rest = parts.slice(i)
854    if (rest.some(part => part === '.' || part === '..')) return undefined
855    if (rest.length > 0) {
856      const names = await disk.names(head)
857      if (names === undefined || names.includes(rest[0] ?? '')) return undefined
858    }
859    return [stat.realPath.replace(/\/$/, ''), ...rest].join('/')
860  }
861  return undefined
862}
863
864// Whether a path lands strictly beneath a scratch root, both through every
865// symlink and with its final name left where it lives: `rm <dir>/link`
866// removes the link where it lives, wherever it points.
867export async function isInScratch(path: string, roots: readonly string[], disk: Disk): Promise<boolean> {
868  const trimmed = path.replace(/\/+$/, '')
869  const cut = trimmed.lastIndexOf('/')
870  const name = trimmed.slice(cut + 1)
871  const followed = await landing(path, disk)
872  let own = followed
873  if (!(name === '' || name === '.' || name === '..' || path.endsWith('/'))) {
874    const parent = await landing(trimmed.slice(0, cut) || '/', disk)
875    own = parent === undefined ? undefined : `${parent.replace(/\/$/, '')}/${name}`
876  }
877  const beneath = (p: string | undefined) => p !== undefined && roots.some(root => p.startsWith(`${root.replace(/\/$/, '')}/`))
878  return beneath(followed) && beneath(own)
879}
880
881// The first write that lands outside the scratch roots, as a finding.
882export async function judgeWrites(writes: readonly Write[], roots: readonly string[], disk: Disk): Promise<Finding | undefined> {
883  for (const w of writes) {
884    const base = `\`${w.writer}\` changes a file's content, location, existence, or metadata`
885    if (roots.length === 0) return { action: `\`${w.writer}\``, reason: `${base}, and no scratch root was found to judge its paths against.` }
886    if (w.path === undefined) return { action: `\`${w.writer}\``, reason: `${base}, and the path \`${w.raw || '(none)'}\` cannot be resolved.` }
887    if (!(await isInScratch(w.path, roots, disk))) return { action: `\`${w.writer}\``, reason: `${base}, and \`${w.raw}\` is outside the scratch roots.` }
888  }
889  return undefined
890}
891
hooks/mods/panes.tsx 147 lines
1// The two dev-team panes' trees, from the views hooks/register.ts keeps. Pure:
2// the elements come in from the hook's $.ui.resolve(e), and nothing here
3// touches `$`. Each row is a keyed Box, so a test finds it by key.
4
5import type { Color, Elements, RenderElement, RenderSurface } from 'claude-code'
6
7import type { BoardItem, BoardLane, BoardView, RunState, RunsView } from '../../types'
8
9export type Draw = Pick<Elements[RenderSurface], 'Box' | 'Text'>
10
11export const RUNS_PANE = 'dev-team-runs'
12export const BOARD_PANE = 'dev-team-board'
13
14const STATE_COLOR: Record<RunState, Color> = {
15  running: 'suggestion',
16  done: 'success',
17  failed: 'error',
18  refused: 'error',
19  'budget-killed': 'warning',
20  escalated: 'warning',
21}
22
23const pad = (n: number): string => String(n).padStart(2, '0')
24
25// A time of day, HH:MM, in the host's zone.
26export const clockOf = (ms: number): string => {
27  const at = new Date(ms)
28  return `${pad(at.getHours())}:${pad(at.getMinutes())}`
29}
30
31export function runsTree({ Box, Text }: Draw, view: RunsView, room: number): RenderElement {
32  if (view.error !== undefined) {
33    return (
34      <Box flexDirection="column">
35        <Text color="error">{view.error}</Text>
36      </Box>
37    )
38  }
39  if (view.rows.length === 0) {
40    return (
41      <Box flexDirection="column">
42        <Text dimColor>No dispatched runs in the logs yet.</Text>
43      </Box>
44    )
45  }
46  // Two lines a run: the run, then its log.
47  const rows = view.rows.slice(0, Math.max(1, Math.floor(room / 2)))
48  return (
49    <Box flexDirection="column">
50      {rows.map((row, i) => (
51        <Box key={`run-${i}`} flexDirection="column">
52          <Text wrap="truncate-end">
53            <Text color={STATE_COLOR[row.state]} bold>
54              {row.state.padEnd(13)}
55            </Text>
56            {` ${row.agent.padEnd(8)} ${row.target} · started ${row.startedAt}`}
57            {row.refusals > 0 ? ` · ${row.refusals} refused call${row.refusals === 1 ? '' : 's'}` : ''}
58          </Text>
59          <Text dimColor wrap="truncate-start">
60            {`  ${row.log}`}
61          </Text>
62        </Box>
63      ))}
64    </Box>
65  )
66}
67
68const LANE_TITLE = { unrefined: 'Unrefined (Lestrade)', review: 'Review (Holmes)', development: 'Development (Watson)' } as const
69
70// How many items a lane holds: the count, or "25+" when the list came back full.
71const countOf = (lane: { limit: number; items: BoardItem[] }): string => (lane.items.length >= lane.limit ? `${lane.limit}+` : String(lane.items.length))
72
73const refOf = (item: BoardItem): string => `${item.number === null ? `item ${item.id}` : `${item.isPr ? 'PR ' : ''}#${item.number}`} ${item.repo ?? ''}`.trimEnd()
74
75// How many items each lane lists under its count.
76export const TOP_ITEMS = 3
77
78function laneRows({ Box, Text }: Draw, name: keyof typeof LANE_TITLE, lane: BoardLane): RenderElement {
79  if ('error' in lane) {
80    return (
81      <Box key={`lane-${name}`} flexDirection="column">
82        <Text>
83          <Text bold>{LANE_TITLE[name]}</Text>
84          <Text color="error">{` could not list: ${lane.error}`}</Text>
85        </Text>
86      </Box>
87    )
88  }
89  const claimed = name === 'development' ? lane.items.filter(item => item.claimedAt !== null).length : 0
90  return (
91    <Box key={`lane-${name}`} flexDirection="column">
92      <Text>
93        <Text bold>{LANE_TITLE[name]}</Text>
94        {` ${countOf(lane)}${claimed > 0 ? ` · ${claimed} claimed` : ''}`}
95      </Text>
96      {lane.items.slice(0, TOP_ITEMS).map(item => (
97        <Text dimColor wrap="truncate-end">
98          {`  ${refOf(item)}  ${item.title ?? ''}`}
99        </Text>
100      ))}
101    </Box>
102  )
103}
104
105export function boardTree(draw: Draw, view: BoardView, cadenceMs: number): RenderElement {
106  const { Box, Text } = draw
107  const board = view.board
108  const claimed = board && !('error' in board.development) ? board.development.items.filter(item => item.claimedAt !== null) : []
109  const status =
110    view.attemptedAt === 0
111      ? 'Fetching the board from The Index.'
112      : `${board && view.boardAt !== undefined ? `Fetched ${clockOf(view.boardAt)}` : 'Not fetched'} · next fetch after ${clockOf(view.attemptedAt + cadenceMs)}`
113  return (
114    <Box flexDirection="column">
115      <Box key="board-status">
116        <Text dimColor>{status}</Text>
117      </Box>
118      {view.error !== undefined && (
119        <Box key="board-error">
120          <Text color="error">{`The last fetch failed: ${view.error}`}</Text>
121        </Box>
122      )}
123      {board && laneRows(draw, 'unrefined', board.unrefined)}
124      {board && laneRows(draw, 'review', board.review)}
125      {board && laneRows(draw, 'development', board.development)}
126      {claimed.length > 0 && (
127        <Box key="claimed" flexDirection="column">
128          <Text bold>Claimed</Text>
129          {claimed.map(item => (
130            <Text wrap="truncate-end">{`  ${refOf(item)}  since ${item.claimedAt}`}</Text>
131          ))}
132        </Box>
133      )}
134      {view.escalated.length > 0 && (
135        <Box key="escalated" flexDirection="column">
136          <Text bold color="warning">
137            Escalated by the breaker
138          </Text>
139          {view.escalated.map(line => (
140            <Text>{`  ${line}`}</Text>
141          ))}
142        </Box>
143      )}
144    </Box>
145  )
146}
147
hooks/mods/runs.ts 140 lines
1// The runs pane's logic, as pure functions. hooks/register.ts holds the hooks
2// and reads the disk; hooks/mods/panes.tsx draws the rows.
3//
4// bin/dispatch-agent.sh writes one log per run in ~/.claude-workbench/dev-team-logs:
5//   <agent>-<item id>-<YYYYMMDD-HHMMSS>.log     an item run
6//   lestrade-sweep-<owner>-<repo>-<stamp>.log  a blocker sweep ("/" written as "-")
7// beside <agent>-<id>.lock (the run's pid) and <agent>-<id>.escalated (the
8// breaker escalated the item after its last run). The pane reads the newest
9// RUN_LIMIT logs only, and each one again only when its size or mtime changed.
10// It reads them with one awk (bin/scan-run-logs.awk) over the changed files,
11// and finds the live runs with one ps. It writes nothing.
12
13export const LOG_DIR = '.claude-workbench/dev-team-logs'
14
15// The most runs the pane shows.
16export const RUN_LIMIT = 8
17
18// How often the runs pane reads the logs while it is open.
19export const RUNS_REFRESH_MS = 15_000
20
21// The view types live in the type contract, types/index.d.ts, which
22// `claude plugin validate` holds every $.state value to.
23import type { DevTeamAgent as Agent, RunRow, RunState } from '../../types'
24
25// What a run's log name says: its agent, what it ran on, when it started, and
26// the pair the lock, the marker and the dispatcher's command line name.
27export type RunLog = { name: string; agent: Agent; target: string; pair: string; startedAt: string }
28
29// The end of a log, as the breaker reads it: its last three lines that are not
30// permission refusals, and how many refusals it holds.
31export type LogEnd = { tail: string[]; refusals: number }
32
33const ITEM_LOG = /^(lestrade|holmes|watson)-([0-9]+)-([0-9]{8})-([0-9]{6})\.log$/
34const SWEEP_LOG = /^lestrade-sweep-(.+)-([0-9]{8})-([0-9]{6})\.log$/
35const MARKER = /^(lestrade|holmes|watson)-([0-9]+)\.escalated$/
36
37const stampOf = (day: string, time: string): string =>
38  `${day.slice(0, 4)}-${day.slice(4, 6)}-${day.slice(6)} ${time.slice(0, 2)}:${time.slice(2, 4)}:${time.slice(4)}`
39
40// The run a log file name records, or undefined for any other file:
41// dispatch-tick.log, a lock, a marker.
42export function runLogOf(name: string): RunLog | undefined {
43  const item = ITEM_LOG.exec(name)
44  if (item) {
45    const [, agent, id, day, time] = item as unknown as [string, Agent, string, string, string]
46    return { name, agent, target: `item ${id}`, pair: `${agent}-${id}`, startedAt: stampOf(day, time) }
47  }
48  const sweep = SWEEP_LOG.exec(name)
49  if (sweep) {
50    const [, slug, day, time] = sweep as unknown as [string, string, string, string]
51    return { name, agent: 'lestrade', target: `sweep ${slug}`, pair: `lestrade-sweep-${slug}`, startedAt: stampOf(day, time) }
52  }
53  return undefined
54}
55
56export type Entry = { name: string; kind: string; mtimeMs: number; size: number }
57export type NewestRun = RunLog & { mtimeMs: number; size: number }
58
59// The newest `limit` run logs in a listing, newest first.
60export function newestRuns(entries: readonly Entry[], limit = RUN_LIMIT): NewestRun[] {
61  const runs: NewestRun[] = []
62  for (const entry of entries) {
63    const log = entry.kind === 'file' ? runLogOf(entry.name) : undefined
64    if (log) runs.push({ ...log, mtimeMs: entry.mtimeMs, size: entry.size })
65  }
66  return runs.sort((a, b) => b.mtimeMs - a.mtimeMs || (a.name < b.name ? 1 : -1)).slice(0, limit)
67}
68
69// The pairs whose item the breaker escalated: one <agent>-<id>.escalated each.
70export function markersOf(entries: readonly Entry[]): string[] {
71  return entries.flatMap(entry => {
72    const marker = MARKER.exec(entry.name)
73    return marker ? [`${marker[1]}-${marker[2]}`] : []
74  })
75}
76
77// The pairs a live dispatcher runs, read from `ps -axo command=`. The
78// dispatcher's run is its wrapper subshell, whose command line is the
79// dispatcher's own: `bash …/dispatch-agent.sh <agent> <target>`. A `--check` or
80// `--mark-escalated` call puts its flag before the agent, so it never matches.
81export function livePairsOf(ps: string): Set<string> {
82  const live = new Set<string>()
83  for (const line of ps.split('\n')) {
84    const run = /dispatch-agent\.sh\s+(lestrade|holmes|watson)\s+(\S+)\s*$/.exec(line)
85    if (!run) continue
86    const [, agent, target] = run as unknown as [string, Agent, string]
87    if (/^[0-9]+$/.test(target)) live.add(`${agent}-${target}`)
88    else if (agent === 'lestrade' && target.includes('/')) live.add(`lestrade-sweep-${target.replaceAll('/', '-')}`)
89  }
90  return live
91}
92
93// The awk program that reads the end of the logs, shipped beside the scripts
94// so bin/test-scan-run-logs.sh runs the real thing.
95export const scanArgv = (root: string, logs: readonly string[]): string[] => ['awk', '-f', `${root}/bin/scan-run-logs.awk`, ...logs]
96
97// The awk run's output, by path. A path that printed nothing is absent.
98export function endsOf(stdout: string): Map<string, LogEnd> {
99  const ends = new Map<string, LogEnd>()
100  let current: LogEnd | undefined
101  for (const line of stdout.split('\n')) {
102    if (line.startsWith('\u001e')) {
103      const tab = line.indexOf('\t')
104      current = { tail: [], refusals: Number(line.slice(1, tab)) || 0 }
105      ends.set(line.slice(tab + 1), current)
106    } else if (current !== undefined && current.tail.length < 3) {
107      current.tail.push(line)
108    }
109  }
110  // The output's own last newline is no line of the log.
111  for (const end of ends.values()) while (end.tail.length > 0 && end.tail[end.tail.length - 1] === '') end.tail.pop()
112  return ends
113}
114
115// A run's state. Only the newest run of a pair can be the one running, or the
116// one the breaker escalated after; an older run of the pair has ended. The end
117// is read as the breaker reads it, from the last lines that are not refusals.
118export function stateOf(facts: { isNewest: boolean; isLive: boolean; isEscalated: boolean; end: LogEnd }): RunState {
119  if (facts.isNewest && facts.isLive) return 'running'
120  if (facts.isNewest && facts.isEscalated) return 'escalated'
121  const tail = facts.end.tail
122  if (tail.some(line => /content filtering policy/i.test(line))) return 'refused'
123  if (tail.some(line => line.includes('Exceeded USD budget'))) return 'budget-killed'
124  if (tail.some(line => /^(API Error|Execution error|Error:)/i.test(line))) return 'failed'
125  return 'done'
126}
127
128// The pane's rows, newest first.
129export function rowsOf(dir: string, runs: readonly NewestRun[], ends: ReadonlyMap<string, LogEnd>, live: ReadonlySet<string>, markers: readonly string[]): RunRow[] {
130  const seen = new Set<string>()
131  return runs.map(run => {
132    const isNewest = !seen.has(run.pair)
133    seen.add(run.pair)
134    const log = `${dir}/${run.name}`
135    const end = ends.get(log) ?? { tail: [], refusals: 0 }
136    const state = stateOf({ isNewest, isLive: live.has(run.pair), isEscalated: markers.includes(run.pair), end })
137    return { agent: run.agent, target: run.target, startedAt: run.startedAt, state, refusals: end.refusals, log }
138  })
139}
140
hooks/mods/scratch.ts 103 lines
1// Each dev-team agent's own scratch folder, as pure functions. hooks/register.ts
2// holds the hooks: tool.call points a bare mktemp at the folder, turn.complete
3// deletes the folder when the agent's run ends with no live child, and
4// session.end deletes every folder still recorded.
5//
6// A bare mktemp is one with no template and no directory option (`mktemp`,
7// `mktemp -d`, `mktemp -dq`), which lands in $TMPDIR, outside every scratch
8// root, and outlives the run. The rewrite adds a template inside the agent's
9// folder, so the folder holds every temporary file and folder the run made.
10//
11// Only a dev-team agent's line is pointed: a sub-agent of a dev-team type
12// (holmes-lens included), and the top-level loop of a `claude -p --agent` run
13// of one. The main session and every other agent keep mktemp as they wrote it.
14//
15// The rewrite is checked, never trusted. The line is read again after it, and
16// the rewrite stands only when every statement reads exactly as before, apart
17// from each bare mktemp gaining the one template word. Anything else (a mktemp
18// the text match missed, or one it found inside a quoted string or a heredoc)
19// leaves the line as it was written, which is how it ran before the mod.
20
21import type { ShellParse, Statement } from './commit-guard'
22
23// mktemp's no-value options. -t, -p and --tmpdir name a directory, so a line
24// with any of them is not bare.
25const FLAGS = /^-[dqu]+$/
26
27export const isBareMktemp = (statement: Statement): boolean =>
28  statement.name === 'mktemp' &&
29  statement.nameAt === 0 &&
30  statement.wrappers.length === 0 &&
31  statement.assignments.length === 0 &&
32  statement.args.every(arg => FLAGS.test(arg))
33
34export const hasBareMktemp = (parse: ShellParse): boolean => parse.statements.some(isBareMktemp)
35
36// The template a bare mktemp gains: the folder, quoted, and mktemp's X run.
37export const templateOf = (folder: string): string => `${folder}/tmp.XXXXXXXX`
38
39// A bare mktemp in the text: the name at a command boundary, its no-value
40// options, and then the end of the command, or a `#` comment after a space.
41const BARE = /(^|[\s;&|(`])mktemp((?:[ \t]+-[dqu]+)*)(?=[ \t]*(?:$|[;&|)`\n])|[ \t]+#)/gm
42
43// A folder the rewrite can quote: an absolute path of plain characters.
44const PLAIN_PATH = /^\/[A-Za-z0-9._@+\/-]+$/
45
46// The line with every bare mktemp the text match finds pointed at `folder`, or
47// undefined when there is nothing to point or the folder cannot be quoted.
48// The caller reads the result again and keeps it only when readsAsPointed.
49export function pointMktemp(line: string, folder: string): string | undefined {
50  if (!PLAIN_PATH.test(folder)) return undefined
51  const rewritten = line.replace(BARE, (_, lead: string, flags: string) => `${lead}mktemp${flags} '${templateOf(folder)}'`)
52  return rewritten === line ? undefined : rewritten
53}
54
55// Whether the rewritten line reads exactly as the line did, apart from each
56// bare mktemp gaining the template word: the same statements, the same words,
57// redirects, heredocs and unknowns, in the same order.
58export function readsAsPointed(before: ShellParse, after: ShellParse, folder: string): boolean {
59  const template = templateOf(folder)
60  if (!hasBareMktemp(before) || JSON.stringify(after.unknowns) !== JSON.stringify(before.unknowns)) return false
61  if (after.statements.length !== before.statements.length) return false
62  const expected = before.statements.map(s => (isBareMktemp(s) ? { ...s, words: [...s.words, template], args: [...s.args, template] } : s))
63  return after.statements.every((s, i) => JSON.stringify(s) === JSON.stringify(expected[i]))
64}
65
66// The scratch root an agent's folder goes in: the first of core's roots that
67// is not ~/.claude/plans, so the session scratchpad when there is one, and
68// ~/Developer/scratchpad when there is not, as the agents' prose orders them.
69export const rootOf = (roots: readonly string[]): string | undefined => roots.find(root => !root.endsWith('/.claude/plans'))
70
71// The folder's name prefix: the agent's bare type (`watson-direct`), in plain
72// characters.
73export const prefixOf = (type: string): string => (type.split(':').pop() ?? 'agent').replace(/[^A-Za-z0-9_-]/g, '') || 'agent'
74
75// The statuses of a loop that is not over: not started, in a turn, held, or
76// between turns until a message wakes it.
77const LIVE: ReadonlySet<string> = new Set(['pending', 'running', 'waiting', 'idle'])
78
79// Whether a run still has a live child: an agent its loop spawned (for the
80// top-level `run`, one with no parent) that is not over. A child may read the
81// run's folder, as Holmes's helpers read his checkout, so the folder stays.
82export const hasLiveChild = (agents: readonly { parentId?: string; status: string }[], agentId: string | undefined): boolean =>
83  agents.some(agent => agent.parentId === agentId && LIVE.has(agent.status))
84
85// Whether `folder` may be deleted: a folder directly or deeper inside one of
86// the scratch roots, never a root itself.
87export const isDeletable = (folder: string, roots: readonly string[]): boolean =>
88  PLAIN_PATH.test(folder) && !folder.split('/').includes('..') && roots.some(root => folder.startsWith(`${root}/`) && folder.length > root.length + 1)
89
90// What leaves room for the engine's own end step out of the end's budget, and
91// the most the sweep's rm may take of it.
92const END_MARGIN_MS = 150
93const END_RM_CAP_MS = 1_000
94
95// The session end's sweep: the recorded folders one rm may delete, the ones
96// under a scratch root, and the time it may take, or undefined when there is
97// nothing to delete or too little time left. A folder left out stays recorded.
98export function sweepOf(folders: readonly string[], roots: readonly string[], remainingMs: number): { doomed: string[]; timeoutMs: number } | undefined {
99  const doomed = folders.filter(folder => isDeletable(folder, roots))
100  const timeoutMs = Math.min(remainingMs, END_RM_CAP_MS) - END_MARGIN_MS
101  return doomed.length === 0 || timeoutMs <= 0 ? undefined : { doomed, timeoutMs }
102}
103
hooks/mods/spawn.ts 291 lines
1// What happens when a session dispatches a dev-team agent, as pure functions:
2// which mode agent a dispatch runs, what model and effort the config gives it,
3// and the dispatch gate's refusal and advisory hint. hooks/register.ts holds
4// every hook that touches `$` and calls these, because the engine follows `$`
5// into no imported function.
6
7import type { EngineInterface } from 'claude-code'
8
9import type { DevTeamEffort } from '../../types'
10
11// workbench-core's $.workbench contract, read off the noun itself: dev-team
12// lists core under `dependencies`, and the engine types the noun on `$`.
13type Workbench = EngineInterface['workbench']
14export type BriefCheck = Awaited<ReturnType<Workbench['briefCheck']>>
15export type BriefSlot = Awaited<ReturnType<Workbench['briefSlots']>>[number]
16export type CallerLane = Awaited<ReturnType<Workbench['callerLane']>>
17
18export const PLUGIN = 'workbench-dev-team'
19
20export type Family = 'watson' | 'holmes' | 'lestrade'
21
22// The agent types a family answers to: the public type, and each mode type
23// bin/compose-agents.sh builds from it. A dispatch of any of them is routed by
24// its token, so a mode type dispatched by name still runs the mode its prompt
25// calls for.
26const FAMILIES: Readonly<Record<Family, readonly string[]>> = {
27  watson: ['watson', 'watson-direct', 'watson-index'],
28  holmes: ['holmes', 'holmes-local', 'holmes-index'],
29  lestrade: ['lestrade', 'lestrade-item', 'lestrade-sweep'],
30}
31
32export function familyOf(subagentType: string | undefined): Family | undefined {
33  if (typeof subagentType !== 'string') return undefined
34  const [plugin, name] = subagentType.split(':')
35  if (plugin !== PLUGIN || name === undefined) return undefined
36  return (Object.keys(FAMILIES) as Family[]).find(family => FAMILIES[family].includes(name))
37}
38
39// The tokens. A token picks a mode only when it is the whole prompt: one
40// non-blank line, and that line the token, which is how workbench-core's
41// briefCheck reads the `item-id` and `repo-sweep` shapes. A brief that only
42// mentions `Item ID: 9`, in its Context or anywhere else, is a brief, so it
43// runs the off-board mode and the agent's own rule that a mention is not a
44// dispatch token still applies. Watson and Holmes take `Item ID: <n>` or one
45// bare id (a plain integer, a UUID, or a PVTI_ id) for The Index mode. Lestrade
46// takes `Item ID: <n>` or a bare integer for Item mode, and
47// `Repo sweep: <owner/repo>` for Sweep mode. Whitespace is the five ASCII
48// characters core's check uses, so an exotic space makes a line no token.
49const SPACE = '[ \\t\\r\\v\\f]*'
50const ITEM_ID = new RegExp(`^${SPACE}Item ID:${SPACE}[0-9]+${SPACE}$`)
51const BARE_INTEGER = new RegExp(`^${SPACE}[0-9]+${SPACE}$`)
52const BARE_ID = new RegExp(
53  `^${SPACE}(?:[0-9]+|[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}|PVTI_[A-Za-z0-9_-]+)${SPACE}$`,
54)
55const REPO_SWEEP = new RegExp(`^${SPACE}Repo sweep:${SPACE}[^ \\t\\r\\v\\f/]+/[^ \\t\\r\\v\\f/]+${SPACE}$`)
56const NONBLANK = /[^ \t\r\v\f\n]/
57
58// Whether the whole prompt is one token: a single non-blank line it matches.
59function isWhole(prompt: string, ...tokens: RegExp[]): boolean {
60  const lines = prompt.split('\n').filter(line => NONBLANK.test(line))
61  return lines.length === 1 && tokens.some(token => token.test(lines[0] ?? ''))
62}
63
64// The item id of a prompt that is the whole token `Item ID: <n>`, or undefined.
65export function itemIdOf(prompt: string): string | undefined {
66  if (!isWhole(prompt, ITEM_ID)) return undefined
67  return /[0-9]+/.exec(prompt.split('\n').find(line => NONBLANK.test(line)) ?? '')?.[0]
68}
69
70// The mode agent a dispatch runs, as a full type. Anything but a whole-prompt
71// token is Direct mode for Watson and Local mode for Holmes, as their files
72// say. Lestrade has no prose mode, so a prompt that is no token stays on the
73// public type, which decides for itself.
74export function modeTypeOf(family: Family, prompt: string): string {
75  switch (family) {
76    case 'watson':
77      return `${PLUGIN}:${isWhole(prompt, ITEM_ID, BARE_ID) ? 'watson-index' : 'watson-direct'}`
78    case 'holmes':
79      return `${PLUGIN}:${isWhole(prompt, ITEM_ID, BARE_ID) ? 'holmes-index' : 'holmes-local'}`
80    case 'lestrade':
81      if (isWhole(prompt, ITEM_ID, BARE_INTEGER)) return `${PLUGIN}:lestrade-item`
82      if (isWhole(prompt, REPO_SWEEP)) return `${PLUGIN}:lestrade-sweep`
83      return `${PLUGIN}:lestrade`
84  }
85}
86
87// What the /config rows give one family's interactive dispatch, read through
88// configTextOf. The values are the person's, so each is checked before it is
89// used: a model is an alias or a full id with an optional bracketed suffix
90// (`claude-opus-5-5[1m]`), and an effort is a level. turn.step refuses a
91// numeric effort from a hook ("a number is internal-only"), so a number is not
92// one. A value outside that shape, and a missing one, leave the knob unset, so
93// the agent file's own value applies. A bad value never blocks a dispatch, as
94// on the scheduled path.
95export type Knobs = { model?: string; effort?: DevTeamEffort }
96
97const MODEL = /^[A-Za-z0-9][A-Za-z0-9._-]*(?:\[[A-Za-z0-9]+\])?$/
98const LEVELS: readonly string[] = ['low', 'medium', 'high', 'xhigh', 'max']
99
100export function knobsOf(configText: string | undefined, family: Family): Knobs {
101  if (configText === undefined) return {}
102  let config: unknown
103  try {
104    config = JSON.parse(configText)
105  } catch {
106    return {}
107  }
108  const entry = (config as { agents?: Record<string, unknown> } | null)?.agents?.[family]
109  if (entry === null || typeof entry !== 'object') return {}
110  const { model, effort } = entry as { model?: unknown; effort?: unknown }
111  const knobs: Knobs = {}
112  if (typeof model === 'string' && MODEL.test(model)) knobs.model = model
113  if (typeof effort === 'string' && LEVELS.includes(effort.toLowerCase())) knobs.effort = effort.toLowerCase() as DevTeamEffort
114  return knobs
115}
116
117// The knobs the review agents read at run time: whether Holmes and Lestrade
118// fan out to helpers, and the model the helpers run on. The mod adds them to
119// the prompt of the three modes that read them, so no agent reads the file:
120// at spawn for an interactive dispatch, and on prompt.submit for the top-level
121// run bin/dispatch-agent.sh starts, which raises no agent.spawn.
122export const CONFIG_MODES: Readonly<Record<string, Family>> = {
123  [`${PLUGIN}:holmes-local`]: 'holmes',
124  [`${PLUGIN}:holmes-index`]: 'holmes',
125  [`${PLUGIN}:lestrade-item`]: 'lestrade',
126}
127
128// The plugin's /config rows, as the config text knobsOf and configLineOf read:
129// one entry per agent, keyed as the old config file keyed it, so
130// `watsonModel` is `agents.watson.model`. A row with no value is left out.
131const ROW_KNOBS = ['model', 'effort', 'fanout', 'lensModel'] as const
132export function configTextOf(options: Readonly<Record<string, unknown>> | undefined): string {
133  const entry = (family: Family) =>
134    Object.fromEntries(
135      ROW_KNOBS.map(knob => [knob, options?.[`${family}${knob.charAt(0).toUpperCase()}${knob.slice(1)}`]] as const).filter(([, value]) => value !== undefined),
136    )
137  return JSON.stringify({ agents: Object.fromEntries((Object.keys(FAMILIES) as Family[]).map(family => [family, entry(family)])) })
138}
139
140// The line's label, which the agents' prose names.
141export const CONFIG_LABEL = 'Dev-team config:'
142
143// The line the mod adds for one family, from the config's text. Only a
144// `false` turns the fan-out off, so a missing or malformed value leaves it on,
145// as the agents' default is. A lensModel outside the model shape is unset, and
146// the helpers then run on the agent's own model.
147export function configLineOf(configText: string | undefined, family: Family): string {
148  let entry: unknown
149  try {
150    entry = (JSON.parse(configText ?? '') as { agents?: Record<string, unknown> } | null)?.agents?.[family]
151  } catch {
152    entry = undefined
153  }
154  const { fanout, lensModel } = entry !== null && typeof entry === 'object' ? (entry as { fanout?: unknown; lensModel?: unknown }) : {}
155  const model = typeof lensModel === 'string' && MODEL.test(lensModel) ? lensModel : undefined
156  return `${CONFIG_LABEL} fanout ${fanout === false ? 'off' : 'on'}; lensModel ${model ?? 'unset'}.`
157}
158
159// The prompt with the config line added once, after a blank line.
160export const withConfigLine = (prompt: string, line: string): string =>
161  prompt.includes(CONFIG_LABEL) ? prompt : `${prompt.replace(/\s+$/, '')}\n\n${line}`
162
163// ── Holmes's helpers run on holmes-lens ──────────────────────────────────────
164
165export const LENS_TYPE = `${PLUGIN}:holmes-lens`
166
167// Holmes in any mode: the public type and its two mode types, by full or bare
168// name. Not holmes-lens, which holds no Agent tool.
169const HOLMES_MODE = /(^|[:/])holmes(-local|-index)?$/iu
170export const isHolmesMode = (type: string | undefined): boolean => type !== undefined && HOLMES_MODE.test(type.trim())
171
172// The type a Holmes spawn runs as: holmes-lens, whatever the call named, or a
173// refusal for a fork or a teammate, which cannot be retyped.
174export function lensSpawnOf(e: { subagentType: string; fork: boolean; isTeammate?: true }): { subagentType: string } | { deny: string } {
175  if (e.subagentType === LENS_TYPE) return { subagentType: LENS_TYPE }
176  if (e.fork || e.isTeammate) return { deny: lensDeny(e.fork ? 'a fork' : 'a teammate') }
177  return { subagentType: LENS_TYPE }
178}
179
180export function lensDeny(what: string): string {
181  return [
182    `🛑 Blocked: Holmes spawned ${what}.`,
183    '',
184    `Helper rule (${PLUGIN}). Every helper Holmes spawns runs on ${LENS_TYPE}, the read-only helper, and ${what} cannot be retyped to it. ` +
185      `Dispatch the helper with subagent_type "${LENS_TYPE}".`,
186  ].join('\n')
187}
188
189// Whether a type is one of the dev-team's agents, the helper included: the
190// types the mod gives a scratch folder and a context budget.
191export const isDevTeamType = (type: string | undefined): boolean => type !== undefined && (familyOf(type) !== undefined || type === LENS_TYPE)
192
193// ── The default-branch check ─────────────────────────────────────────────────
194
195export const DEFAULT_BRANCHES: readonly string[] = ['main', 'master', 'trunk']
196
197// A brief's Workdir: the path, and whether the slot records a workspace
198// decision. The path runs to the first " (", and the slot records a decision
199// when the text after the path names a branch or a worktree, as
200// `/repo (branch: main, Mike chose main)` does. Undefined when the brief has
201// no Workdir line.
202export function workdirOf(prompt: string): { path: string; isRecorded: boolean } | undefined {
203  const line = prompt.split('\n').find(l => /^[ \t]*Workdir:/.test(l))
204  if (line === undefined) return undefined
205  const value = line.replace(/^[ \t]*Workdir:/, '').trim()
206  const cut = value.indexOf(' (')
207  const path = (cut === -1 ? value : value.slice(0, cut)).trim()
208  return { path, isRecorded: /\b(branch|worktree)\b/i.test(value.slice(path.length)) }
209}
210
211export function branchDeny(path: string, branch: string): string {
212  return [
213    `🛑 Blocked: a Watson Direct-mode dispatch onto ${branch}, with no branch recorded in Workdir:.`,
214    '',
215    `Workspace check (${PLUGIN}). The repo at ${path} is on its default branch, ${branch}, and the brief's Workdir: names only the path. ` +
216      'Ask the human which branch the work goes on, through AskUserQuestion, and propose a branch name. ' +
217      `Then record the answer beside the path, for example "Workdir: ${path} (branch: fix/short-name)", or "(branch: ${branch}, <who> chose to work on ${branch})" when the human picks ${branch}, and dispatch again.`,
218  ].join('\n')
219}
220
221export function branchUnreadDeny(path: string): string {
222  return [
223    '🛑 Blocked: a Watson Direct-mode dispatch whose branch could not be read.',
224    '',
225    `Workspace check (${PLUGIN}). git did not answer for ${path}, so the check cannot tell whether the repo is on its default branch. ` +
226      'Ask the human which branch the work goes on, record it beside the path in Workdir:, for example "(branch: fix/short-name)", and dispatch again.',
227  ].join('\n')
228}
229
230// Whether the dispatch gate judges a dispatch: a main-session dispatch while
231// orchestrator mode is on, as hooks/agent-dispatch-gate.sh decides it. A
232// sub-agent's dispatch and a top-level `--agent` run's are exempt.
233export const isGated = (lane: CallerLane, isOn: boolean): boolean => lane === 'main' && isOn
234
235// The refusal for a brief missing slots: the gate's human line, then what it
236// tells the model, from the same slot records the check read.
237export function denyOf(missing: readonly string[], slots: readonly BriefSlot[]): string {
238  const catalogue = slots.map(slot => `${slot.header} (${slot.description})`).join(', ')
239  return [
240    `🛑 Blocked: an Agent dispatch without a complete brief. Missing: ${missing.join(', ')}.`,
241    '',
242    `Dispatch gate (${PLUGIN}). Every Agent dispatch from the main session uses the six-slot brief, research included. ` +
243      `Slots: ${catalogue}. Add the missing slots and dispatch again. ` +
244      'The dev-team specialists and the brief they expect are in /workbench-dev-team:orchestrate. ' +
245      'To dispatch without the brief in this session, the human can run /orchestrator off.',
246  ].join('\n')
247}
248
249// The refusal when the gate judges a dispatch but cannot check its brief:
250// workbench-core's briefCheck rejected or the gate itself failed. Core is a
251// declared dependency, so that is a fault, and the gate fails closed.
252// /orchestrator off stands the gate down, so the refusal never locks the
253// human out.
254export function uncheckedDeny(): string {
255  return [
256    '🛑 Blocked: an Agent dispatch whose brief could not be checked.',
257    '',
258    `Dispatch gate (${PLUGIN}). The gate judges this dispatch, and workbench-core's brief check did not answer, so the gate refuses rather than let an unchecked brief through. ` +
259      'Dispatch again. If the check keeps failing, workbench-core needs repair. ' +
260      'To dispatch without the brief check in this session, the human can run /orchestrator off.',
261  ].join('\n')
262}
263
264// The advisory note on a complete brief that dictates method: a fenced block, a
265// shell command on its own line, or three or more numbered steps. It never
266// blocks. Each pattern is the bash gate's, line by line, over its five ASCII
267// whitespace characters.
268const WS = ' \\t\\r\\v\\f'
269const FENCE = new RegExp(`^[${WS}]*\`\`\``)
270const COMMAND = new RegExp(
271  `^[${WS}]*(\\$[${WS}]+)?(git|gh|npm|npx|yarn|pnpm|composer|php|python3?|pytest|cargo|rustc|bash|zsh|sed|awk|grep|rg|jq|cp|mv|rm|mkdir|chmod|ln|curl|docker)[${WS}]+[^${WS}]`,
272)
273const STEP = new RegExp(`^[${WS}]{0,3}[0-9]+[.)][${WS}]+[^${WS}]`)
274
275export function hintOf(check: BriefCheck, prompt: string): string | undefined {
276  if (check.shape !== 'brief' || !check.isComplete) return undefined
277  const lines = prompt.split('\n')
278  const markers: string[] = []
279  if (lines.some(line => FENCE.test(line))) markers.push('a fenced code block')
280  if (lines.some(line => COMMAND.test(line))) markers.push('a shell command on its own line')
281  const steps = lines.filter(line => STEP.test(line)).length
282  if (steps >= 3) markers.push(`${steps} numbered steps`)
283  if (markers.length === 0) return undefined
284  return (
285    `📐 Dispatch hint (advisory, nothing was blocked): this brief carries ${markers.join(', ')}. ` +
286    'A brief states the outcome and lets the sub-agent pick the method. The sub-agent has the repo in front of it and you do not. ' +
287    'Prefer moving that detail into Done when: as an observable result, or into Constraints: as a hard limit. ' +
288    'Send it as-is if the detail is genuinely a constraint rather than a recipe.'
289  )
290}
291
types/index.d.ts 79 lines
1// The type contract of workbench-dev-team's hooks module: the values it keeps
2// in $.state for the session, each one per agent. `claude plugin validate` holds every $.state key
3// the module names to this file.
4//
5// The module adds no noun to $. It builds on workbench-core's $.workbench, which
6// the engine types on $ because .claude-plugin/plugin.json lists workbench-core
7// under "dependencies".
8
9// ── The runs pane (hooks/mods/runs.ts) ──────────────────────────────────────
10
11export type DevTeamAgent = 'lestrade' | 'holmes' | 'watson'
12
13//   running        the dispatcher's process for the run is alive
14//   done           the run ended with no error line
15//   failed         the run ended on an API, execution or other error line
16//   refused        the API refused the run's output (the content filter)
17//   budget-killed  the run hit its USD budget cap
18//   escalated      the breaker escalated the item after this run, its last
19export type RunState = 'running' | 'done' | 'failed' | 'refused' | 'budget-killed' | 'escalated'
20
21// One run. `target` is "item <id>" or "sweep <owner>-<repo>", `startedAt` the
22// stamp in its log's name, and `refusals` the log's "Permission denied:"
23// lines, one per tool call a rule, hook or prompt refused.
24export type RunRow = { agent: DevTeamAgent; target: string; startedAt: string; state: RunState; refusals: number; log: string }
25
26// What the runs pane draws: the rows, newest first, or why there are none.
27export type RunsView = { rows: RunRow[]; error?: string }
28
29// ── The board pane (hooks/mods/board.ts) ────────────────────────────────────
30
31export type BoardItem = { id: number; number: number | null; isPr: boolean; repo: string | null; title: string | null; claimedAt: string | null }
32
33// A lane: its items, at most `limit` of them, or the list tool's own error.
34export type BoardLane = { limit: number; items: BoardItem[] } | { error: string }
35
36export type Board = { unrefined: BoardLane; review: BoardLane; development: BoardLane }
37
38// What the board pane draws:
39//   attemptedAt  when the last fetch started (0: never); the cadence counts from it
40//   board        the last board fetched, and boardAt, when that fetch started
41//   error        why the last fetch failed, when it did
42//   escalated    the items the breaker escalated, read from the log folder
43export type BoardView = { attemptedAt: number; board?: Board; boardAt?: number; error?: string; escalated: string[] }
44
45// An effort the config may give an agent: a level, as a hook may set a model
46// request's `effort`.
47export type DevTeamEffort = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
48
49declare module 'claude-code' {
50  interface PluginState {
51    'workbench-dev-team': {
52      // The effort each dev-team sub-agent runs at, keyed by its agentId: set
53      // when agent.spawn starts it, read by every turn.step of its loop.
54      effort: StateFamily<DevTeamEffort>
55      // The type each dev-team sub-agent runs as, holmes-lens included, keyed
56      // by its agentId: set when agent.spawn starts it. The scratch folder, the
57      // context budget and the holmes-lens rule read it.
58      agentType: StateFamily<string>
59      // Each dev-team agent's scratch folder, keyed by its agentId, or `run`
60      // for the top-level loop of a `claude -p --agent` run: made at its first
61      // bare mktemp, deleted and set to "" when its run completes.
62      scratch: StateFamily<string>
63      // Every scratch folder the mod made and has not deleted yet. A family
64      // cannot be listed, so session.end reads this to delete what each run's
65      // end left behind for a child that was still live.
66      scratchFolders: string[]
67      // Whether the human was told this run passed its context budget, keyed
68      // as scratch is, so the notice comes once per run.
69      overBudget: StateFamily<boolean>
70      // What the runs pane draws: the newest dispatched runs, read from the
71      // dispatch logs while the pane is open.
72      runs: RunsView
73      // What the board pane draws: The Index's three lanes as last fetched,
74      // when the last fetch started, and the items the breaker escalated.
75      board: BoardView
76    }
77  }
78}
79