brain-seat delegation for Claude Code: tiering, verification, background debrief and eval, dashboard


Brain-seat delegation for Claude Code. The main model (the "brain") hands a task card to a worker subagent with one tool call, dispatch. The mod does the rest. It writes the brief, cuts a git worktree and picks the model tier. It spawns the worker, holding it in a queue while two others run. When the worker reports, the mod checks the report against the repo itself (branch, sha, changed files, scope, gate, PR) and hands the brain one verdict line with what to do next. At a quiet stop it can write a debrief in the background, and it can run your eval when origin/main moves. It is a Claude Code mod (a plugin of function hooks). It needs nothing from any other repo: no scripts, no harness, no network. Version 0.5.0, MIT.
git clone https://github.com/nelsben/chassis-delegation.git ~/chassis-delegation
claude --plugin-dir ~/chassis-delegation. Repeat the flag to load several.plugin-authoring skill; the skill's first paragraph names this session's hot-reload folder, ~/.claude/dev-mods/<session-id>/. Clone the mod into a child of it as a real copy (git clone https://github.com/nelsben/chassis-delegation.git ~/.claude/dev-mods/<session-id>/chassis-delegation; the watcher does not follow a symlink). When the turn ends, Claude Code asks "Enable hot reloading for this session?": answer Enable for this session. The mod loads then, and reloads after each later change to that folder. This is the way in for the VS Code extension, which takes no --plugin-dir flag.CLAUDE_CODE_PLUGIN_DIRS in your shell, or in the env block of ~/.claude/settings.json. It applies to every project and every new process. Point it at a checkout of your own that you update deliberately, never at a session's hot-reload folder, and not while some session also hot-loads the mod, or that session loads it twice./delegation there gets Claude Code's own "no command with that name" answer. Hot reload needs no restart./delegation setup inside Claude Code, in the repo you work on. It checks the repo and the tools the mod needs (see Setup), names the exact fix for each one that fails, and writes nothing until every required check holds. Run the fixes, then run it again. Once they hold it scaffolds the repo and asks you for your first task in a sentence. /delegation init is the bare scaffold, with no checks. Either way it adds four things and prints what it wrote. It never overwrites a file:agents/tasks/README.md, describing the card format;agents/tasks/OPS-000-sample.md (bare /delegation init only; setup does not write it);.chassis-delegation.json, holding every key with its default and a _comment per key (setup also fills in gateMap, and baseRef or cardDir when it found the need);.delegation/ line in .gitignore.If Claude Code says an update is pending, restart the session once before dispatching; init itself needs no re-run.
In a repo that is itself a plugin (it holds .claude-plugin/plugin.json), init puts the cards under docs/cards/ instead and writes "cardDir": "docs/cards" into the config, because the engine reads every agents/*.md of a plugin as a subagent definition. --replay still reads cards from agents/tasks/ at the base commit (the allowlist's git shapes name that folder).
To check the folder on the new machine, run tests/selfcheck.sh. It runs claude plugin validate and claude plugin test, then prints the --plugin-dir line to use.
Upgrade between turns, with no worker running. These steps are written so the agent in the session that loads the mod can follow them.
/delegation. From 0.5.0 on, one line reads mod: chassis-delegation <version> loaded from <folder>. An older version has no such line: the folder is the session's hot-reload clone, ~/.claude/dev-mods/<session-id>/chassis-delegation, or the one named by --plugin-dir or CLAUDE_CODE_PLUGIN_DIRS.git -C <folder> pull --ff-only origin main
If the pull refuses with "unrelated histories", the clone was made before 2026-10-05, when this repository was republished with a fresh history. Delete the folder and clone it again:
git clone https://github.com/nelsben/chassis-delegation.git <folder>
A folder that is a plain copy, not a git clone, is replaced the same way.
--plugin-dir or CLAUDE_CODE_PLUGIN_DIRS folder needs Claude Code restarted./delegation names the new version. Then run /delegation setup in the repo: in a repo already set up it changes no file, checks the repo, and says what it would have set.What the brain will see:
scope: reads as prose stops before the worktree and the spawn; dispatch again with --scope <globs>, or write globs on the card. Before, the worker ran and the brief was refused afterwards.spend=<usd> from spendByTier (economy 2, standard 6, frontier 15 dollars), and the worker is told it. A run that ends past it with no report gets one wrap-up message; at twice it the attempt is over-spend. Set spend: on a card or spendByTier in the repo file to change it; 0 means no ceiling.~ marks the old figure, the session's cost growth, when no usage came back.(1 live: A; 1 queued: B); a second dispatch of a queued task answers already queued since HH:MM; a queued task that is refused keeps its place and the refusal is a row.no-report respawns at the same tier. Finished work found in a worktree gives next=verify sha=…: run /dispatch <ID> --verify <sha>.--base <sha> is written into the brief as base=, and the verifier diffs a stacked task against it.tier: opus is frontier)./delegation setup and the card tool. Say a task in a sentence; Claude writes the card, shows the brief, and dispatches when you say go.repo=here leaves the card folder out of the delta, so nothing needs committing before a dispatch.haiku, which Claude Code 2.1.294 and later resolves to Claude Haiku 5.5 on the Anthropic API; on Bedrock, Vertex and Foundry it still resolves to Haiku 4.5. The mod toasts the change once: haiku now resolves to claude-haiku-5-5 (was …).What to do:
.chassis-delegation.json has to change; a key you never set takes its default.<root>/.delegation/briefs/<ID>.brief.md before its next dispatch to get one.--scope.Clone again (step 2: the history changed), then read CHANGELOG.md from your version up.
A live view of what delegation is doing and spending, in two places:
delegation · 2 live · 1 queued · $4.12 · a sparkline of the last 30 minutes of spend, and a [ details ] button that opens the pane. The whole band is a hover scope: hover it and a card opens beneath the row with the first three blocks below. Turn the band off with the dashboardBand setting in /config (default on); /delegation dashboard still opens the pane./delegation dashboard or the band's [ details ] and never on its own. It draws all four blocks, then "Across sessions" (the six cross-session sparklines over the last 14 sessions) and the open items.The four blocks:
<repo>-<ID>, or under worktreeRoot): task, model (a chip in the model's colour and its name), state (live 04:12, queued #2, verdict owed, or the verdict with its icon), worktree folder and branch, tokens, cost, and attempt over budget.sonnet · $3.10 · 412k tok · 4 verified.What is live and what is not. The spend line samples the session's cost every 15 seconds while a worker is live or queued and every 60 seconds otherwise (the last 240 points are kept, and survive a reload), and the live and queued counts and the live mm:ss clock follow the engine's agent list. A worker the mod spawns itself does not run the mod's per-step hooks (public issue #22), so a worker's tokens and cost, and the by-model bars, update when its run ends, not during it. The tokens column is the total the run reported; the dollar figure is the worker's own cost.
A screen that shows no mod panes (the VS Code extension today) answers /delegation dashboard with the same blocks as markdown text instead: the headline, the spend over the session, the worktree table and spend by model.
| Name | Who calls it | What it does | ||||||
|---|---|---|---|---|---|---|---|---|
/delegation | you type it | shows the delegation state and where the config came from | ||||||
/delegation setup | you type it | checks the repo and tools, names the fixes, then scaffolds and asks for your first task (see Setup) | ||||||
/delegation init | you type it | the bare scaffold, no checks (step 4 above) | ||||||
/delegation dashboard | you type it | opens the live dashboard pane (see Dashboard); nothing opens it unasked | ||||||
| `/dispatch <ID> [--dry-run\ | --scope\ | --forbid\ | --replay\ | --base\ | --here\ | --force-overlap]` | you type it | dispatches a card; the model can also run it through the tool. --here shares the session's own checkout (see repo=here) |
/dispatch <ID> --verify <sha> | you type it, or the brain after a work present row | spawns nothing: runs the verifier on the work already in the task's worktree at that sha (see Look before you respawn under How a report is verified) | ||||||
mcp__chassis-delegation__dispatch | the model, on its own | the same dispatch, as a tool | ||||||
mcp__chassis-delegation__card | the model, on its own | you say a task in a sentence; the model looks at the repo, calls this with the title, why, done-when, scope globs and red test; it writes the card, runs the dry run and returns the one-line summary and the brief header. Dispatch when you say go (the dispatch tool, or card again with dispatch: true) | ||||||
mcp__chassis-delegation__init | the model, on its own | the same scaffold as /delegation init, as a tool | ||||||
mcp__chassis-delegation__setup | the model, on its own | the same checks and scaffold as /delegation setup, as a tool (no input) |
A user skill or command named delegation or dispatch under ~/.claude/skills or ~/.claude/commands shadows the mod's commands; remove it.
/delegation setup (or the setup tool) checks the repo and prints one line per check, ✓ or ✗ for the required ones and · for advice, with the fix on the line under a failing one. The mod cannot run the fixes itself (its host-command allowlist has no git init, commit, remote or install), so the brain or you run them, then run setup again. It is idempotent.
Required:
git init -b main);package.json scripts.test (not npm's placeholder; pnpm test or yarn test by lockfile), pytest, cargo test or go test ./...;git, and the stack's runner (node and its package manager, python3 and pytest, cargo, go);~/.claude/skills/{delegation,dispatch} or ~/.claude/commands/{delegation,dispatch}.md shadowing the mod.Advice (never blocks): the mode (origin/main resolves: worktree mode; no remote: repo=here, setup writes "baseRef" and cards dispatch with --here); a lockfile for the brief's install step; a nested, gitignored child repo (start Claude Code inside it to delegate there); a plugin repo (cards under docs/cards/); gh on PATH (a report's pr= claim is checked only with it); uncommitted changes in worktree mode (a worker's worktree lacks them); the background debrief spending an agent at a quiet stop.
When every required check holds, setup runs the init scaffold (the card folder's README, .chassis-delegation.json, the .gitignore line; no sample card), writes .chassis-delegation.json with gateMap: {"test": "<detected>"} (an existing config is left as it is and setup prints what it would have set), prints one config: line per key it set in a fresh config (gateMap.test, and baseRef or cardDir when set), and ends by asking for the first task:
Set up. Tell Claude your first task in a sentence, for example: "add a function that reads a file header and returns its size, with a unittest". Claude writes the card, shows you the brief, and dispatches when you say go.
The example follows the detected gate. Through the setup tool the result adds Ask the person for the first task, then call the card tool. Setup prints init's file lines but not init's own "Next:" line. Nothing needs committing before a dispatch: a worktree dispatch reads the card from the main checkout, and in repo=here mode the card folder is in the always-applied ignore= set. A verified verdict means the report matches git, not that the work is right: read the diff. /delegation in a root with no config adds not set up here: run /delegation setup.
/delegation setup once. Then tell Claude what you want in a sentence ("add a function that reads a file header and returns its size, with a unittest"). Claude looks at the repo, calls the card tool, and shows you the card's path, a one-line summary and the brief header:wrote …/agents/tasks/OPS-1-add-a-function-that-reads.md OPS-1 · standard → sonnet · scope game_decompiler/, tests/ · gate test · red: python3 -m unittest tests.test_rom · 2 attempts · $6 ceiling
[[brief v=1 task=OPS-1 subtask=main purpose=build tier=standard model=sonnet …]]
Say go and Claude dispatches OPS-1.
dispatch tool (or the card tool again with dispatch: true); you can also type /dispatch OPS-1 yourself. The tool reports what it did, one line per step:chassis-delegation: OPS-1 attempt 1/2 verified · sonnet · $0.41 · next=accept
A failed check names its first failed claim, and next= tells the brain what to do:
chassis-delegation: OPS-1 attempt 1/2 refuted on files (files= does not match the actual delta …) · sonnet · $0.38 · next=resume agent=agent-7
/delegation prints what is running, pending, queued and owed, plus the last verdicts and where the config came from.The brain does the typing. You see:
tier=standard → sonnet (brief) · attempt 1/2. A briefed spawn's notice ends with its attempt of the budget, so a retry's spend is never a surprise. The tier comes from the brief header's tier=, else the caller's model, else the classifier. The classifier is called only when there is no header and no caller model, and the debug log names the source: T-7: tier=economy picked by the brief header's tier= (no classify call);<n> workers · $<usd> · ctx <pct>%;debrief running in the background, T1 63/63, started queued OPS-3 (waited 4 min);A compaction keeps the delegation loop's position. While anything runs or is owed, the system prompt carries a short "Delegation state" section, at most 40 lines.
Each card is one file, agents/tasks/<ID>-<slug>.md: YAML frontmatter, then the spec in markdown. The card tool writes cards from a sentence (the next free id for the domain's prefix, a slug from the title, scope given as globs, a gate id that exists in gateMap, a red test or none), and refuses with the fix when a field is wrong, writing nothing. You can still write one by hand. <ID> is <PREFIX>-<number>[letter], for example OPS-12, BE-101 or FE-7b.
id: OPS-12 title: One line that says what done looks like domain: ops # one of domains; the branch is agent/<domain>/<id> tier: standard # economy | standard | frontier | premium, or a model name (haiku | sonnet | opus | fable) as its tier; anything else runs at standard with a warning status: queued # /dispatch takes queued or claimed; template, merged … are refused scope: [src/feature/, docs/feature.md] forbid: [src/secrets/] red_test: npm test -- feature.test.ts gate: test # ids in gateMap, comma-separated budget: 2-attempts # spawns + resumes before it comes back to you spend: 4 # optional: dollars one attempt may spend (else spendByTier; 0 = no ceiling) repo: here # optional: no worktree, the worker shares this checkout (see repo=here)
## Why … ## Done when …
Glob rules for scope and forbid:
* crosses folders;** is the same as *;? is one character;/ means everything under the folder.A card whose scope is prose is caught at dispatch (GH-103). /dispatch and the tool write the brief, print its header and the prose scope, and stop before the worktree and the spawn:
stopped: the card scope is prose; pass --scope <globs> (or scope on the tool) and dispatch again
Dispatch again with --scope <globs> (the tool's scope, which must be globs and replaces the card's scope in the header); the brief written the first time is reused. --dry-run behaves as before.
A budget that is not <n>-attempts, such as a chassis frontier-60m, falls back to defaultBudget, and /dispatch says so once: warning: budget "frontier-60m" is not <n>-attempts; using the default 3. A spawn whose header carries such a budget logs the same line to debug. The frontier part is never taken as the tier: the card's tier stays the tier.
Brief. /dispatch writes <root>/.delegation/briefs/<ID>.brief.md. It starts with one header line:
[[brief v=1 task=<ID> subtask=main purpose=build tier=<tier> model=<alias> scope=<globs> forbid=<globs> red_test="<cmd>" gate=<ids> spend=<usd> budget=<n>-attempts report=chassis.report.v1]]
spend= is the per-attempt ceiling in dollars (GH-106): the card's spend: when it has one, else the tier's entry in spendByTier; 0 writes none. The body's Rules carry the line "Spend: about $<spend> for this attempt. Do what the card asks and no more; when you are near it, stop and hand back what you have with the report line." See Spend ceiling below.
The body comes from hooks/templates/brief.md. It tells the worker to:
| Lockfile | Install step |
|---|---|
package-lock.json | npm ci |
pnpm-lock.yaml | pnpm i --frozen-lockfile |
yarn.lock | yarn install --immutable |
requirements.txt | pip install -r requirements.txt |
.delegation/<ID>/red-<attempt>.txt in the worktree (a new file for each attempt);The repo file's briefExtra adds lines for the repo. An existing brief is reused, never overwritten (and it decides the mode: a reused repo=here brief dispatches with no worktree, whatever the flags say). Two more fields are optional: scope_globs= and forbid_globs= replace a prose scope or forbid. A brief whose scope= is prose and that has no scope_globs= is not refused: the verifier marks the scope claim unchecked (add scope_globs= to check it), checks every other claim, and the verdict is unverified at worst. Only a brief with no scope= at all is refused. Other header fields are optional too:
repo=none marks a task with no git repo, and repo=<dir> names the folder to verify;repo=here marks a task worked in the session's own checkout, no worktree (see repo=here);base=<ref> names where the delta starts when the worker commits (GH-10): a branch, HEAD~2, a sha. It beats baseRef and the origin/main chain. A value that is not a git ref (one starting with -, or a a..b range) makes the brief refused. /dispatch <ID> --base <sha> (and the tool's base) writes it whenever the base is not the default origin/main, in worktree mode and repo=here alike (GH-105): a card stacked on an unpushed sibling is cut from that sha and judged on what the worker added to it, not on the sibling's files. A reused brief keeps its own header; when it has no base= the dispatch says so (note: the reused brief has no base=; …) rather than rewriting it. --verify <sha> takes its delta from the same base=. The verdict's scope line names the base it diffed from (every changed path since <base> is within …);ignore=<globs> (repo=here only) names paths taken off the delta before scope and files are checked.A
hooks/register.ts 3193 lines1// chassis-delegation — brain-seat delegation for Claude Code: part 1 (SPEC v1 +
2// amendment 1), part 2 (the dispatch tool, quiet verdicts, the clean-stop
3// debrief, the eval trigger, compaction state, the scheduler) and part 5
4// (standalone: native verification, a built-in tier map, a per-repo config
5// file, `/delegation init`, portable briefs, a native git guard, a built-in
6// debrief). The hooks stay thin: every decision is a pure function in ./lib,
7// and every host command passes ./lib/allow.ts first. Nothing here calls the
8// chassis scripts: the mod works in a repo that has only agents/tasks/ cards.
9import type { AgentSpawnResult, EngineInterface, PluginOptions, Register, TurnUsage } from 'claude-code'
10
11import type { BandItem, DelegationVerdict, DelegationWorker, QueuedSpawn } from './types'
12import { blocksFor, openItemLines } from './lib/dashboard'
13import { bandTree } from './lib/band'
14import { paneTree, type PaneTable } from './lib/pane'
15import { dashboardText, isActive, liveView, pickScheme, sampleEvery, sessionRecords, spendSeries, type LiveView, type SpendPoint } from './lib/live'
16import { metricsFromRecords } from './lib/metrics'
17import { checkArgv, refusedLine, type AllowConfig } from './lib/allow'
18import {
19 attemptsFor,
20 escalationSource,
21 holdSource,
22 judgedShas,
23 laneFields,
24 lineageResumes,
25 nextAttempt,
26 patchRecord,
27 priorRedHashes,
28 taskLabel,
29 type AttemptRecord,
30 type Lane,
31} from './lib/attempts'
32import {
33 amendNeedsApproval,
34 amendScopeInsideForbid,
35 amendInsideForbidLine,
36 forbidCovering,
37 scopeInsideForbidWarning,
38 scopeInsideForbid,
39 appendAmends,
40 budgetWarning,
41 effectiveList,
42 parseAmends,
43 type Amend,
44 extractAmendBlocks,
45 extractReport,
46 findBriefPath,
47 inlineHeaderMissing,
48 lacksLine,
49 noBriefLine,
50 parseBudget,
51 parseHeader,
52 parseReport,
53 scopeOverlap,
54 spendOf,
55 type BriefHeader,
56} from './lib/brief'
57import { addTurn, ceilingState, liveWorker, overSpendLine, round4, spendCeiling, warnText, OVER_SPEND_NEXT, type Spend } from './lib/cost'
58import {
59 addFriction,
60 breadcrumbPath,
61 builtInDebriefPrompt,
62 cleanStop,
63 countLines,
64 debriefPathOf,
65 debriefPrompt,
66 debriefSkillPath,
67 debriefSource,
68 debriefToast,
69 DEFAULT_DEBRIEF_AGENT,
70 DEFAULT_DEBRIEF_COOLDOWN_HOURS,
71 DEFAULT_DEBRIEF_IDLE_MINUTES,
72 DEFAULT_DEBRIEF_MIN_EVENTS,
73 DEFAULT_DEBRIEF_MIN_NEW_LINES,
74 frictionFacts,
75 isCorrection,
76 parseWatermark,
77 watermarkPath,
78 type FrictionEvent,
79} from './lib/cleanstop'
80import { appendInstructions, compactBlock, composeSection, emptySnapshot, isEmptyState, owedFrom, prsFrom, renderState, SECTION_ID, type RecentVerdict, type StateSnapshot } from './lib/compaction'
81import {
82 agentTypeFor,
83 branchName,
84 briefDirFor,
85 briefFileName,
86 budgetAttempts,
87 cardMatches,
88 cardNamesFromLsTree,
89 checkDomain,
90 checkStatus,
91 currentBranchArgv,
92 defaultBriefDir,
93 DISPATCH_TOOL,
94 fetchArgv,
95 headShaArgv,
96 hereIgnore,
97 inlineBriefFileName,
98 inlineBriefText,
99 installStep,
100 lsTreeArgv,
101 cardDirPath,
102 overlapRefusal,
103 overlapWarning,
104 parseCard,
105 parseDispatchArgs,
106 parseDispatchTool,
107 renderBrief,
108 renderHeader,
109 scopeLooksProse,
110 PROSE_SCOPE_STOP,
111 scratchpadFor,
112 showCardArgv,
113 spawnDescription,
114 spawnPrompt,
115 worktreeAddArgv,
116 worktreePath,
117 type Card,
118 type DispatchArgs,
119} from './lib/dispatch'
120import {
121 DEFAULT_EVAL_IDLE_MINUTES,
122 DEFAULT_EVAL_LIVE_MAX_USD,
123 DEFAULT_SESSION_USD_CAP,
124 evalFailureRow,
125 evalRunnerPrompt,
126 evalToast,
127 extractEvalBlock,
128 failingLines,
129 liveEligible,
130 parseEvalBlock,
131 parseSha,
132 revParseArgv,
133 shouldEval,
134 type EvalTier,
135} from './lib/evaltrigger'
136import { gateTemplatesOf, resolveGateRuns } from './lib/gates'
137import { gitWrites, guardDeny, joinDir, parseGuardBranches } from './lib/gitguard'
138import { handbackMessages, workerSaid } from './lib/handback'
139import { INIT_FILES, PLUGIN_MANIFEST, initPlan, initText } from './lib/init'
140import { CARD_TOOL, cardFromFields, cardSummary, sayGo } from './lib/card'
141import { PATH_TOOLS, ROOT_MARKERS, SHADOWS, allRequiredHold, configLines, detectGate, foundOf, handoverText, BRAIN_HANDOVER, scaffoldConfig, setupChecks, setupText, wouldSet, type SetupProbe } from './lib/setup'
142import { globList, globRoot, isNotWorkTree, noRepoVerdict, reportFiles, scopeCheck, workTreeArgv } from './lib/norepo'
143import { deliveryFor, parseVerbosity, quietLine, shortNext, type Rendered, type Verbosity } from './lib/quiet'
144import { ignoreWithCards, isGitRef, mergeConfig, parseRepoConfig, REPO_CONFIG_FILE, settingsLayer, type RepoConfig } from './lib/repoconfig'
145import { alreadyQueuedDeny, alreadyQueuedPart, drainRefusalRow, enqueue, hasSlot, isFinalDeny, isQueuedDeny, promptKey, queuedDeny, queuedIndex, queuedText, startingOthers, waitedMinutes } from './lib/scheduler'
146import {
147 CLASSIFIER_LABELS,
148 classifierText,
149 fableRequested,
150 finalAlias,
151 tierOf,
152 tierWarning,
153 needsClassifier,
154 noticeText,
155 pickTier,
156 tierPickLine,
157 type Alias,
158 type Tier,
159 type TierPick,
160} from './lib/tier'
161import {
162 advise,
163 amendApprovalLine,
164 amendedLine,
165 amendMalformedLine,
166 amendPendingLine,
167 branchDelta,
168 budgetDenyMessage,
169 contextBlock,
170 filesPathFor,
171 isFailing,
172 isProveResume,
173 isWorkPresentDeny,
174 newestRed,
175 probeWork,
176 resumeText,
177 scratchFor,
178 syntheticReport,
179 verdictLine,
180 verifyAdvice,
181 workPresentDeny,
182 workPresentLine,
183 type Advice,
184 type Verdict,
185 type WorkPresent,
186} from './lib/verify'
187import { briefContract, briefWantsRed, verifyCardless, verifyNative, type RedEvidence } from './lib/verify-native'
188import { driftMessage, newAgentTypes, parseCandidateIds, shouldClearStatus, statusText } from './lib/watch'
189import { MOD_VERSION } from './lib/version'
190
191type Host = EngineInterface
192
193const PLUGIN = 'chassis-delegation'
194const WORKERS = { plugin: 'chassis-delegation', key: 'workers' } as const
195const STATUS = { plugin: 'chassis-delegation', key: 'status' } as const
196const LAST_VERDICT = { plugin: 'chassis-delegation', key: 'lastVerdict' } as const
197const QUEUE = { plugin: 'chassis-delegation', key: 'queue' } as const
198const SPEND = { plugin: 'chassis-delegation', key: 'spend' } as const
199/** GH-112: the pane `/delegation dashboard` and the band's `[ details ]` open. */
200const DASH_PANE = 'delegation-dash'
201
202const K = {
203 tasks: (task: string) => `delegation.tasks.${task}`,
204 adhoc: 'delegation.adhoc',
205 spawn: (key: string) => `delegation.spawn.${key}`,
206 agent: (agentId: string) => `delegation.agent.${agentId}`,
207 name: (name: string) => `delegation.agentName.${name}`,
208 alias: (alias: string) => `delegation.alias.${alias}`,
209 agentTypes: 'delegation.agentTypes',
210 answered: 'delegation.models.answered',
211 replay: (briefPath: string) => `delegation.replay.${briefPath}`,
212 /** A worker's own spend this attempt, from its turn usage (GH-106). */
213 cost: (agentId: string) => `delegation.cost.${agentId}`,
214 /** Set once an attempt's verdict row is posted to the conversation. */
215 posted: (attemptKey: string) => `delegation.posted.${attemptKey}`,
216 /** This session's verdict lines, for the compaction block and the system prompt section (2E). */
217 recent: (sessionId: string) => `delegation.recent.${sessionId}`,
218 /** The background debrief of a session (2C). */
219 debrief: (sessionId: string) => `delegation.debrief.${sessionId}`,
220 /** The friction the mod saw in a session (5E): corrections, denials, refutes. */
221 friction: (sessionId: string) => `delegation.friction.${sessionId}`,
222 /** origin/main when the last T1 eval runner started (2D): never twice per sha. */
223 lastSha: 'delegation.eval.lastSha',
224 /** The eval runner that is out. */
225 evalInflight: 'delegation.eval.inflight',
226 /** The T2 run of a session: once a session. */
227 live: (sessionId: string) => `delegation.eval.live.${sessionId}`,
228 /** Every eval block the runners reported. */
229 evals: 'delegation.evals',
230}
231
232const PROBE_EVERY_MS = 24 * 60 * 60 * 1000
233const STATUS_EVERY_MS = 60 * 1000
234const PROMPT_CAP = 20000
235const MIN_MS = 60 * 1000
236const HOUR_MS = 60 * MIN_MS
237const RECENT_CAP = 50
238/** An eval runner out longer than this is taken as gone (a T1 run takes minutes). */
239const EVAL_STALE_MS = 3 * HOUR_MS
240const EVALS_CAP = 100
241const GATE_TIMEOUT_MS = 9 * 60 * 1000
242/** The advice kinds that leave something for the brain (or the person) to do. */
243const OWED_KINDS = ['resume', 'respawn', 'verify', 'exhausted', 'check', 'fix-brief']
244const LEDGER_FILE = '.delegation/ledger.jsonl'
245
246/** `delegation.debrief.<session>`: the background debrief (2C, 5E). */
247type DebriefRecord = { lastAt: number; watermark: number; lines: number; sessionId: string; mode?: 'breadcrumbs' | 'events'; builtIn?: boolean; agentId?: string; finishedAt?: number; path?: string; denied?: string }
248/** `delegation.friction.<session>`: every friction event's count, and the last 100. */
249type FrictionRecord = { total: number; events: FrictionEvent[] }
250/** `delegation.eval.inflight`: the eval runner that is out (2D). */
251type EvalInflight = { agentId?: string; tier: EvalTier; sha: string; at: number; sessionId: string }
252/** One `delegation.evals` row. */
253type EvalEntry = { tier: EvalTier; sha: string; total?: number; pass?: number; fail?: number; result?: string; failing: string[]; at: number; agentId: string; sessionId: string }
254
255/**
256 * GH-16: what a repo=here attempt record carries besides the attempt: `here`,
257 * the checkout it shares (the session root), and `files`, the files= its
258 * hand-back claimed (written before the verify, so a sibling's verify sees them).
259 */
260type HereFields = { here?: string; files?: string[] }
261/**
262 * GH-1: `requestedAlias: 'fable'` when fable was asked for and opus spawned
263 * (item 8); `adhoc: true` on the attempt a cardless hand-back recorded under
264 * its report's task= (item 2).
265 */
266type IssueOneFields = { requestedAlias?: Alias; adhoc?: true }
267type HereRecord = AttemptRecord & HereFields & IssueOneFields
268/** GH-16: another repo=here card in flight in the same checkout: a record without a verdict. */
269type InFlight = { label: string; scope: string[]; files: string[] }
270
271/** What the mod keeps per spawned worker under `delegation.spawn.<key>` (key = the spawn's tool_use_id). */
272type SpawnRecord = {
273 key: string
274 task: string
275 subtask: string
276 adhoc: boolean
277 /** The worker's current attempt (a resume moves it on). */
278 attempt: number
279 lineage: number
280 tier: Tier
281 alias: string
282 budget: number
283 purpose: string
284 briefPath?: string
285 /** Why a spawn with no brief file has none (GH-6): what its inline header lacks, or why its brief could not be written. */
286 noBrief?: string
287 prompt: string
288 description: string
289 subagentType: string
290 cwd?: string
291 agentId?: string
292 /** The model id the alias resolved to at spawn; a resume keeps it. */
293 resolvedModel?: string
294 usdAtStart?: number
295 /** GH-106: the brief's `spend=` ceiling in dollars (absent: none). */
296 spend?: number
297 verdictAttempt?: number
298 verdictBlock?: string
299 lastFailed?: boolean
300 /** A `/dispatch --replay` run, from `base`. */
301 replay?: boolean
302 base?: string
303 /** ms since the epoch at spawn. */
304 at?: number
305 /** The one-line row of the last verdict (2B). */
306 verdictLine?: string
307 /** GH-16: a repo=here worker: the checkout it shares (the session root). */
308 here?: string
309}
310
311const laneOf = (s: { replay?: boolean; base?: string }): Lane | undefined =>
312 s.replay ? { replay: true, ...(s.base ? { base: s.base } : {}) } : undefined
313
314/** `<task>:<lineage>:<attempt>` (the subtask and a replay lane folded into the task) — one verdict per key. */
315const attemptKey = (s: SpawnRecord): string =>
316 `${taskLabel(s.task, s.subtask)}${s.replay ? `@replay${s.base ? `-${s.base}` : ''}` : ''}:${s.lineage}:${s.attempt}`
317
318type Config = {
319 autoEscalate: boolean
320 applyAmends: boolean
321 defaultBudget: number
322 probeModels: boolean
323 candidateIds: string[]
324 briefDir: string
325 verbosity: Verbosity
326 autoDebrief: boolean
327 debriefIdleMs: number
328 debriefMinNewLines: number
329 debriefMinEvents: number
330 debriefCooldownMs: number
331 debriefAgent: string
332 /** On only with an evalCommand configured (5E). */
333 autoEval: boolean
334 evalLive: boolean
335 evalLiveMaxUsd: number
336 sessionUsdCap: number
337 evalIdleMs: number
338 evalAgent: string
339 ledgerFile: boolean
340 gitGuard: boolean
341 /** GH-112: the band above the prompt. */
342 dashboardBand: boolean
343 guardBranches: string[]
344 // defaults < .chassis-delegation.json < settings (5B)
345 gateMap: Record<string, string>
346 gateTemplates: string[][]
347 agentTypes: Record<string, string>
348 tierMap: Record<Tier, Alias>
349 evalCommand: string
350 evalLiveCommand: string
351 briefTemplate: string
352 briefExtra: string
353 maxWorkers: number
354 domains: string[]
355 worktreeRoot: string
356 /** GH-12: the card folder, relative to the root. */
357 cardDir: string
358 /** GH-16: the delta's base when a brief names none ('' = origin/main → main → origin/master → master). */
359 baseRef: string
360 /** GH-16: globs always subtracted from a repo=here delta. */
361 ignore: string[]
362 /** GH-106: dollars one attempt may spend, per tier; 0 = no ceiling. */
363 spendByTier: Record<'economy' | 'standard' | 'frontier', number>
364}
365
366function readConfig(o: PluginOptions, repo: RepoConfig): { cfg: Config; errors: string[] } {
367 const str = (k: string) => (typeof o[k] === 'string' ? (o[k] as string) : '')
368 const budget = typeof o.defaultBudget === 'number' && o.defaultBudget >= 1 ? Math.floor(o.defaultBudget) : 3
369 const num = (k: string, fallback: number): number => {
370 const v = o[k]
371 return typeof v === 'number' && Number.isFinite(v) && v >= 0 ? v : fallback
372 }
373 const settings = settingsLayer(o)
374 const eff = mergeConfig(repo, settings.config)
375 return {
376 errors: settings.errors,
377 cfg: {
378 autoEscalate: o.autoEscalate === true,
379 applyAmends: o.applyAmends !== false,
380 defaultBudget: budget,
381 probeModels: o.probeModels === true,
382 candidateIds: parseCandidateIds(str('candidateIds')),
383 briefDir: str('briefDir'),
384 verbosity: parseVerbosity(o.verdictVerbosity),
385 autoDebrief: o.autoDebrief !== false,
386 debriefIdleMs: num('debriefIdleMinutes', DEFAULT_DEBRIEF_IDLE_MINUTES) * MIN_MS,
387 debriefMinNewLines: Math.floor(num('debriefMinNewLines', DEFAULT_DEBRIEF_MIN_NEW_LINES)),
388 debriefMinEvents: Math.max(1, Math.floor(num('debriefMinEvents', DEFAULT_DEBRIEF_MIN_EVENTS))),
389 debriefCooldownMs: num('debriefCooldownHours', DEFAULT_DEBRIEF_COOLDOWN_HOURS) * HOUR_MS,
390 debriefAgent: str('debriefAgent') || DEFAULT_DEBRIEF_AGENT,
391 autoEval: eff.autoEval && eff.evalCommand.trim() !== '',
392 evalLive: o.evalLive === true,
393 evalLiveMaxUsd: num('evalLiveMaxUsd', DEFAULT_EVAL_LIVE_MAX_USD),
394 sessionUsdCap: num('sessionUsdCap', DEFAULT_SESSION_USD_CAP),
395 evalIdleMs: num('evalIdleMinutes', DEFAULT_EVAL_IDLE_MINUTES) * MIN_MS,
396 evalAgent: str('evalAgent') || DEFAULT_DEBRIEF_AGENT,
397 ledgerFile: o.ledgerFile !== false,
398 gitGuard: o.gitGuard !== false,
399 dashboardBand: o.dashboardBand !== false,
400 guardBranches: parseGuardBranches(str('guardBranches')),
401 gateMap: eff.gateMap,
402 gateTemplates: gateTemplatesOf(eff.gateMap),
403 agentTypes: eff.agentTypes,
404 tierMap: eff.tierMap,
405 evalCommand: eff.evalCommand,
406 evalLiveCommand: eff.evalLiveCommand,
407 briefTemplate: eff.briefTemplate,
408 briefExtra: eff.briefExtra,
409 maxWorkers: eff.maxWorkers,
410 domains: eff.domains,
411 worktreeRoot: eff.worktreeRoot,
412 cardDir: eff.cardDir,
413 baseRef: eff.baseRef,
414 ignore: eff.ignore,
415 spendByTier: eff.spendByTier,
416 },
417 }
418}
419
420// Module memory: a reload starts it over; the store and $.state stay.
421let options: PluginOptions = {}
422let repoLayer: RepoConfig = {}
423/** The repo file's text as last read: undefined = not read yet, null = no file. */
424let repoText: string | null | undefined
425let cfg: Config = readConfig({}, {}).cfg
426const waiting = new Set<string>() // Agent tool calls still awaiting their result
427const handbacks = new Map<string, string>() // agentId → a SubagentHandback message a tool.call hook saw (the transcript is the live source: GH-2)
428const finalizing = new Map<string, Promise<Rendered>>() // attempt key → its verify, so a second fire waits on the first
429const posted = new Set<string>() // attempt keys whose verdict row went to the conversation
430const offered = new Set<string>()
431let seedingTypes = false
432let delegated = false
433/** GH-106: the turn id a subagent's `turn.start` carried (the engine gives none today), by agentId. */
434const turnOf = new Map<string, string>()
435let lastUsd: number | undefined
436let lastCostChangeAt = 0
437let statusTimer: { cancel: () => void } | undefined
438let sampleTimer: { cancel: () => void } | undefined
439/** GH-112: `git worktree list --porcelain`, kept for 15 s so a redraw does not run git. */
440let worktreeCache: { at: number; text: string } | undefined
441/** GH-112: every task's attempt records, kept 15 s so a redraw does not read the whole store; the mod's own record writes clear it at once. */
442let recordsCache: { at: number; records: AttemptRecord[] } | undefined
443const RECORDS_TTL_MS = 15_000
444let probeTimer: { cancel: () => void } | undefined
445let spawnSeq = 0
446// Part 2: the main loop's turn state, the idle timers, the scheduler's slots.
447// A (re)load does not know whether a main turn is running, so it assumes one is
448// until the next main-loop turn.complete: nothing runs in the background before.
449let inTurn = true
450let idleTimer: { cancel: () => void } | undefined
451let evalTimer: { cancel: () => void } | undefined
452let debriefBusy = false
453let evalBusy = false
454let slotSeq = 0
455const starting = new Map<number, string>() // slot token → the task label of a spawn holding a slot before $.agent.list shows it
456const drainReserved = new Map<string, number>() // promptKey → the slot token the drain reserved for that spawn
457/** GH-107: the last refusal posted for a queued task that keeps its place; the same refusal is not posted again on every drain. */
458const drainRefusalSeen = new Map<string, string>()
459// promptKey → what to record once the spawn hook sees the agent id of a debrief or eval runner the mod started
460const runnerStarted = new Map<string, (agentId: string) => Promise<void>>()
461const selfDecided = new Set<string>() // promptKeys of spawns spawnSelf has decided: the hook passes them through
462const selfIds = new Map<string, string>() // promptKey → the agent id the hook saw for a spawnSelf spawn
463let lockTail: Promise<unknown> = Promise.resolve()
464let ledgerTail: Promise<unknown> = Promise.resolve()
465let frictionTail: Promise<unknown> = Promise.resolve()
466
467const allowCfg = (): AllowConfig => ({ gateTemplates: cfg.gateTemplates, domains: cfg.domains, ...(cfg.worktreeRoot ? { worktreeRoot: cfg.worktreeRoot } : {}) })
468
469// ---- small host helpers (fail-open) ---------------------------------------
470function debug($: Host, text: string) {
471 try {
472 $.ui.log(`${PLUGIN}: ${text}`, { to: 'debug' })
473 } catch {
474 // nothing to do
475 }
476}
477async function storeGet<T>($: Host, key: string): Promise<T | undefined> {
478 try {
479 return (await $.store.get(key)) as T | undefined
480 } catch {
481 return undefined
482 }
483}
484async function storeSet($: Host, key: string, value: unknown): Promise<void> {
485 // a task's attempt records changed: the dashboard's copy is stale (GH-112)
486 if (key.startsWith(K.tasks(''))) recordsCache = undefined
487 try {
488 await $.store.set(key, value)
489 } catch (err) {
490 debug($, `store.set ${key} failed: ${String(err)}`)
491 }
492}
493async function exists($: Host, path: string): Promise<boolean> {
494 try {
495 return await $.fs.exists(path)
496 } catch {
497 return false
498 }
499}
500async function now($: Host): Promise<number> {
501 try {
502 return await $.clock.now()
503 } catch {
504 return Date.now()
505 }
506}
507async function sessionUsd($: Host): Promise<number | undefined> {
508 try {
509 return (await $.session.usage()).cost?.usd
510 } catch {
511 return undefined
512 }
513}
514
515type RunOut = { ok: true; exitCode: number; stdout: string; stderr: string } | { ok: false; why: string }
516
517/** The ONLY way the mod runs a host command: a refused argv never reaches `$.process.run`. */
518async function run($: Host, argv: string[], init?: { cwd?: string; env?: Record<string, string>; stdin?: string; timeoutMs?: number }): Promise<RunOut> {
519 const check = checkArgv(argv, allowCfg())
520 if (!check.ok) {
521 try {
522 $.ui.log(refusedLine(argv, check.reason))
523 } catch {
524 // the refusal stands either way
525 }
526 return { ok: false, why: `refused: ${check.reason}` }
527 }
528 try {
529 const r = await $.process.run(argv, init)
530 return { ok: true, exitCode: r.exitCode, stdout: r.stdout, stderr: r.stderr }
531 } catch (err) {
532 return { ok: false, why: String(err) }
533 }
534}
535
536// ---- part 5B: the repo file ---------------------------------------------------
537/** Reads `<root>/.chassis-delegation.json` (when it changed) and remakes the config: defaults < repo file < settings. */
538async function loadRepoConfig($: Host): Promise<void> {
539 let root: string
540 try {
541 root = await $.session.root()
542 } catch {
543 return
544 }
545 const text = (await readText($, `${root}/${REPO_CONFIG_FILE}`)) ?? null
546 if (text === repoText) return
547 repoText = text
548 const parsed = text === null ? { config: {}, errors: [] } : parseRepoConfig(text)
549 repoLayer = parsed.config
550 const next = readConfig(options, repoLayer)
551 cfg = next.cfg
552 const errors = [...parsed.errors, ...next.errors]
553 for (const err of errors) debug($, `${REPO_CONFIG_FILE}: ${err}`)
554 if (errors.length > 0) $.ui.toast(`${PLUGIN}: ${REPO_CONFIG_FILE}: ${errors[0]}${errors.length > 1 ? ` (+${errors.length - 1} more in the debug log)` : ''}`)
555}
556
557// ---- $.state for the band/pane (part 4 reads it) ---------------------------
558async function setWorkers($: Host, fn: (list: DelegationWorker[]) => DelegationWorker[]) {
559 try {
560 const { value } = await $.state.get(WORKERS)
561 await $.state.set(WORKERS, fn(value ?? []).slice(-50))
562 } catch {
563 // state is a view; the store is the record
564 }
565 redraw($)
566}
567async function setLastVerdict($: Host, v: DelegationVerdict) {
568 try {
569 await $.state.set(LAST_VERDICT, v)
570 } catch {
571 // as above
572 }
573 redraw($)
574}
575
576// ---- GH-112: the live dashboard (the band above the prompt and the pane) -----------
577function redraw($: Host) {
578 try {
579 $.ui.invalidate('ui.render')
580 } catch {
581 // a view only
582 }
583}
584
585async function allRecords($: Host): Promise<AttemptRecord[]> {
586 // the store is shared with every session on the machine, so a write elsewhere shows within RECORDS_TTL_MS
587 const t = await now($)
588 if (recordsCache && t - recordsCache.at < RECORDS_TTL_MS) return recordsCache.records
589 let keys: string[]
590 try {
591 keys = await $.store.keys()
592 } catch {
593 return []
594 }
595 const out: AttemptRecord[] = []
596 for (const key of keys) {
597 if (!key.startsWith(K.tasks(''))) continue
598 const list = await storeGet<AttemptRecord[]>($, key)
599 if (Array.isArray(list)) out.push(...list)
600 }
601 recordsCache = { at: t, records: out }
602 return out
603}
604
605async function readSpend($: Host): Promise<SpendPoint[]> {
606 try {
607 const { value } = await $.state.get(SPEND)
608 return Array.isArray(value) ? value : []
609 } catch {
610 return []
611 }
612}
613
614/** `git worktree list --porcelain`, from the cache while it is under 15 s old. */
615async function worktreeList($: Host, root: string, t: number): Promise<string> {
616 if (worktreeCache && t - worktreeCache.at < 15_000) return worktreeCache.text
617 const out = await run($, ['git', '-C', root, 'worktree', 'list', '--porcelain'])
618 const text = out.ok && out.exitCode === 0 ? out.stdout : ''
619 worktreeCache = { at: t, text }
620 return text
621}
622
623/** Everything the band and the pane draw: the session's facts gathered, the model made by live.ts. */
624async function liveModel($: Host): Promise<{ view: LiveView; since: number; records: AttemptRecord[]; running: { label: string; alias: string; tier: string; at: number }[] }> {
625 const t = await now($)
626 const live = await liveState($)
627 const queue = await readQueue($)
628 let since = 0
629 let usd = 0
630 try {
631 const u = await $.session.usage()
632 since = u.startedAt ?? 0
633 usd = u.cost?.usd ?? 0
634 } catch {
635 // unknown: the series and the records still draw
636 }
637 let root = ''
638 try {
639 root = await $.session.root()
640 } catch {
641 // no repo: no worktree rows
642 }
643 const records = await allRecords($)
644 const liveAt: Record<string, number> = {}
645 for (const w of live.workers) liveAt[w.label] = w.spawn.at ?? t
646 const view = liveView({
647 root,
648 ...(cfg.worktreeRoot ? { worktreeRoot: cfg.worktreeRoot } : {}),
649 porcelain: root ? await worktreeList($, root, t) : '',
650 records,
651 queue,
652 liveAt,
653 now: t,
654 budget: cfg.defaultBudget,
655 since,
656 usd,
657 series: await readSpend($),
658 owed: live.pending.length,
659 })
660 return { view, since, records, running: live.workers.map(w => ({ label: w.label, alias: w.spawn.alias, tier: w.spawn.tier, at: w.spawn.at ?? t })) }
661}
662
663/** One sample of the session's dollars; the next waits 15 s while a worker is live or queued, else 60 s. */
664async function sampleSpend($: Host): Promise<void> {
665 let active = false
666 try {
667 const t = await now($)
668 const live = await liveState($)
669 active = live.workers.length + live.pending.length + live.queued > 0
670 const usd = await sessionUsd($)
671 if (usd !== undefined) await $.state.set(SPEND, spendSeries(await readSpend($), { t, usd }))
672 redraw($)
673 } catch (err) {
674 debug($, `spend not sampled: ${String(err)}`)
675 }
676 scheduleSample($, sampleEvery(active))
677}
678function scheduleSample($: Host, ms: number) {
679 sampleTimer?.cancel()
680 try {
681 sampleTimer = $.clock.after(ms, () => void sampleSpend($))
682 } catch {
683 sampleTimer = undefined
684 }
685}
686
687// ---- status line ------------------------------------------------------------
688async function refreshStatus($: Host) {
689 if (!delegated) return
690 let running = 0
691 try {
692 running = (await $.agent.list()).filter(a => a.status === 'running').length
693 } catch {
694 running = 0
695 }
696 let pct: number | undefined
697 let usd: number | undefined
698 try {
699 const u = await $.session.usage()
700 pct = u.context.percent
701 usd = u.cost?.usd
702 } catch {
703 // leave them unknown
704 }
705 const t = await now($)
706 if (usd !== lastUsd) {
707 lastUsd = usd
708 lastCostChangeAt = t
709 }
710 if (shouldClearStatus({ running, now: t, lastCostChangeAt })) {
711 $.ui.status(undefined)
712 statusTimer?.cancel()
713 statusTimer = undefined
714 delegated = false
715 try {
716 await $.state.set(STATUS, '')
717 } catch {
718 // view only
719 }
720 return
721 }
722 const liveNow = (await liveState($)).workers
723 const text = statusText({ running, usd, pct }) + (liveNow.length > 0 ? ` · (${liveNow.length} live: ${(await liveLabels($, liveNow)).join(', ')})` : '')
724 $.ui.status(text)
725 try {
726 await $.state.set(STATUS, text)
727 } catch {
728 // view only
729 }
730 if (!statusTimer) {
731 try {
732 statusTimer = $.clock.every(STATUS_EVERY_MS, () => {
733 void refreshStatus($)
734 // a worker killed without a turn.complete frees its slot here
735 void drainQueue($)
736 })
737 } catch {
738 statusTimer = undefined
739 }
740 }
741}
742
743// ---- part 2: live state ($.agent.list × the store) ------------------------------
744async function agentList($: Host): Promise<{ id: string; status: string; description: string }[]> {
745 try {
746 return (await $.agent.list()).map(a => ({ id: a.id, status: a.status, description: a.description }))
747 } catch {
748 return []
749 }
750}
751async function sessionIdOf($: Host): Promise<string> {
752 try {
753 return await $.session.id()
754 } catch {
755 return ''
756 }
757}
758async function readText($: Host, path: string): Promise<string | undefined> {
759 try {
760 return await $.fs.read(path)
761 } catch {
762 return undefined
763 }
764}
765async function readQueue($: Host): Promise<QueuedSpawn[]> {
766 try {
767 const { value } = await $.state.get(QUEUE)
768 return Array.isArray(value) ? value : []
769 } catch {
770 return []
771 }
772}
773async function writeQueue($: Host, queue: QueuedSpawn[]): Promise<void> {
774 try {
775 await $.state.set(QUEUE, queue)
776 } catch (err) {
777 debug($, `queue not written: ${String(err)}`)
778 }
779}
780
781type Live = {
782 /** Agents the mod's spawn hook recorded that $.agent.list says are running (ad hoc ones included). */
783 running: number
784 /** Briefed workers running with their verdict still out: what holds a scheduler slot. */
785 workers: { label: string; agentId: string; spawn: SpawnRecord }[]
786 /** Briefed workers that handed back and whose verdict is not in, and verifies in flight. */
787 pending: { task: string; agentId?: string }[]
788 queued: number
789}
790
791const isOpen = (s: SpawnRecord): boolean => s.verdictAttempt !== s.attempt
792
793/** The session's delegation as it stands; `exclude` is an agent whose turn just ended. */
794async function liveState($: Host, exclude?: string): Promise<Live> {
795 const out: Live = { running: 0, workers: [], pending: [], queued: 0 }
796 for (const a of await agentList($)) {
797 if (a.id === exclude) continue
798 const spawn = await spawnByAgent($, a.id)
799 if (!spawn) continue
800 const label = taskLabel(spawn.task, spawn.subtask)
801 if (a.status === 'running') {
802 out.running += 1
803 if (!spawn.adhoc && isOpen(spawn)) out.workers.push({ label, agentId: a.id, spawn })
804 } else if (a.status === 'completed' && !spawn.adhoc && isOpen(spawn)) out.pending.push({ task: label, agentId: a.id })
805 }
806 for (const key of finalizing.keys()) {
807 const label = key.split(':')[0] ?? key
808 if (!out.pending.some(p => p.task === label)) out.pending.push({ task: label })
809 }
810 out.queued = (await readQueue($)).length
811 return out
812}
813
814// ---- part 2F: the scheduler --------------------------------------------------------
815/** One at a time: the slot count and the queue are read and written under this lock. */
816function withLock<T>(fn: () => Promise<T>): Promise<T> {
817 const chained = lockTail.then(fn, fn)
818 lockTail = chained.then(
819 () => undefined,
820 () => undefined,
821 )
822 return chained
823}
824
825/** A slot for a briefed spawn (a token, released when the spawn hook ends), or the spawn queued. */
826async function claimSlot($: Host, item: QueuedSpawn, label: string): Promise<{ token: number } | { deny: string }> {
827 return withLock(async () => {
828 const live = await liveState($)
829 // the token the dispatch path already holds for this very task is not another worker (GH-5)
830 const others = startingOthers(starting.values(), label)
831 const holders = [...(await liveLabels($, live.workers)), ...others]
832 const queue = await readQueue($)
833 const held = queuedIndex(queue, item.task, item.subtask ?? 'main')
834 if (hasSlot(live.workers.length, others.length, cfg.maxWorkers)) {
835 // a spawn of a task that is queued takes the queued place (GH-101)
836 if (held >= 0) await writeQueue($, queue.filter((_, i) => i !== held))
837 slotSeq += 1
838 starting.set(slotSeq, label)
839 return { token: slotSeq }
840 }
841 const entered = enqueue(queue, item)
842 if (entered.existing) return { deny: alreadyQueuedDeny(label, entered.existing.at, entered.existing.position) }
843 await writeQueue($, entered.queue)
844 debug($, `queued ${label}: ${holders.length} workers hold the ${cfg.maxWorkers} slots`)
845 return { deny: queuedDeny(label, holders, entered.queue.map(q => taskLabel(q.task, q.subtask ?? 'main'))) }
846 })
847}
848
849/** Starts the head of the queue while a slot is free; `exclude` is the worker whose turn just ended. */
850async function drainQueue($: Host, exclude?: string): Promise<void> {
851 for (;;) {
852 const picked = await withLock(async () => {
853 const queue = await readQueue($)
854 const head = queue[0]
855 if (!head) return undefined
856 const live = await liveState($, exclude)
857 if (!hasSlot(live.workers.length, starting.size, cfg.maxWorkers)) return undefined
858 slotSeq += 1
859 const token = slotSeq
860 starting.set(token, taskLabel(head.task, head.subtask ?? 'main'))
861 drainReserved.set(promptKey(head), token)
862 return { head, token }
863 })
864 if (!picked) return
865 const { head, token } = picked
866 let res: { agentId?: string; deny?: string }
867 try {
868 // model omitted: the spawn hook picks it from the brief, as for any spawn
869 res = await spawnSelf($, { prompt: head.prompt, description: head.description, subagentType: head.subagentType, ...(head.cwd ? { cwd: head.cwd } : {}) })
870 } catch (err) {
871 res = { deny: String(err) }
872 }
873 // the spawn hook took the reservation and released the slot; if it never ran, release it here
874 if (drainReserved.get(promptKey(head)) === token) drainReserved.delete(promptKey(head))
875 starting.delete(token)
876 const label = taskLabel(head.task, head.subtask ?? 'main')
877 if (res.deny !== undefined) {
878 const deny = res.deny
879 // work present is final too: the work needs --verify, and a respawn would be refused on every drain (the status timer runs one a minute)
880 const dropped = isFinalDeny(deny) || isWorkPresentDeny(deny)
881 $.ui.toast(`queued ${label} not started: ${deny}`)
882 if (dropped) {
883 drainRefusalSeen.delete(label)
884 await appendRow($, drainRefusalRow(label, deny, true))
885 await withLock(async () => writeQueue($, (await readQueue($)).filter(q => !(q.task === head.task && (q.subtask ?? 'main') === (head.subtask ?? 'main')))))
886 // the row behind it may start
887 continue
888 }
889 // GH-107: the head keeps its place (position 1); the brain reads a row, not a toast, but the same refusal only once
890 if (drainRefusalSeen.get(label) !== deny) {
891 drainRefusalSeen.set(label, deny)
892 await appendRow($, drainRefusalRow(label, deny, false))
893 }
894 return
895 }
896 drainRefusalSeen.delete(label)
897 // the head leaves the queue only now that its spawn has succeeded
898 await withLock(async () => writeQueue($, (await readQueue($)).filter(q => !(q.task === head.task && (q.subtask ?? 'main') === (head.subtask ?? 'main')))))
899 $.ui.toast(`started queued ${label} (waited ${waitedMinutes(head.at, await now($))} min)`)
900 }
901}
902
903// ---- part 2B: where a verdict goes -------------------------------------------------
904function logDebug($: Host, text: string) {
905 try {
906 $.ui.log(text, { to: 'debug' })
907 } catch {
908 // nothing to do
909 }
910}
911
912/** Appends a user-role row the brain reads; refused, the row goes to the transcript log and a toast. */
913async function appendRow($: Host, row: string) {
914 try {
915 await $.session.append({ message: { type: 'user', content: [{ type: 'text', text: row }] } })
916 } catch (err) {
917 // The model does not read a log line: the toast tells the person to look.
918 debug($, `row not appended: ${String(err)}`)
919 $.ui.log(row)
920 $.ui.toast(row.split('\n')[0] ?? row)
921 }
922}
923
924/** This session's verdict lines (2E reads them). */
925async function noteRecent($: Host, entry: RecentVerdict) {
926 const sid = await sessionIdOf($)
927 if (!sid) return
928 const list = (await storeGet<RecentVerdict[]>($, K.recent(sid))) ?? []
929 await storeSet($, K.recent(sid), [...list, entry].slice(-RECENT_CAP))
930}
931
932// ---- GH-106: what each worker spends ----------------------------------------------------
933/** A worker's own spend this attempt: the sum of its turns' usage; reset when the attempt moves on (a resume). */
934type CostRecord = { attempt: number; spend: Spend; warned?: true; stopped?: true }
935
936async function workerCost($: Host, agentId: string, attempt: number): Promise<Spend | undefined> {
937 const rec = await storeGet<CostRecord>($, K.cost(agentId))
938 return rec && rec.attempt === attempt ? rec.spend : undefined
939}
940
941/** `BE-310 $3.10` for each live worker: its own running cost, the label alone while none is measured. */
942async function liveLabels($: Host, workers: readonly { label: string; agentId: string; spawn: SpawnRecord }[]): Promise<string[]> {
943 const out: string[] = []
944 for (const w of workers) out.push(liveWorker(w.label, (await workerCost($, w.agentId, w.spawn.attempt))?.usd))
945 return out
946}
947
948/**
949 * A worker's turn ended: add its usage to the attempt's own cost, then hold it
950 * to its brief's ceiling: one wrap-up message when the cost first reaches
951 * `spend=`, and at twice it the over-spend verdict (and the turn ended, if the
952 * engine gave us its id). A turn that hands back a report is left to the verifier.
953 * True when the wrap-up message went out: the worker carries on, so this turn's end is not judged.
954 */
955async function trackSpend($: Host, spawn: SpawnRecord, e: { agentId: string; usage?: TurnUsage; answer: string }): Promise<boolean> {
956 if (!e.usage) return false
957 const prior = await storeGet<CostRecord>($, K.cost(e.agentId))
958 const rec: CostRecord = prior && prior.attempt === spawn.attempt ? prior : { attempt: spawn.attempt, spend: undefined as unknown as Spend }
959 const spend = addTurn(rec.spend, { ...e.usage, model: e.usage.model || spawn.resolvedModel })
960 const next: CostRecord = { ...rec, spend }
961 await storeSet($, K.cost(e.agentId), next)
962 const ceiling = spawn.spend
963 if (spawn.adhoc || ceiling === undefined || ceiling <= 0 || spend.usd === null) return false
964 if (spawn.verdictAttempt === spawn.attempt || extractReport(e.answer) !== undefined) return false
965 const state = ceilingState(spend.usd, ceiling)
966 if (state === 'stop' && !next.stopped) {
967 await storeSet($, K.cost(e.agentId), { ...next, stopped: true })
968 await overSpend($, spawn, spend.usd, ceiling)
969 return false
970 } else if (state === 'warn' && !next.warned) {
971 await storeSet($, K.cost(e.agentId), { ...next, warned: true })
972 try {
973 const sent = await $.session.send({ to: { agentId: e.agentId }, text: warnText(spend.usd, ceiling) })
974 if (!sent.isDelivered) debug($, `${taskLabel(spawn.task, spawn.subtask)}: spend warning not delivered: ${sent.reason}`)
975 return sent.isDelivered
976 } catch (err) {
977 debug($, `${taskLabel(spawn.task, spawn.subtask)}: spend warning not sent: ${String(err)}`)
978 }
979 }
980 return false
981}
982
983/** Twice the ceiling: the attempt's verdict is over-spend; the row says where to look; no escalation (never a failing verdict). */
984async function overSpend($: Host, spawnIn: SpawnRecord, usd: number, ceiling: number): Promise<void> {
985 const t = await now($)
986 const label = taskLabel(spawnIn.task, spawnIn.subtask)
987 const own = round4(usd)
988 let judged: AttemptRecord | undefined
989 if (!spawnIn.adhoc) {
990 const lane = laneOf(spawnIn)
991 const records = patchRecord(await loadAttempts($, spawnIn.task), spawnIn.subtask, spawnIn.attempt, { verdict: 'over-spend', usd: own, verdictAt: t }, lane)
992 await storeSet($, K.tasks(spawnIn.task), records)
993 judged = attemptsFor(records, spawnIn.subtask, lane).find(r => r.attempt === spawnIn.attempt)
994 }
995 const row = overSpendLine({ label, task: spawnIn.task, attempt: spawnIn.attempt, budget: spawnIn.budget, usd: own, spend: ceiling })
996 const latest = (await storeGet<SpawnRecord>($, K.spawn(spawnIn.key))) ?? spawnIn
997 await storeSet($, K.spawn(spawnIn.key), { ...latest, verdictAttempt: spawnIn.attempt, lastFailed: false, verdictBlock: row, verdictLine: row })
998 await setWorkers($, list => list.map(w => (w.task === spawnIn.task && w.subtask === spawnIn.subtask && w.attempt === spawnIn.attempt ? { ...w, verdict: 'over-spend' } : w)))
999 const next = OVER_SPEND_NEXT(spawnIn.task)
1000 await setLastVerdict($, { task: label, attempt: spawnIn.attempt, verdict: 'over-spend', next, at: t, text: row })
1001 await noteRecent($, { task: label, attempt: spawnIn.attempt, verdict: 'over-spend', line: row, at: t, owed: `${label}: ${next}` })
1002 await appendLedger($, { ...(judged ?? { task: spawnIn.task, subtask: spawnIn.subtask, attempt: spawnIn.attempt }), verdict: 'over-spend', usd: own, next, sessionId: await sessionIdOf($) })
1003 await appendRow($, row)
1004 // the worker's running turn, when the engine gave us its id (a subagent's turn.start carries none today)
1005 const turnId = spawnIn.agentId ? turnOf.get(spawnIn.agentId) : undefined
1006 if (turnId) {
1007 try {
1008 await $.turn.abort({ turnId })
1009 } catch (err) {
1010 debug($, `${label}: turn ${turnId} not aborted: ${String(err)}`)
1011 }
1012 }
1013}
1014
1015// ---- part 5E: friction the mod can see -------------------------------------------------
1016/** One correction, denial or refute, counted for the debrief's friction signal. */
1017function noteFriction($: Host, ev: FrictionEvent): Promise<unknown> {
1018 frictionTail = frictionTail.then(
1019 async () => {
1020 const sid = await sessionIdOf($)
1021 if (!sid) return
1022 const rec = (await storeGet<FrictionRecord>($, K.friction(sid))) ?? { total: 0, events: [] }
1023 await storeSet($, K.friction(sid), { total: rec.total + 1, events: addFriction(rec.events, ev) })
1024 },
1025 () => undefined,
1026 )
1027 return frictionTail
1028}
1029
1030// ---- part 5A: the ledger file -------------------------------------------------------------
1031/** One JSON line per judged attempt in `<root>/.delegation/ledger.jsonl` (append-only; `ledgerFile` off skips it). */
1032function appendLedger($: Host, row: Record<string, unknown>): Promise<unknown> {
1033 if (!cfg.ledgerFile) return Promise.resolve()
1034 ledgerTail = ledgerTail.then(
1035 async () => {
1036 const path = `${await $.session.root()}/${LEDGER_FILE}`
1037 const prev = (await readText($, path)) ?? ''
1038 await $.fs.write(path, `${prev}${prev === '' || prev.endsWith('\n') ? '' : '\n'}${JSON.stringify(row)}\n`)
1039 },
1040 () => undefined,
1041 )
1042 return ledgerTail.catch(err => debug($, `ledger line not written: ${String(err)}`))
1043}
1044
1045// ---- part 2E: the delegation state -------------------------------------------------
1046async function snapshot($: Host): Promise<StateSnapshot> {
1047 const s = emptySnapshot()
1048 const sid = await sessionIdOf($)
1049 const live = await liveState($)
1050 s.running = live.workers.map(w => ({ task: w.label, tier: w.spawn.tier, agentId: w.agentId, ...(w.spawn.at !== undefined ? { at: w.spawn.at } : {}) }))
1051 s.pending = live.pending
1052 s.queued = (await readQueue($)).map((q, i) => ({ task: taskLabel(q.task, q.subtask ?? 'main'), position: i + 1 }))
1053 if (!sid) return s
1054 const recent = (await storeGet<RecentVerdict[]>($, K.recent(sid))) ?? []
1055 s.recent = recent.map(r => ({ line: r.line, at: r.at }))
1056 const busy = new Set([...s.running.map(r => r.task), ...s.pending.map(p => p.task), ...s.queued.map(q => q.task)])
1057 s.owed = owedFrom(recent, busy)
1058 s.prs = prsFrom(recent)
1059 const debrief = await storeGet<DebriefRecord>($, K.debrief(sid))
1060 if (debrief?.lastAt !== undefined) {
1061 s.debrief = { at: debrief.lastAt, ...(debrief.agentId ? { agentId: debrief.agentId } : {}), ...(debrief.path ? { path: debrief.path } : {}), ...(debrief.finishedAt !== undefined ? { finishedAt: debrief.finishedAt } : {}) }
1062 }
1063 const inflight = await storeGet<EvalInflight>($, K.evalInflight)
1064 const last = ((await storeGet<EvalEntry[]>($, K.evals)) ?? []).filter(e => e.sessionId === sid).at(-1)
1065 if (inflight && inflight.sessionId === sid) s.eval = { tier: inflight.tier, sha: inflight.sha, at: inflight.at, running: true, ...(inflight.agentId ? { agentId: inflight.agentId } : {}) }
1066 else if (last) s.eval = { tier: last.tier, sha: last.sha, at: last.at, ...(last.pass !== undefined ? { pass: last.pass } : {}), ...(last.total !== undefined ? { total: last.total } : {}) }
1067 return s
1068}
1069
1070async function recipeDir($: Host): Promise<string | undefined> {
1071 const sid = await sessionIdOf($)
1072 if (!sid) return undefined
1073 const pad = scratchpadFor(await $.session.root(), sid)
1074 return (await exists($, pad)) ? pad : undefined
1075}
1076
1077// ---- part 2C/2D: idle timers --------------------------------------------------------
1078function cancelIdle() {
1079 idleTimer?.cancel()
1080 evalTimer?.cancel()
1081 idleTimer = undefined
1082 evalTimer = undefined
1083}
1084
1085/** (Re)starts the idle timers: the debrief after debriefIdleMinutes, the eval after evalIdleMinutes. */
1086function armIdle($: Host) {
1087 cancelIdle()
1088 try {
1089 if (cfg.autoDebrief) idleTimer = $.clock.after(cfg.debriefIdleMs, () => void maybeDebrief($))
1090 if (cfg.autoEval) evalTimer = $.clock.after(cfg.evalIdleMs, () => void maybeEval($))
1091 } catch (err) {
1092 debug($, `idle timers not armed: ${String(err)}`)
1093 }
1094}
1095
1096/**
1097 * 2C + 5E: at a clean stop, a background agent writes the debrief. The
1098 * friction signal is the harness breadcrumb file when it exists (lines past its
1099 * `.done` watermark), else the events the mod saw (corrections, denials,
1100 * refutes) since the last debrief. The instructions are the person's
1101 * `~/.claude/commands/debrief.md` when it exists, else hooks/templates/debrief.md.
1102 */
1103async function maybeDebrief($: Host) {
1104 if (debriefBusy || inTurn || !cfg.autoDebrief) return
1105 debriefBusy = true
1106 try {
1107 const sid = await sessionIdOf($)
1108 let home: string | undefined
1109 try {
1110 home = await $.env.get('HOME')
1111 } catch {
1112 home = undefined
1113 }
1114 if (!sid || !home) return debug($, 'no debrief: the session id or HOME is unknown')
1115 const live = await liveState($)
1116 const prev = await storeGet<DebriefRecord>($, K.debrief(sid))
1117 const crumbs = await readText($, breadcrumbPath(home, sid))
1118 const friction = (await storeGet<FrictionRecord>($, K.friction(sid))) ?? { total: 0, events: [] }
1119 const mode: 'breadcrumbs' | 'events' = crumbs !== undefined ? 'breadcrumbs' : 'events'
1120 const lines = mode === 'breadcrumbs' ? countLines(crumbs ?? '') : friction.total
1121 const watermark = mode === 'breadcrumbs' ? parseWatermark(await readText($, watermarkPath(home, sid))) : prev?.mode === 'events' ? prev.lines : 0
1122 const t = await now($)
1123 const verdict = cleanStop({
1124 inTurn,
1125 running: live.running,
1126 pending: live.pending.length,
1127 queued: live.queued,
1128 lines,
1129 watermark,
1130 minNewLines: mode === 'breadcrumbs' ? cfg.debriefMinNewLines : cfg.debriefMinEvents,
1131 now: t,
1132 cooldownMs: cfg.debriefCooldownMs,
1133 ...(prev ? { lastAt: prev.lastAt, ...(mode === 'breadcrumbs' && prev.mode !== 'events' ? { lastWatermark: prev.watermark } : {}) } : {}),
1134 })
1135 if (!verdict.ok) return debug($, `no debrief: ${verdict.why}${mode === 'events' ? ' (friction events)' : ''}`)
1136 if (inTurn) return
1137 const source = debriefSource(home, $.plugin.root, await exists($, debriefSkillPath(home)))
1138 // claimed before the spawn: never twice for this watermark, whatever the spawn does
1139 const record: DebriefRecord = { lastAt: t, watermark, lines, sessionId: sid, mode, builtIn: source.builtIn }
1140 await storeSet($, K.debrief(sid), record)
1141 let prompt: string
1142 if (source.builtIn) {
1143 const recent = ((await storeGet<RecentVerdict[]>($, K.recent(sid))) ?? []).filter(r => prev === undefined || r.at > prev.lastAt)
1144 const fresh = friction.events.slice(-Math.max(0, Math.min(friction.events.length, lines - watermark)))
1145 prompt = builtInDebriefPrompt(source.path, sid, await $.session.root(), [...frictionFacts(fresh), ...recent.map(r => `verdict: ${r.line}`)])
1146 } else prompt = debriefPrompt(source.path, sid)
1147 const recordAgent = async (agentId: string) => storeSet($, K.debrief(sid), { ...record, agentId })
1148 runnerStarted.set(promptKey({ prompt }), recordAgent)
1149 let res: { agentId?: string; deny?: string }
1150 try {
1151 res = await spawnSelf($, { subagentType: cfg.debriefAgent, model: 'sonnet', description: 'debrief', prompt })
1152 } catch (err) {
1153 res = { deny: String(err) }
1154 }
1155 runnerStarted.delete(promptKey({ prompt }))
1156 if (res.deny !== undefined) {
1157 await storeSet($, K.debrief(sid), { ...record, denied: res.deny })
1158 return debug($, `debrief not started: ${res.deny}`)
1159 }
1160 if (res.agentId) await recordAgent(res.agentId)
1161 $.ui.toast('debrief running in the background')
1162 } finally {
1163 debriefBusy = false
1164 }
1165}
1166
1167/** The eval runner of this session still out (its agent not listed as finished). */
1168async function evalOut($: Host, sid: string): Promise<EvalInflight | undefined> {
1169 const inflight = await storeGet<EvalInflight>($, K.evalInflight)
1170 if (!inflight || inflight.sessionId !== sid) return undefined
1171 if ((await now($)) - inflight.at > EVAL_STALE_MS) return undefined
1172 const listed = (await agentList($)).find(a => a.id === inflight.agentId)
1173 return listed && listed.status !== 'running' ? undefined : inflight
1174}
1175
1176/** 2D: origin/main moved and the session is idle → one eval runner per sha (only with an evalCommand, 5E). */
1177async function maybeEval($: Host) {
1178 if (evalBusy || inTurn || !cfg.autoEval) return
1179 evalBusy = true
1180 try {
1181 const sid = await sessionIdOf($)
1182 const live = await liveState($)
1183 const inFlight = (await evalOut($, sid)) !== undefined
1184 const base = { autoEval: cfg.autoEval, inTurn, running: live.running, pending: live.pending.length, queued: live.queued, inFlight }
1185 // the cheap checks first: a busy session never runs git
1186 const busy = shouldEval({ ...base, sha: 'idle', lastSha: undefined })
1187 if (!busy.ok) return debug($, `no eval: ${busy.why}`)
1188 const root = await $.session.root()
1189 const r = await run($, revParseArgv(root), { cwd: root, timeoutMs: 20000 })
1190 const sha = r.ok && r.exitCode === 0 ? parseSha(r.stdout) : undefined
1191 const lastSha = await storeGet<string>($, K.lastSha)
1192 const verdict = shouldEval({ ...base, ...(sha ? { sha } : {}), ...(lastSha ? { lastSha } : {}) })
1193 if (!verdict.ok || !sha) return debug($, `no eval: ${verdict.ok ? 'origin/main unknown' : verdict.why}`)
1194 if (inTurn) return
1195 await storeSet($, K.lastSha, sha) // never twice per sha, whatever the runner does
1196 await startEvalRunner($, 'T1', sha, root, sid)
1197 } finally {
1198 evalBusy = false
1199 }
1200}hooks/lib/dashboard.ts 317 lines1// Part 4A/4B and 3D, the drawing's data: the band's lines, the pane's six
2// blocks and their sparklines, the open items, and the status tool's text.
3// The trees themselves are band.tsx and pane.tsx. Pure: no `$`.
4import type { BandItem, DashboardBlock } from '../types'
5import { usdText } from './cost'
6import { SERIES, type Metrics } from './metrics'
7
8export type { BandItem } from '../types'
9
10const ARROW = '▸'
11
12/** `04:12` (mm:ss), `1:02:05` past the hour. */
13export function elapsed(ms: number): string {
14 const s = Math.max(0, Math.floor(ms / 1000))
15 const two = (n: number) => String(n).padStart(2, '0')
16 const h = Math.floor(s / 3600)
17 const m = Math.floor((s % 3600) / 60)
18 return h > 0 ? `${h}:${two(m)}:${two(s % 60)}` : `${two(m)}:${two(s % 60)}`
19}
20
21/** Cut to `columns` cells, the last one an ellipsis. */
22export function fit(text: string, columns: number): string {
23 const cells = [...text]
24 if (cells.length <= columns) return text
25 return columns <= 1 ? cells.slice(0, Math.max(0, columns)).join('') : `${cells.slice(0, columns - 1).join('')}…`
26}
27
28const ADVICE: Record<string, string> = {
29 resume: 'resume advised',
30 respawn: 'respawn advised',
31 exhausted: 'budget exhausted',
32 check: 'check by hand',
33 'fix-brief': 'fix the brief',
34}
35
36/**
37 * A verdict that waits on a decision, in one clause: the verdict with its
38 * claim (`refuted on scope`; a no-repo verdict keeps its why), then the
39 * advice, or that an amend needs approval.
40 */
41export function decisionText(r: { verdict: string; reason?: string; advice?: string; approval?: boolean; noRepo?: boolean }): string {
42 const claim = r.reason ? (r.noRepo ? `: ${r.reason}` : ` ${r.reason.replace(/ \(.*$/, '')}`) : ''
43 const verdict = `${r.verdict}${r.noRepo ? ' (no repo)' : ''}${claim}`
44 const advice = r.approval ? 'amend needs approval' : (r.advice && ADVICE[r.advice]) || r.advice
45 return advice ? `${verdict} · ${advice}` : verdict
46}
47
48export type BandLine = { key: string; text: string; clear?: boolean }
49
50/** Room the band keeps at a line's end for its `[ clear ]`. */
51export const CLEAR_COLUMNS = 10
52
53/**
54 * One line per item, sized to the band's columns. The running workers'
55 * fields are padded to one another so the numbers line up.
56 */
57export function bandLines(items: readonly BandItem[], now: number, columns: number): BandLine[] {
58 const running = items.filter((i): i is Extract<BandItem, { kind: 'running' }> => i.kind === 'running')
59 const width = (f: (i: Extract<BandItem, { kind: 'running' }>) => string) => Math.max(0, ...running.map(i => [...f(i)].length))
60 const taskW = width(i => i.task)
61 const tierW = width(i => `${i.tier}/${i.alias}`)
62 const timeW = width(i => elapsed(now - i.at))
63 const usdW = width(i => usdText(i.usd === undefined ? undefined : { usd: i.usd }))
64 return items.map((item, n): BandLine => {
65 switch (item.kind) {
66 case 'running': {
67 const cost = usdText(item.usd === undefined ? undefined : { usd: item.usd })
68 const text = [
69 `${ARROW} ${item.task.padEnd(taskW)}`,
70 `${item.tier}/${item.alias}`.padEnd(tierW),
71 elapsed(now - item.at).padStart(timeW),
72 cost.padEnd(usdW),
73 item.phase,
74 ].join(' · ')
75 return { key: `run-${item.task}-${n}`, text: fit(text, columns) }
76 }
77 case 'decision':
78 return { key: `decide-${item.task}-${n}`, text: fit(`${ARROW} ${item.task} · ${item.text}`, columns) }
79 case 'queued':
80 return { key: `queued-${item.task}-${n}`, text: fit(`${ARROW} queued: ${item.task} (pos ${item.position})`, columns) }
81 case 'eval-failed': {
82 const score = item.total !== undefined && item.pass !== undefined ? ` ${item.pass}/${item.total}` : ' (no result block)'
83 return { key: 'eval-failed', text: fit(`${ARROW} eval ${item.tier}${score} — see transcript`, Math.max(1, columns - CLEAR_COLUMNS)), clear: true }
84 }
85 case 'runner':
86 return { key: `runner-${item.what}-${n}`, text: fit(`${ARROW} ${item.what} running · ${elapsed(now - item.at)}`, columns) }
87 }
88 })
89}
90
91/** The pane's "Open items": every band item with the exact next action. */
92export function openItemLines(items: readonly BandItem[], now: number): string[] {
93 return items.map(item => {
94 switch (item.kind) {
95 case 'running':
96 return item.phase === 'verifying'
97 ? `${item.task}: verifying ${item.tier}/${item.alias} ${elapsed(now - item.at)} — the verdict lands on its own`
98 : `${item.task}: running ${item.tier}/${item.alias} ${elapsed(now - item.at)} — wait for the hand-back`
99 case 'decision':
100 return `${item.task}: ${item.next}`
101 case 'queued':
102 return `${item.task}: queued (position ${item.position}) — starts when a worker slot frees`
103 case 'eval-failed': {
104 const score = item.total !== undefined && item.pass !== undefined ? ` ${item.pass}/${item.total}` : ''
105 return `eval ${item.tier}${score} at ${item.sha.slice(0, 8)}: read the failing names in the transcript, then press clear in the band`
106 }
107 case 'runner':
108 return `${item.what}: running ${elapsed(now - item.at)} — nothing to do`
109 }
110 })
111}
112
113// ---- the pane's blocks ------------------------------------------------------------
114
115export type Block = DashboardBlock
116
117const plural = (n: number, one: string, many = `${one}s`) => `${n} ${n === 1 ? one : many}`
118const STEPS_PER_DISPATCH = 10
119
120/**
121 * The six blocks: a heading, one line of numbers for this session, and the
122 * series over `history` (the last 14 sessions, oldest first, this one last).
123 */
124export function blocksFor(input: {
125 current: Metrics
126 history: readonly Metrics[]
127 rateLimits?: readonly { kind: string; percentUsed: number }[]
128 owedRows?: readonly string[]
129}): Block[] {
130 const m = input.current
131 const series = (key: string) => {
132 const s = SERIES.find(x => x.key === key)
133 return s ? input.history.map(h => s.value(h)) : []
134 }
135 const window = input.rateLimits?.find(r => r.kind === 'five_hour') ?? input.rateLimits?.[0]
136 const v = m.verdicts
137 return [
138 {
139 key: 'steps',
140 heading: 'Steps saved',
141 number: `${plural(m.dispatches, 'dispatch', 'dispatches')} · ~${m.dispatches * STEPS_PER_DISPATCH} hand steps saved`,
142 detail: [],
143 series: series('steps'),
144 alt: 'dispatches per session',
145 },
146 {
147 key: 'verdicts',
148 heading: 'Verdicts',
149 number: `${v.verified} verified · ${v.unverified} unverified · ${v.refuted} refuted${v.falseRefuted > 0 ? ` (${v.falseRefuted} false)` : ''}`,
150 detail: [],
151 series: series('verdicts'),
152 alt: 'refuted share per session',
153 },
154 {
155 key: 'tiers',
156 heading: 'Tiers',
157 number: `haiku ${m.spawns.haiku} · sonnet ${m.spawns.sonnet} · opus ${m.spawns.opus} · resume ${m.escalations.resume} · respawn ${m.escalations.respawn}`,
158 detail: [],
159 series: series('tiers'),
160 alt: 'share of spawns on the cheapest tier that verified, per session',
161 },
162 {
163 key: 'owed',
164 heading: 'Owed work that ran itself',
165 number: `${plural(m.debriefs, 'debrief')} · ${plural(m.evals.run, 'eval')}${m.evals.failed > 0 ? ` (${m.evals.failed} failed)` : ''}`,
166 detail: [...(input.owedRows ?? [])],
167 series: series('owed'),
168 alt: 'debriefs and evals per session',
169 },
170 {
171 key: 'compaction',
172 heading: 'Compaction',
173 number: `${plural(m.compactions, 'compaction')} served, state block attached`,
174 detail: [],
175 series: series('compaction'),
176 alt: 'compactions per session',
177 },
178 {
179 key: 'spend',
180 heading: 'Spend',
181 number: `$${m.usd.toFixed(2)} on workers · ${window ? `${window.kind} window ${window.percentUsed}%` : 'window –'}`,
182 detail: [],
183 series: series('spend'),
184 alt: 'worker dollars per session',
185 },
186 ]
187}
188
189// ---- sparklines ------------------------------------------------------------------
190
191export const SPARK_POINTS = 14
192export const SPARK_HEIGHT = 16
193/** CSS pixels per column the Svg takes, and the most columns it spans. */
194export const PX_PER_COLUMN = 8
195export const SPARK_MAX_COLUMNS = 60
196
197const r1 = (n: number) => Math.round(n * 10) / 10
198
199/** The values drawn: nulls dropped, the last 14 kept. */
200export const sparkValues = (values: readonly (number | null)[]): number[] =>
201 values.filter((v): v is number => v !== null && Number.isFinite(v)).slice(-SPARK_POINTS)
202
203/**
204 * The polyline's points over a `width` × `height` box, 1 px inside it: zero at
205 * the bottom, the largest value at the top; one value is a flat line across.
206 */
207export function sparkPoints(values: readonly (number | null)[], width: number, height: number): string {
208 const v = sparkValues(values)
209 if (v.length === 0) return ''
210 const pad = 1
211 const max = Math.max(0, ...v)
212 const y = (n: number) => r1(max > 0 ? pad + (1 - n / max) * (height - 2 * pad) : height - pad)
213 if (v.length === 1) return `${pad},${y(v[0] as number)} ${r1(width - pad)},${y(v[0] as number)}`
214 const step = (width - 2 * pad) / (v.length - 1)
215 return v.map((n, i) => `${r1(pad + i * step)},${y(n)}`).join(' ')
216}
217
218/**
219 * The Svg sparkline: one polyline, no fill, 2 px, stroked in the surface's
220 * foreground (`currentColor`, and the CSS system colour `CanvasText` where the
221 * markup's style is kept), the viewBox sized to the column. Undefined when
222 * there is nothing to draw.
223 */
224export function sparkSvg(values: readonly (number | null)[], columns: number): { source: string; width: number; height: number } | undefined {
225 const width = Math.max(1, Math.min(columns, SPARK_MAX_COLUMNS)) * PX_PER_COLUMN
226 const height = SPARK_HEIGHT
227 const points = sparkPoints(values, width, height)
228 if (!points) return undefined
229 const source =
230 `<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 ${width} ${height}" width="${width}" height="${height}">` +
231 '<style>svg{color-scheme:light dark}polyline{stroke:CanvasText}</style>' +
232 `<polyline points="${points}" fill="none" stroke="currentColor" stroke-width="2" stroke-linejoin="round" stroke-linecap="round"/>` +
233 '</svg>'
234 return { source, width, height }
235}
236
237const BLOCKS = '▁▂▃▄▅▆▇█'
238/** The terminal's default colour (bit 24 alone), foreground and background. */
239const DEFAULT_COLOUR = 0x01000000
240
241/** The Raster sparkline: one row of block glyphs, one per session, in the terminal's own colours. */
242export function sparkCells(values: readonly (number | null)[]): { cells: string; columns: number; rows: 1; glyphs: string } | undefined {
243 const v = sparkValues(values)
244 if (v.length === 0) return undefined
245 const max = Math.max(0, ...v)
246 const glyphs = v.map(n => BLOCKS[max > 0 ? Math.round((n / max) * (BLOCKS.length - 1)) : 0] as string)
247 const words = new Uint32Array(v.length * 3)
248 glyphs.forEach((g, i) => {
249 words[i * 3] = g.codePointAt(0) as number
250 words[i * 3 + 1] = DEFAULT_COLOUR
251 words[i * 3 + 2] = DEFAULT_COLOUR
252 })
253 return { cells: base64(littleEndian(words)), columns: v.length, rows: 1, glyphs: glyphs.join('') }
254}
255
256/** u32 words as little-endian bytes, whatever the host's order. */
257function littleEndian(words: Uint32Array): Uint8Array {
258 const bytes = new Uint8Array(words.length * 4)
259 words.forEach((w, i) => {
260 bytes[i * 4] = w & 0xff
261 bytes[i * 4 + 1] = (w >>> 8) & 0xff
262 bytes[i * 4 + 2] = (w >>> 16) & 0xff
263 bytes[i * 4 + 3] = (w >>> 24) & 0xff
264 })
265 return bytes
266}
267
268const B64 = 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/'
269
270/** Standard padded base64 (the environment's Uint8Array may have no toBase64). */
271export function base64(bytes: Uint8Array): string {
272 let out = ''
273 for (let i = 0; i < bytes.length; i += 3) {
274 const a = bytes[i] as number
275 const b = bytes[i + 1]
276 const c = bytes[i + 2]
277 const n = (a << 16) | ((b ?? 0) << 8) | (c ?? 0)
278 out += B64[(n >> 18) & 63]
279 out += B64[(n >> 12) & 63]
280 out += b === undefined ? '=' : B64[(n >> 6) & 63]
281 out += c === undefined ? '=' : B64[n & 63]
282 }
283 return out
284}
285
286// ---- 3D: the status tool -----------------------------------------------------------
287
288export const STATUS_TOOL = {
289 name: 'status',
290 description:
291 'The chassis-delegation state, as the system prompt section renders it: running workers, pending verdicts, queued spawns and what is owed, then the last 10 verdict lines and the last eval and debrief. Call it instead of reading the store. Pass open: true only when the person asks to see the Delegation pane.',
292 inputSchema: {
293 type: 'object',
294 properties: {
295 open: { type: 'boolean', description: 'Also open the Delegation pane for the person (only when they asked to see it)' },
296 },
297 additionalProperties: false,
298 },
299} as const
300
301/** The name the model calls it by: `mcp__<plugin>__<name>`. */
302export const STATUS_TOOL_NAME = 'mcp__chassis-delegation__status'
303
304export const STATUS_VERDICTS = 10
305const NOTHING = 'Delegation state (chassis-delegation): nothing running, nothing owed.'
306
307/** The compose section's text (or that nothing runs), then the tails: 10 verdicts, the last eval, the last debrief. */
308export function statusToolText(section: string | undefined, tails: { verdicts: readonly string[]; eval?: string; debrief?: string }): string {
309 const verdicts = tails.verdicts.slice(-STATUS_VERDICTS)
310 return [
311 section ?? NOTHING,
312 ...(verdicts.length > 0 ? [`Last ${STATUS_VERDICTS} verdicts:`, ...verdicts.map(l => `- ${l}`)] : [`Last ${STATUS_VERDICTS} verdicts: none this session`]),
313 `Last eval: ${tails.eval ?? 'none'}`,
314 `Last debrief: ${tails.debrief ?? 'none this session'}`,
315 ].join('\n')
316}
317hooks/lib/band.tsx 62 lines1// Part 4A, as GH-112 wires it: the band above the prompt, one row, drawn only
2// while a worker is live, a spawn is queued or a verdict is owed. The whole
3// band is a hover scope: hovering it reveals the card (the KPI tiles, the spend
4// chart, the worktree table) beneath the row. `[ details ]` opens the pane.
5// No `$` here: register.ts resolves the element table and hands in the handler.
6import type { RenderNode } from 'claude-code'
7
8import type { BandLine } from './dashboard'
9import { sparkCells, sparkSvg } from './dashboard'
10import { bandSummary, recentSpend, type LiveView, type Scheme } from './live'
11import { liveBlocks, type PaneTable } from './pane'
12
13export const DASH_SCOPE = 'chassis-delegation-dash'
14
15/** Lines past the band's rows fold into one `… n more` line; the Button keeps the last row. */
16export function capLines(lines: readonly BandLine[], maxRows: number): BandLine[] {
17 const room = Math.max(1, maxRows - 1)
18 if (lines.length <= room) return [...lines]
19 const keep = Math.max(0, room - 1)
20 return [...lines.slice(0, keep), { key: 'more', text: `▸ … ${lines.length - keep} more — see details` }]
21}
22
23function sparkline(p: PaneTable, v: LiveView): RenderNode | null {
24 const values = recentSpend(v.series, v.now)
25 if (values.length < 2) return null
26 const { Box } = p.t
27 if (p.surface === 'terminal') {
28 const cells = sparkCells(values)
29 if (!cells) return null
30 const { Raster } = p.t
31 return (
32 <Box flexDirection="row">
33 <Raster key="band-spark" columns={cells.columns} rows={1} cells={cells.cells} />
34 </Box>
35 )
36 }
37 const svg = sparkSvg(values, 14)
38 if (!svg) return null
39 const { Svg } = p.t
40 return (
41 <Box flexDirection="row">
42 <Svg source={svg.source} alt={`Session spend, last 30 minutes: ${values.map(n => `$${n.toFixed(2)}`).join(', ')}`} width={svg.width} height={svg.height} />
43 </Box>
44 )
45}
46
47export function bandTree(p: PaneTable, v: LiveView, scheme: Scheme, columns: number, onDetails: () => void): RenderNode {
48 const { Box, Text, Button } = p.t
49 return (
50 <Box key="dash-band" flexDirection="column" hover={{ scope: DASH_SCOPE }}>
51 <Box flexDirection="row" gap={1} hover={{ scope: DASH_SCOPE }}>
52 <Text wrap="truncate-end">{bandSummary(v)} ·</Text>
53 {sparkline(p, v)}
54 <Button key="details" label="details" hotkey="d" variant="primary" onPress={onDetails} />
55 </Box>
56 <Box key="dash-card" flexDirection="column" display="none" hover={{ scope: DASH_SCOPE, display: 'flex' }}>
57 {liveBlocks(p, v, scheme, Math.max(10, columns - 2), false)}
58 </Box>
59 </Box>
60 )
61}
62hooks/lib/pane.tsx 234 lines1// Part 4B: the Delegation pane, as a tree. Opened only by the band's
2// `[ details ]`, `/delegation`, or the status tool with `open: true`. Six
3// blocks (a heading, this session's number line, a sparkline over the last 14
4// sessions), then "Open items", then `[ close ]`, the one control. Sparklines
5// are an Svg on vscode, desktop and mobile and a Raster on the terminal, which
6// has no Svg. No `$` here: register.ts resolves the table and hands in the
7// close handler.
8import type { Elements, RenderNode } from 'claude-code'
9
10import type { DashboardBlock } from '../types'
11import { sparkCells, sparkSvg, sparkValues } from './dashboard'
12import { chartCells, chartSvg, modelColor, modelLabel, tokText, verdictMark, type LiveView, type Scheme, type WorktreeRow } from './live'
13
14/** The surface's table, narrowed: the terminal draws a Raster, the rest an Svg. */
15export type PaneTable =
16 | { surface: 'terminal'; t: Elements['terminal'] }
17 | { surface: 'desktop'; t: Elements['desktop'] }
18 | { surface: 'vscode'; t: Elements['vscode'] }
19 | { surface: 'mobile'; t: Elements['mobile'] }
20
21export type PaneView = {
22 blocks: readonly DashboardBlock[]
23 /** Every item the band would show, with its next action. */
24 open: readonly string[]
25 /** The pane body's width in cells (`e.props.bodyColumns`). */
26 columns: number
27 /** GH-112: the live model; absent, the pane draws only the cross-session blocks. */
28 live?: LiveView
29 scheme?: Scheme
30}
31
32const INDENT = 2
33
34function spark(p: PaneTable, b: DashboardBlock, columns: number): RenderNode {
35 const { Text } = p.t
36 const shown = sparkValues(b.series)
37 if (shown.length === 0) return <Text>{' '.repeat(INDENT)}no sessions measured yet</Text>
38 const alt = `${b.alt}, last ${shown.length} session${shown.length === 1 ? '' : 's'}: ${shown.map(v => +v.toFixed(2)).join(', ')}`
39 if (p.surface === 'terminal') {
40 const cells = sparkCells(b.series)
41 if (!cells) return <Text>{' '.repeat(INDENT)}no sessions measured yet</Text>
42 const { Box, Raster } = p.t
43 return (
44 <Box flexDirection="row" paddingLeft={INDENT}>
45 <Raster key={`spark-${b.key}`} columns={cells.columns} rows={1} cells={cells.cells} />
46 </Box>
47 )
48 }
49 const svg = sparkSvg(b.series, Math.max(1, columns - INDENT))
50 if (!svg) return <Text>{' '.repeat(INDENT)}no sessions measured yet</Text>
51 const { Box, Svg } = p.t
52 return (
53 <Box flexDirection="row" paddingLeft={INDENT}>
54 <Svg source={svg.source} alt={alt} width={svg.width} height={svg.height} />
55 </Box>
56 )
57}
58
59// ---- GH-112: the live blocks, drawn by the band's card (the first three, compact) and the pane (all four) ----
60
61const pad = (text: string, w: number) => (text.length >= w ? text.slice(0, Math.max(0, w - 1)) + (text.length > w ? '…' : '') : text.padEnd(w))
62const usdCell = (n: number | undefined) => (n === undefined ? '$–' : `$${n.toFixed(2)}`)
63
64/** Stat tiles: live workers, queued, spend this session (the hero), verified on first attempt. */
65function kpiRow(p: PaneTable, v: LiveView): RenderNode {
66 const { Box, Text } = p.t
67 const tile = (key: string, label: string, value: string, hero?: boolean) => (
68 <Box key={key} flexDirection="column" borderStyle="round" paddingX={1}>
69 <Text dimColor>{label}</Text>
70 <Text bold>{value}</Text>
71 {hero ? <Text dimColor>whole session</Text> : null}
72 </Box>
73 )
74 return (
75 <Box key="dash-kpis" flexDirection="row" gap={1}>
76 {tile('kpi-live', 'live', String(v.live))}
77 {tile('kpi-queued', 'queued', String(v.queued))}
78 {tile('kpi-spend', 'spend', `$${v.usd.toFixed(2)}`, true)}
79 {tile('kpi-first', 'verified 1st try', `${v.firstTry.n} of ${v.firstTry.m}`)}
80 </Box>
81 )
82}
83
84/** Cumulative session dollars over time: an Svg line on desktop, vscode and mobile; a Raster on the terminal. */
85function spendChart(p: PaneTable, v: LiveView, columns: number): RenderNode {
86 const { Box, Text } = p.t
87 const alt = `Session spend over time: $${v.usd.toFixed(2)} now, ${v.series.length} samples, ${v.events.length} spawn and verdict ticks`
88 let chart: RenderNode = <Text dimColor>collecting samples…</Text>
89 if (p.surface === 'terminal') {
90 const c = chartCells(v.series, v.events, v.now, Math.max(8, columns - 2))
91 if (c) {
92 const { Raster } = p.t
93 chart = <Raster key="spend-raster" columns={c.columns} rows={c.rows} cells={c.cells} />
94 }
95 } else {
96 const svg = chartSvg(v.series, v.events, v.now, Math.max(8, columns - 2))
97 if (svg) {
98 const { Svg } = p.t
99 chart = <Svg source={svg.source} alt={alt} width={svg.width} height={svg.height} isInteractive />
100 }
101 }
102 return (
103 <Box key="dash-spend" flexDirection="column">
104 <Text bold>Spend over time</Text>
105 <Text dimColor>cumulative dollars, this session · ticks: spawn ┬ verdict ┴ · hover for time, $ and event</Text>
106 {chart}
107 </Box>
108 )
109}
110
111function stateCell(p: PaneTable, r: WorktreeRow, w: number): RenderNode {
112 const { Box, Text } = p.t
113 const mark = verdictMark(r.verdict)
114 return (
115 <Box flexDirection="row" width={w}>
116 {mark ? <Text color={mark.color}>{mark.icon} </Text> : null}
117 <Text wrap="truncate-end">{r.state}</Text>
118 </Box>
119 )
120}
121
122/** One row per task: task · model (a chip in the model's colour, its name in ink) · state · worktree · tokens · cost · attempt. */
123function worktreeTable(p: PaneTable, v: LiveView, scheme: Scheme, columns: number, max: number): RenderNode {
124 const { Box, Text } = p.t
125 const rows = v.rows.slice(0, max)
126 const w = (f: (r: WorktreeRow) => string, floor: number, cap: number) => Math.min(cap, Math.max(floor, ...rows.map(r => [...f(r)].length)))
127 const taskW = w(r => r.task, 4, 14)
128 const modelW = w(r => (r.model === r.family || r.model === '–' ? r.family : `${r.family} ${r.model}`), 5, 28) + 2
129 const stateW = w(r => r.state, 5, 16) + 2
130 const folderW = w(r => (r.branch ? `${r.folder} ${r.branch}` : r.folder), 8, 40)
131 const wide = columns >= taskW + modelW + stateW + folderW + 26
132 return (
133 <Box key="dash-worktrees" flexDirection="column">
134 <Text bold>Worktrees</Text>
135 {rows.length === 0 ? <Text dimColor>no worker has run this session</Text> : null}
136 {rows.length > 0 ? (
137 <Box flexDirection="row">
138 <Text dimColor>{pad('task', taskW + 1)}</Text>
139 <Text dimColor>{pad('model', modelW + 1)}</Text>
140 <Text dimColor>{pad('state', stateW + 1)}</Text>
141 {wide ? <Text dimColor>{pad('worktree · branch', folderW + 1)}</Text> : null}
142 <Text dimColor>{pad('tok', 7)}</Text>
143 <Text dimColor>{pad('cost', 8)}</Text>
144 <Text dimColor>att</Text>
145 </Box>
146 ) : null}
147 {rows.map(r => {
148 const c = modelColor(r.family, scheme)
149 const name = r.model === r.family || r.model === '–' ? r.family : `${r.family} ${r.model}`
150 return (
151 <Box flexDirection="row">
152 <Text wrap="truncate-end">{pad(r.task, taskW + 1)}</Text>
153 <Box flexDirection="row" width={modelW + 1}>
154 {c ? <Text color={c}>■ </Text> : <Text dimColor>■ </Text>}
155 <Text wrap="truncate-end">{name}</Text>
156 </Box>
157 {stateCell(p, r, stateW + 1)}
158 {wide ? <Text wrap="truncate-end">{pad(r.branch ? `${r.folder} ${r.branch}` : r.folder, folderW + 1)}</Text> : null}
159 <Text>{pad(r.tokens === undefined ? '–' : tokText(r.tokens), 7)}</Text>
160 <Text>{pad(usdCell(r.usd), 8)}</Text>
161 <Text>{r.attempt}</Text>
162 </Box>
163 )
164 })}
165 {v.rows.length > max ? <Text dimColor>… {v.rows.length - max} more in the pane</Text> : null}
166 <Text dimColor>tokens and cost update when a run ends</Text>
167 </Box>
168 )
169}
170
171/** Up to three bars plus "other", each direct-labelled `sonnet · $3.10 · 412k tok · 4 verified`. */
172function modelBars(p: PaneTable, v: LiveView, scheme: Scheme): RenderNode {
173 const { Box, Text } = p.t
174 const shown = v.byModel.filter(m => m.family !== 'other' || m.usd > 0 || m.tokens > 0)
175 const top = Math.max(0.0001, ...shown.map(m => m.usd))
176 return (
177 <Box key="dash-models" flexDirection="column">
178 <Text bold>Spend by model</Text>
179 {shown.map(m => {
180 const c = modelColor(m.family, scheme)
181 const cells = m.usd > 0 ? Math.max(1, Math.round((m.usd / top) * 20)) : 0
182 return (
183 <Box flexDirection="row" gap={1}>
184 <Box width={20}>{c ? <Text color={c}>{'█'.repeat(cells)}</Text> : <Text dimColor>{'█'.repeat(cells)}</Text>}</Box>
185 <Text wrap="truncate-end">{modelLabel(m)}</Text>
186 </Box>
187 )
188 })}
189 </Box>
190 )
191}
192
193/** The card (compact: tiles, chart, worktrees) and the pane (all four) draw these. */
194export function liveBlocks(p: PaneTable, v: LiveView, scheme: Scheme, columns: number, full: boolean): RenderNode {
195 const { Box } = p.t
196 return (
197 <Box flexDirection="column" gap={1}>
198 {kpiRow(p, v)}
199 {spendChart(p, v, columns)}
200 {worktreeTable(p, v, scheme, columns, full ? 50 : 6)}
201 {full ? modelBars(p, v, scheme) : null}
202 </Box>
203 )
204}
205
206export function paneTree(p: PaneTable, view: PaneView, onClose: () => void): RenderNode {
207 const { Box, Text, Button } = p.t
208 const gap = ' '.repeat(INDENT)
209 return (
210 <Box flexDirection="column" gap={1}>
211 {view.live ? liveBlocks(p, view.live, view.scheme ?? 'dark', view.columns, true) : null}
212 <Text bold>Across sessions</Text>
213 {view.blocks.length === 0 ? <Text>No delegation measured yet.</Text> : null}
214 {view.blocks.map(b => (
215 <Box flexDirection="column">
216 <Text bold>{b.heading}</Text>
217 <Text wrap="truncate-end">{gap + b.number}</Text>
218 {b.detail.map(d => (
219 <Text wrap="truncate-end">{gap + d}</Text>
220 ))}
221 {spark(p, b, view.columns)}
222 </Box>
223 ))}
224 <Box flexDirection="column">
225 <Text bold>Open items</Text>
226 {view.open.length === 0 ? <Text>{gap}nothing open</Text> : view.open.map(o => <Text wrap="wrap">{gap + o}</Text>)}
227 </Box>
228 <Box flexDirection="row">
229 <Button key="close" label="close" role="dismiss" onPress={onClose} />
230 </Box>
231 </Box>
232 )
233}
234hooks/lib/live.ts 484 lines1// GH-112: the live delegation model behind the band and the pane. Pure: no `$`.
2// register.ts gathers the raw facts (the session cost, the attempt records, the
3// queue, `git worktree list --porcelain`) and hands them in; everything the
4// drawing shows is made here: the spend series, the events, the worktree rows,
5// spend by model, the chart geometry (an Svg source, or Raster cells).
6import type { AttemptRecord } from './attempts'
7import { taskLabel } from './attempts'
8import { base64, elapsed } from './dashboard'
9import { worktreePath } from './paths'
10
11// ---- the spend series -------------------------------------------------------------
12
13export type SpendPoint = { t: number; usd: number }
14export const SPEND_CAP = 240
15export const SAMPLE_LIVE_MS = 15_000
16export const SAMPLE_IDLE_MS = 60_000
17
18/** How often to sample the session's dollars: fast while a worker is live or queued, slow otherwise. */
19export const sampleEvery = (active: boolean): number => (active ? SAMPLE_LIVE_MS : SAMPLE_IDLE_MS)
20
21/** The series with one more sample, the last `cap` kept; a sample that is no finite number is dropped. */
22export function spendSeries(prev: readonly SpendPoint[] | undefined, point: SpendPoint, cap = SPEND_CAP): SpendPoint[] {
23 const list = Array.isArray(prev) ? prev : []
24 if (!Number.isFinite(point.usd) || !Number.isFinite(point.t)) return [...list]
25 return [...list, point].slice(-cap)
26}
27
28// ---- model families ---------------------------------------------------------------
29
30export type Family = 'haiku' | 'sonnet' | 'opus' | 'other'
31export const FAMILIES: readonly Family[] = ['haiku', 'sonnet', 'opus', 'other']
32
33/** The family of an alias or a resolved model id; anything not haiku, sonnet or opus is "other". */
34export function familyOf(model: string | undefined): Family {
35 const m = (model ?? '').toLowerCase()
36 if (m.includes('haiku')) return 'haiku'
37 if (m.includes('sonnet')) return 'sonnet'
38 if (m.includes('opus')) return 'opus'
39 return 'other'
40}
41
42const recordFamily = (r: Pick<AttemptRecord, 'alias' | 'resolvedModel'>): Family => {
43 const f = familyOf(r.alias)
44 return f === 'other' ? familyOf(r.resolvedModel) : f
45}
46
47const round4 = (n: number) => Math.round(n * 10000) / 10000
48
49export type ModelSpend = { family: Family; usd: number; tokens: number; verified: number }
50
51/** Dollars and tokens per family from the records' own spend, and how many cards each verified. */
52export function spendByModel(records: readonly AttemptRecord[]): ModelSpend[] {
53 const out = new Map<Family, { usd: number; tokens: number; cards: Set<string> }>(FAMILIES.map(f => [f, { usd: 0, tokens: 0, cards: new Set<string>() }]))
54 for (const r of records) {
55 const row = out.get(recordFamily(r)) as { usd: number; tokens: number; cards: Set<string> }
56 row.usd += r.usd ?? 0
57 row.tokens += r.tokens ?? 0
58 if (r.verdict === 'verified') row.cards.add(taskLabel(r.task, r.subtask))
59 }
60 return FAMILIES.map(family => {
61 const row = out.get(family) as { usd: number; tokens: number; cards: Set<string> }
62 return { family, usd: round4(row.usd), tokens: row.tokens, verified: row.cards.size }
63 })
64}
65
66/** Records of this session: those that began at or after `since`. */
67export const sessionRecords = (records: readonly AttemptRecord[], since: number): AttemptRecord[] => records.filter(r => r.at >= since)
68
69// ---- events -----------------------------------------------------------------------
70
71export type LiveEvent = { t: number; kind: 'spawn' | 'verdict'; task: string; family: Family; verdict?: string }
72
73/** Spawn and verdict times this session, oldest first (a spawn before a verdict at the same time). */
74export function events(records: readonly AttemptRecord[], since: number): LiveEvent[] {
75 const out: LiveEvent[] = []
76 for (const r of records) {
77 if (r.at < since || r.kind === 'verify') continue
78 const task = taskLabel(r.task, r.subtask)
79 const family = recordFamily(r)
80 out.push({ t: r.at, kind: 'spawn', task, family })
81 if (r.verdict !== 'pending' && r.verdictAt !== undefined) out.push({ t: r.verdictAt, kind: 'verdict', task, family, verdict: r.verdict })
82 }
83 return out.sort((a, b) => a.t - b.t || (a.kind === b.kind ? 0 : a.kind === 'spawn' ? -1 : 1))
84}
85
86// ---- verdict marks and colours ----------------------------------------------------
87
88export type Scheme = 'light' | 'dark'
89
90/** Light or dark from what the render input reports of the surface; dark when it reports none. */
91export function pickScheme(e: unknown): Scheme {
92 const seen = (o: unknown): string | undefined => {
93 if (!o || typeof o !== 'object') return undefined
94 for (const k of ['colorScheme', 'theme', 'scheme', 'appearance']) {
95 const v = (o as Record<string, unknown>)[k]
96 if (typeof v === 'string') return v.toLowerCase()
97 }
98 return undefined
99 }
100 const v = seen(e) ?? seen((e as { props?: unknown } | undefined)?.props) ?? seen((e as { viewport?: unknown } | undefined)?.viewport)
101 return v?.includes('light') ? 'light' : 'dark'
102}
103
104/** A model family's chip colour; "other" has none (the surface's secondary ink). */
105export const MODEL_COLORS: Record<Exclude<Family, 'other'>, Record<Scheme, string>> = {
106 haiku: { light: '#2a78d6', dark: '#3987e5' },
107 sonnet: { light: '#eb6834', dark: '#d95926' },
108 opus: { light: '#1baf7a', dark: '#199e70' },
109}
110export const modelColor = (family: Family, scheme: Scheme): string | undefined => (family === 'other' ? undefined : MODEL_COLORS[family][scheme])
111
112export const SURFACE: Record<Scheme, string> = { light: '#fcfcfb', dark: '#1a1a19' }
113
114export type Mark = { icon: string; color: string; word: string }
115
116/** Verdict status: an icon, a colour and the word, never a colour alone. Undefined for a state that is no verdict. */
117export function verdictMark(verdict: string | undefined): Mark | undefined {
118 switch (verdict) {
119 case 'verified':
120 return { icon: '✓', color: '#0ca30c', word: 'verified' }
121 case 'unverified':
122 case 'over-spend':
123 return { icon: '!', color: '#fab219', word: verdict }
124 case 'refuted':
125 case 'no-report':
126 case 'refused':
127 return { icon: '✗', color: '#d03b3b', word: verdict }
128 default:
129 return undefined
130 }
131}
132
133// ---- the worktree rows ------------------------------------------------------------
134
135export type RowKind = 'live' | 'queued' | 'verdict' | 'owed' | 'disk'
136export type WorktreeRow = {
137 task: string
138 family: Family
139 /** The resolved model id when the record has one, else the alias. */
140 model: string
141 state: string
142 kind: RowKind
143 verdict?: string
144 /** The worktree's folder name, or an en dash when it has none on disk. */
145 folder: string
146 branch: string
147 tokens?: number
148 usd?: number
149 /** `n/b`: the latest attempt over the budget. */
150 attempt: string
151}
152
153export type WorktreeInput = {
154 root: string
155 worktreeRoot?: string
156 /** `git worktree list --porcelain`. */
157 porcelain: string
158 /** This session's attempt records, all tasks. */
159 records: readonly AttemptRecord[]
160 /** Queued spawns, front first. */
161 queue: readonly { task: string; subtask?: string }[]
162 /** Task label to the ms since the epoch its live worker started or resumed. */
163 liveAt: Readonly<Record<string, number>>
164 now: number
165 /** The attempt budget shown after the slash. */
166 budget: number
167}
168
169/** The worktrees `git worktree list --porcelain` names: path and branch. */
170export function parsePorcelain(text: string): { path: string; branch: string }[] {
171 const out: { path: string; branch: string }[] = []
172 for (const block of text.split(/\n\s*\n/)) {
173 let path = ''
174 let branch = ''
175 for (const line of block.split('\n')) {
176 if (line.startsWith('worktree ')) path = line.slice('worktree '.length).trim()
177 else if (line.startsWith('branch ')) branch = line.slice('branch '.length).trim().replace(/^refs\/heads\//, '')
178 }
179 if (path) out.push({ path, branch })
180 }
181 return out
182}
183
184const base = (p: string) => p.slice(p.lastIndexOf('/') + 1)
185
186/** One row per task this session, per queued task, and per worktree on disk under the mod's naming. */
187export function worktreeRows(input: WorktreeInput): WorktreeRow[] {
188 const prefix = worktreePath(input.root, '', false, input.worktreeRoot)
189 const disk = new Map<string, { path: string; branch: string }>()
190 for (const w of parsePorcelain(input.porcelain)) {
191 if (!w.path.startsWith(prefix) || w.path.length === prefix.length) continue
192 disk.set(w.path.slice(prefix.length).replace(/-replay$/, ''), w)
193 }
194 const byLabel = new Map<string, AttemptRecord[]>()
195 for (const r of [...input.records].sort((a, b) => a.at - b.at)) {
196 const label = taskLabel(r.task, r.subtask)
197 byLabel.set(label, [...(byLabel.get(label) ?? []), r])
198 }
199 const queued = input.queue.map(q => taskLabel(q.task, q.subtask ?? 'main'))
200 const rows: WorktreeRow[] = []
201 const seen = new Set<string>()
202 const labels = [...byLabel.keys(), ...queued.filter(l => !byLabel.has(l))]
203 for (const label of labels) {
204 seen.add(label)
205 const recs = byLabel.get(label) ?? []
206 const latest = recs.reduce<AttemptRecord | undefined>((a, r) => (a === undefined || r.attempt >= a.attempt ? r : a), undefined)
207 const id = (latest?.task ?? label.split('/')[0]) as string
208 const wt = disk.get(id)
209 const usds = recs.flatMap(r => (r.usd === undefined ? [] : [r.usd]))
210 const toks = recs.flatMap(r => (r.tokens === undefined ? [] : [r.tokens]))
211 const resolved = [...recs].reverse().find(r => r.resolvedModel)?.resolvedModel
212 const q = queued.indexOf(label)
213 let state: string
214 let kind: RowKind
215 let verdict: string | undefined
216 if (input.liveAt[label] !== undefined) {
217 state = `live ${elapsed(input.now - (input.liveAt[label] as number))}`
218 kind = 'live'
219 } else if (q >= 0) {
220 state = `queued #${q + 1}`
221 kind = 'queued'
222 } else if (latest && latest.verdict !== 'pending') {
223 state = latest.verdict
224 kind = 'verdict'
225 verdict = latest.verdict
226 } else {
227 state = 'verdict owed'
228 kind = 'owed'
229 }
230 rows.push({
231 task: label,
232 family: latest ? recordFamily(latest) : 'other',
233 model: latest ? (resolved ?? latest.alias) : '–',
234 state,
235 kind,
236 ...(verdict ? { verdict } : {}),
237 folder: wt ? base(wt.path) : '–',
238 branch: wt?.branch ?? '',
239 ...(toks.length > 0 ? { tokens: toks.reduce((a, b) => a + b, 0) } : {}),
240 ...(usds.length > 0 ? { usd: round4(usds.reduce((a, b) => a + b, 0)) } : {}),
241 attempt: `${latest?.attempt ?? 0}/${input.budget}`,
242 })
243 }
244 for (const [id, wt] of disk) {
245 if (seen.has(id) || [...seen].some(l => l.split('/')[0] === id)) continue
246 rows.push({ task: id, family: 'other', model: '–', state: 'on disk', kind: 'disk', folder: base(wt.path), branch: wt.branch, attempt: `0/${input.budget}` })
247 }
248 const rank: Record<RowKind, number> = { live: 0, owed: 1, queued: 2, verdict: 3, disk: 4 }
249 return rows.map((r, i) => ({ r, i })).sort((a, b) => rank[a.r.kind] - rank[b.r.kind] || a.i - b.i).map(x => x.r)
250}
251
252// ---- the view both drawings read --------------------------------------------------
253
254export type LiveView = {
255 now: number
256 live: number
257 queued: number
258 owed: number
259 usd: number
260 series: SpendPoint[]
261 events: LiveEvent[]
262 rows: WorktreeRow[]
263 byModel: ModelSpend[]
264 firstTry: { n: number; m: number }
265}
266
267/** Cards judged on their first attempt this session: how many verified. */
268export function firstTry(records: readonly AttemptRecord[]): { n: number; m: number } {
269 const first = new Map<string, AttemptRecord>()
270 for (const r of records) if (r.attempt === 1 && r.kind !== 'verify') first.set(taskLabel(r.task, r.subtask), r)
271 const judged = [...first.values()].filter(r => r.verdict !== 'pending')
272 return { n: judged.filter(r => r.verdict === 'verified').length, m: judged.length }
273}
274
275export type ViewInput = WorktreeInput & { since: number; usd: number; series: readonly SpendPoint[]; owed: number }
276
277export function liveView(i: ViewInput): LiveView {
278 const records = sessionRecords(i.records, i.since)
279 return {
280 now: i.now,
281 live: Object.keys(i.liveAt).length,
282 queued: i.queue.length,
283 owed: i.owed,
284 usd: i.usd,
285 series: [...i.series],
286 events: events(records, i.since),
287 rows: worktreeRows({ ...i, records }),
288 byModel: spendByModel(records),
289 firstTry: firstTry(records),
290 }
291}
292
293/** The band shows while a worker is live, a spawn is queued or a verdict is owed. */
294export const isActive = (v: Pick<LiveView, 'live' | 'queued' | 'owed'>): boolean => v.live + v.queued + v.owed > 0
295
296// ---- the band's line --------------------------------------------------------------
297
298const dollars = (n: number) => `$${n.toFixed(2)}`
299
300export function bandSummary(v: LiveView): string {
301 return ['delegation', `${v.live} live`, `${v.queued} queued`, ...(v.owed > 0 ? [`${v.owed} owed`] : []), dollars(v.usd)].join(' · ')
302}
303
304/** The last 30 minutes of cumulative dollars, thinned to at most `n` values: the band's sparkline. */
305export function recentSpend(series: readonly SpendPoint[], now: number, windowMs = 30 * 60_000, n = 14): number[] {
306 const inside = series.filter(p => p.t >= now - windowMs)
307 if (inside.length <= n) return inside.map(p => p.usd)
308 return Array.from({ length: n }, (_, i) => (inside[Math.round((i * (inside.length - 1)) / (n - 1))] as SpendPoint).usd)
309}
310
311/** `1.2k`, `412k`, `3.4M`. */
312export function tokText(n: number): string {
313 if (n >= 1_000_000) return `${+(n / 1_000_000).toFixed(1)}M`
314 if (n >= 1000) return `${Math.round(n / 100) / 10 >= 100 ? Math.round(n / 1000) : +(n / 1000).toFixed(1)}k`
315 return String(n)
316}
317
318/** `sonnet · $3.10 · 412k tok · 4 verified`. */
319export const modelLabel = (m: ModelSpend): string => `${m.family} · ${dollars(m.usd)} · ${tokText(m.tokens)} tok · ${m.verified} verified`
320
321// ---- the dashboard as text (a screen that shows no panes) --------------------------
322
323const BARS = '▁▂▃▄▅▆▇█'
324
325/** `▁▃▅█`: values scaled to eight glyph heights; empty under two values. */
326export function sparkText(values: readonly number[]): string {
327 if (values.length < 2) return ''
328 const lo = Math.min(...values)
329 const hi = Math.max(...values)
330 return values.map(v => BARS[hi === lo ? 0 : Math.round(((v - lo) / (hi - lo)) * 7)]).join('')
331}
332
333const cell = (s: string) => s.replace(/\|/g, '\\|')
334
335/**
336 * The dashboard as markdown, for a screen that shows no mod panes (the VS Code
337 * extension): the headline, the spend over the session, the worktree table and
338 * spend by model.
339 */
340export function dashboardText(v: LiveView): string {
341 const ft = v.firstTry.m > 0 ? ` · ${v.firstTry.n} of ${v.firstTry.m} verified on the first attempt` : ''
342 const lines = [`**Delegation** · ${v.live} live · ${v.queued} queued${v.owed > 0 ? ` · ${v.owed} owed` : ''} · ${dollars(v.usd)} this session${ft}`]
343 const series = v.series
344 if (series.length >= 2) {
345 const first = series[0] as SpendPoint
346 const mins = Math.max(1, Math.round((v.now - first.t) / 60_000))
347 lines.push(`Spend: ${dollars(first.usd)} → ${dollars(v.usd)} over the last ${mins} min ${sparkText(recentSpend(series, v.now, Number.MAX_SAFE_INTEGER, 24))}`)
348 }
349 lines.push('')
350 if (v.rows.length === 0) lines.push('No worktrees or tasks this session.')
351 else {
352 lines.push('| Task | Model | State | Worktree | Tokens | Cost | Attempt |', '| --- | --- | --- | --- | --- | --- | --- |')
353 for (const r of v.rows) {
354 const mark = r.verdict ? verdictMark(r.verdict) : undefined
355 const state = mark ? `${mark.icon} ${r.state}` : r.state
356 const where = r.folder === '–' ? '–' : `${r.folder}${r.branch ? ` · ${r.branch}` : ''}`
357 lines.push(`| ${cell(r.task)} | ${cell(r.family === 'other' ? r.model : `${r.family} (${r.model})`)} | ${cell(state)} | ${cell(where)} | ${r.tokens !== undefined ? tokText(r.tokens) : '–'} | ${r.usd !== undefined ? dollars(r.usd) : '–'} | ${r.attempt} |`)
358 }
359 }
360 if (v.byModel.length > 0) lines.push('', `By model: ${v.byModel.map(modelLabel).join('; ')}`)
361 lines.push('', 'Tokens and cost for a worker update when its run ends.')
362 return lines.join('\n')
363}
364
365// ---- the spend chart --------------------------------------------------------------
366
367type Geo = { t0: number; t1: number; max: number }
368const geo = (series: readonly SpendPoint[], now: number): Geo | undefined => {
369 if (series.length === 0) return undefined
370 const t0 = (series[0] as SpendPoint).t
371 const t1 = Math.max(now, (series.at(-1) as SpendPoint).t, t0 + 1)
372 return { t0, t1, max: Math.max(0.01, ...series.map(p => p.usd)) }
373}
374
375const esc = (s: string) => s.replace(/&/g, '&').replace(/</g, '<').replace(/>/g, '>')
376const r1 = (n: number) => Math.round(n * 10) / 10
377const agoText = (ms: number) => (ms < 60_000 ? 'now' : `${Math.round(ms / 60_000)}m ago`)
378
379/** The nearest event to `t`, within `within` ms. */
380export function nearestEvent(evs: readonly LiveEvent[], t: number, within = 30_000): LiveEvent | undefined {
381 let best: LiveEvent | undefined
382 for (const e of evs) if (Math.abs(e.t - t) <= within && (best === undefined || Math.abs(e.t - t) < Math.abs(best.t - t))) best = e
383 return best
384}
385
386export const CHART_HEIGHT = 96
387const TICKS = 10
388
389/**
390 * The Svg: cumulative session dollars as a 2 px line over a light area in the
391 * surface's text ink (CanvasText: the total is no model's colour), 1 px event
392 * ticks on the time axis (short for a spawn, long for a verdict), no second y
393 * axis. Interactive: a crosshair and a tooltip (time, dollars so far, the
394 * nearest event) at the pointer, by `:hover` and `<title>`.
395 */
396export function chartSvg(series: readonly SpendPoint[], evs: readonly LiveEvent[], now: number, columns: number): { source: string; width: number; height: number } | undefined {
397 const g = geo(series, now)
398 if (!g) return undefined
399 const width = Math.max(8, Math.min(columns, 120)) * 8
400 const height = CHART_HEIGHT
401 const top = 6
402 const base = height - TICKS - 2
403 const x = (t: number) => r1(1 + ((t - g.t0) / (g.t1 - g.t0)) * (width - 2))
404 const y = (usd: number) => r1(base - (usd / g.max) * (base - top))
405 const pts = series.map(p => `${x(p.t)},${y(p.usd)}`)
406 const line = series.length === 1 ? `1,${y((series[0] as SpendPoint).usd)} ${width - 1},${y((series[0] as SpendPoint).usd)}` : pts.join(' ')
407 const area = `M1,${base} L${line.replace(/ /g, ' L')} L${series.length === 1 ? width - 1 : x((series.at(-1) as SpendPoint).t)},${base} Z`
408 const ticks = evs
409 .filter(e => e.t >= g.t0 && e.t <= g.t1)
410 .map(e => `<line x1="${x(e.t)}" x2="${x(e.t)}" y1="${base}" y2="${base + (e.kind === 'verdict' ? TICKS : TICKS / 2)}" stroke="currentColor" stroke-width="1"/>`)
411 .join('')
412 const step = series.length > 1 ? (width - 2) / (series.length - 1) : width
413 const hits = series
414 .map((p, i) => {
415 const ev = nearestEvent(evs, p.t)
416 const tip = `${agoText(now - p.t)} · ${dollars(p.usd)} so far${ev ? ` · ${ev.kind} ${ev.task} ${ev.family}${ev.verdict ? ` ${ev.verdict}` : ''}` : ''}`
417 return `<g class="h"><rect x="${r1(x(p.t) - step / 2)}" y="0" width="${r1(step)}" height="${height}" fill="transparent"/><line class="x" x1="${x(p.t)}" x2="${x(p.t)}" y1="${top}" y2="${base}" stroke="currentColor" stroke-width="1"/><title>${esc(tip)}</title></g>`
418 })
419 .join('')
420 const source =
421 `<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 ${width} ${height}" width="${width}" height="${height}" color="CanvasText">` +
422 '<style>svg{color-scheme:light dark}.h .x{opacity:0}.h:hover .x{opacity:.6}</style>' +
423 `<path d="${area}" fill="currentColor" fill-opacity="0.12" stroke="none"/>` +
424 `<polyline points="${line}" fill="none" stroke="currentColor" stroke-width="2" stroke-linejoin="round" stroke-linecap="round"/>` +
425 `<line x1="1" x2="${width - 1}" y1="${base}" y2="${base}" stroke="currentColor" stroke-width="1" stroke-opacity="0.4"/>` +
426 ticks +
427 hits +
428 '</svg>'
429 return { source, width, height }
430}
431
432const BLOCKS = ' ▁▂▃▄▅▆▇█'
433const DEFAULT_COLOUR = 0x01000000
434
435/**
436 * The terminal's chart: a Raster, one column per `columns`, `rows` high, the
437 * area under the line filled with block glyphs in the terminal's own ink, and
438 * one row of event ticks beneath (`┬` a spawn, `┴` a verdict). No Svg here.
439 */
440export function chartCells(series: readonly SpendPoint[], evs: readonly LiveEvent[], now: number, columns: number, rows = 4): { cells: string; columns: number; rows: number; glyphs: string[] } | undefined {
441 const g = geo(series, now)
442 if (!g) return undefined
443 const cols = Math.max(4, Math.min(columns, 120))
444 const total = rows + 1
445 const valueAt = (c: number): number => {
446 const t = g.t0 + (c / (cols - 1)) * (g.t1 - g.t0)
447 let v = 0
448 for (const p of series) if (p.t <= t) v = p.usd
449 return series.length === 1 || t < (series[0] as SpendPoint).t ? (series[0] as SpendPoint).usd : v
450 }
451 const grid: string[][] = Array.from({ length: total }, () => Array.from({ length: cols }, () => ' '))
452 for (let c = 0; c < cols; c++) {
453 const eighths = Math.round((valueAt(c) / g.max) * rows * 8)
454 for (let r = 0; r < rows; r++) {
455 const fill = Math.max(0, Math.min(8, eighths - (rows - 1 - r) * 8))
456 ;(grid[r] as string[])[c] = BLOCKS[fill] as string
457 }
458 }
459 for (const e of evs) {
460 if (e.t < g.t0 || e.t > g.t1) continue
461 const c = Math.min(cols - 1, Math.round(((e.t - g.t0) / (g.t1 - g.t0)) * (cols - 1)))
462 const cur = (grid[rows] as string[])[c]
463 ;(grid[rows] as string[])[c] = e.kind === 'verdict' || cur === '┴' ? '┴' : '┬'
464 }
465 const glyphs = grid.map(r => r.join(''))
466 const words = new Uint32Array(cols * total * 3)
467 grid.forEach((row, r) =>
468 row.forEach((ch, c) => {
469 const i = (r * cols + c) * 3
470 words[i] = ch.codePointAt(0) as number
471 words[i + 1] = DEFAULT_COLOUR
472 words[i + 2] = DEFAULT_COLOUR
473 }),
474 )
475 const bytes = new Uint8Array(words.length * 4)
476 words.forEach((w, i) => {
477 bytes[i * 4] = w & 0xff
478 bytes[i * 4 + 1] = (w >>> 8) & 0xff
479 bytes[i * 4 + 2] = (w >>> 16) & 0xff
480 bytes[i * 4 + 3] = (w >>> 24) & 0xff
481 })
482 return { cells: base64(bytes), columns: cols, rows: total, glyphs }
483}
484hooks/lib/metrics.ts 189 lines1// Part 4C: the dashboard's per-session numbers, `delegation.metrics.<sessionId>`
2// in the store, updated by the hooks that already handle each event, and the
3// six series the pane's sparklines draw over the last 14 sessions. Pure: no `$`.
4import { aliasOf } from './tier'
5
6type ByAlias = { haiku: number; sonnet: number; opus: number }
7
8export type Metrics = {
9 /** ms since the epoch when this session first counted anything. */
10 firstSeen: number
11 /** Dispatches made through the `dispatch` tool or `/dispatch` (spawned or queued). */
12 dispatches: number
13 verdicts: { verified: number; unverified: number; refuted: number; falseRefuted: number }
14 /** Briefed spawns by alias. */
15 spawns: ByAlias
16 escalations: { resume: number; respawn: number }
17 /** Background debriefs started. */
18 debriefs: number
19 evals: { run: number; failed: number }
20 /** `session.compact` hooks served (the state block attached). */
21 compactions: number
22 /** Dollars the briefed workers spent (3B), summed as their turns complete. */
23 usd: number
24 /** Verified verdicts by alias: which tiers proved enough (the Tiers sparkline). */
25 verifiedBy: ByAlias
26}
27
28export const emptyMetrics = (firstSeen: number): Metrics => ({
29 firstSeen,
30 dispatches: 0,
31 verdicts: { verified: 0, unverified: 0, refuted: 0, falseRefuted: 0 },
32 spawns: { haiku: 0, sonnet: 0, opus: 0 },
33 escalations: { resume: 0, respawn: 0 },
34 debriefs: 0,
35 evals: { run: 0, failed: 0 },
36 compactions: 0,
37 usd: 0,
38 verifiedBy: { haiku: 0, sonnet: 0, opus: 0 },
39})
40
41export type MetricEvent =
42 | { kind: 'dispatch' }
43 | { kind: 'spawn'; alias: string; respawn?: boolean }
44 | { kind: 'resume' }
45 | { kind: 'verdict'; verdict: string; alias: string; falseRefute?: boolean }
46 | { kind: 'debrief' }
47 | { kind: 'eval' }
48 | { kind: 'eval-failed' }
49 | { kind: 'compaction' }
50 | { kind: 'usd'; usd: number }
51
52const tierKey = (alias: string): keyof ByAlias | undefined => {
53 const a = aliasOf(alias)
54 return a === 'haiku' || a === 'sonnet' || a === 'opus' ? a : undefined
55}
56
57/** A stored object read back whole: fields an older release did not write start at zero. */
58export function fillMetrics(m: Partial<Metrics> | undefined, now: number): Metrics {
59 const base = emptyMetrics(m?.firstSeen ?? now)
60 if (!m) return base
61 return {
62 ...base,
63 ...m,
64 verdicts: { ...base.verdicts, ...m.verdicts },
65 spawns: { ...base.spawns, ...m.spawns },
66 escalations: { ...base.escalations, ...m.escalations },
67 evals: { ...base.evals, ...m.evals },
68 verifiedBy: { ...base.verifiedBy, ...m.verifiedBy },
69 }
70}
71
72/** One event applied to the session's metrics (created at `now` when absent). */
73export function applyMetric(prev: Partial<Metrics> | undefined, ev: MetricEvent, now: number): Metrics {
74 const m = fillMetrics(prev, now)
75 switch (ev.kind) {
76 case 'dispatch':
77 return { ...m, dispatches: m.dispatches + 1 }
78 case 'spawn': {
79 const k = tierKey(ev.alias)
80 const spawns = k ? { ...m.spawns, [k]: m.spawns[k] + 1 } : m.spawns
81 const escalations = ev.respawn ? { ...m.escalations, respawn: m.escalations.respawn + 1 } : m.escalations
82 return { ...m, spawns, escalations }
83 }
84 case 'resume':
85 return { ...m, escalations: { ...m.escalations, resume: m.escalations.resume + 1 } }
86 case 'verdict': {
87 if (ev.verdict !== 'verified' && ev.verdict !== 'unverified' && ev.verdict !== 'refuted') return m
88 const verdicts = { ...m.verdicts, [ev.verdict]: m.verdicts[ev.verdict] + 1, falseRefuted: m.verdicts.falseRefuted + (ev.falseRefute ? 1 : 0) }
89 const k = tierKey(ev.alias)
90 const verifiedBy = ev.verdict === 'verified' && k ? { ...m.verifiedBy, [k]: m.verifiedBy[k] + 1 } : m.verifiedBy
91 return { ...m, verdicts, verifiedBy }
92 }
93 case 'debrief':
94 return { ...m, debriefs: m.debriefs + 1 }
95 case 'eval':
96 return { ...m, evals: { ...m.evals, run: m.evals.run + 1 } }
97 case 'eval-failed':
98 return { ...m, evals: { ...m.evals, failed: m.evals.failed + 1 } }
99 case 'compaction':
100 return { ...m, compactions: m.compactions + 1 }
101 case 'usd':
102 return Number.isFinite(ev.usd) && ev.usd > 0 ? { ...m, usd: Math.round((m.usd + ev.usd) * 1e6) / 1e6 } : m
103 }
104}
105
106/**
107 * A refute that was false: a refuted attempt whose task later verified at
108 * the SAME report sha (the code did not change; a resume or a re-verify by
109 * hand), or verified because an amend made the difference.
110 */
111export function isFalseRefute(prev: { verdict: string; reportSha?: string } | undefined, cur: { verdict: string; reportSha?: string; amended: boolean }): boolean {
112 if (!prev || prev.verdict !== 'refuted' || cur.verdict !== 'verified') return false
113 if (cur.amended) return true
114 return prev.reportSha !== undefined && prev.reportSha !== '' && prev.reportSha === cur.reportSha
115}
116
117/** `delegation.metricsIndex`: every session the metrics have counted, oldest first. */
118export type SessionIndex = { id: string; firstSeen: number }[]
119
120const INDEX_CAP = 200
121
122/** Notes a session once, in first-seen order (the oldest dropped past the cap). */
123export function noteSession(idx: SessionIndex, id: string, firstSeen: number): SessionIndex {
124 if (idx.some(s => s.id === id)) return idx
125 return [...idx, { id, firstSeen }].sort((a, b) => a.firstSeen - b.firstSeen).slice(-INDEX_CAP)
126}
127
128export const SPARK_SESSIONS = 14
129
130/** The ids of the last 14 sessions by first-seen time, oldest first. */
131export const lastSessions = (idx: SessionIndex, n = SPARK_SESSIONS): string[] =>
132 [...idx].sort((a, b) => a.firstSeen - b.firstSeen).slice(-n).map(s => s.id)
133
134const CHEAP_FIRST: (keyof ByAlias)[] = ['haiku', 'sonnet', 'opus']
135
136/**
137 * The six blocks' series, one value per session (`null` where the session
138 * has nothing to measure, so the line does not read a zero into it).
139 *
140 * Tiers: the share of the session's briefed spawns that went to the cheapest
141 * tier any of its verdicts verified on (all spawns on sonnet, sonnet verified:
142 * 1; half on opus while sonnet verified: 0.5).
143 */
144export const SERIES: readonly { key: string; value: (m: Metrics) => number | null }[] = [
145 { key: 'steps', value: m => m.dispatches },
146 {
147 key: 'verdicts',
148 value: m => {
149 const total = m.verdicts.verified + m.verdicts.unverified + m.verdicts.refuted
150 return total > 0 ? m.verdicts.refuted / total : null
151 },
152 },
153 {
154 key: 'tiers',
155 value: m => {
156 const cheapest = CHEAP_FIRST.find(k => m.verifiedBy[k] > 0)
157 const total = m.spawns.haiku + m.spawns.sonnet + m.spawns.opus
158 return cheapest && total > 0 ? m.spawns[cheapest] / total : null
159 },
160 },
161 { key: 'owed', value: m => m.debriefs + m.evals.run },
162 { key: 'compaction', value: m => m.compactions },
163 { key: 'spend', value: m => m.usd },
164]
165
166/**
167 * GH-112: this session's numbers read from its attempt records (the dashboard's
168 * "Across sessions" blocks draw the current session from what the store holds).
169 */
170export function metricsFromRecords(
171 records: readonly { task: string; subtask: string; alias: string; kind: string; verdict: string; usd?: number }[],
172 firstSeen: number,
173): Metrics {
174 const m = emptyMetrics(firstSeen)
175 m.dispatches = new Set(records.map(r => `${r.task}/${r.subtask}`)).size
176 for (const r of records) {
177 const a = tierKey(r.alias)
178 if (r.kind === 'spawn' && a) m.spawns[a] += 1
179 if (r.kind === 'resume') m.escalations.resume += 1
180 m.usd += r.usd ?? 0
181 if (r.verdict === 'verified') {
182 m.verdicts.verified += 1
183 if (a) m.verifiedBy[a] += 1
184 } else if (r.verdict === 'unverified') m.verdicts.unverified += 1
185 else if (r.verdict === 'refuted') m.verdicts.refuted += 1
186 }
187 return m
188}
189hooks/lib/allow.ts 165 lines1// The host-command allowlist. HARD RULE: `$.process.run` runs with no
2// permission prompt, so every argv the mod would run passes `checkArgv` first
3// and a refused argv never runs. Pure: no `$`.
4//
5// Allowed, and nothing else (part 5A: no chassis script is on the list):
6// git [-C <dir>] diff|merge-base|rev-parse|status|log … (no --output, --ext-diff, --textconv; the verifier's delta is `diff --name-status -M`)
7// git [-C <dir>] worktree list [--porcelain|-v|--verbose|-z]
8// git [-C <dir>] fetch [-q] origin main (exact)
9// git -C <root> worktree add -q -b agent/<domain>/<id>[-replay] <worktree> <origin/main|7-40 hex sha>
10// (<domain>: the configured domains, by default
11// frontend|backend|ops|dispatcher|cross|shared;
12// <worktree>: <root>-<id>[-replay], or
13// <worktreeRoot>/<repo name>-<id>[-replay];
14// the literal -replay suffix on both or neither)
15// git -C <root> ls-tree --name-only <sha> agents/tasks/ (exact; /dispatch --replay)
16// git -C <root> show <sha>:agents/tasks/<id>-<name>.md (exact; /dispatch --replay)
17// gh pr list|view … (no --web)
18// claude plugin validate|test <absolute folder> (exact; a no-repo brief's gate)
19// a gate-map command, word for word, `{files}` and `{worktree}` filled by absolute paths
20// (./gates.ts holds what a map may name)
21// <id> is a dispatchable task id, <PREFIX>-<number>[letter]: BE-101, OPS-195b.
22// Notably refused: sf, curl, git push|commit|reset|checkout, any other git
23// global option (-c, --exec-path, --git-dir …), gh pr merge|create, and any
24// script or package command the gate map does not name exactly.
25
26import { fillPlaceholders, gateCommandTemplates, matchesAnyTemplate } from './gates'
27import { worktreePath } from './paths'
28
29export const DEFAULT_DOMAINS = ['frontend', 'backend', 'ops', 'dispatcher', 'cross', 'shared'] as const
30
31export type Check = { ok: true } | { ok: false; reason: string }
32
33/** What the allowlist reads of the config: the gate templates (gateTemplatesOf(gateMap)), the domains, the worktree root. */
34export type AllowConfig = { gateTemplates?: readonly (readonly string[])[]; domains?: readonly string[]; worktreeRoot?: string }
35
36const GIT_READ = ['diff', 'merge-base', 'rev-parse', 'status', 'log']
37const GIT_WRITEY_FLAGS = /^(--output(=|$)|--ext-diff$|--textconv$|-O)/
38const TASK_ID = /^[A-Z][A-Z0-9]*-\d+[a-z]?$/
39const TASK_ID_SRC = '[A-Z][A-Z0-9]*-\\d+[a-z]?'
40const SHA = /^[0-9a-f]{7,40}$/
41const SHOW_CARD = new RegExp(`^[0-9a-f]{7,40}:agents/tasks/${TASK_ID_SRC}-[A-Za-z0-9._-]+\\.md$`)
42const REPLAY = '-replay'
43
44const refuse = (reason: string): Check => ({ ok: false, reason })
45const ok: Check = { ok: true }
46
47export const isTaskId = (s: string): boolean => TASK_ID.test(s) && s.length <= 64
48export const isSha = (s: string): boolean => SHA.test(s)
49export const isBaseRef = (s: string): boolean => s === 'origin/main' || SHA.test(s)
50
51export function checkArgv(argv: readonly string[], allow: AllowConfig = {}): Check {
52 const [cmd, ...rest] = argv
53 if (cmd === undefined) return refuse('empty argv')
54 if (cmd === 'git') return checkGit(rest, allow)
55 if (cmd === 'gh') return checkGh(rest)
56 if (cmd === 'claude') return checkClaude(rest)
57 if (matchesAnyTemplate(argv, allow.gateTemplates ?? [])) return ok
58 return refuse(`${cmd} is not on the list (not git, gh, claude plugin, nor a gate-map command word for word)`)
59}
60
61function checkGit(args: readonly string[], allow: AllowConfig): Check {
62 let i = 0
63 const dirs: string[] = []
64 while (args[i] === '-C') {
65 const dir = args[i + 1]
66 if (!dir || dir.startsWith('-')) return refuse('git -C needs a directory')
67 dirs.push(dir)
68 i += 2
69 }
70 const sub = args[i]
71 if (sub === undefined) return refuse('git with no subcommand')
72 if (sub.startsWith('-')) return refuse(`git global option ${sub} is not on the list`)
73 const tail = args.slice(i + 1)
74 if (GIT_READ.includes(sub)) {
75 const bad = tail.find(a => GIT_WRITEY_FLAGS.test(a))
76 return bad ? refuse(`git ${sub} ${bad} is not on the list`) : ok
77 }
78 if (sub === 'ls-tree') {
79 const [flag, sha, path, ...more] = tail
80 const exact = dirs.length === 1 && flag === '--name-only' && SHA.test(sha ?? '') && path === 'agents/tasks/' && more.length === 0
81 return exact ? ok : refuse(`git ls-tree ${tail.join(' ')} is not the exact shape (-C <root> ls-tree --name-only <sha> agents/tasks/)`)
82 }
83 if (sub === 'show') {
84 const exact = dirs.length === 1 && tail.length === 1 && SHOW_CARD.test(tail[0] as string) && !(tail[0] as string).includes('..')
85 return exact ? ok : refuse(`git show ${tail.join(' ')} is not the exact shape (-C <root> show <sha>:agents/tasks/<id>-<name>.md)`)
86 }
87 if (sub === 'fetch') {
88 const exact = tail.join(' ')
89 return exact === '-q origin main' || exact === 'origin main' ? ok : refuse(`git fetch ${exact} is not the exact shape (fetch [-q] origin main)`)
90 }
91 if (sub === 'worktree') {
92 const action = tail[0]
93 if (action === 'list') {
94 const bad = tail.slice(1).find(a => !['--porcelain', '-v', '--verbose', '-z'].includes(a))
95 return bad ? refuse(`git worktree list ${bad} is not on the list`) : ok
96 }
97 if (action === 'add') return checkWorktreeAdd(dirs, tail.slice(1), allow)
98 return refuse(`git worktree ${action ?? ''} is not on the list`.trim())
99 }
100 return refuse(`git ${sub} is not on the list`)
101}
102
103function checkWorktreeAdd(dirs: readonly string[], args: readonly string[], allow: AllowConfig): Check {
104 if (dirs.length !== 1) return refuse('git worktree add needs exactly one -C <root>')
105 const root = (dirs[0] as string).replace(/\/+$/, '')
106 if (args.length !== 5 || args[0] !== '-q' || args[1] !== '-b') {
107 return refuse('git worktree add must be exactly: -q -b agent/<domain>/<id> <worktree> <base>')
108 }
109 const [, , branch, path, base] = args as [string, string, string, string, string]
110 const domains: readonly string[] = allow.domains && allow.domains.length > 0 ? allow.domains : DEFAULT_DOMAINS
111 const m = new RegExp(`^agent/([A-Za-z0-9_-]+)/(${TASK_ID_SRC})(${REPLAY})?$`).exec(branch)
112 if (!m || !domains.includes(m[1] as string)) return refuse(`branch ${branch} is not agent/<${domains.join('|')}>/<id>[-replay]`)
113 const want = worktreePath(root, m[2] as string, m[3] !== undefined, allow.worktreeRoot)
114 if (path !== want) return refuse(`worktree path ${path} is not ${want}`)
115 if (!isBaseRef(base)) return refuse(`base ${base} is not origin/main or a 7-40 hex sha`)
116 return ok
117}
118
119function checkGh(args: readonly string[]): Check {
120 if (args[0] !== 'pr' || (args[1] !== 'list' && args[1] !== 'view')) return refuse(`gh ${args.slice(0, 2).join(' ')} is not on the list`)
121 const bad = args.find(a => a === '--web' || a === '-w')
122 return bad ? refuse(`gh pr ${args[1]} ${bad} is not on the list`) : ok
123}
124
125/** An absolute folder with no `..` segment and nothing a shell would read. */
126const ABS_FOLDER = /^\/[A-Za-z0-9._@%+=:,\/-]*$/
127
128/** `claude plugin validate <folder>` and `claude plugin test <folder>`, exactly, the folder absolute. */
129function checkClaude(args: readonly string[]): Check {
130 const [noun, verb, folder, ...more] = args
131 const shape = 'claude plugin validate|test <absolute folder>'
132 if (noun !== 'plugin' || (verb !== 'validate' && verb !== 'test')) return refuse(`claude ${args.slice(0, 2).join(' ')} is not on the list (only ${shape})`)
133 if (folder === undefined || more.length > 0) return refuse(`claude plugin ${verb} must be exactly: ${shape}`)
134 if (!ABS_FOLDER.test(folder) || folder.split('/').includes('..')) return refuse(`claude plugin ${verb} ${folder} is not an absolute folder`)
135 return ok
136}
137
138/** Whether a gate-map command may stand in the map at all (./gates.ts holds the rules). */
139export function checkGateCommand(cmd: string): Check {
140 const t = gateCommandTemplates(cmd)
141 return 'reason' in t ? refuse(t.reason) : ok
142}
143
144export const refusedLine = (argv: readonly string[], reason: string): string =>
145 `chassis-delegation: refused argv ${JSON.stringify(argv)} (${reason})`
146
147const SAMPLE_FILES = '/sample/.delegation/T-1/files.txt'
148const SAMPLE_TREE = '/sample/worktree'
149
150/**
151 * A gate-map entry at config load: the shape rules, then every part of it with
152 * `{files}` and `{worktree}` filled by a sample absolute path through the
153 * allowlist itself, so what would be refused at verify time is refused now.
154 */
155export function checkGateEntry(cmd: string): Check {
156 const t = gateCommandTemplates(cmd)
157 if ('reason' in t) return refuse(t.reason)
158 for (const words of t.templates) {
159 const argv = fillPlaceholders(words, SAMPLE_FILES, SAMPLE_TREE)
160 const c = checkArgv(argv, { gateTemplates: t.templates })
161 if (!c.ok) return refuse(`${c.reason.includes('not an absolute folder') ? `${c.reason}; write {worktree}` : c.reason}`)
162 }
163 return ok
164}
165hooks/lib/attempts.ts 122 lines1// Attempt records: what the mod keeps per task in `$.store` under
2// `delegation.tasks.<task-id>` (an array, appended per attempt). The eval's
3// data source. Pure: no `$`.
4import type { Tier, TierSource } from './tier'
5
6/**
7 * `work-present` (GH-104): a respawn or an auto-resume found finished work in
8 * the worker's worktree (commits ahead of the base, a clean tree) and did not
9 * spawn; `/dispatch <id> --verify <sha>` judges it, and the record takes that
10 * verdict.
11 */
12/**
13 * `over-spend` (GH-106): the worker's own cost reached twice its brief's
14 * `spend=` before it handed back; work may be present in its worktree. Not a
15 * failing verdict: it never escalates the tier.
16 */
17export type AttemptVerdict = 'pending' | 'verified' | 'unverified' | 'refuted' | 'no-report' | 'refused' | 'work-present' | 'over-spend'
18/** `verify` (GH-104): no worker ran; the attempt is the work found in the worktree, judged by `--verify`. */
19export type AttemptKind = 'spawn' | 'resume' | 'verify'
20
21export type AttemptRecord = {
22 task: string
23 subtask: string
24 /** 1-based, counted per task + subtask across spawns and resumes. */
25 attempt: number
26 kind: AttemptKind
27 /** The attempt number of the spawn this attempt descends from (a resume keeps its spawn's). */
28 lineage: number
29 tier: Tier
30 alias: string
31 source?: TierSource
32 resolvedModel?: string
33 verdict: AttemptVerdict
34 /** The report's own `gate=` claim. */
35 reportGate?: string
36 /**
37 * GH-104: the sha the attempt's report named (as written), or a
38 * `work-present` / `verify` attempt's branch head. A judged one is not new
39 * work: a respawn past it is not stopped by the look at the worktree.
40 */
41 sha?: string
42 /** The red evidence the report named (`red=`, as written) and read in the worker's tree (GH-20). */
43 red?: string
44 /** sha-256 of that file's bytes: a later attempt naming the same bytes is refuted on red. */
45 redHash?: string
46 /** The worker's own cost from its turn usage (GH-106); else the session's cost growth between spawn (or resume) and verdict, marked by `usdApprox`. */
47 usd?: number
48 /** True when `usd` is the session delta, not the worker's own cost. */
49 usdApprox?: true
50 tokens?: number
51 /** ms since the epoch at spawn (or resume). */
52 at: number
53 verdictAt?: number
54 agentId?: string
55 toolUseId?: string
56 briefPath?: string
57 purpose?: string
58 /** A `/dispatch --replay` run of an already-merged card, from `base`. */
59 replay?: true
60 base?: string
61}
62
63/** Which attempts count together: real runs, or the replays from one base commit. */
64export type Lane = { replay?: boolean; base?: string }
65
66const inLane = (r: AttemptRecord, lane?: Lane): boolean =>
67 Boolean(r.replay) === Boolean(lane?.replay) && (!r.replay || (r.base ?? '') === (lane?.base ?? ''))
68
69const FAILING: readonly AttemptVerdict[] = ['refuted', 'no-report']
70
71export const attemptsFor = (records: readonly AttemptRecord[], subtask: string, lane?: Lane): AttemptRecord[] =>
72 records.filter(r => r.subtask === subtask && inLane(r, lane)).sort((a, b) => a.attempt - b.attempt)
73
74export const nextAttempt = (records: readonly AttemptRecord[], subtask: string, lane?: Lane): number =>
75 attemptsFor(records, subtask, lane).reduce((n, r) => Math.max(n, r.attempt), 0) + 1
76
77export const lineageResumes = (records: readonly AttemptRecord[], subtask: string, lineage: number, lane?: Lane): number =>
78 attemptsFor(records, subtask, lane).filter(r => r.lineage === lineage && r.kind === 'resume').length
79
80export const recordFailed = (r: AttemptRecord): boolean => FAILING.includes(r.verdict) || r.reportGate === 'fail'
81
82/**
83 * GH-104: what earns a respawn the next tier: a refuted report, or a
84 * `gate=fail` the verifier confirmed (the verdict held it: verified). A
85 * no-report is a reporting defect, not a capability one, and never escalates.
86 */
87export const escalatesOn = (verdict: string, reportGate?: string): boolean => verdict === 'refuted' || (verdict === 'verified' && reportGate === 'fail')
88
89/** The tier a respawn escalates from: the last attempt's, when that attempt was refuted or confirmed gate=fail (GH-104). */
90export function escalationSource(records: readonly AttemptRecord[], subtask: string, lane?: Lane): Tier | undefined {
91 const last = attemptsFor(records, subtask, lane).at(-1)
92 return last && escalatesOn(last.verdict, last.reportGate) ? last.tier : undefined
93}
94
95/**
96 * GH-104: the tier a respawn after a no-report is held at: that attempt's
97 * own, never one up (and never below it, should the brief's tier be lower).
98 */
99export function holdSource(records: readonly AttemptRecord[], subtask: string, lane?: Lane): Tier | undefined {
100 const last = attemptsFor(records, subtask, lane).at(-1)
101 return last && last.verdict === 'no-report' ? last.tier : undefined
102}
103
104/** GH-104: the shas the verifier already judged for this task + subtask (pending and work-present attempts are not judged). */
105export const judgedShas = (records: readonly AttemptRecord[], subtask: string, lane?: Lane): string[] =>
106 attemptsFor(records, subtask, lane).flatMap(r => (r.sha && r.verdict !== 'pending' && r.verdict !== 'work-present' ? [r.sha] : []))
107
108/** The red hashes of the attempts before `attempt` (same subtask and lane): what a fresh red file must differ from. */
109export const priorRedHashes = (records: readonly AttemptRecord[], subtask: string, attempt: number, lane?: Lane): { attempt: number; hash: string }[] =>
110 attemptsFor(records, subtask, lane).flatMap(r => (r.attempt < attempt && r.redHash ? [{ attempt: r.attempt, hash: r.redHash }] : []))
111
112/** Replaces the record with the same subtask + attempt by `patch` applied to it. */
113export function patchRecord(records: readonly AttemptRecord[], subtask: string, attempt: number, patch: Partial<AttemptRecord>, lane?: Lane): AttemptRecord[] {
114 return records.map(r => (r.subtask === subtask && r.attempt === attempt && inLane(r, lane) ? { ...r, ...patch } : r))
115}
116
117/** The fields a replay attempt carries. */
118export const laneFields = (lane?: Lane): Pick<AttemptRecord, 'replay' | 'base'> =>
119 lane?.replay ? { replay: true, ...(lane.base ? { base: lane.base } : {}) } : {}
120
121export const taskLabel = (task: string, subtask: string): string => (subtask === 'main' ? task : `${task}/${subtask}`)
122hooks/lib/brief.ts 373 lines1// Brief header + amend parsing and report-block extraction. Pure: no `$`.
2//
3// The grammar is chassis.brief.v1 and chassis.report.v1 (the README's
4// "Contracts"; the same grammar the original shell verifier read): a value runs to
5// the next RECOGNIZED field marker, so a value may hold spaces (red_test=cd x
6// && npx vitest run y); a value written in double quotes runs to its closing
7// quote, so field-looking text inside the quotes stays in the value.
8
9import { globMatches } from './verify-native'
10
11export const BRIEF_FIELDS = [
12 'v', 'task', 'subtask', 'purpose', 'tier', 'model', 'scope', 'forbid', 'scope_globs', 'forbid_globs',
13 'red_test', 'gate', 'spend', 'budget', 'report', 'repo', 'base', 'ignore', 'context', 'note',
14] as const
15
16/** `red=` (GH-20): the red test's failing output, a path relative to the worker's tree, or `none`. */
17export const REPORT_FIELDS = ['v', 'task', 'subtask', 'branch', 'pr', 'sha', 'gate', 'red', 'files', 'tokens', 'note'] as const
18
19export type BriefHeader = {
20 /** The header as written, `[[brief` to `]]`. */
21 raw: string
22 fields: Record<string, string>
23 task?: string
24 subtask: string
25 purpose: string
26 tier?: string
27 model?: string
28 gate?: string
29 budget?: string
30 /** GH-106: the per-attempt spend ceiling in dollars, as written. */
31 spend?: string
32 repo?: string
33}
34
35export type Report = Partial<Record<(typeof REPORT_FIELDS)[number], string>>
36
37export type Amend = { ops: string[]; reason: string }
38export type AmendParse = { amends: Amend[]; malformed?: string }
39
40const escapeRe = (s: string) => s.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')
41
42/**
43 * Splits `k=v k2=v two words k3="quoted v"` into fields. A field starts at a
44 * `key=` that follows whitespace (or the start); an unquoted value ends at the
45 * next whitespace + known key + `=`; the LAST field (`note=` by contract) runs
46 * to the end.
47 */
48export function parseFields(body: string, known: readonly string[]): Record<string, string> {
49 const out: Record<string, string> = {}
50 const keys = [...known].sort((a, b) => b.length - a.length).map(escapeRe).join('|')
51 const marker = new RegExp(`(?:^|\\s)(${keys})=`, 'g')
52 const lead = /^\s*([A-Za-z_][A-Za-z0-9_]*)=/
53 let rest = body
54 while (rest.length > 0) {
55 const m = lead.exec(rest)
56 if (!m) break
57 const key = m[1] as string
58 let after = rest.slice(m[0].length)
59 let value: string
60 const quoted = after.startsWith('"') ? /^"([\s\S]*?)"(?=\s|$)/.exec(after) : null
61 if (quoted) {
62 value = quoted[1] as string
63 after = after.slice(quoted[0].length)
64 } else if (key === 'note') {
65 value = after
66 after = ''
67 } else {
68 marker.lastIndex = 0
69 const next = marker.exec(after)
70 const cut = next ? next.index : after.length
71 value = after.slice(0, cut)
72 after = after.slice(cut)
73 }
74 if (!(key in out)) out[key] = value.trim()
75 rest = after
76 }
77 return out
78}
79
80/** The first `[[brief …]]` header in the text: it ends at the first `]]` that closes a line. */
81export function parseHeader(text: string): BriefHeader | undefined {
82 const m = /\[\[brief(\s[\s\S]*?)\]\](?=[ \t]*(?:\r?\n|$))/.exec(text)
83 if (!m) return undefined
84 const fields = parseFields((m[1] as string).replace(/\s+/g, ' ').trim(), BRIEF_FIELDS)
85 const opt = (k: string) => (fields[k] !== undefined && fields[k] !== '' ? fields[k] : undefined)
86 return {
87 raw: m[0],
88 fields,
89 task: opt('task'),
90 subtask: opt('subtask') ?? 'main',
91 purpose: opt('purpose') ?? 'build',
92 tier: opt('tier'),
93 model: opt('model'),
94 gate: opt('gate'),
95 budget: opt('budget'),
96 spend: opt('spend'),
97 repo: opt('repo'),
98 }
99}
100
101/**
102 * What a header still lacks to stand as its own brief (GH-6): a task, a
103 * scope (`scope=` or `scope_globs=`) and a gate (`gate=` or `repo=none`).
104 * Empty when it is complete.
105 */
106export function inlineHeaderMissing(h: BriefHeader): string[] {
107 const has = (k: string) => (h.fields[k] ?? '').trim() !== ''
108 return [
109 ...(h.task ? [] : ['task=']),
110 ...(has('scope') || has('scope_globs') ? [] : ['scope= (or scope_globs=)']),
111 ...(h.gate || h.repo === 'none' ? [] : ['gate= (or repo=none)']),
112 ]
113}
114
115/** The unverified note of a spawn that has no brief file; `why` says what kept an inline header from becoming one. */
116export const noBriefLine = (why?: string): string => `note: no brief file named in the prompt${why ? `; ${why}` : ''}; verify skipped`
117
118/** `why` for an incomplete inline header: the fields it lacks. */
119export const lacksLine = (missing: readonly string[]): string => `the inline header lacks ${missing.join(' and ')}`
120
121/** The first absolute `….brief.md` path the text names. */
122export function findBriefPath(text: string): string | undefined {
123 const m = /(?:^|[\s"'`(<[])(\/[^\s"'`<>()[\]]*?\.brief\.md)(?![A-Za-z0-9_-])/.exec(text)
124 return m ? (m[1] as string) : undefined
125}
126
127/**
128 * `[[amend v=1 scope+=… scope-=… forbid+=… forbid-=… ignore+=… reason=…]]`
129 * blocks, the verifier's grammar (plus GH-16's `ignore+=`, globs a repo=here
130 * delta drops): line-anchored, applied in file order; one bad block makes the
131 * whole set malformed (the verifier then applies none).
132 */
133export function parseAmends(text: string): AmendParse {
134 const blocks: string[] = []
135 let buf: string | undefined
136 for (const rawLine of text.split(/\r?\n/)) {
137 const line = rawLine.trim()
138 if (buf !== undefined) {
139 buf += ' ' + line
140 if (line.includes(']]')) {
141 if (!line.endsWith(']]')) return { amends: [], malformed: 'a block has trailing text after its closing brackets' }
142 blocks.push(buf)
143 buf = undefined
144 }
145 continue
146 }
147 if (line.startsWith('[[amend')) {
148 if (line.includes(']]')) {
149 if (!line.endsWith(']]')) return { amends: [], malformed: 'a block has trailing text after its closing brackets' }
150 blocks.push(line)
151 } else buf = line
152 }
153 }
154 if (buf !== undefined) return { amends: [], malformed: 'an amend block was opened and never closed' }
155 const amends: Amend[] = []
156 for (const block of blocks) {
157 const body = block.slice('[[amend'.length, -2).replace(/\t/g, ' ')
158 const at = body.indexOf(' reason=')
159 const reason = at >= 0 ? body.slice(at + ' reason='.length).trim() : ''
160 const ahead = at >= 0 ? body.slice(0, at) : body
161 let version = false
162 const ops: string[] = []
163 for (const tok of ahead.split(/\s+/).filter(Boolean)) {
164 if (tok === 'v=1') version = true
165 else if (tok.startsWith('v=')) return { amends: [], malformed: `unknown amend version '${tok}'` }
166 else if (/^(scope|forbid)[+-]=/.test(tok) || tok.startsWith('ignore+=')) {
167 if (tok.split('=').slice(1).join('=') === '') return { amends: [], malformed: `empty value in operation '${tok}'` }
168 ops.push(tok)
169 } else return { amends: [], malformed: `unrecognized token '${tok}' (allowed: v, scope+=, scope-=, forbid+=, forbid-=, ignore+=, reason)` }
170 }
171 if (!version) return { amends: [], malformed: "a block is missing 'v=1'" }
172 if (ops.length === 0) return { amends: [], malformed: 'a block carries no scope/forbid/ignore operation' }
173 if (!reason) return { amends: [], malformed: "a block is missing a non-empty 'reason='" }
174 amends.push({ ops, reason })
175 }
176 return { amends }
177}
178
179/** One `[[amend v=1 …]]` block from a worker's hand-back, as the brief will carry it. */
180export type HandbackAmend = Amend & { block: string }
181
182/**
183 * The amend blocks a hand-back carries, in order, read the way the verifier
184 * reads a brief: a block starts its own line (leading whitespace allowed) and
185 * may wrap until its closing `]]`; a block quoted mid-line is prose, never an
186 * amend. A wrapped block is joined with single spaces, as the verifier
187 * normalizes it. Each block is checked alone with `parseAmends`: a malformed
188 * one is set aside with its reason (appending it would make the verifier
189 * discard every amendment in the brief). A block repeated in the text is
190 * taken once.
191 */
192export function extractAmendBlocks(text: string): { amends: HandbackAmend[]; malformed: { block: string; why: string }[] } {
193 const raw: string[] = []
194 const malformed: { block: string; why: string }[] = []
195 let buf: string | undefined
196 for (const rawLine of text.split(/\r?\n/)) {
197 const line = rawLine.trim()
198 if (buf !== undefined) {
199 buf += ' ' + line
200 if (line.includes(']]')) {
201 if (line.endsWith(']]')) raw.push(buf)
202 else malformed.push({ block: buf, why: 'a block has trailing text after its closing brackets' })
203 buf = undefined
204 }
205 continue
206 }
207 if (!line.startsWith('[[amend')) continue
208 if (!line.includes(']]')) buf = line
209 else if (line.endsWith(']]')) raw.push(line)
210 else malformed.push({ block: line, why: 'a block has trailing text after its closing brackets' })
211 }
212 if (buf !== undefined) malformed.push({ block: buf, why: 'an amend block was opened and never closed' })
213 const amends: HandbackAmend[] = []
214 for (const block of raw) {
215 if (amends.some(a => a.block === block) || malformed.some(m => m.block === block)) continue
216 const parsed = parseAmends(block)
217 const one = parsed.amends[0]
218 if (parsed.malformed !== undefined || !one) malformed.push({ block, why: parsed.malformed ?? 'not an amend block' })
219 else amends.push({ block, ops: one.ops, reason: one.reason })
220 }
221 return { amends, malformed }
222}
223
224/**
225 * True for a block that takes anything away (`scope-=` or `forbid-=`), or that
226 * hides paths from the check (`ignore+=`): it is never appended on the
227 * worker's word. `scope+=` (a path the worker had to touch) and `forbid+=`
228 * (less freedom) apply on their own.
229 */
230export const amendNeedsApproval = (a: Amend, forbid: readonly string[] = []): boolean =>
231 a.ops.some(op => /^(scope|forbid)-=|^ignore\+=/.test(op)) || amendScopeInsideForbid(a, forbid).length > 0
232
233/**
234 * The first forbid glob that entirely covers a scope entry, or undefined.
235 * Forbid wins over scope, so such an entry can never be touched. Conservative:
236 * a literal path is covered when the forbid matches it; a glob when its literal
237 * prefix (up to the first wildcard) is non-empty, the forbid matches that prefix,
238 * and the forbid ends in a wildcard or `/` (so it matches every extension of it).
239 */
240export function forbidCovering(entry: string, forbid: readonly string[]): string | undefined {
241 const wild = entry.search(/[*?[]/)
242 const prefix = wild < 0 ? entry : entry.slice(0, wild)
243 if (prefix === '') return undefined
244 return forbid.find(f => f !== '' && globMatches(prefix, f) && (wild < 0 || /[*/]$/.test(f)))
245}
246
247/** The scope entries a forbid glob entirely covers (see forbidCovering). */
248export const scopeInsideForbid = (scope: readonly string[], forbid: readonly string[]): string[] =>
249 scope.filter(e => forbidCovering(e, forbid) !== undefined)
250
251/**
252 * GH-16: the first pair of scope globs, one from each card, that overlap: one
253 * glob's literal prefix matches the other glob (forbidCovering's conservative
254 * rule, tried both ways). Undefined when no pair does. Two repo=here workers
255 * whose scopes overlap would both claim the same paths of the shared checkout.
256 */
257export function scopeOverlap(mine: readonly string[], theirs: readonly string[]): { mine: string; theirs: string } | undefined {
258 for (const a of mine) {
259 for (const b of theirs) {
260 if (a === '' || b === '') continue
261 if (forbidCovering(a, [b]) !== undefined || forbidCovering(b, [a]) !== undefined) return { mine: a, theirs: b }
262 }
263 }
264 return undefined
265}
266
267/** Each `scope+=` path of the block that lies inside a forbid glob, with that glob. */
268export function amendScopeInsideForbid(a: Amend, forbid: readonly string[]): { path: string; forbid: string }[] {
269 const out: { path: string; forbid: string }[] = []
270 for (const op of a.ops) {
271 if (!op.startsWith('scope+=')) continue
272 for (const path of op.slice('scope+='.length).split(',').filter(Boolean)) {
273 const f = forbidCovering(path, forbid)
274 if (f !== undefined) out.push({ path, forbid: f })
275 }
276 }
277 return out
278}
279
280/** The verdict row for a `scope+=` held back because it lies inside a forbid. */
281export const amendInsideForbidLine = (p: { path: string; forbid: string }): string =>
282 `amend needs approval: scope+=${p.path} lies inside forbid ${p.forbid}; scope+= alone does nothing, shrink the forbid`
283
284/** The dispatch warning for a scope entry a forbid covers. */
285export const scopeInsideForbidWarning = (entry: string, forbid: string): string =>
286 `warning: scope entry ${entry} is inside forbid ${forbid} and can never be touched; shrink the forbid (forbid-=) to allow it`
287
288/**
289 * Appends each block to the brief on its own line, in order (append-only: the
290 * header and body are never rewritten), skipping a block the brief already
291 * holds, the brief's own blocks read the same way as the hand-back's.
292 */
293export function appendAmends(brief: string, blocks: readonly string[]): { text: string; added: string[] } {
294 const held = new Set(briefAmendBlocks(brief))
295 let text = brief
296 const added: string[] = []
297 for (const block of blocks) {
298 if (held.has(block)) continue
299 text = (text === '' || text.endsWith('\n') ? text : text + '\n') + block + '\n'
300 held.add(block)
301 added.push(block)
302 }
303 return { text, added }
304}
305
306/** Every line-anchored amend block text in a brief, malformed or not. */
307function briefAmendBlocks(brief: string): string[] {
308 const found = extractAmendBlocks(brief)
309 return [...found.amends.map(a => a.block), ...found.malformed.map(m => m.block)]
310}
311
312/** A comma list after the amendments for `key` (`scope`, `forbid` or `ignore`), entries compared literally. */
313export function effectiveList(base: string, amends: readonly Amend[], key: 'scope' | 'forbid' | 'ignore'): string {
314 let list = base.split(',').filter(Boolean)
315 for (const amend of amends) {
316 for (const op of amend.ops) {
317 const add = op.startsWith(`${key}+=`)
318 const del = op.startsWith(`${key}-=`)
319 if (!add && !del) continue
320 const items = op.slice(key.length + 2).split(',').filter(Boolean)
321 list = add ? [...list, ...items.filter(i => !list.includes(i))] : list.filter(i => !items.includes(i))
322 }
323 }
324 return list.join(',')
325}
326
327/**
328 * The LAST `[[report …]]` block of a hand-back. It ends at the first `]]`
329 * that closes a line (so a note holding `]]` mid-line stays whole), else at
330 * the first `]]`.
331 */
332export function extractReport(text: string): string | undefined {
333 const start = text.lastIndexOf('[[report')
334 if (start < 0) return undefined
335 const tail = text.slice(start)
336 const lineEnd = /^\[\[report[\s\S]*?\]\](?=[ \t]*(?:\r?\n|$))/.exec(tail)
337 if (lineEnd) return lineEnd[0]
338 const close = tail.indexOf(']]')
339 return close >= 0 ? tail.slice(0, close + 2) : undefined
340}
341
342export function parseReport(block: string): Report {
343 const body = block.replace(/^\[\[report/, '').replace(/\]\]$/, '').replace(/\s+/g, ' ').trim()
344 return parseFields(body, REPORT_FIELDS) as Report
345}
346
347/** `budget=<n>-attempts` → n (n ≥ 1); anything else (`frontier-60m`, absent) → the fallback. */
348export function parseBudget(value: string | undefined, fallback: number): number {
349 const m = /^(\d+)-attempts$/.exec((value ?? '').trim())
350 const n = m ? Number(m[1]) : NaN
351 return Number.isInteger(n) && n >= 1 ? n : fallback
352}
353
354/**
355 * GH-1 item 6: the line for a budget written but not in the grammar (a
356 * chassis `frontier-60m`, `0-attempts`), which parseBudget quietly replaces
357 * by the fallback; undefined when the budget is absent or well formed. The
358 * tier part of a `<tier>-<minutes>m` budget is never taken as the tier.
359 */
360export function budgetWarning(value: string | undefined, fallback: number): string | undefined {
361 const v = (value ?? '').trim()
362 if (v === '' || parseBudget(v, 0) >= 1) return undefined
363 return `warning: budget "${v}" is not <n>-attempts; using the default ${fallback}`
364}
365
366/** GH-106: `spend=<usd>` → dollars (0 or more); absent or off the grammar → undefined (no ceiling). */
367export function spendOf(h: Pick<BriefHeader, 'spend'> | undefined): number | undefined {
368 const v = h?.spend?.trim().replace(/^\$/, '')
369 if (v === undefined || !/^\d+(\.\d+)?$/.test(v)) return undefined
370 const n = Number(v)
371 return Number.isFinite(n) && n > 0 ? n : undefined
372}
373hooks/lib/cost.ts 129 lines1// Part 3B: what a worker cost, from the usage each of its turns reported.
2// Every `turn.complete` that carries the worker's agentId adds its TurnUsage;
3// the dollars come from a built-in price table keyed by model id prefix.
4// Pure: no `$`.
5
6export type Tokens = { in: number; out: number; cacheRead: number; cacheWrite: number }
7/** Dollars per million tokens. */
8export type Price = { in: number; out: number }
9
10/** Cache reads cost 10% of the input price; cache writes 125%. */
11export const CACHE_READ = 0.1
12export const CACHE_WRITE = 1.25
13
14const TABLE: readonly [string, Price][] = [
15 ['opus-5-5', { in: 4, out: 20 }],
16 ['opus-5', { in: 5, out: 25 }],
17 ['sonnet-5-5', { in: 2, out: 10 }],
18 ['sonnet-5', { in: 3, out: 15 }],
19 // Haiku 5.5 bills $0.50 / $2.50 for a request whose prompt passes 100K tokens. Usage
20 // arrives summed per turn (a worker's whole run), so that tier cannot be applied:
21 // for a long Haiku 5.5 run this figure is a floor.
22 ['haiku-5-5', { in: 0.1, out: 0.5 }],
23 ['haiku-4-5', { in: 1, out: 5 }],
24]
25
26/** The table, longest prefix first, so `opus-5-5` is tried before `opus-5`. */
27export const PRICES: readonly [string, Price][] = [...TABLE].sort((a, b) => b[0].length - a[0].length)
28
29/**
30 * The price of a model id: the part from `claude-` on (a provider's
31 * `us.anthropic.` prefix dropped) starts with a table prefix followed by the
32 * end, `-`, `[`, `@` or `:`. Unknown: undefined.
33 */
34export function priceFor(model: string | undefined): Price | undefined {
35 if (!model) return undefined
36 const lower = model.toLowerCase()
37 const at = lower.indexOf('claude-')
38 const id = at >= 0 ? lower.slice(at + 'claude-'.length) : lower
39 for (const [prefix, price] of PRICES) {
40 if (id === prefix || (id.startsWith(prefix) && /^[-[@:]/.test(id.slice(prefix.length)))) return price
41 }
42 return undefined
43}
44
45/** The four counts of a TurnUsage (ModelUsage), as the store keeps them. */
46export type UsageLike = { input_tokens: number; output_tokens: number; cache_read_input_tokens?: number | null; cache_creation_input_tokens?: number | null; model?: string }
47
48export const tokensOf = (u: UsageLike): Tokens => ({
49 in: u.input_tokens || 0,
50 out: u.output_tokens || 0,
51 cacheRead: u.cache_read_input_tokens || 0,
52 cacheWrite: u.cache_creation_input_tokens || 0,
53})
54
55export const usdOf = (t: Tokens, p: Price): number =>
56 (t.in * p.in + t.out * p.out + t.cacheRead * p.in * CACHE_READ + t.cacheWrite * p.in * CACHE_WRITE) / 1e6
57
58/**
59 * What an attempt has cost so far: the tokens summed over its turns, and the
60 * dollars, `null` once any turn ran on a model the table does not price
61 * (tokens only, shown `usd=?`).
62 */
63export type Spend = { tokens: Tokens; usd: number | null; turns: number; models: string[] }
64
65export function addTurn(acc: Spend | undefined, u: UsageLike): Spend {
66 const t = tokensOf(u)
67 const price = priceFor(u.model)
68 const prev: Spend = acc ?? { tokens: { in: 0, out: 0, cacheRead: 0, cacheWrite: 0 }, usd: 0, turns: 0, models: [] }
69 return {
70 tokens: {
71 in: prev.tokens.in + t.in,
72 out: prev.tokens.out + t.out,
73 cacheRead: prev.tokens.cacheRead + t.cacheRead,
74 cacheWrite: prev.tokens.cacheWrite + t.cacheWrite,
75 },
76 usd: prev.usd === null || price === undefined ? null : prev.usd + usdOf(t, price),
77 turns: prev.turns + 1,
78 models: u.model && !prev.models.includes(u.model) ? [...prev.models, u.model] : prev.models,
79 }
80}
81
82/** Four decimals, as the store keeps dollars. */
83export const round4 = (n: number): number => Math.round(n * 10000) / 10000
84
85/** `$1.75`; `$?` for tokens on an unpriced model; `$–` while nothing is measured. */
86export function usdText(s: { usd: number | null } | undefined): string {
87 if (!s) return '$–'
88 return s.usd === null ? '$?' : `$${s.usd.toFixed(2)}`
89}
90
91export const totalTokens = (t: Tokens): number => t.in + t.out + t.cacheRead + t.cacheWrite
92
93// ---- GH-106: the per-attempt spend ceiling ---------------------------------------
94
95/** Config `spendByTier` default: dollars one attempt may spend, per tier; 0 = no ceiling. */
96export const DEFAULT_SPEND_BY_TIER: Readonly<Record<'economy' | 'standard' | 'frontier', number>> = { economy: 2, standard: 6, frontier: 15 }
97
98/** The ceiling a tier gets when the card names none: the tier's entry, premium and unknown tiers the frontier's. */
99export const spendForTier = (tier: string, table: Readonly<Record<string, number>>): number =>
100 table[tier] ?? table.frontier ?? DEFAULT_SPEND_BY_TIER.frontier
101
102/** Dollars as the ceiling lines write them: `$2`, `$6.50`. */
103export const dollars = (n: number): string => `$${Number.isInteger(n) ? n : n.toFixed(2)}`
104
105/** Where a worker stands against its ceiling: under, at (warn once), or over twice it (stop). */
106export type CeilingState = 'under' | 'warn' | 'stop'
107export const ceilingState = (usd: number, spend: number): CeilingState => (usd >= 2 * spend ? 'stop' : usd >= spend ? 'warn' : 'under')
108
109/** The one message a worker gets when it first crosses its ceiling. */
110export const warnText = (usd: number, spend: number): string =>
111 `chassis-delegation: you have spent about $${usd.toFixed(2)} of a ${dollars(spend)} ceiling; wrap up now and hand back with the report line`
112
113/** `next=` of the over-spend row. */
114export const OVER_SPEND_NEXT = (task: string): string => `check the worktree (work may be present: /dispatch ${task} --verify <sha>)`
115
116/** The over-spend row: `chassis-delegation: T-1 attempt 1/3 over-spend · $4.20 of $2 · next=…`. */
117export const overSpendLine = (v: { label: string; task: string; attempt: number; budget: number; usd: number; spend: number }): string =>
118 `chassis-delegation: ${v.label} attempt ${v.attempt}/${v.budget} over-spend · $${v.usd.toFixed(2)} of ${dollars(v.spend)} · next=${OVER_SPEND_NEXT(v.task)}`
119
120/** `T-6 $3.10` — a live worker with its running cost (the label alone while nothing is measured). */
121export const liveWorker = (label: string, usd: number | null | undefined): string => (usd === undefined || usd === null ? label : `${label} $${usd.toFixed(2)}`)
122
123/** The ceiling /dispatch writes: the card's `spend:` when it is a number (0 = none), else the tier's. */
124export function spendCeiling(card: string | undefined, tier: string, table: Readonly<Record<string, number>>): number {
125 const v = card?.trim().replace(/^\$/, '')
126 if (v !== undefined && /^\d+(\.\d+)?$/.test(v)) return Number(v)
127 return spendForTier(tier, table)
128}
129hooks/lib/cleanstop.ts 127 lines1// The clean-stop detector (SPEC part 2C): when the session has gone quiet with
2// friction worth writing down, a background agent runs the /debrief skill.
3// Pure: no `$`.
4
5const H = 60 * 60 * 1000
6
7export const DEFAULT_DEBRIEF_IDLE_MINUTES = 20
8export const DEFAULT_DEBRIEF_MIN_NEW_LINES = 25
9export const DEFAULT_DEBRIEF_COOLDOWN_HOURS = 6
10export const DEFAULT_DEBRIEF_AGENT = 'general-purpose'
11
12/** Lines as `wc -l` counts them (newlines), the count the debrief skill writes into `.done`. */
13export const countLines = (text: string): number => {
14 let n = 0
15 for (let i = 0; i < text.length; i += 1) if (text.charCodeAt(i) === 10) n += 1
16 return n
17}
18
19/** The `.done` watermark: the line count the last debrief covered; absent, empty (a legacy touch) or junk is 0. */
20export function parseWatermark(text: string | undefined): number {
21 const m = /^\s*(\d+)\s*$/.exec(text ?? '')
22 return m ? Number(m[1]) : 0
23}
24
25const harness = (home: string): string => `${home.replace(/\/+$/, '')}/.claude/harness/breadcrumbs`
26export const breadcrumbPath = (home: string, sessionId: string): string => `${harness(home)}/${sessionId}.jsonl`
27export const watermarkPath = (home: string, sessionId: string): string => `${harness(home)}/${sessionId}.done`
28export const debriefSkillPath = (home: string): string => `${home.replace(/\/+$/, '')}/.claude/commands/debrief.md`
29
30export type CleanStopInput = {
31 /** A main-loop turn is running. */
32 inTurn: boolean
33 /** Agents this mod spawned (or saw spawned) that `$.agent.list()` says are running. */
34 running: number
35 /** Attempts handed back whose verdict is not in yet. */
36 pending: number
37 /** Spawns waiting in the scheduler queue (2F). */
38 queued: number
39 /** Lines in the session's breadcrumb file. */
40 lines: number
41 /** The `.done` watermark (0 when absent). */
42 watermark: number
43 minNewLines: number
44 now: number
45 /** When the last background debrief of this session started. */
46 lastAt?: number
47 cooldownMs: number
48 /** The watermark the last background debrief started at. */
49 lastWatermark?: number
50}
51
52const plural = (n: number, one: string, many: string): string => `${n} ${n === 1 ? one : many}`
53
54/**
55 * A clean stop: no turn, no worker running, no verdict pending, nothing
56 * queued, at least `minNewLines` breadcrumb lines past the watermark, the last
57 * debrief older than the cooldown, and never twice for the same watermark.
58 */
59export function cleanStop(i: CleanStopInput): { ok: true } | { ok: false; why: string } {
60 if (i.inTurn) return { ok: false, why: 'a turn is running' }
61 if (i.running > 0) return { ok: false, why: `${plural(i.running, 'worker', 'workers')} running` }
62 if (i.pending > 0) return { ok: false, why: `${plural(i.pending, 'verdict', 'verdicts')} pending` }
63 if (i.queued > 0) return { ok: false, why: `${plural(i.queued, 'spawn', 'spawns')} queued` }
64 const fresh = i.lines - i.watermark
65 if (fresh < i.minNewLines) return { ok: false, why: `only ${Math.max(0, fresh)} new breadcrumb lines (need ${i.minNewLines})` }
66 if (i.lastWatermark !== undefined && i.lastWatermark === i.watermark) return { ok: false, why: `already debriefed at watermark ${i.watermark}` }
67 if (i.lastAt !== undefined && i.now - i.lastAt < i.cooldownMs) {
68 return { ok: false, why: `last debrief ${((i.now - i.lastAt) / H).toFixed(1)}h ago (cooldown ${i.cooldownMs / H}h)` }
69 }
70 return { ok: true }
71}
72
73export const debriefPrompt = (skillPath: string, sessionId: string): string =>
74 `Run the /debrief skill exactly as written in ${skillPath}. Session id ${sessionId}. Write only what the skill allows.`
75
76/** The debrief file the agent's answer names: `…/harness/debriefs/<date>-<slug>.json` or `<root>/.delegation/debriefs/…`. */
77export function debriefPathOf(answer: string): string | undefined {
78 const m = /(?:~|\/)[^\s`'"()<>]*\/(?:harness|\.delegation)\/debriefs\/[^\s`'"()<>]+\.json/.exec(answer)
79 return m ? m[0] : undefined
80}
81
82// ---- part 5E: the debrief without the harness ------------------------------------
83
84/** Friction events past the last debrief a debrief needs, when no breadcrumb file exists. */
85export const DEFAULT_DEBRIEF_MIN_EVENTS = 5
86export const FRICTION_CAP = 100
87
88/**
89 * The instructions the debrief agent follows: the person's own
90 * `~/.claude/commands/debrief.md` when it exists, else the mod's built-in
91 * `hooks/templates/debrief.md` (the same JSON schema; it writes under `<root>/.delegation/`).
92 */
93export const debriefSource = (home: string, pluginRoot: string, skillExists: boolean): { path: string; builtIn: boolean } =>
94 skillExists ? { path: debriefSkillPath(home), builtIn: false } : { path: `${pluginRoot.replace(/\/+$/, '')}/hooks/templates/debrief.md`, builtIn: true }
95
96export const builtInDebriefPrompt = (templatePath: string, sessionId: string, root: string, facts: readonly string[]): string =>
97 [
98 `Run the debrief exactly as written in ${templatePath}.`,
99 `Session id ${sessionId}. Repo root ${root}.`,
100 facts.length > 0 ? 'What chassis-delegation saw since the last debrief:' : 'chassis-delegation saw no friction events beyond the verdicts below.',
101 ...facts.map(f => `- ${f}`),
102 'Write only what it allows.',
103 ].join('\n')
104
105/** One thing that went wrong in the session, as the mod saw it. */
106export type FrictionEvent = { kind: 'correction' | 'denial' | 'refuted'; detail: string; at: number }
107
108const CORRECTION_WORDS = ['no', 'stop', "don't", 'dont', 'actually']
109
110/** A prompt that corrects: its first word is no, stop, don't or actually, in any case. */
111export function isCorrection(text: string): boolean {
112 const first = (text.trim().split(/\s+/)[0] ?? '').toLowerCase().replace(/[\u2019]/g, "'").replace(/[.,!?:;]+$/, '')
113 return CORRECTION_WORDS.includes(first)
114}
115
116export const addFriction = (list: readonly FrictionEvent[], ev: FrictionEvent, cap = FRICTION_CAP): FrictionEvent[] => [...list, ev].slice(-cap)
117
118const KIND_TEXT: Record<FrictionEvent['kind'], string> = { correction: 'correction', denial: 'tool denied', refuted: 'refuted' }
119
120/** The facts line by line, as the built-in debrief's prompt carries them. */
121export const frictionFacts = (events: readonly FrictionEvent[]): string[] => events.map(e => `${KIND_TEXT[e.kind]}: ${e.detail.replace(/\s+/g, ' ').slice(0, 200)}`)
122
123export const debriefToast = (answer: string): string => {
124 const path = debriefPathOf(answer)
125 return path ? `debrief written: ${path}` : 'debrief finished'
126}
127hooks/lib/compaction.ts 100 lines1// Session health across compaction (SPEC part 2E): the delegation state as
2// lines, the instructions a compaction is given, and the system prompt's
3// "Delegation state" section. Pure: no `$`.
4
5export type StateSnapshot = {
6 /** Briefed workers `$.agent.list()` says are running. */
7 running: { task: string; tier: string; agentId: string; at?: number }[]
8 /** Workers that handed back and whose verdict is not in yet. */
9 pending: { task: string; agentId?: string }[]
10 /** This session's verdict lines, oldest first. */
11 recent: { line: string; at: number }[]
12 /** The scheduler queue, head first. */
13 queued: { task: string; position: number }[]
14 /** Advice the brain (or Ben) has not acted on: `<task>: <next>`. */
15 owed: string[]
16 debrief?: { at: number; agentId?: string; path?: string; finishedAt?: number }
17 eval?: { tier: string; sha: string; pass?: number; total?: number; at: number; running?: boolean; agentId?: string }
18 /** PRs the workers' reports named. */
19 prs: string[]
20}
21
22/** One `delegation.recent.<session>` row: a verdict as the one line has it, and what it leaves owed. */
23export type RecentVerdict = { task: string; attempt: number; verdict: string; line: string; at: number; owed?: string; pr?: string }
24
25/** What is owed: each task's latest verdict that left advice undone, unless the task is running, pending or queued again. */
26export function owedFrom(recent: readonly RecentVerdict[], busy: ReadonlySet<string>): string[] {
27 const latest = new Map<string, RecentVerdict>()
28 for (const r of recent) latest.set(r.task, r)
29 return [...latest.values()].filter(r => r.owed !== undefined && !busy.has(r.task)).map(r => r.owed as string)
30}
31
32/** The PRs the reports named (`pr=` other than none), each once. */
33export const prsFrom = (recent: readonly RecentVerdict[]): string[] =>
34 [...new Set(recent.map(r => r.pr ?? '').filter(pr => pr !== '' && pr !== 'none'))]
35
36export const emptySnapshot = (): StateSnapshot => ({ running: [], pending: [], recent: [], queued: [], owed: [], prs: [] })
37
38export const isEmptyState = (s: StateSnapshot): boolean => s.running.length === 0 && s.pending.length === 0 && s.queued.length === 0 && s.owed.length === 0
39
40export const SECTION_ID = 'chassis-delegation:state'
41export const HEADER = 'Delegation state (chassis-delegation):'
42export const MAX_LINES = 40
43const RECENT = 5
44
45export const KEEP_VERBATIM =
46 'KEEP VERBATIM: (1) the delegation state below; (2) the release recipe pointer (session scratchpad FOLDnn scripts); (3) the brief and report contracts `[[brief v=1 …]]` / `[[report v=1 …]]`; (4) every card held for Ben and every item owed by Ben.'
47
48/** `2026-10-03 14:00Z` */
49export const stamp = (ms: number): string => `${new Date(ms).toISOString().slice(0, 16).replace('T', ' ')}Z`
50
51/**
52 * The state, one line per fact, `HEADER` first, at most 40 lines: running
53 * workers, pending verdicts, queued spawns and what is owed come first (they
54 * are what the brain acts on), then the last 5 verdicts, then the last debrief
55 * and eval and the PRs reports named.
56 */
57export function renderState(s: StateSnapshot): string[] {
58 const live = [
59 ...s.running.map(r => `- running: ${r.task} ${r.tier} agent ${r.agentId}${r.at !== undefined ? ` since ${stamp(r.at)}` : ''}`),
60 ...s.pending.map(p => `- pending verdict: ${p.task}${p.agentId ? ` agent ${p.agentId}` : ''}`),
61 ...s.queued.map(q => `- queued: ${q.task} (position ${q.position})`),
62 ...s.owed.map(o => `- owed: ${o}`),
63 ]
64 const history = [
65 ...s.recent.slice(-RECENT).map(r => `- verdict: ${r.line}`),
66 ...(s.debrief
67 ? [`- last debrief: ${stamp(s.debrief.at)}${s.debrief.agentId ? ` agent ${s.debrief.agentId}` : ''} — ${s.debrief.path ?? (s.debrief.finishedAt ? 'finished' : 'running')}`]
68 : []),
69 ...(s.eval
70 ? [
71 s.eval.running
72 ? `- eval running: ${s.eval.tier} at ${s.eval.sha.slice(0, 8)}${s.eval.agentId ? ` agent ${s.eval.agentId}` : ''}`
73 : `- last eval: ${s.eval.tier} ${s.eval.pass ?? '?'}/${s.eval.total ?? '?'} at ${s.eval.sha.slice(0, 8)} (${stamp(s.eval.at)})`,
74 ]
75 : []),
76 ...(s.prs.length > 0 ? [`- PRs named in reports: ${s.prs.join(', ')}`] : []),
77 ]
78 const room = MAX_LINES - 1
79 // History keeps at least its verdicts when live lines would crowd it out.
80 const keepHistory = Math.min(history.length, Math.max(RECENT, room - live.length))
81 const keepLive = Math.min(live.length, room - keepHistory)
82 const liveShown = live.length > keepLive ? [...live.slice(0, keepLive - 1), `- … ${live.length - keepLive + 1} more`] : live
83 return [HEADER, ...liveShown, ...history.slice(0, keepHistory)]
84}
85
86export const appendInstructions = (existing: string | undefined, block: string): string => (existing ?? '') + '\n' + block
87
88/** What a compaction is told to keep: KEEP VERBATIM, the release recipe pointer, the state. */
89export function compactBlock(s: StateSnapshot, scratchpad?: string): string {
90 const recipe = scratchpad ? [`Release recipe: the FOLDnn scripts in ${scratchpad}`] : []
91 const state = isEmptyState(s) && s.recent.length === 0 && !s.debrief && !s.eval ? [`${HEADER} nothing running, nothing owed.`] : renderState(s)
92 return [KEEP_VERBATIM, ...recipe, ...state].join('\n')
93}
94
95/** The system prompt's section: present only while something runs, waits or is owed. */
96export function composeSection(s: StateSnapshot): { id: string; text: string; scope: 'session' } | undefined {
97 if (isEmptyState(s)) return undefined
98 return { id: SECTION_ID, text: renderState(s).join('\n'), scope: 'session' }
99}
100