Graph engineering for the harness flow. A local MCP server owns the node graph, routes execution across vendor CLIs or the session, adjudicates results against…

English · 한국어
Graph engineering for the harness flow, mediated by a local MCP server. The main session orchestrates; the graph engine owns state, routing, execution, and adjudication.
Sibling to harness, not a replacement. harness owns the six-stage reasoning contract. graph owns who executes a node, and whether the result survives contact with the worktree.
The static Workflow engine had to spawn a transport subagent per node just to drive a vendor CLI through Bash. A subagent waiting on a long-running process can only poll, and cache-read scales with turn count. Measured on one real run, same unit of work:
| shell turns in the node | cache-read tokens |
|---|---|
| 4 | 227,323 |
| 12 | 1,281,319 |
| 17 | 1,965,390 |
Across that run the transport layer cost more than the reasoning layer. An MCP call is one turn and blocks, so a node cannot poll — the failure mode is removed structurally rather than discouraged in a prompt.
Install the marketplace plugin and use graph:install to verify the connection:
/plugin install graph@newkayak12-claude-skills
trophy rides along. From this version, the first interactive session after you install or update this plugin installs trophy (achievements) once, in user scope, if you don't have it. Nothing is sent until you say yes; uninstalling trophy is respected (it is never reinstalled). To opt out beforehand:
mkdir -p ~/.claude/plugins/.newkayak12-trophy-ride.done. Needssh(Windows without one is not covered).
For a source checkout, graph:install can instead merge this direct registration:
{
"mcpServers": {
"graph-engineering": { "command": "node", "args": ["<plugin root>/mcp/broker.mjs"] }
}
}
Zero runtime dependencies, Node 18+.
Optional: graph:install can run develop:like-my-code to write .claude/conventions/; the broker then tells every plan, setgoal, implement, and test node, on any vendor, to read and follow the relevant files.
.claude/conventions/** (setgoal folds them into acceptance); install offers develop:like-my-code for a reference project/graph-live pane (Claude Code 2.1.292+, early access, interactive sessions): a read-only view of this folder's newest graph run in the teams 0.46.0 style. Tabs Flow / Nodes (keys 1-2), a stage rail, the Next / Now / Blocked line, one row per subgoal (impl → test → gate) and a gates bar; Nodes is a To do / Doing / Done board. The mod only reads the run files; files over 4 MiB are skipped. No band of its own: graph runs already show in the harness pipeline band. English by default, Korean with Claude Code's language. Mod tests 20; real Claude Code captures EN/KO.acquireLock set its deadline from Date.now(), so a forward wall-clock jump (wake from sleep, an NTP step) made a waiter throw LockTimeoutError at once instead of waiting out its timeout. The deadline now uses performance.now(); stale-lock mtime checks stay on the wall clock. The default timeout is unchanged, nothing imported from teams. Tests: graph suite 146 (from 145).ownerGone judged a claim by its owner's pid and boot, so a running claim stamped under the broker's own pid was never reclaimed, even when this process no longer held its ticket. activeNodes now maps each held node to its ticket, and an own-pid claim with a matching boot is reclaimed (abandoned) only when this process does not hold that ticket; a held claim is never reclaimed. The existing boot-id branch (a claim from another boot is reclaimed though its pid is alive) gets its first repo test, with a live-pid control that stays running and a skip where /proc/sys/kernel/random/boot_id is absent; deleting the boot comparison makes the test fail. No default changed, nothing imported from teams. Tests: graph suite 145 (from 143).writeFileSync), so a reader could see half a file, and a lock that timed out let the write go ahead unlocked. New mcp/store.mjs (graph's own; nothing imported from teams): mutateRun takes the lock, reads the run fresh, runs fn, and writes to a temp file then renames, all under one lock. A lock timeout (GRAPH_LOCK_TIMEOUT_MS, 5000) throws LockTimeoutError; a live owner's lock is never taken; a dead owner's lock is broken only under a separate <lock>.steal lock. Every broker write goes through it, and open-nodes.json is written the same way. Claim before run: graph_run claims the node with a ticket before it starts the adapter, so a second call on the same node is refused instead of running the vendor twice; a result or interruption whose ticket no longer holds is recorded result_superseded, not applied; retrying a subgoal or spec also retires its running nodes. Ported from teams (each re-checked as a defect here, each with a test that failed before): #1 a retry after a spec retry no longer waits on the dead generation; #2 report waits on live nodes; #6 changed_files claims with spaces, non-ASCII, notes, globs or renames match git; #7 a claimed file that exists but is git-ignored is not contradicted; #8 the torn read above; #9 an author stage_ok:false with no reason and passing checks goes on to be judged; #10 a rejection with no reason gets one from its checks/evidence; #11 malformed vendor JSON gets one fresh attempt. Not ported: #3 (finished ≠ delivered) only adds a settled field to runState's return, and graph already shows a settled failure in its node states (unreachable) and the report, so it adds nothing graph lacks; #4 (test goes to the non-implementing vendor) and #5 (author runs last under either allocation) are routing policy, and graph routes by the explicit or ordered allocation its user set. Also not ported, as features or policy rather than defects: daemon/taskmanager-only fixes (e332ed4 dispatchSettled/fold_deferred, 2917e2e, 5cbfb19, 1d4bb37); 0a34817 (graph has no autoReassign, and spec problems already reach graph_retry); c182b99/453ff06 (graph has no goal threshold; the rest is cards and pins); 1fccf7f (graph has no draft/review stages); 4983843, fd78c4a, e8086b7, dd610ef, 8406e45, fe1ed28, ff1e8df m1/m2/m4/m5/m8/m10/m12 and f0ee114 (teams features); 9d359b0 capacityNotice/routing_blocked_capacity (manager parking); 094de8c (a cost policy that changes the gate contract); edd29fc (requireRunnable already refuses duplicates); ecd8c81/32f05eb (--verify Bash, a feature). Race and crash repairs: the steal is race-free: <lock>.steal is a link()ed file carrying its owner, a dead holder is succeeded through <lock>.steal.<key> and never removed by name, so one link() wins (same scheme as teams, graph's own code). A throw after the claim (prompt write, adapter, outcome) releases the claim: the node goes back to pending with its pre-claim fields, ledger claim_failed. A claim records the owner's pid and boot id, and a running node is reclaimed only when that process is dead or the claim came from another boot; elapsed time is used only for old run files with no owner, so a live adapter run longer than 10 minutes keeps its node and its result applies. Tests: new test-store (24) and test-ports (13), graph suite 143.## What Claude Does / What You Do table. Documentation only; no broker change.teams/scripts/bench/) drove this engine's code through real sessions and hit three defects the unit suite never had; both stable bench runs reproduced the first. (1) A driving session that reports itself as claude-opus-5[1m] — a context variant the fresh-agent picker does not list — blocked at plan with native host cannot select model; the host's own model is now selectable by definition. (2) The execution default names a tier (sonnet) and hosts declare ids (claude-sonnet-5); the check compared strings, so every implement node could be vendor-failure with zero failed nodes, and the only way out was a second graph_open. Tiers now resolve against the declared list, and an undeclared tier runs on the host model with the substitution written into the routing reason. (3) A gate:goal that rejected the assembled result was never re-judged after the subgoal it blamed was retried — the run wedged with the fix in place; a subgoal retry now opens gate:goal:N over the live gates and moves report behind it, and the retried subgoal's briefing carries the goal gate's reason and gaps (and a failed test's checks, which a retry used to lose). Tests ported with the fixes. Bench: one run of this version's predecessor on a 4-package request completed 9/9 at 6× the cost and time of a plain session; teams' manager on the same request is a different topology (four runs plus a manager) and is read separately there.report hung behind gate:goal, and a subgoal that ran out of retries left the run blocked forever - the subgoals that HAD passed were never reported. Edges now come in two kinds: deps is a data dependency (the dep must be done) and after is order-only, Make's | prerequisite (the dep must have finished, not passed). report hangs off the goal gate with after. When graph_retry finds the budget gone it settles the failure instead of returning a dead end: every node that needed the dead node's output through a data edge becomes unreachable, transitively and with the reason; a node that had already failed downstream is final too; order-only edges do not propagate. The goal gate goes unreachable, report becomes ready, and the run ends complete with the partial account written by the report node. A plain failed with retries left settles nothing - the report cannot run ahead of a retry. A spec may give a subgoal after: [id], checked for self/dangling/cycle like deps. graph_retry returns unreachable[] and the next ready nodes when it declines; graph_status shows after and counts unreachable. Also fixed: a report rebuilt after a spec retry (report:2) was briefed without the whole run, and a critique's blocking / problems never reached the goal gate or the report.in_progress on dispatch, completed on the verdict — with the text lines kept as the fallback for hosts without one. Same small vocabulary either way (node_id, vendor/model, state, short reason), and the same rule holds: a line you cannot write from the verdict is a line you do not write.report to the peer vendor but left the driver's output template narrating the run, so a graph node's payload had nowhere to go and a driver under context pressure would quietly re-narrate. The template now ends with a ### Report section relaying that node verbatim, and the mandates say plainly that the report node writes the run's account. The loop also prints one line per node as it goes — node_id, vendor/model, state, short reason — so a long run is visible while it runs; the line is written from the verdict the driver already holds, never from an opened payload. Duplication that had regrown against references/ is gone again. QA: Usefulness/Authoring/Output-Quality PASS, MCP NONE, Weight OK; eval delta +0.42 (12/12 vs 7/12), above the +0.375 baseline, re-measured after the edits with no regression. The driver is also told plainly that the broker does not serialize concurrent implement nodes — readyNodes returns every dependency-satisfied node, so keeping two of them off one worktree is the caller's job.report routes to the vendor that did not drive the run, so a run's account of itself is not written by its own driver; it stays reasoning work (read-only sandbox, host_model if it falls back to the host). graph_status now takes no run_id: it lists every run in a directory with state, counts, the node running right now with its vendor and elapsed seconds, and the last node to finish — so a lead that lost the id, or a second operator looking in, can find the run without the transcript. A running node in the single-run view carries the same detail. First full cross-vendor E2E on 1.5.7 completed 8/8 with Codex implementing and testing under danger-full-access and changed_files_verified: true.changed_files claim against git's relative output and failed every truthful node that used the briefing's own paths. Claude gains danger-full-access (--permission-mode bypassPermissions) as a per-run opt-in, matching Codex; defaults are unchanged and still answer to the project's permission settings.isolated: true requires an actual worktree rather than a user's stated wish for one.cwd so a restarted client can find the run, and the last duplication the weight check found is gone (briefing rules and the ranking sentence live in references/ only).references/, and names its trigger phrases.skills/orchestrate/references/; the skill keeps the loop, the mandates, and the verdict contract (275 -> 186 lines).graph_retry that is going to be rejected no longer clears capacity exclusions first, and an ordinary probe failure is never laundered into a capacity exclusion.graph_retry({reset_capacity:true}) with no interrupted node to name.graph:orchestrate defaults to the active Codex or Claude session for every stage, using its native tools and current model. External vendor routing remains optional; no Codex CLI or fixed model is required.graph:orchestrate now keeps reasoning, gates, and reporting on the Claude session while requiring Codex for implement/test. The node's selected model is forwarded into both Codex readiness probes, with model-scoped probe caching, so a bad global default cannot reject a run that names a working model explicitly.graph:orchestrate: ordinary skill-driven runs now open with vendor: "auto", candidates: ["codex"], so the bundled adapter is actually used when its readiness probe passes and falls back visibly to self when unavailable. Direct callers that omit candidates retain the quiet auto behavior introduced in v1.0.1.graph_open takes model (a run-level default) and policy, a per-stage override map keyed by stage name plus the optional gate:goal — each entry may set vendor, candidates, sandbox, model. This expresses the harness contract directly: reasoning on a strong model, execution on whatever can actually write on this host. graph_next reports the chosen model per ready node, and an explicit graph_run({model}) still wins for one call. Verified by a reduced-scale codex E2E that reached report with artifacts checked independently of the run's own verdicts.vendor: "auto" no longer enrols every registered vendor as a candidate; the candidate list is empty by default, so an unnamed run degrades to self through the existing path. The codex vendor, its adapter, and the readiness probe are unchanged and still route when named (vendor: "codex") or listed in candidates.broker plugin name and split lifecycle from execution: graph:install connects/verifies it, while graph:orchestrate drives the graph. Engine code and on-disk run format remain the existing implementation.| tool | purpose |
|---|---|
graph_open | throw a raw request in; the broker builds the flow as a node graph on disk |
graph_next | ask which nodes are ready, and how each is routed |
graph_run | the routed vendor executes one node; blocks; returns a one-line verdict |
graph_submit | record a node the orchestrator executed itself; same adjudication |
graph_retry | open a fresh attempt, carrying rejection feedback — a subgoal, or the spec itself. Budget gone: settles the failure, returns unreachable[] and the now-ready report |
graph_status | compact run state; omit run_id for every run in a directory and what is running right now; full:true only for one node at a time |
The goal-spec, subgoal acceptance, upstream handoffs, prior rejection feedback, changed-file lists and evidence all stay in the graph on disk. Tools return {node_id, stage, vendor, state, stage_ok} and a short reason. Two consequences:
setgoal expands the graph inside the broker. The spec it produces never passes through the caller; the per-subgoal implement/test/gate nodes and their dependencies are derived from it server-side.graph_run does not accept a prompt — passing one is impossible on purpose.This is what makes a long loop possible. If node results accumulated in the orchestrator's context, a graph with retries would exhaust it and the loop would die before the work did.
plan -> setgoal -> critique, then per subgoal implement -> test -> gate with subgoal dependencies mapped onto gate nodes, then gate:goal:1 -> report.
Edges come in two kinds. deps is a data dependency: the node consumes what the dep produced, so the dep must be done. after is order-only — Make's | prerequisite: the node must not start before the dep has finished, but it does not need the dep to have passed. report hangs off the goal gate with after, which is what lets it write the account of a failure. A spec may give a subgoal after: ["U1"] alongside deps.
Failure becomes settled at exactly one point: when graph_retry finds the retry budget gone. Until then a failed node is a retry waiting to happen and nothing downstream is written off. Once settled, every node that needed the dead node through a data edge becomes unreachable, transitively and with the reason (unreachable: gate:U1:2 is unreachable); a downstream node that had already failed is final too; order-only edges do not propagate. The goal gate goes unreachable, report becomes ready, and the run ends complete with a report that names what shipped and what did not. Only a run with no report node at all — setgoal never produced a spec — still ends blocked.
A rejected subgoal gets a new attempt rather than a re-run node: the failed attempt stays in the graph as evidence, its still-pending nodes are retired as skipped, and anything that waited on the old gate is rewired to the new one.
When critique rejects the spec, retrying one subgoal fixes nothing — the decomposition itself is in question. graph_retry with no subgoal_id reopens setgoal and critique with the critique's problems as feedback and retires the subgoal graph the rejected spec produced; the rebuilt graph hangs off the live critique node.
A judging node can fail with stage_ok: true. There, stage_ok means "the judging itself worked" and the verdict is accept / verified / sound. Reading only stage_ok once let a rejected subgoal flow downstream as if it had passed, which made the gate decorative. Both graph_run and graph_submit refuse a node whose deps are unmet, that is already finished, or that does not exist — the ordering is enforced, not advisory.
vendor: "auto" (default) tries each candidate in order and falls back to self. A bare direct call has no candidates. graph:orchestrate instead opens with vendor: "auto", allocation: "balanced", host_vendor, host_model, native_models. Reasoning prefers the driving AI/model; Implement/Test prefer the other vendor using Claude sonnet or Codex gpt-5.6-sol. If only one vendor is available, it can fill both roles with fresh contexts and selectable models. Fable/Astra are never inherited from the driving session automatically; they require an explicit model request. Hosts must declare their native model capabilities honestly. A Codex session uses native agents instead of nested Codex CLI; external vendors pass readiness probes.
Balanced ranking considers stage preference, assigned/running work, completion counts, and execution errors. It is a deterministic heuristic, not learned performance or cost prediction. Explicit policies override it. No token/spending cap is added. The lead passes scoped artifact paths and submits compact results; it does not perform every role in its own conversation. Each task's Implement/Test/Gate shares the same working directory and code snapshot. See skills/orchestrate/SKILL.md for the routing, artifact, and retry contract and the broker's current enforcement limits. A **named
hooks/mod.tsx 320 lines1import { atom, read, update } from 'claude-code'
2import type { Register } from 'claude-code'
3
4import type { GraphNode, GraphRun, Snap } from '../types'
5import { MARK, TINT, board, bar, cap, fmt, isKorean, rail, tabs } from './draw'
6import type { Mark } from './draw'
7
8// Read only: the mod lists and reads the broker's run files and writes nothing.
9const snap = atom({ plugin: 'graph', key: 'snap' } as const, null as Snap | null)
10const view = atom({ plugin: 'graph', key: 'view' } as const, 'flow' as 'flow' | 'nodes')
11const lang = atom({ plugin: 'graph', key: 'lang' } as const, 'en' as 'en' | 'ko')
12
13const PANE = 'graph-live'
14const DIR = '.harness-run/broker/runs'
15const FILE = /^[0-9a-f-]+\.json$/ // not the `<file>.<pid>.<uuid>.tmp` the store renames from
16const MAX_BYTES = 4 * 1024 * 1024 // $.fs.read rejects more
17const LIVE_MS = 2 * 60 * 60 * 1000
18const TICK_MS = 3000
19const BAR = 24
20const COL_MAX = 6
21
22const en = {
23 tabFlow: 'Flow', tabNodes: 'Nodes',
24 now: 'Now', next: 'Next', blockedAt: 'Blocked at', allDone: 'All done', nothing: 'Nothing to run',
25 stateRunning: 'running', stateBlocked: 'blocked', stateComplete: 'finished',
26 headLine: '{state} · {done}/{total}', gates: '{done}/{total} gates', goalGate: 'goal gate {pct}%',
27 stagePlan: 'Plan', stageSetgoal: 'SetGoal', stageCritique: 'Critique', stageBuild: 'Build', stageGate: 'Goal gate', stageReport: 'Report',
28 nodeImplement: 'impl', nodeTest: 'test', nodeGate: 'gate',
29 attempt: '(attempt {n})', again: '#{n}',
30 noRun: 'No graph run in this folder.', tooBig: 'This run file is too large to show.',
31 more: '+{n} more', cmdDesc: 'Open the graph live pane', paneOpened: 'pane opened',
32 colTodo: 'To do', colDoing: 'Doing', colDone: 'Done',
33}
34
35// one table, two languages; the type keeps the keys identical
36export const STRINGS: Record<'en' | 'ko', Record<keyof typeof en, string>> = {
37 en,
38 ko: {
39 tabFlow: '흐름', tabNodes: '노드',
40 now: '지금', next: '다음', blockedAt: '막힌 곳', allDone: '모두 끝남', nothing: '실행할 것 없음',
41 stateRunning: '진행 중', stateBlocked: '막힘', stateComplete: '완료',
42 headLine: '{state} · {done}/{total}', gates: '관문 {done}/{total}', goalGate: '목표 관문 {pct}%',
43 stagePlan: '계획', stageSetgoal: '목표 설정', stageCritique: '비평', stageBuild: '구현', stageGate: '목표 관문', stageReport: '보고',
44 nodeImplement: '구현', nodeTest: '테스트', nodeGate: '관문',
45 attempt: '({n}번째 시도)', again: '#{n}',
46 noRun: '이 폴더에 그래프 실행이 없습니다.', tooBig: '실행 파일이 너무 커서 보여줄 수 없습니다.',
47 more: '+{n}건 더', cmdDesc: '그래프 실행 현황 창 열기', paneOpened: '창을 열었습니다',
48 colTodo: '대기', colDoing: '진행', colDone: '완료',
49 },
50}
51type Key = keyof typeof en
52
53const nodeMark = (state: string): Mark =>
54 state === 'done' ? 'done' : state === 'running' ? 'running' : state === 'failed' ? 'failed' : 'pending'
55
56// graph.mjs settled(): an order-only dep that can no longer change
57const settled = (n: GraphNode) =>
58 n.state === 'done' || n.state === 'skipped' || n.state === 'unreachable' || (n.state === 'failed' && n.final === true)
59
60// graph.mjs unmetDeps(), recomputed here: what still holds a pending node back
61function unmet(run: GraphRun, n: GraphNode): string[] {
62 const get = (id: string) => run.nodes.find(x => x.node_id === id)
63 const data = n.deps.filter(d => get(d)?.state !== 'done')
64 const order = (n.after ?? []).filter(d => {
65 const dep = get(d)
66 return !dep || !settled(dep)
67 })
68 const live =
69 n.stage === 'report' && data.length === 0 && order.length === 0
70 ? run.nodes.filter(x => x !== n && x.stage !== 'report' && (x.state === 'running' || (x.state === 'pending' && unmet(run, x).length === 0)))
71 : []
72 return [...data, ...order, ...live.map(x => x.node_id)]
73}
74
75const readyOf = (run: GraphRun) => run.nodes.filter(n => n.state === 'pending' && unmet(run, n).length === 0)
76
77// graph.mjs runState()
78function runState(run: GraphRun): 'running' | 'blocked' | 'complete' {
79 if (run.nodes.some(n => n.stage === 'report' && n.state === 'done')) return 'complete'
80 const isRunning = run.nodes.some(n => n.state === 'running')
81 if (run.routing_blocked && !isRunning) return 'blocked'
82 return readyOf(run).length === 0 && !isRunning ? 'blocked' : 'running'
83}
84
85const sgOf = (n: GraphNode) => (typeof n.subgoal_id === 'string' && n.subgoal_id !== '' ? n.subgoal_id : null)
86
87// a person's name for a node, never its raw id
88function label(s: (k: Key, v?: Record<string, string | number>) => string, n: GraphNode, how: 'attempt' | 'again') {
89 const sg = sgOf(n)
90 const name = sg !== null ? `${sg} ${s(`node${cap(n.stage)}` as Key)}` : s(`stage${cap(n.stage)}` as Key)
91 if (sg === null && n.attempt <= 1) return name
92 return `${name} ${s(how, { n: n.attempt })}`
93}
94
95function stageMark(nodes: GraphNode[]): Mark {
96 if (nodes.length === 0) return 'pending'
97 if (nodes.some(n => n.state === 'failed')) return 'failed'
98 if (nodes.every(n => n.state === 'done')) return 'done'
99 return nodes.some(n => n.state === 'running' || n.state === 'done') ? 'running' : 'pending'
100}
101
102function model(run: GraphRun) {
103 const nodes = run.nodes.filter(n => n.state !== 'skipped')
104 const last = (stage: string) => nodes.filter(n => n.stage === stage && sgOf(n) === null).at(-1)
105 const single = (stage: string): Mark => {
106 const n = last(stage)
107 return n === undefined ? 'pending' : stageMark([n])
108 }
109 const rail: { key: string; state: Mark }[] = [
110 { key: 'Plan', state: single('plan') },
111 { key: 'Setgoal', state: single('setgoal') },
112 { key: 'Critique', state: single('critique') },
113 { key: 'Build', state: stageMark(nodes.filter(n => sgOf(n) !== null)) },
114 { key: 'Gate', state: single('gate') },
115 { key: 'Report', state: single('report') },
116 ]
117 const state = runState(run)
118 // nothing running yet (a session run never marks a node running): the next stage is the one in progress
119 if (state === 'running' && !rail.some(g => g.state === 'running' || g.state === 'failed')) {
120 const first = rail.find(g => g.state === 'pending')
121 if (first !== undefined) first.state = 'running'
122 }
123 const subgoals = run.spec?.subgoals ?? [...new Set(nodes.map(sgOf).filter((x): x is string => x !== null))].map(id => ({ id, title: '' }))
124 const gateOf = (id: string) => nodes.filter(n => n.stage === 'gate' && sgOf(n) === id).at(-1)
125 const pct = nodes.filter(n => n.stage === 'gate' && sgOf(n) === null && typeof n.result?.match_pct === 'number').at(-1)?.result?.match_pct
126 const failed = nodes.filter(n => n.state === 'failed')
127 return {
128 state,
129 nodes,
130 rail,
131 subgoals,
132 total: subgoals.length,
133 done: subgoals.filter(g => gateOf(g.id)?.state === 'done').length,
134 pct,
135 running: nodes.filter(n => n.state === 'running'),
136 ready: readyOf(run),
137 blocked: failed.find(n => n.final !== true) ?? failed[0] ?? nodes.find(n => n.state === 'pending'),
138 nodeOf: (id: string, stage: string) => nodes.filter(n => n.stage === stage && sgOf(n) === id).at(-1),
139 }
140}
141type Model = ReturnType<typeof model>
142
143// the sentence the pane leads with
144function focus(s: (k: Key, v?: Record<string, string | number>) => string, m: Model) {
145 if (m.state === 'complete') return { kind: 'done', head: s('now'), text: s('allDone'), color: 'success' }
146 if (m.state === 'blocked') {
147 const reason = m.blocked?.result?.reason
148 const what = m.blocked === undefined ? s('nothing') : label(s, m.blocked, 'attempt')
149 return { kind: 'blocked', head: s('blockedAt'), text: reason ? `${what}: ${reason}` : what, color: 'warning' }
150 }
151 const one = m.running[0] ?? m.ready[0]
152 const head = m.running.length > 0 ? s('now') : s('next')
153 return { kind: 'run', head, text: one === undefined ? s('nothing') : label(s, one, 'attempt'), color: 'claude' }
154}
155
156const titleOf = (run: GraphRun) => run.request.split('\n')[0]!.replace(/^\[[^\]]*\]\s*/, '').slice(0, 80)
157
158export const register: Register = on => {
159 let isRunning = false
160
161 on('session.start', async ($, e, next) => {
162 if ((await $.session.surfaces()).length === 0) return next(e)
163
164 // the newest run file; skipped when it is unchanged, kept when it does not parse
165 async function tick() {
166 if (isRunning) return
167 isRunning = true
168 try {
169 const now = await $.clock.now()
170 const prev = await read($, snap)
171 let entries: { name: string; kind: string; size: number; mtimeMs: number }[]
172 try {
173 entries = await $.fs.list(DIR)
174 } catch {
175 if (prev !== null) await update($, snap, () => null)
176 return
177 }
178 const newest = entries.filter(f => f.kind === 'file' && FILE.test(f.name)).sort((a, b) => b.mtimeMs - a.mtimeMs)[0]
179 if (newest === undefined) {
180 if (prev !== null) await update($, snap, () => null)
181 return
182 }
183 const fresh = now - newest.mtimeMs <= LIVE_MS
184 if (newest.size > MAX_BYTES) {
185 await update($, snap, () => ({ run: null, big: true, mtimeMs: newest.mtimeMs, size: newest.size, live: false }))
186 return
187 }
188 let run: GraphRun
189 if (prev !== null && !prev.big && prev.run !== null && prev.mtimeMs === newest.mtimeMs && prev.size === newest.size) {
190 run = prev.run
191 } else {
192 try {
193 run = JSON.parse(await $.fs.read(`${DIR}/${newest.name}`)) as GraphRun
194 if (!Array.isArray(run.nodes)) return
195 } catch {
196 return // a failed read or parse leaves the last good view
197 }
198 }
199 const live = fresh && runState(run) !== 'complete'
200 if (prev !== null && prev.run === run && prev.live === live) return
201 await update($, snap, () => ({ run, big: false, mtimeMs: newest.mtimeMs, size: newest.size, live }))
202 } catch {
203 // an unreadable folder leaves the last view as it was
204 } finally {
205 isRunning = false
206 }
207 }
208
209 // Claude Code's own language setting, read once per session
210 try {
211 const language = (await $.settings.read()).language
212 await update($, lang, () => (isKorean(language) ? 'ko' : 'en'))
213 } catch {
214 // unreadable settings: English
215 }
216 await $.command.register({ name: 'graph-live', description: STRINGS[await read($, lang)].cmdDesc })
217 await tick()
218 $.clock.every(TICK_MS, tick)
219
220 return next(e)
221 })
222
223 on('command.run', { command: 'graph-live' }, async $ => {
224 await $.ui.open({ id: PANE, title: 'Graph' })
225 return { text: STRINGS[await read($, lang)].paneOpened }
226 })
227
228 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
229 const ui = $.ui.resolve(e)
230 const { Box, Text } = ui
231 const kind = await read($, view)
232 const lang$ = await read($, lang)
233 const s = (key: Key, vars?: Record<string, string | number>) => fmt(STRINGS[lang$][key], vars)
234 const sn = await read($, snap)
235 if (sn === null || (sn.run === null && !sn.big)) return <Box flexDirection="column"><Text>{s('noRun')}</Text></Box>
236 if (sn.run === null) return <Box flexDirection="column"><Text color="warning">{s('tooBig')}</Text></Box>
237
238 const m = model(sn.run)
239 const f = focus(s, m)
240 const stateWord = s(`state${cap(m.state)}` as Key)
241 const gatesText = `${s('gates', { done: m.done, total: m.total })}${m.pct === undefined ? '' : ` · ${s('goalGate', { pct: m.pct })}`}`
242 const mark = (n: GraphNode | undefined, word: string) => {
243 const st = n === undefined ? 'pending' : nodeMark(n.state)
244 return <Text color={TINT[st]}>{`${MARK[st]} ${word}`}</Text>
245 }
246
247 let body
248 if (kind === 'flow') {
249 body = (
250 <Box flexDirection="column">
251 <Box flexDirection="column" marginTop={1}>
252 <Text color={f.color} bold wrap="truncate-end">{`${f.kind === 'blocked' ? '⚠' : '▶'} ${f.head} · ${f.text}`}</Text>
253 </Box>
254 <Box flexDirection="column" marginY={1}>
255 {m.subgoals.slice(0, COL_MAX).map(g => {
256 const imp = m.nodeOf(g.id, 'implement')
257 const tst = m.nodeOf(g.id, 'test')
258 const gat = m.nodeOf(g.id, 'gate')
259 const joined = (n: GraphNode | undefined) => (n !== undefined && n.state !== 'pending' ? ' ━ ' : ' ┄ ')
260 return (
261 <Box key={g.id}>
262 <Box flexShrink={1}><Text wrap="truncate-end">{`${g.id} ${g.title ?? ''}`.trim()}</Text></Box>
263 <Box flexShrink={0} marginLeft={2}>
264 {mark(imp, s('nodeImplement'))}
265 <Text dimColor>{joined(tst)}</Text>
266 {mark(tst, s('nodeTest'))}
267 <Text dimColor>{joined(gat)}</Text>
268 {mark(gat, s('nodeGate'))}
269 </Box>
270 </Box>
271 )
272 })}
273 {m.subgoals.length > COL_MAX && <Text dimColor>{` ${s('more', { n: m.subgoals.length - COL_MAX })}`}</Text>}
274 </Box>
275 {bar(ui, m.done, m.total, BAR, gatesText)}
276 </Box>
277 )
278 } else {
279 const card = (n: GraphNode, i: number) => {
280 const st = n.state === 'unreachable' ? 'pending' : nodeMark(n.state)
281 return (
282 <Box key={`n${i}`} flexDirection="column">
283 <Text wrap="truncate-end"><Text color={TINT[st]}>{MARK[st]}</Text>{` ${label(s, n, 'again')}`}</Text>
284 {n.state === 'failed' && n.result?.reason && <Text color="error" wrap="truncate-end">{` ${n.result.reason}`}</Text>}
285 </Box>
286 )
287 }
288 const cols = [
289 { label: s('colTodo'), tint: 'inactive', items: m.nodes.filter(n => n.state !== 'running' && n.state !== 'done') },
290 { label: s('colDoing'), tint: 'claude', items: m.nodes.filter(n => n.state === 'running') },
291 { label: s('colDone'), tint: 'success', items: m.nodes.filter(n => n.state === 'done') },
292 ]
293 body = (
294 <Box flexDirection="column">
295 {board(ui, cols, card)}
296 {bar(ui, m.done, m.total, BAR, gatesText)}
297 </Box>
298 )
299 }
300 return (
301 <Box flexDirection="column" borderStyle="round" borderColor="claude" paddingX={1}>
302 <Box>
303 <Box flexShrink={1}><Text bold wrap="truncate-end">{titleOf(sn.run)}</Text></Box>
304 <Box flexShrink={0} marginLeft={2}>
305 <Text color="claude">{s('headLine', { state: stateWord, done: m.done, total: m.total })}</Text>
306 </Box>
307 </Box>
308 <Box marginY={1} justifyContent="space-between">
309 {tabs(ui, [{ key: 'flow', label: s('tabFlow') }, { key: 'nodes', label: s('tabNodes') }], kind, key => update($, view, () => key as 'flow' | 'nodes'))}
310 <Text dimColor>{sn.run.run_id.slice(0, 8)}</Text>
311 </Box>
312 <Box flexDirection="column">
313 {m.rail.length > 0 && rail(ui, m.rail.map(g => ({ label: s(`stage${g.key}` as Key), state: g.state })))}
314 </Box>
315 {body}
316 </Box>
317 )
318 })
319}
320hooks/draw.tsx 95 lines1/*
2 * Drawing kit for mods: stage rail, progress bar, tabs, board.
3 * Origin: teams/hooks/mod.tsx (teams 0.46.0). Copied per plugin because a mod imports only its own
4 * plugin's files; copies may drift, so change one and diff the other.
5 */
6export type Mark = 'done' | 'running' | 'pending' | 'failed'
7
8export const MARK: Record<Mark, string> = { done: '✔', running: '●', pending: '○', failed: '✘' }
9// the stage rail: a dot per stage, a solid rail up to where the run is, dotted after
10export const DOT: Record<Mark, string> = { done: '●', running: '◉', pending: '○', failed: '✘' }
11export const TINT: Record<Mark, string> = { done: 'success', running: 'claude', pending: 'inactive', failed: 'error' }
12
13// a key the table lacks formats to ''
14export const fmt = (text: string | undefined, vars: Record<string, string | number> = {}) =>
15 (text ?? '').replace(/\{(\w+)\}/g, (_, k: string) => String(vars[k] ?? ''))
16
17export const cap = (s: string) => s.charAt(0).toUpperCase() + s.slice(1)
18
19// Claude Code's `language` setting
20export const isKorean = (language: unknown) => typeof language === 'string' && /^(ko|korean|한국어)/i.test(language)
21
22// filled cells of a bar `width` wide, clamped to 0..width
23export const cells = (done: number, total: number, width: number) =>
24 total > 0 ? Math.min(width, Math.max(0, Math.round((width * done) / total))) : 0
25
26// between two stages: solid once the next stage has started, dotted before
27export const connector = (nextState: Mark) => (nextState !== 'pending' ? ' ━━ ' : ' ┄┄ ')
28
29// the resolved components from $.ui.resolve(e)
30export type Ui = { Box: any; Text: any; Button: any }
31
32export function rail(ui: Ui, stages: { label: string; state: Mark }[]) {
33 const { Box, Text } = ui
34 return (
35 <Box flexWrap="wrap">
36 {stages.map((g, i) => {
37 const next = stages[i + 1]
38 return (
39 <Box key={`g${i}`}>
40 <Text color={TINT[g.state]} bold={g.state === 'running'}>{`${DOT[g.state]} ${g.label}`}</Text>
41 {next !== undefined && (
42 <Text color={next.state !== 'pending' ? 'success' : 'inactive'}>{connector(next.state)}</Text>
43 )}
44 </Box>
45 )
46 })}
47 </Box>
48 )
49}
50
51export function bar(ui: Ui, done: number, total: number, width: number, suffix?: string) {
52 const { Box, Text } = ui
53 const filled = cells(done, total, width)
54 return (
55 <Box>
56 <Text color="success">{'━'.repeat(filled)}</Text>
57 <Text color="inactive">{'─'.repeat(width - filled)}</Text>
58 {suffix !== undefined && <Text bold>{` ${suffix}`}</Text>}
59 </Box>
60 )
61}
62
63// plain tabs with hotkeys 1-9: the selected one in full strength with a dot, the rest dim
64export function tabs(ui: Ui, items: { key: string; label: string }[], current: string, onPick: (key: string) => void) {
65 const { Box, Button } = ui
66 return (
67 <Box columnGap={3}>
68 {items.map((one, i) => (
69 <Button key={one.key} plain hotkey={String(i + 1)} dimColor={current !== one.key}
70 label={`${current === one.key ? '● ' : ''}${one.label}`} onPress={() => onPick(one.key)} />
71 ))}
72 </Box>
73 )
74}
75
76// columns of bordered cards, each with a count in its title
77export function board<T>(
78 ui: Ui,
79 cols: { label: string; tint: string; items: T[] }[],
80 render: (item: T, i: number) => unknown,
81) {
82 const { Box, Text } = ui
83 const width = `${Math.floor(100 / Math.max(cols.length, 1))}%`
84 return (
85 <Box>
86 {cols.map(col => (
87 <Box key={col.label} flexDirection="column" width={width} borderStyle="round" borderColor={col.tint} paddingX={1}>
88 <Text bold color={col.tint}>{`${col.label} ${col.items.length}`}</Text>
89 {col.items.map(render)}
90 </Box>
91 ))}
92 </Box>
93 )
94}
95types/index.d.ts 32 lines1// the parts of a broker run file (.harness-run/broker/runs/<runId>.json) the mod reads
2export type GraphNode = {
3 node_id: string
4 stage: string
5 deps: string[]
6 after?: string[]
7 state: string
8 attempt: number
9 subgoal_id?: string | null
10 final?: boolean
11 result?: { stage_ok?: boolean; reason?: string; match_pct?: number } | null
12}
13export type GraphRun = {
14 run_id: string
15 request: string
16 routing_blocked?: boolean
17 spec?: { goal?: string; subgoals?: { id: string; title?: string }[] } | null
18 nodes: GraphNode[]
19}
20// what the tick keeps: the newest run, or that its file is too large to read
21export type Snap = { run: GraphRun | null; big: boolean; mtimeMs: number; size: number; live: boolean }
22
23declare module 'claude-code' {
24 interface PluginState {
25 graph: {
26 snap: Snap | null
27 view: 'flow' | 'nodes'
28 lang: 'en' | 'ko'
29 }
30 }
31}
32