Build and maintain a verified Figma component library from real code with resumable, schema-validated workflows.

Build and maintain a Figma component library from a codebase.
Requires Node 24. Run node scripts/lab_setup.ts install playwright once to install the pinned Node dependencies and Chromium into the shared cache. See TYPESCRIPT.md for plain-copy setup, checks and isolated comparison harnesses.
Extract, model, render are separate on purpose. Writing source straight into Figma gives a one-shot script that cannot re-run, cannot diff against the source later, and cannot feed anything but Figma. components.json and tokens.json are the contract; Figma is one renderer.
Component source, token source and usage source vary separately. A Single Directory Component site has no tokens in configuration at all — they live in stylesheets. Conflating the axes forces one site down another's path.
| Site | Components | Config path | Tokens |
|---|---|---|---|
| Drupal authoring site | 69 Drupal authoring bundles | config/default | 97 planned variables from authored Sass |
| Site Studio site A | 146 Site Studio + 6 custom | config/packages (declared in settings) | 129 custom style entities |
| Site Studio site B | 101 Site Studio + 3 custom | config/packages (declared in settings) | 176 custom style entities |
| Site Studio site C | 168 Site Studio + 4 custom | config/sitestudio (declared in settings) | 166 custom style entities |
| Paragraphs site | 43 Paragraph types | config/default | 113 base tokens via Sass source map |
| Paragraphs site, compiled-CSS branch | 43 Paragraph types | config/default | 94 authored custom properties |
The Paragraphs site also has 13 custom Single Directory Components, but only 6 are invoked by a paragraph template - they are a partial rendering layer, not the component source. It was recorded as a 13-component Single Directory Component site until 2026-08-31; that profile came from a bug, not the site. See references/strategies/README.md.
Run design-lab:init once on a new machine. It settles everything about you and the machine, asking before it installs or changes anything: where runs live (next to each project as PROJECT/design/<date>, or in ~/.design/<project>/<date>; runs are personal and never committed), one shared Playwright and its browser for capture, the baseline packages, the Figma runner, and the Claude Code setting that would otherwise make runs ask you to approve commands. lab_setup.ts check reports the same without changing anything. Everything about one site belongs to preflight, at the start of each run.
Use design-lab:run for a complete library. It creates a run folder outside the repository (where, design-lab:init decided) with its project.json, records the repository commit and every strategy decision, validates artifacts before rendering, and can resume from the first incomplete phase. Use a narrower skill only when the request names a single phase.
Use design-lab:figma-build to build an earlier run's capture and plan into a new, empty Figma file: a new run folder beside the others, with its own pane, verification and report, in minutes rather than the hours a fresh capture takes. The earlier run is left as it was.
Both open the design-lab pane beside the conversation as they start (below); nobody has to ask for it.
Builds write into Figma through the design-lab runner, a Figma development plugin that must be imported once per machine by hand — Figma offers no command-line install. To get the steps with this machine's paths filled in, ask:
Using design-lab's references/relay.md, give me concise steps to install the design-lab runner plugin in Figma desktop, including the absolute path to its manifest on this machine.
| Skill | Does |
|---|---|
design-lab:init | once per machine: where runs live, capture tools, the Figma runner, Claude Code settings |
design-lab:run | end-to-end, resumable workflow and completion gate |
design-lab:figma-build | an earlier run's capture and plan, built into a new empty Figma file as a new run, verified and scored |
design-lab:detect | which strategies apply |
design-lab:inventory | components + fields + slots + source defects -> components.json |
design-lab:usage | verified anonymous example addresses + placement counts + tiers |
design-lab:capture | measures and photographs each component on a running site, per breakpoint |
design-lab:tokens | colour, spacing, type per breakpoint, each with its code name -> tokens.json |
design-lab:plan | reviewable build proposal with variant arithmetic and hard refusals |
design-lab:figma-foundation | variable collections, modes, scopes, code syntax. Once per file |
design-lab:figma-component | one named atomic component transaction — variants, properties, bindings, documentation card, build record |
design-lab:figma-index | the Getting Started page: inventory, linked index, coverage, known gaps. Refresh after every component |
design-lab:verify | checks the whole file against the base expectations; every gap ends as a fix or a recorded waiver |
design-lab:evaluate | the last step of every run: scores it into scorecard.json, a self-contained HTML report (coverage, accuracy against the live site, time and tokens, repeatability) and the fixed completion message |
scripts/workflow.ts is the deterministic front door. Its init, identity, detect, select, preflight, connect, extract, usage, plan, variables, approve, target, register, record, validate, status, and watch commands write atomically and keep artifact hashes in the project manifest. The schemas in schemas/ are the machine-readable contracts; references/library-standard.md is the canonical product definition.
Component extractors cover Site Studio, SDCs, Paragraphs, and combined Drupal authoring vocabularies (block_content + Paragraphs). Token extractors cover Site Studio styles, theme-loaded CSS custom properties, Sass source maps, and source-authored Sass. Combined Drupal extraction also writes render-evidence.json, a bounded map from each authoring bundle to its existing Twig, SDC, stylesheet, root-class, and referenced-field evidence. That evidence includes deterministic styleFacts parsed from the component's own Sass: root and nested-part declarations stay separate, retain token/literal provenance, and give the model the visual facts it needs without asking it to rediscover every stylesheet rule.
Drupal database usage is deterministic too: extract_drupal_usage.ts reads the running DDEV project, preserves placements and structural references as separate measures, writes a validated usage.json, and merges tiers into components.json. Planning hard-stops when detection found usage evidence but it was neither measured nor explicitly waived as degraded. Waivers are human decisions: the workflow requires both a named decider and a reason, and final verification refuses any unresolved prerequisite phase. Registering one component receipt also cannot complete the component phase; valid non-failing receipts must cover the approved plan.
Capture: scaffold_configs.ts writes a config per component and names the ones a human must finish; capture_all.ts measures the box model and typography per breakpoint and takes element-scoped screenshots. Playwright is not vendored: the pinned packages and Chromium come from the shared cache that lab_setup.ts install fills, whatever the current working directory. DESIGN_LAB_BROWSER_EXECUTABLE selects an installed browser when a Playwright-managed Chromium download is unavailable. See skills/capture/SKILL.md.
Token sources are ranked by evidence. A substantial theme-loaded custom-property layer states runtime intent and wins. Source-authored Sass is next, then a recovered Sass source map. Weak or unloaded CSS never outranks the theme sources merely because similarly named files exist.
figma-component owns one component transaction on purpose. run may process many transactions in one model session, but each becomes durable only after its build record is validated. Every recorded assertion must explicitly pass; skipped fidelity work and an empty assertion object remain incomplete. A partial run therefore resumes safely instead of pretending the library is done.
verify is the one that runs last and the one that should have existed first. Every other skill reports on its own step, so a library can pass all of them and still be half a library — which is exactly what happened on the Paragraphs site: four empty Foundations pages, 36 of 43 components missing, no documentation links anywhere, and sixteen variables whose Dev Mode names existed nowhere in the codebase. Nothing was looking at the whole.
Planned: drift.
Every site's font trouble has been its own, so fonts are a step of each run, not of setup. When the build is ready to write, workflow.ts connect asks Figma desktop which fonts it can draw with and writes the run's font plan, fonts.json. For each text stack the site renders, it works out the family the visitor really sees: a family declared with no source is skipped, as the browser skips it; icon fonts and generic families are set aside. Each family is then available to Figma, or missing with a stand-in chosen by genre (the same design under its other macOS name when Figma desktop leaves a system font out of its list, such as Courier New for Courier, which needs nothing installed; a metric-compatible clone where one exists, such as Arimo for Arial) and the route to the real font: the Adobe Fonts kit to activate, the open-licence files to install, or the client's desktop files or the foundry's trial for a commercial family, never the site's web font files. The run never stops for a font. The build draws each weight in the face the site really serves (a font-weight: 500 rule serving a Semibold file is drawn Semibold), matches style names generously, accepts trial and web family names, and uses a variable font's weight axis when no named style fits. Verify and the completion message name each stand-in as a decision, not as a failure. workflow.ts fonts --project <run> writes the plan again from the font list recorded at the last connect (or, before the first, without one), and workflow.ts report fonts --project <run> shows it. The step writes fonts.json, figma/available-fonts.json (Figma's list) and, for a site using Adobe Fonts, fonts-kits.json (the kit's families, read once) in the run folder.
design-lab:run and design-lab:figma-build open a pane beside the transcript as they start, and it stays open for the whole run. Typed as slash commands, the pane opens at any window width; asked for in words, Claude Code places a pane it was not asked for only in a wide terminal (144 columns). If the project's newest run is finished, the pane waits for the new run's folder rather than showing the old recap; if it is unfinished, that is the run being resumed and it shows at once. /design-lab:watch [run folder] opens the same pane by hand, after it was closed or for a run started elsewhere. The pane shows the run as five stages (Preflight, Discovery, Build, Verify, Report), one row each with its result, and opens the stage under way beneath its row: the preflight checklist, ticked off as preflight proves each check; the build's steps done of total and whether the runner is connected. A word in the header says where the run stands (for example Building, Needs you, Failed, Done or Done · needs review), and a card above the stages says what the person has to do, or what went wrong. It also keeps a status line such as design-lab: steps 112/158 · runner connected · 41m in that session. With no folder it shows this project's newest run, found by the convention design-lab:init chose from the folder the session is in, and moves to each newer run there as it starts; before the first run it opens and waits for one. Outside any project it falls back to the run workflow.ts init, preflight or connect last recorded in ~/.design-lab/active-run.json. Figma is first touched when the build is ready to write: workflow.ts connect asks the person to open the target file and start the runner, and the pane shows that as the one thing that needs them, with no button, until the runner connects. When the runner later stops asking for steps, the pane says what to do in Figma desktop and offers one button, Resume run, which puts the resume request in the prompt box. When the benchmark has written completion.md, a toast says the run is done, the status line clears, and the pane leads with the verdict (when verification left blocking or major problems open), the run's coverage, accuracy, time and tokens, and links to the Figma library, the benchmark report and the run's files; its Recap starts folded and shows the full completion message when Show is pressed. /design-lab:recap [run folder] shows any finished run's completion message again, with no Claude turn.
The pane is a Claude Code mod (hooks/mod/), which needs a Claude Code that loads mods (2.1.286 does; 2.1.284 does not) and draws in the terminal and the desktop app's Code tab. The desktop app runs its own copy of Claude Code, updated separately from the app; the pane draws there as a sidebar once that copy loads mods, and until then the command answers with text. It only reads the run folder and the pointer (preflight-checks.json, which workflow.ts preflight rewrites as each check starts and settles, is in the run folder); it writes nothing and never reads the runner token. Everything works without it: where the mod is not loaded, or nothing draws (claude -p, the VS Code chat panel), the same command prints the same summary from workflow.ts watch. Mod tests run with node design-lab/scripts/test-mod.ts, which invokes claude plugin test on an isolated mod-only test copy.
references/prior-art.md — read first. Look for an existing Figma file and existing tooling before extracting anythingreferences/model.md — the universal model and the provenance rulereferences/variant-policy.md — the decision that makes or breaks the libraryreferences/findability.md — how anyone finds a component in a 146-component filereferences/tokens-and-variables.md — code syntax, and why the Figma name is not the tokenreferences/defaults.md — which variant goes first, and the evidence for itreferences/verification.md — assert numbers, do not eyeball 146 componentsreferences/benchmark.md — run checklist and fixed opening prompt; every run ends with design-lab:evaluatereferences/completion-message.md — the fixed reply design-lab:run ends withreferences/build-records.md — the idempotency and resume contractreferences/strategies/README.md — per-strategy mapping and counting trapsThe corpus and scoreboard use explicit per-person locations from ~/.claude/design-lab.json, or the file named by DESIGN_LAB_CONFIG:
{
"corpus": "/path/to/corpus",
"scoreboard": {
"ledger": "/path/to/ledger.jsonl",
"dashboard": "/path/to/dashboard.html"
}
}
All three keys are required. The commands never create or edit this configuration. scoreboard.ts record appends to the ledger and redraws the dashboard from the whole ledger; --open shows it.
node scripts/corpus.ts freeze --run /path/to/finished-run --label site-a
node scripts/corpus.ts list
node scripts/tier1.ts --all --out /tmp/property-results.json
node scripts/tier1.ts --run /path/to/run --out /tmp/property-results.json
node scripts/tier2.ts --site site-a --file-key SCRATCH_FILE_KEY
node scripts/scoreboard.ts record --run /path/to/evaluated-run --tier 2 --site site-a
node scripts/scoreboard.ts rows
Freezing copies the run and writes a manifest with artifact hashes and producer identity. An existing label is refused; use a new label to record a corpus refresh. Older manifests may have no producer commit; that absence stays explicit. Both component-id and older machine-name measurement files are supported.
Tier 1 rebuilds the same trees as figma_build.ts init, in a temporary directory, and compares their resolved breakpoint properties with the saved measurements. Its output is JSON; one summary per site goes to standard error. Geometry uses tree layout arithmetic, not font shaping or Figma rendering. Geometry needing font shaping says unmeasured. Older captures have styled inline descendants rather than exact character ranges, so text run counts identify distinct measured inline styles and flag flattening; their basis is recorded beside each result. Captured states absent from the default responsive tree are reported as unmeasured. This comparison does not change the builder or verify's gate.
Tier 2 requires Figma open with the runner. It creates a new workspace under the site's replays/ directory, clears the designated scratch file, rebuilds with frozen images, waits for the build and verification dumps, assembles measurements, verifies, and scores. The runner stays open between evaluations; for --all, have it open in each scratch file. The image step has no site or public fallback requests: uncached sources remain failures. For --all, add a scratchFileKey to each site's corpus.json; file keys are checked before any replay starts. --timeout limits the wait per site. Failed evaluations keep the workspace and its evidence for inspection. Verification findings do not prevent writing the scorecard. New build receipts and the master-matches-capture gate use the corrected comparison metric; historical scores retain their original metric for comparison. A report is produced even on failed verification, but the manifest records failed quality until the shared completion gate passes.
The benchmark and ledger share run_metrics.ts; neither imports the other. Missing metrics remain unmeasured or null. Recording an evaluation is a separate explicit step, so replay does not silently append a ledger row.
Variable collections default to one shared <Brand> Core, with slash groups for domains and width modes shared by invariant and responsive values. Additional collections require an independent mode axis or documented publishing/ownership boundary; inventory size alone never causes a split.
hooks/mod/register.tsx 719 lines1// design-lab's pane: where a run is and when it is done, from the files the run writes. It reads
2// only run folders (and, to find them, the person's design-lab settings, the project's runs folder
3// and the active-run pointer), writes nothing to disk, and asks nothing of the person except, when
4// the runner has stopped, a button that puts the resume request in the prompt.
5
6import { atom, read, update } from 'claude-code'
7import type { CommandRunInput, EngineInterface, Register, Timer } from 'claude-code'
8
9import type { Summary as RunSummary } from '../../src/protocol'
10type Summary = RunSummary<'mod'>
11import {
12 afterFill, barOf, CHECK_COLORS, CHECK_MARKS, checkMessage, compactOf, durationOf, elapsedOf, failureOf, needsYouOf, parseJson,
13 percentOf, PHASE_DONE, phaseLabel, plainOf, RESUME_PROMPT, isIdle, runnerLine, stagesOf, statusOf, stepsLine,
14 summaryOf, toneOf, verdictOf,
15} from './model'
16import { ancestors, base, MARKERS, parent } from './locate'
17
18const PANE = 'design-lab'
19const COMMAND = 'design-lab:watch'
20const RECAP_COMMAND = 'design-lab:recap'
21// The skills that start or resume a run: each opens the pane by itself, so nobody has to know
22// about design-lab:watch to see where a run is.
23export const RUN_SKILLS = ['design-lab:run', 'design-lab:figma-build'] as const
24import { POLL_MS } from '../../src/protocol.ts'
25export { POLL_MS } from '../../src/protocol.ts'
26// The runner log is tailed only while it is small enough to read whole every poll.
27const LOG_READ_LIMIT = 1024 * 1024
28
29const runAtom = atom({ plugin: 'design-lab', key: 'run' } as const, null)
30const summaryAtom = atom({ plugin: 'design-lab', key: 'summary' } as const, null)
31const alarmedAtom = atom({ plugin: 'design-lab', key: 'alarmed' } as const, false)
32// The runs folder the pane follows when the person named no run: it moves to each newer run there.
33const followAtom = atom({ plugin: 'design-lab', key: 'follow' } as const, null)
34// A finished run the pane passes over while a run skill is starting the next one, so the pane
35// never shows the last run's recap as if it were the new run.
36const skipAtom = atom({ plugin: 'design-lab', key: 'skip' } as const, null)
37// Whether a finished run's full completion message is open under its figures.
38// Named by run, so opening one run's recap never opens the next run's.
39const recapOpenAtom = atom({ plugin: 'design-lab', key: 'recapOpen' } as const, null)
40
41// The files read under a run folder, and nothing else there.
42export const RUN_FILES = {
43 project: 'project.json',
44 phaseLog: 'phase-log.jsonl',
45 progress: 'figma/progress.json',
46 runnerLog: 'figma/runner.log',
47 completion: 'benchmark/completion.md',
48 scorecard: 'benchmark/scorecard.json',
49 preflightChecks: 'preflight-checks.json',
50 verifyReport: 'verify-report.json',
51} as const
52
53// The files a finished run links to, when they exist: a rebuild has no plan or components of its
54// own unless it copied them.
55export const ARTIFACT_FILES = {
56 report: 'benchmark/report.html',
57 verifyReport: 'verify-report.json',
58 plan: 'plan.json',
59 components: 'components.json',
60 scorecard: 'benchmark/scorecard.json',
61} as const
62
63// The bar's fill by the color its Text would take, and its empty track.
64const BAR_COLORS: Record<string, string> = { cyan: '#4fa8d6', green: '#3fb36b', yellow: '#d6b44f', red: '#d65f5f' }
65const BAR_TRACK = '#8888884d'
66
67/** A rounded progress bar, wide and short so it scales to the pane's width. */
68function barSvg(share: number, fill: string): string {
69 const filled = share > 0 ? `<rect width="${Math.max(12, share)}" height="12" rx="6" fill="${fill}"/>` : ''
70 return `<svg xmlns="http://www.w3.org/2000/svg" width="1000" height="12" viewBox="0 0 1000 12">`
71 + `<rect width="1000" height="12" rx="6" fill="${BAR_TRACK}"/>${filled}</svg>`
72}
73
74let timer: Timer | undefined
75// The last JSON that parsed, so a file caught half-written keeps the last good reading.
76const lastGood = new Map<string, unknown>()
77
78async function readText($: EngineInterface, path: string): Promise<string | undefined> {
79 try {
80 return await $.fs.read(path)
81 } catch {
82 return undefined
83 }
84}
85
86async function readJson($: EngineInterface, path: string): Promise<unknown> {
87 const value = parseJson(await readText($, path))
88 if (value !== undefined) lastGood.set(path, value)
89 return value ?? lastGood.get(path)
90}
91
92async function readLog($: EngineInterface, path: string): Promise<string | undefined> {
93 try {
94 const stat = await $.fs.stat(path)
95 return stat.size <= LOG_READ_LIMIT ? await $.fs.read(path) : undefined
96 } catch {
97 return undefined
98 }
99}
100
101export async function summarise($: EngineInterface, run: string): Promise<Summary> {
102 const at = (file: string) => `${run}/${file}`
103 const raw = {
104 project: await readJson($, at(RUN_FILES.project)),
105 phaseLog: await readText($, at(RUN_FILES.phaseLog)),
106 progress: await readJson($, at(RUN_FILES.progress)),
107 preflightChecks: await readJson($, at(RUN_FILES.preflightChecks)),
108 verifyReport: await readJson($, at(RUN_FILES.verifyReport)),
109 runnerLog: await readLog($, at(RUN_FILES.runnerLog)),
110 ...((await $.fs.exists(at(RUN_FILES.completion)))
111 ? { completion: await readText($, at(RUN_FILES.completion)), scorecard: await readJson($, at(RUN_FILES.scorecard)),
112 present: await presentOf($, run) }
113 : {}),
114 }
115 return summaryOf(run, raw, await $.clock.now())
116}
117
118/** Which of the run's artifact files exist, so the pane links only to those. */
119async function presentOf($: EngineInterface, run: string): Promise<string[]> {
120 const found: string[] = []
121 for (const path of Object.values(ARTIFACT_FILES)) if (await exists($, `${run}/${path}`)) found.push(path)
122 return found
123}
124
125let refreshing = false
126
127async function refresh($: EngineInterface): Promise<void> {
128 // A slow file system must not stack refreshes: skip a tick while the last one is still reading.
129 if (refreshing) return
130 refreshing = true
131 try {
132 await refreshNow($)
133 } finally {
134 refreshing = false
135 }
136}
137
138// Nothing awaits a refresh the timer or an event starts, so a failure there would be an unhandled
139// rejection. It is told once, as a toast, and again only after a refresh worked in between.
140let refreshFailed = false
141function refreshInBackground($: EngineInterface): void {
142 void refresh($).then(
143 () => { refreshFailed = false },
144 () => { if (!refreshFailed) { refreshFailed = true; reportFailure($, 'refreshing the pane') } },
145 )
146}
147
148async function refreshNow($: EngineInterface): Promise<void> {
149 const follow = await read($, followAtom)
150 if (follow) {
151 const newest = await newestRun($, follow)
152 if (newest && newest !== (await read($, skipAtom)) && newest !== (await read($, runAtom))) {
153 await update($, runAtom, () => newest)
154 await update($, skipAtom, () => null)
155 await update($, alarmedAtom, () => false)
156 }
157 }
158 const run = await read($, runAtom)
159 if (!run) return
160 const summary = await summarise($, run)
161 const before = await read($, summaryAtom)
162 if (JSON.stringify(before) !== JSON.stringify(summary)) await update($, summaryAtom, () => summary)
163 $.ui.status(statusOf(summary, await $.clock.now()))
164 // Done: say so once, the moment the recap appears.
165 if (summary.hasRecap && before && before.found && !before.hasRecap) {
166 // Say what the header says: a run that left problems open is done but needs review.
167 const review = verdictOf(summary.findings) ? ', with verification problems to review' : ''
168 $.ui.toast(`design-lab: ${summary.siteLabel ?? 'the run'} is done${review}. The recap is in the design-lab pane.`, { timeoutMs: 10_000 })
169 }
170 // The watchdog: once per transition, never again until the run needs nothing from the person.
171 // It says what the Needs you card says.
172 const need = needsYouOf(summary)
173 const alarmed = await read($, alarmedAtom)
174 if (need && !alarmed) {
175 $.ui.toast(`design-lab needs you: ${need.message}`, { timeoutMs: 10_000 })
176 await update($, alarmedAtom, () => true)
177 } else if (!need && alarmed) {
178 await update($, alarmedAtom, () => false)
179 }
180}
181
182function watch($: EngineInterface): void {
183 timer?.cancel()
184 timer = $.clock.every(POLL_MS, () => refreshInBackground($))
185}
186
187/** The run a command names, or with none: the newest in this project's runs folder, which the pane
188 * then follows; else the run the machine-wide pointer names. */
189async function runOf($: EngineInterface, args: string): Promise<string | { missing: string; follow?: string }> {
190 const given = args.trim()
191 if (given) return given.startsWith('/') ? given.replace(/\/+$/, '') : `${await $.session.cwd()}/${given}`
192 const folder = await runsFolder($, await $.session.cwd())
193 if (folder) {
194 const newest = await newestRun($, folder)
195 return newest ?? { missing: `No design-lab run yet in ${folder}. The pane shows the first one as soon as it starts.`, follow: folder }
196 }
197 const home = await $.env.get('HOME')
198 const pointer = parseJson(home ? await readText($, `${home}/.design-lab/active-run.json`) : undefined)
199 const workspace = typeof pointer === 'object' && pointer !== null ? (pointer as { workspace?: unknown }).workspace : undefined
200 if (typeof workspace === 'string') return workspace
201 return { missing: (await settings($)).convention
202 ? 'No design-lab run found for this folder. Give the run folder: /design-lab:watch <run folder>'
203 : 'design-lab is not set up on this machine yet: run design-lab:init once. Or give the run folder: /design-lab:watch <run folder>' }
204}
205
206// Which run to watch when the person names none: the newest in this project's runs folder, by
207// the convention design-lab:init recorded (src/lab-config.ts holds the same rule).
208
209async function exists($: EngineInterface, path: string): Promise<boolean> {
210 return $.fs.exists(path).catch(() => false)
211}
212
213async function isDir($: EngineInterface, path: string): Promise<boolean> {
214 const found = await $.fs.stat(path, { resolve: false }).catch(() => undefined)
215 return found?.kind === 'dir'
216}
217
218async function hasMarker($: EngineInterface, folder: string): Promise<boolean> {
219 for (const marker of MARKERS) if (await isDir($, `${folder}/${marker}`)) return true
220 return false
221}
222
223/** A path with every link followed, as the scripts resolve it; the path itself when it cannot be. */
224async function real($: EngineInterface, path: string): Promise<string> {
225 const found = await $.fs.stat(path, { resolve: true }).catch(() => undefined)
226 return (found?.realPath ?? path).replace(/\/+$/, '') || '/'
227}
228
229/** The person's design-lab settings, or {} before design-lab:init has run. */
230async function settings($: EngineInterface): Promise<{ convention?: string; home?: string }> {
231 const rawHome = await $.env.get('HOME')
232 const home = rawHome ? await real($, rawHome) : undefined
233 const path = (await $.env.get('DESIGN_LAB_CONFIG')) ?? (home ? `${home}/.claude/design-lab.json` : undefined)
234 const text = path ? await $.fs.read(path).catch(() => undefined) : undefined
235 const value = parseJson(typeof text === 'string' ? text : undefined) as { runs?: { convention?: unknown } } | undefined
236 const convention = typeof value?.runs?.convention === 'string' ? value.runs.convention : undefined
237 return { convention, home }
238}
239
240/** The project folder for a session in `cwd`: above worktrees/ for PROJECT/worktrees/<name>,
241 * else the nearest folder above the repository holding plans/, analysis-reports/ or design/. */
242async function projectFolder($: EngineInterface, cwd: string, home?: string): Promise<string | undefined> {
243 const stop = (folder: string) => folder === '/' || folder === home
244 const chain = ancestors(cwd)
245 let repo: string | undefined
246 for (const folder of chain) if (await exists($, `${folder}/.git`)) { repo = folder; break }
247 if (!repo) {
248 for (const folder of chain) {
249 if (stop(folder)) return undefined
250 if (await isDir($, `${folder}/worktrees`) || await hasMarker($, folder)) return folder
251 }
252 return undefined
253 }
254 if (base(parent(repo)) === 'worktrees') return parent(parent(repo))
255 for (const folder of ancestors(parent(repo))) {
256 if (stop(folder)) return undefined
257 if (await hasMarker($, folder)) return folder
258 }
259 return undefined
260}
261
262/** Where this project's runs live, or undefined when design-lab:init has not chosen. */
263async function runsFolder($: EngineInterface, sessionCwd: string): Promise<string | undefined> {
264 const { convention, home } = await settings($)
265 const cwd = await real($, sessionCwd)
266 const project = await projectFolder($, cwd, home)
267 if (convention === 'project') return project ? `${project}/design` : undefined
268 if (convention === 'home' && home) {
269 let repo: string | undefined
270 for (const folder of ancestors(cwd)) if (await exists($, `${folder}/.git`)) { repo = folder; break }
271 return `${home}/.design/${base(project ?? repo ?? cwd)}`
272 }
273 return undefined
274}
275
276// When each run began, read once per run folder: a run's start never changes, so a poll lists the
277// runs folder and reads only the project.json of runs it has not seen.
278const startedAt = new Map<string, string>()
279
280/** The newest run in a runs folder: latest start (project.json createdAt), then folder name. */
281async function newestRun($: EngineInterface, folder: string): Promise<string | undefined> {
282 const entries = await $.fs.list(folder).catch(() => [])
283 let best: { created: string; name: string; path: string } | undefined
284 for (const entry of entries) {
285 if (entry.kind !== 'dir') continue
286 const path = `${folder}/${entry.name}`
287 let created = startedAt.get(path)
288 if (created === undefined) {
289 const text = await $.fs.read(`${path}/project.json`).catch(() => undefined)
290 const value = (parseJson(typeof text === 'string' ? text : undefined) as { createdAt?: unknown } | undefined)?.createdAt
291 if (typeof value !== 'string') continue
292 startedAt.set(path, value)
293 created = value
294 }
295 if (!best || created > best.created || (created === best.created && entry.name > best.name)) {
296 best = { created, name: entry.name, path }
297 }
298 }
299 return best?.path
300}
301
302/** A run skill is starting: follow this project's runs folder and open the pane beside the
303 * conversation. A newest run with no recap yet is the one being resumed, so it shows at once; a finished
304 * one is passed over until the new run's folder appears. Opened from the person's own slash
305 * command, Claude Code places the pane at any width; from the Skill tool, only in a wide window. */
306async function followRunSkill($: EngineInterface): Promise<void> {
307 const folder = await runsFolder($, await $.session.cwd())
308 // Not set up yet: the skill sends the person to design-lab:init first.
309 if (!folder) return
310 const newest = await newestRun($, folder)
311 const summary = newest ? await summarise($, newest) : undefined
312 const resuming = newest && summary?.found && !summary.hasRecap ? newest : null
313 await update($, followAtom, () => folder)
314 await update($, skipAtom, () => resuming ? null : newest ?? null)
315 await update($, runAtom, () => resuming)
316 await update($, summaryAtom, () => resuming ? summary! : null)
317 await update($, alarmedAtom, () => false)
318 watch($)
319 if (resuming) refreshInBackground($)
320 if ((await $.session.surfaces()).length > 0) await $.ui.open({ id: PANE, title: 'design-lab' }).catch(() => undefined)
321}
322
323async function resume($: EngineInterface, run: string): Promise<void> {
324 const text = RESUME_PROMPT(run)
325 const step = afterFill(await $.prompt.fill({ text, mode: 'replace' }).catch(() => undefined))
326 if (step === 'submit') await $.prompt.submit({ text }).catch(() => undefined)
327 if (step === 'explain') $.ui.toast('design-lab: close the open dialog, then press resume again.')
328}
329
330/** A hook whose registration catches a failure says so once, as a toast; the engine logs the error itself. */
331function reportFailure($: EngineInterface, what: string): void {
332 $.ui.toast(`design-lab: ${what} failed; claude --debug has the error`, { timeoutMs: 10_000 })
333}
334
335/** A gating command hook must answer, so a failure becomes the command's answer rather than a hang. */
336function answerFailure($: EngineInterface, command: string): { text: string } {
337 reportFailure($, command)
338 return { text: `design-lab could not answer ${command}; claude --debug has the error.` }
339}
340
341async function watchCommand($: EngineInterface, e: CommandRunInput): Promise<{ text: string }> {
342 const run = await runOf($, e.args)
343 // Named, a run is watched as named; found by convention, the pane follows the runs folder.
344 const follow = e.args.trim() ? null : typeof run === 'string'
345 ? (await runsFolder($, await $.session.cwd())) ?? null : run.follow ?? null
346 await update($, followAtom, () => follow)
347 await update($, skipAtom, () => null)
348 if (typeof run !== 'string') {
349 if (!follow) return { text: run.missing }
350 await update($, runAtom, () => null)
351 await update($, summaryAtom, () => null)
352 watch($)
353 if ((await $.session.surfaces()).length > 0) await $.ui.open({ id: PANE, title: 'design-lab', focus: true })
354 return { text: run.missing }
355 }
356 await update($, runAtom, () => run)
357 await update($, alarmedAtom, () => false)
358 await refresh($)
359 watch($)
360 const summary = (await read($, summaryAtom)) ?? (await summarise($, run))
361 if ((await $.session.surfaces()).length === 0) return { text: plainOf(summary) }
362 // The person asked for it: bring it to the front, over any other pane already open.
363 const opened = await $.ui.open({ id: PANE, title: 'design-lab', focus: true })
364 return { text: opened.isPlaced ? `Watching ${summary.siteLabel ?? run}.` : plainOf(summary) }
365}
366
367async function recapCommand($: EngineInterface, e: CommandRunInput): Promise<{ text: string }> {
368 const run = await runOf($, e.args)
369 if (typeof run !== 'string') return { text: run.missing }
370 const summary = await summarise($, run)
371 if (!summary.found) return { text: plainOf(summary) }
372 return { text: summary.recap ?? `${summary.siteLabel ?? run} has no recap yet: the run has not finished its benchmark.` }
373}
374
375export const register: Register = on => {
376 on('session.start', async ($, e, next) => {
377 // After a reload the run is still in state; pick the watch back up.
378 if ((await read($, runAtom)) || (await read($, followAtom))) {
379 watch($)
380 refreshInBackground($)
381 }
382 return next(e)
383 })
384
385 on('session.end', async (_$, e, next) => {
386 timer?.cancel()
387 timer = undefined
388 return next(e)
389 })
390
391 // A gating hook: its registration's .catch answers in its place when the work throws.
392 on('command.run', { command: COMMAND }, ($, e) => watchCommand($, e))
393 .catch(($) => answerFailure($, COMMAND))
394
395 // Typed as a slash command: the person asked, so the pane opens at any width. A failure here is
396 // reported, and the command still runs through next.
397 for (const command of RUN_SKILLS) {
398 on('command.run', { command }, async ($, e, next) => {
399 await followRunSkill($)
400 return next(e)
401 }).catch(($, e, next) => {
402 reportFailure($, command)
403 return next(e)
404 })
405 }
406
407 // Called through the Skill tool (the person asked in their own words): the same, unasked. A typed
408 // command raises this too, after command.run; following again then changes nothing.
409 on('skill.prompt', async ($, e, next) => {
410 if ((RUN_SKILLS as readonly string[]).includes(e.skill)) {
411 await followRunSkill($).catch(() => reportFailure($, e.skill))
412 }
413 return next(e)
414 })
415
416 on('command.run', { command: RECAP_COMMAND }, ($, e) => recapCommand($, e))
417 .catch(($) => answerFailure($, RECAP_COMMAND))
418
419 // The recap's output row, drawn as the Markdown it is, so its links are links.
420 on('ui.render', { component: 'CommandOutput', props: { command: RECAP_COMMAND } }, async ($, e, next) => {
421 if (!e.props.text) return next(e)
422 const { Markdown } = $.ui.resolve(e)
423 return <Markdown text={e.props.text} />
424 })
425
426 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
427 const elements = $.ui.resolve(e)
428 const { Box, Text, Button, Markdown } = elements
429 // Drawn surfaces (desktop, editor, phone) draw the bar as a vector; the terminal as cells.
430 const Svg = e.surface !== 'terminal' && 'Svg' in elements ? elements.Svg : undefined
431 const summary = await read($, summaryAtom)
432 const run = await read($, runAtom)
433 // The pane's ✕ sits on its first row: start one row lower so it never covers text.
434 const follow = await read($, followAtom)
435 if (!summary || !run) {
436 const skip = await read($, skipAtom)
437 return <Box marginTop={1} flexDirection="column" gap={1}>
438 <Text bold>design-lab</Text>
439 <Text dimColor>{follow
440 ? skip
441 ? 'Waiting for the new design-lab run to start. It appears here as soon as its folder is made.'
442 : `No design-lab run yet in ${follow}. It appears here as soon as one starts.`
443 : 'No design-lab run is being watched.'}</Text>
444 </Box>
445 }
446 if (!summary.found) return <Box marginTop={1}><Text>{plainOf(summary)}</Text></Box>
447 const now = await $.clock.now()
448 const columns = e.props.bodyColumns
449 // The card around the bar takes margin 2, border 2 and padding 2; the rest is slack.
450 const width = Math.max(4, columns - 10)
451 const tone = toneOf(summary)
452 const stages = stagesOf(summary)
453 const need = needsYouOf(summary)
454 const failure = failureOf(summary)
455 const runner = summary.runner
456 const preflight = summary.preflight
457 const scores = summary.scores
458 const finished = summary.hasRecap
459 const verdict = finished ? verdictOf(summary.findings) : null
460 const recapOpen = (await read($, recapOpenAtom)) === run
461 const time = finished ? durationOf(scores?.workingSeconds ?? null) : elapsedOf(summary.startedAt, now)
462 const name = run.split('/').pop()
463 // Finished with figures, the Time tile holds the time, so the subtitle does not repeat it.
464 const subtitle = finished && scores ? [name] : [name, time && `${time} ${finished ? 'working time' : 'elapsed'}`]
465 // Two tiles side by side need about 52 columns; narrower, they stack.
466 const narrowTiles = columns < 52
467 // Narrower still, a stage's steps go one per line.
468 const narrowSteps = columns < 40
469
470 // One figure in a quietly bordered tile: only the value carries color, so the tiles never
471 // compete with a card that asks for attention.
472 const tile = (key: string, label: string, value: string | null, details: (string | null)[], color: string | undefined) => (
473 <Box key={`tile-${key}`} flexDirection="column" flexGrow={1} flexShrink={1} width={narrowTiles ? '100%' : '50%'}
474 borderStyle="round" borderDimColor paddingX={1}>
475 <Text dimColor>{label}</Text>
476 <Text bold color={color}>{value ?? '–'}</Text>
477 {details.filter(Boolean).map((detail, i) => <Text key={`${key}-${i}`} dimColor wrap="wrap">{detail}</Text>)}
478 </Box>
479 )
480 // Coverage and accuracy are judgments: green only when nothing is missing, yellow otherwise.
481 const judged = (share: number | null) => share === null ? undefined : share >= 100 ? 'green' : 'yellow'
482 const progress = (done: number, total: number, color: string) => {
483 if (Svg) {
484 const share = Math.round(Math.min(1, Math.max(0, total > 0 ? done / total : 0)) * 1000)
485 return <Svg alt={`${done} of ${total} steps`} source={barSvg(share, BAR_COLORS[color] ?? BAR_COLORS.cyan!)} />
486 }
487 const bar = barOf(done, total, width)
488 return <Text wrap="truncate-end"><Text color={color}>{bar.filled}</Text><Text dimColor>{bar.empty}</Text></Text>
489 }
490 // A stage's phases as ticks: done, failed, under way (or waiting on the person), still to come.
491 const steps = (phases: { name: string; status: string }[], stopped: boolean) => {
492 const ticks = phases.map(phase => {
493 const done = PHASE_DONE.has(phase.status)
494 const failed = phase.status === 'failed'
495 const current = phase.name === summary.current || phase.status === 'running'
496 const color = done ? 'green' : failed ? 'red' : current ? (stopped ? 'yellow' : 'cyan') : undefined
497 const mark = done ? '✓' : failed ? '✗' : current ? (stopped ? '!' : '▸') : '○'
498 return { key: phase.name, color, current, quiet: !done && !failed && !current, text: `${mark} ${phaseLabel(phase.name)}` }
499 })
500 if (narrowSteps) {
501 return <Box flexDirection="column">{ticks.map(t =>
502 <Text key={t.key} color={t.color} dimColor={t.quiet} bold={t.current} wrap="truncate-end">{t.text}</Text>)}</Box>
503 }
504 return <Text wrap="wrap">{ticks.map((t, i) =>
505 <Text key={t.key} color={t.color} dimColor={t.quiet} bold={t.current}>{i > 0 ? ' ' : ''}{t.text}</Text>)}</Text>
506 }
507
508 // What the open stage shows beneath its row: only the detail that stage needs.
509 const detail = (id: string, state: string, phases: { name: string; status: string }[]) => {
510 if (id === 'preflight') {
511 return preflight?.checks
512 ? preflight.checks.map(check => {
513 const message = checkMessage(check)
514 return (
515 <Box key={`check-${check.id}`} flexDirection="column">
516 <Text dimColor={check.status === 'waiting'} color={CHECK_COLORS[check.status]}>
517 {CHECK_MARKS[check.status] ?? '·'} {check.label}
518 </Text>
519 {message && <Text dimColor={check.status === 'checking'}> {message}</Text>}
520 </Box>
521 )
522 })
523 : <Text dimColor>Checking the site, the tools and the Figma file.</Text>
524 }
525 if (id === 'build') {
526 const stopped = state === 'stopped'
527 const building = runner && runner.state === 'building' && runner.stepsTotal
528 // Stopped, the Needs you card says what to do; the step count would only repeat the bar.
529 const line = !runner || stopped || state === 'failed' ? null
530 : runner.stepsTotal && (runner.state === 'building' || runner.state === 'done')
531 ? `${runner.stepsDone ?? 0} of ${runner.stepsTotal} steps` : stepsLine(runner)
532 return (
533 <Box flexDirection="column" gap={1}>
534 {steps(phases, stopped)}
535 {(building || line) && runner && (
536 <Box flexDirection="column">
537 {building && progress(runner.stepsDone ?? 0, runner.stepsTotal!, stopped ? 'yellow' : state === 'failed' ? 'red' : 'cyan')}
538 {line && <Text dimColor wrap="truncate-end">{line}</Text>}
539 </Box>
540 )}
541 {runner && (
542 <Text wrap="truncate-end">
543 <Text color={runner.connected || runner.state === 'done' ? 'green' : isIdle(runner) ? 'gray' : 'yellow'}>● </Text>
544 <Text dimColor={isIdle(runner)}>{runnerLine(runner)}</Text>
545 </Text>
546 )}
547 </Box>
548 )
549 }
550 if (id === 'verify') return <Text dimColor>Checking the file against the design-lab standard.</Text>
551 if (id === 'report') return <Text dimColor>Scoring the run and writing the report.</Text>
552 return steps(phases, state === 'stopped')
553 }
554
555 const MARK: Record<string, string> = { done: '✓', flagged: '!', active: '▸', stopped: '!', failed: '✗', pending: '○', reused: '↺' }
556 const COLOR: Record<string, string | undefined> = {
557 done: 'green', flagged: 'yellow', active: 'cyan', stopped: 'yellow', failed: 'red', pending: undefined, reused: undefined,
558 }
559 // One card opens: a failure first, then a stop for the person, then the first stage under way.
560 const opened = stages.find(stage => stage.state === 'failed') ?? stages.find(stage => stage.state === 'stopped')
561 ?? stages.find(stage => stage.state === 'active')
562
563 return (
564 <Box flexDirection="column" marginTop={1} gap={1}>
565 {/* Header: what the run is, and where it stands in one colored word. */}
566 <Box flexDirection="column">
567 <Box flexDirection="row" justifyContent="space-between" alignItems="center">
568 <Text bold wrap="truncate-end">{summary.siteLabel ?? run}</Text>
569 <Text bold inverse color={tone.color}>{`\u00a0${tone.label}\u00a0`}</Text>
570 </Box>
571 <Text dimColor wrap="truncate-end">{subtitle.filter(Boolean).join(' · ')}</Text>
572 </Box>
573
574 {/* What went wrong, above everything else. */}
575 {failure && (
576 <Box key="failed" flexDirection="column" borderStyle="round" borderColor="red" paddingX={1}>
577 <Text bold color="red">Failed</Text>
578 <Text>{failure.text}</Text>
579 </Box>
580 )}
581
582 {/* The one thing the person has to do. */}
583 {need && (
584 <Box key="needs-you" flexDirection="column" borderStyle="round" borderColor="yellow" paddingX={1}>
585 <Text bold color="yellow">Needs you</Text>
586 <Text>{need.message}</Text>
587 {need.canResume && <Text dimColor>When the runner is open again, press Resume run.</Text>}
588 {summary.facts.figmaUrl && <Markdown text={`[Open the Figma file ↗](${summary.facts.figmaUrl})`} />}
589 {need.canResume && (
590 <Box marginTop={1}>
591 <Button key="resume" label="Resume run" onPress={() => void resume($, run)} />
592 </Box>
593 )}
594 </Box>
595 )}
596
597 {/* The five stages, in order: one row each, the open one with its detail beneath. */}
598 <Box flexDirection="column">
599 {stages.map((stage, i) => {
600 const lit = stage.state === 'active' || stage.state === 'stopped' || stage.state === 'failed'
601 const color = COLOR[stage.state]
602 return (
603 <Box key={`stage-${stage.id}`} flexDirection="column">
604 <Box flexDirection="row" justifyContent="space-between" gap={1}>
605 <Box flexShrink={0}>
606 <Text bold={lit} dimColor={stage.state === 'pending' || stage.state === 'reused'}
607 color={lit || stage.state === 'flagged' ? color : undefined}>
608 <Text color={color}>{MARK[stage.state]}</Text> {i + 1} {stage.label}
609 </Text>
610 </Box>
611 {stage.note ? <Text color={stage.noteColor} dimColor={!stage.noteColor && !lit}
612 wrap={stage.id === 'verify' ? 'wrap' : 'truncate-end'}>{stage.note}</Text> : null}
613 </Box>
614 {opened === stage && (
615 // Stopped, the card stays quiet: the Needs you card is the only yellow box.
616 <Box key={`card-${stage.id}`} flexDirection="column" marginLeft={2} marginY={1} paddingX={1} borderStyle="round"
617 borderColor={stage.state === 'stopped' ? undefined : color} borderDimColor>
618 {detail(stage.id, stage.state, stage.phases)}
619 </Box>
620 )}
621 </Box>
622 )
623 })}
624 </Box>
625
626 {/* Finished: the verdict, the report's figures, then what the run made and where it is. */}
627 {finished && (
628 <Box flexDirection="column" gap={1}>
629 {verdict && (
630 <Box key="verdict" flexDirection="column" borderStyle="round" borderColor="yellow" paddingX={1}>
631 <Text bold color="yellow">Verification found problems</Text>
632 <Text>{verdict}</Text>
633 <Markdown text={`[Open the verification findings](${fileUrl(run, ARTIFACT_FILES.verifyReport)})`} />
634 </Box>
635 )}
636 {scores && (() => {
637 const coverage = percentOf(scores.built, scores.eligible)
638 const accuracy = percentOf(scores.withinTolerance, scores.widths)
639 const missing = scores.built !== null && scores.eligible !== null && scores.eligible > scores.built
640 ? `${scores.eligible - scores.built} not built` : null
641 return (
642 <Box flexDirection="column" gap={1}>
643 <Box flexDirection={narrowTiles ? 'column' : 'row'} gap={1}>
644 {tile('coverage', 'Coverage', coverage !== null ? `${coverage}%` : null,
645 [scores.built !== null && scores.eligible !== null ? `${scores.built} of ${scores.eligible} buildable` : null, missing],
646 judged(coverage))}
647 {tile('accuracy', 'Accuracy', accuracy !== null ? `${accuracy}%` : null,
648 [scores.withinTolerance !== null && scores.widths !== null ? `${scores.withinTolerance} of ${scores.widths} widths within tolerance` : null],
649 judged(accuracy))}
650 </Box>
651 <Box flexDirection={narrowTiles ? 'column' : 'row'} gap={1}>
652 {tile('time', 'Time', durationOf(scores.workingSeconds),
653 [scores.buildSeconds !== null ? `Figma build ${durationOf(scores.buildSeconds)}` : null], 'magenta')}
654 {tile('tokens', 'Tokens', compactOf(scores.tokens),
655 [scores.toolCalls !== null ? `${scores.toolCalls} tool calls` : null], 'blue')}
656 </Box>
657 </Box>
658 )
659 })()}
660 {(() => {
661 const links = artifactsOf(run, summary.facts.figmaUrl, new Set(summary.present))
662 return (links.main || links.files) && (
663 <Box flexDirection="column" gap={1}>
664 {links.main && (
665 <Box flexDirection="column">
666 <Text bold dimColor>Results</Text>
667 <Markdown text={links.main} />
668 </Box>
669 )}
670 {links.files && (
671 <Box flexDirection="column">
672 <Text bold dimColor>Run files</Text>
673 <Markdown text={links.files} />
674 </Box>
675 )}
676 </Box>
677 )
678 })()}
679 {summary.recap && (
680 <Box flexDirection="column">
681 <Box flexDirection="row" justifyContent="space-between" alignItems="center">
682 <Text bold dimColor>Recap</Text>
683 <Button key="recap" label={recapOpen ? 'Hide' : 'Show'} onPress={() => void update($, recapOpenAtom, open => open === run ? null : run)} />
684 </Box>
685 {recapOpen && <Markdown text={summary.recap} />}
686 </Box>
687 )}
688 </Box>
689 )}
690 </Box>
691 )
692 })
693}
694
695/** A file in the run folder as a link: each path segment encoded, so a space, #, ? or bracket in a
696 * folder name cannot end or break the link. */
697export function fileUrl(run: string, path: string): string {
698 const encode = (segment: string) => encodeURIComponent(segment).replace(/[!'()*]/g, c => `%${c.charCodeAt(0).toString(16).toUpperCase()}`)
699 return `file://${`${run}/${path}`.split('/').map(encode).join('/')}`
700}
701
702/** What a finished run made, as links: the two that matter (the Figma file and the benchmark
703 * report), then the run's own files. Only files that exist are listed. */
704export function artifactsOf(run: string, figmaUrl: string | null, present: Set<string>): { main: string | null; files: string | null } {
705 const main = [
706 figmaUrl ? `**[Open the Figma library ↗](${figmaUrl})**` : null,
707 present.has(ARTIFACT_FILES.report) ? `**[Open the benchmark report](${fileUrl(run, ARTIFACT_FILES.report)})**` : null,
708 ].filter(Boolean)
709 const files = ([
710 ['Verification findings', ARTIFACT_FILES.verifyReport],
711 ['Build plan', ARTIFACT_FILES.plan],
712 ['Components', ARTIFACT_FILES.components],
713 ['Scorecard', ARTIFACT_FILES.scorecard],
714 ] as const).filter(([, path]) => present.has(path)).map(([label, path]) => `[${label}](${fileUrl(run, path)})`)
715 // No bullets: Markdown indents them unevenly. The two results stand one per line (a hard break
716 // is two trailing spaces); the run files share one line.
717 return { main: main.length ? main.join(' \n') : null, files: files.length ? files.join(' · ') : null }
718}
719src/protocol.ts 210 lines1/** Self-contained values: safe in the Node server, Figma stripping and the no-Node mod host.
2 * The only imports are erased types generated from our JSON schemas. */
3import type { Progress } from './generated/progress.ts';
4import type { RunnerStep } from './generated/runner-step.ts';
5export const PORT = 8765;
6export const WAIT_MS = 5000;
7export const RETRY_MS = WAIT_MS;
8export const HEARTBEAT_SECONDS = 10;
9export const SERVER_FRESH_MS = 3 * HEARTBEAT_SECONDS * 1000;
10export const RUNNER_ABSENT_MS = 120_000;
11export const POLL_MS = WAIT_MS;
12export type ProgressState = Progress['state'];
13export type StepKind = RunnerStep['kind'];
14export const PROGRESS_STATES = {
15 waiting: 'waiting',
16 preflight: 'preflight',
17 building: 'building',
18 done: 'done',
19 failed: 'failed',
20} as const satisfies { [State in ProgressState]: State };
21export const STEP_KINDS = {
22 wait: 'wait',
23 done: 'done',
24 check: 'check',
25 dump: 'dump',
26 use_figma: 'use_figma',
27 upload: 'upload',
28 screenshot: 'screenshot',
29 skip: 'skip',
30} as const satisfies { [Kind in StepKind]: Kind };
31export const isProgressState = (value: string): value is ProgressState =>
32 Object.values(PROGRESS_STATES).some((state) => state === value);
33export const isStepKind = (value: string): value is StepKind =>
34 Object.values(STEP_KINDS).some((kind) => kind === value);
35
36/** Finite workflow vocabulary; dynamic artifact paths remain generated typed maps. */
37export const PHASE_NAMES = [
38 'init',
39 'discovery',
40 'inventory',
41 'usage',
42 'capture',
43 'tokens',
44 'plan',
45 'variables',
46 'preflight',
47 'connect',
48 'foundation',
49 'components',
50 'index',
51 'verify',
52 'benchmark',
53] as const;
54export type PhaseName = (typeof PHASE_NAMES)[number];
55export const REGISTRABLE_PHASES = [
56 'usage',
57 'capture',
58 'foundation',
59 'components',
60 'index',
61 'verify',
62] as const satisfies readonly PhaseName[];
63export type RegistrablePhase = (typeof REGISTRABLE_PHASES)[number];
64export const RECORDABLE_PHASES = [...REGISTRABLE_PHASES, 'benchmark'] as const satisfies readonly PhaseName[];
65export type RecordablePhase = (typeof RECORDABLE_PHASES)[number];
66export const PHASE_STATUSES = [
67 'pending',
68 'running',
69 'complete',
70 'failed',
71 'waived',
72 'awaiting-approval',
73 'approved',
74 'waiting',
75 'stopped',
76 'invalidated',
77] as const;
78export type PhaseStatus = (typeof PHASE_STATUSES)[number];
79export type CheckStatus = 'done' | 'checking' | 'needs-you' | 'failed' | 'waiting';
80export interface ChecklistDocument {
81 pass: string;
82 at: string;
83 ready: boolean | null;
84 goAheadAt?: string | null;
85 checks: (Omit<Check, 'status'> & { status: CheckStatus; at: string })[];
86}
87export interface Handshake {
88 ok: boolean;
89 at?: string;
90 failure?: string;
91 runnerConnected?: boolean;
92 fileKey?: string | null;
93 fileName?: string | null;
94 fileKeyMatches?: boolean;
95 empty?: boolean;
96 onlyPreflightCover?: boolean;
97 writable?: boolean;
98 pluginData?: boolean;
99 connectionOnly?: boolean;
100 coverPageId?: string;
101 coverId?: string;
102 font?: string;
103 fontLoaded?: boolean;
104 fonts?: Record<string, string[]> | null;
105 server?: { pid: number | null; started: boolean };
106 install?: { folder: string; manifest: string; version: string; firstInstall: boolean; updated: boolean };
107 instructions?: string[];
108 outdated?: boolean;
109 runnerVersion?: string;
110}
111
112// DESIGN_LAB_MOD_CONTRACT_BEGIN
113// reused: copied from an earlier run (design-lab:figma-build), not run again here
114export type Phase = { name: string; status: string; reused: boolean };
115
116// What the phases recorded that the pane says beside each stage; each null until recorded.
117export type Facts = {
118 found: number | null;
119 toBuild: number | null;
120 built: number | null;
121 expected: number | null;
122 figmaUrl: string | null;
123};
124
125// The verification report's open findings by severity, the checks it passed, and those it waived.
126export type Findings = { blocker: number; major: number; minor: number; passed: number; waived: number };
127
128export type Runner = {
129 state: ProgressState | 'connecting';
130 stepsDone: number | null;
131 stepsTotal: number | null;
132 stepKind: StepKind | null;
133 message: string | null;
134 serverAlive: boolean;
135 connected: boolean;
136 lastSeenMs: number | null;
137};
138
139export type Check = {
140 id: string;
141 label: string;
142 status: string;
143 message: string | null;
144 dependsOn: string[];
145};
146
147// The headline figures of a finished run, from its scorecard: each null when the scorer left it out.
148export type Scores = {
149 built: number | null;
150 eligible: number | null;
151 withinTolerance: number | null;
152 widths: number | null;
153 workingSeconds: number | null;
154 buildSeconds: number | null;
155 buildSteps: number | null;
156 tokens: number | null;
157 toolCalls: number | null;
158 blockers: number | null;
159 majors: number | null;
160};
161
162type ModSummary = {
163 workspace: string;
164 found: boolean;
165 siteLabel: string | null;
166 phases: Phase[];
167 current: string | null;
168 // checks: the list preflight last wrote, or null when it has none newer than the recorded phase
169 preflight: { status: string; at: string | null; checks: Check[] | null } | null;
170 runner: Runner | null;
171 blocker: string | null;
172 // the run is waiting for the person to do something it will notice by itself (start the runner)
173 waiting: string | null;
174 log: string[];
175 hasRecap: boolean;
176 recap: string | null;
177 scores: Scores | null;
178 facts: Facts;
179 findings: Findings | null;
180 startedAt: string | null;
181 // the last message the phase log holds for each failed phase, by phase name
182 phaseErrors: Record<string, string>;
183 // the run's artifact files that exist, run-relative (read only once the recap is written)
184 present: string[];
185};
186
187// DESIGN_LAB_MOD_CONTRACT_END
188
189/** Both readers describe one run, but the CLI keeps its historical file-path/seconds
190 * projection; the pane adds presentation facts and recap text. The surface parameter
191 * makes these differences explicit without weakening either contract. */
192type CliRunner = Omit<Runner, 'lastSeenMs'> & { lastSeenSeconds: number | null };
193type CliSummary =
194 | { found: false; workspace: string; recap?: never }
195 | {
196 found: true;
197 workspace: string;
198 siteLabel?: string;
199 phases: Omit<Phase, 'reused'>[];
200 nextPhase: string | null;
201 preflightChecks: Check[] | null;
202 runner: CliRunner | null;
203 blocker: string | null;
204 waiting: string | null;
205 recap: string | null;
206 startedAt: string | null;
207 elapsedSeconds: number | null;
208 };
209export type Summary<Surface extends 'cli' | 'mod' = 'cli'> = Surface extends 'mod' ? ModSummary : CliSummary;
210hooks/mod/model.ts 704 lines1// What a design-lab run looks like from the files it writes, with no engine calls: the same
2// reading `workflow.ts watch` does, so the pane and the text fallback agree.
3
4import type { Check, Facts, Findings, Phase, Runner, Scores, Summary as RunSummary } from '../../src/protocol';
5import type { Summary as ManifestSummary } from '../../types';
6type Summary = RunSummary<'mod'>;
7type Assignable<Expected, Actual extends Expected> = Actual;
8// Both directions enforce the generated host contract without runtime dependencies.
9export type HostSummaryMatchesProtocol = Assignable<Summary, ManifestSummary>;
10export type ProtocolSummaryMatchesHost = Assignable<ManifestSummary, Summary>;
11
12// protocol.ts has no runtime dependencies, so these values are safe in the mod host.
13import { SERVER_FRESH_MS, RUNNER_ABSENT_MS, isProgressState, isStepKind } from '../../src/protocol.ts';
14export { SERVER_FRESH_MS, RUNNER_ABSENT_MS } from '../../src/protocol.ts';
15export const LOG_LINES = 8;
16// The Markdown element draws at most 10,000 characters.
17export const RECAP_LIMIT = 9_500;
18
19const DONE = new Set(['complete', 'approved', 'waived']);
20
21/** The files the mod reads, as text; a missing or unreadable file is undefined. */
22export type Raw = {
23 project?: unknown;
24 phaseLog?: string;
25 progress?: unknown;
26 runnerLog?: string;
27 completion?: string;
28 scorecard?: unknown;
29 preflightChecks?: unknown;
30 verifyReport?: unknown;
31 // the artifact files found in the run folder, run-relative
32 present?: string[];
33};
34
35export function parseJson(text: string | undefined): unknown {
36 if (text === undefined) return undefined;
37 try {
38 return JSON.parse(text);
39 } catch {
40 return undefined;
41 }
42}
43
44function record(value: unknown): Record<string, unknown> {
45 return typeof value === 'object' && value !== null ? (value as Record<string, unknown>) : {};
46}
47
48function text(value: unknown): string | null {
49 return typeof value === 'string' ? value : null;
50}
51
52function count(value: unknown): number | null {
53 return typeof value === 'number' ? value : null;
54}
55
56function msSince(stamp: string | null, nowMs: number): number | null {
57 if (!stamp) return null;
58 const at = Date.parse(stamp);
59 return Number.isNaN(at) ? null : nowMs - at;
60}
61
62/** The phase log's lines that parse; a half-written last line is skipped. */
63export function entriesOf(log: string | undefined): Record<string, unknown>[] {
64 if (!log) return [];
65 return log.split('\n').flatMap((line) => {
66 const value = parseJson(line.trim() || undefined);
67 return typeof value === 'object' && value !== null ? [value] : [];
68 });
69}
70
71export function tailOf(log: string | undefined, lines = LOG_LINES): string[] {
72 if (!log) return [];
73 return log
74 .split('\n')
75 .filter((line) => line.trim() !== '')
76 .slice(-lines);
77}
78
79export function runnerOf(progress: unknown, nowMs: number): Runner | null {
80 const p = record(progress);
81 if (!('state' in p)) return null;
82 const serverAge = msSince(text(p.at), nowMs);
83 const serverAlive = serverAge !== null && serverAge <= SERVER_FRESH_MS;
84 const seenAge = msSince(text(p.lastSeen), nowMs);
85 const asked = seenAge !== null && seenAge <= RUNNER_ABSENT_MS;
86 const connected = serverAlive && (asked || p.inflight === true);
87 return {
88 state: typeof p.state === 'string' && isProgressState(p.state) ? p.state : 'waiting',
89 stepsDone: count(p.stepsDone),
90 stepsTotal: count(p.stepsTotal),
91 stepKind: typeof p.stepKind === 'string' && isStepKind(p.stepKind) ? p.stepKind : null,
92 message: text(p.message),
93 serverAlive,
94 connected,
95 lastSeenMs: seenAge,
96 };
97}
98
99export function summaryOf(workspace: string, raw: Raw, nowMs: number): Summary {
100 const project = record(raw.project);
101 if (raw.project === undefined) {
102 return {
103 workspace,
104 found: false,
105 siteLabel: null,
106 phases: [],
107 current: null,
108 preflight: null,
109 runner: null,
110 blocker: null,
111 waiting: null,
112 log: [],
113 hasRecap: false,
114 recap: null,
115 scores: null,
116 facts: { found: null, toBuild: null, built: null, expected: null, figmaUrl: null },
117 findings: null,
118 startedAt: null,
119 phaseErrors: {},
120 present: [],
121 };
122 }
123 const recorded: Phase[] = Object.entries(record(project.phases)).map(([name, value]) => ({
124 name,
125 status: text(record(value).status) ?? 'pending',
126 reused: Boolean(record(value).from),
127 }));
128 const finished = recapIsCurrent(project, raw);
129 // Preflight records its phase only once it passes, so until then a run that has not reached the
130 // build is still in preflight. A rebuild copies preflight from its source and never runs it here.
131 const preflighting =
132 !recorded.some((phase) => phase.name === 'preflight') &&
133 !recorded.some((phase) => phase.reused) &&
134 !finished &&
135 !recorded.some((phase) => BUILD_PHASES.has(phase.name) && phase.status !== 'pending');
136 // In the order the run takes them, not the order project.json happens to hold them.
137 const phases = [...recorded, ...(preflighting ? [{ name: 'preflight', status: 'running', reused: false }] : [])].sort(
138 (a, b) => flowRank(a.name) - flowRank(b.name),
139 );
140 // A phase copied from an earlier run never takes the current phase, whatever status it copied.
141 const own = phases.filter((phase) => !phase.reused);
142 const running = own.find((phase) => phase.status === 'running');
143 const due = own.find((phase) => !DONE.has(phase.status));
144 const runner = runnerOf(raw.progress, nowMs);
145 const entries = entriesOf(raw.phaseLog);
146 const last = entries[entries.length - 1];
147 const phaseErrors: Record<string, string> = {};
148 for (const phase of phases) {
149 if (phase.status !== 'failed') continue;
150 const said = entries.filter((entry) => entry.phase === phase.name && typeof entry.message === 'string').pop();
151 if (said) phaseErrors[phase.name] = said.message as string;
152 }
153 // The open blocker: the newest entry stopped the run for the person, and the runner has not
154 // come back since.
155 const blocker = last && last.status === 'stopped' && !runner?.connected ? text(last.message) : null;
156 // Waiting on the person for something the run notices by itself (the runner starting at the
157 // build's connection): what to do, with nothing to press.
158 const waiting = last && last.status === 'waiting' && !runner?.connected ? text(last.message) : null;
159 // While the build waits for the person to start the runner, the runner is awaited, not idle.
160 const shown = waiting && runner && runner.state === 'waiting' ? { ...runner, state: 'connecting' as const } : runner;
161 const preflight = record(record(project.phases).preflight);
162 const preflightAt = preflight.from ? null : text(preflight.updatedAt);
163 const checks = checksOf(raw.preflightChecks, preflight);
164 return {
165 workspace,
166 found: true,
167 siteLabel: text(record(project.run).siteLabel),
168 phases,
169 current: (running ?? due)?.name ?? null,
170 preflight:
171 preflight.status || checks ? { status: text(preflight.status) ?? 'running', at: preflightAt, checks } : null,
172 runner: shown,
173 blocker,
174 waiting,
175 log: tailOf(raw.runnerLog),
176 hasRecap: finished,
177 recap: finished ? recapOf(raw.completion!) : null,
178 scores: finished ? scoresOf(raw.scorecard) : null,
179 facts: factsOf(project, finished ? raw.completion : undefined),
180 findings: findingsOf(raw.verifyReport, text(project.createdAt)),
181 startedAt: preflightAt ?? text(project.createdAt),
182 phaseErrors,
183 present: raw.present ?? [],
184 };
185}
186
187export const CHECK_MARKS: Record<string, string> = {
188 done: '✓',
189 checking: '▸',
190 'needs-you': '!',
191 failed: '✗',
192 waiting: '·',
193};
194export const CHECK_COLORS: Record<string, string | undefined> = {
195 done: 'green',
196 checking: 'cyan',
197 'needs-you': 'yellow',
198 failed: 'red',
199};
200
201/** The checklist preflight last wrote, or null when there is none or it is older than the recorded
202 * preflight phase (a run that passed preflight before the checklist existed). */
203export function checksOf(document: unknown, phase: Record<string, unknown>): Check[] | null {
204 const d = record(document);
205 if (!Array.isArray(d.checks)) return null;
206 const written = Date.parse(text(d.at) ?? '');
207 const passed = Date.parse(phase.from ? '' : (text(phase.updatedAt) ?? ''));
208 if (phase.status === 'complete' && written < passed) return null;
209 return d.checks.flatMap((value) => {
210 const c = record(value);
211 const id = text(c.id);
212 return id === null
213 ? []
214 : [
215 {
216 id,
217 label: text(c.label) ?? id,
218 status: text(c.status) ?? 'waiting',
219 message: text(c.message),
220 dependsOn: Array.isArray(c.dependsOn) ? c.dependsOn.filter((x): x is string => typeof x === 'string') : [],
221 },
222 ];
223 });
224}
225
226/** Preflight passed and every check is done: the group can fold to one line. */
227export function preflightPassed(summary: Summary): boolean {
228 const p = summary.preflight;
229 return p?.status === 'complete' && (p.checks ?? []).every((check) => check.status === 'done');
230}
231
232/** What a check says beside its label: only while it is working or needs something. */
233export function checkMessage(check: Check): string | null {
234 return ['checking', 'needs-you', 'failed'].includes(check.status) ? check.message : null;
235}
236
237/** HH:MM on the person's clock. */
238export function clockOf(stamp: string | null): string | null {
239 if (!stamp) return null;
240 const at = new Date(stamp);
241 if (Number.isNaN(at.getTime())) return null;
242 return `${String(at.getHours()).padStart(2, '0')}:${String(at.getMinutes()).padStart(2, '0')}`;
243}
244
245/** The recap belongs to this build: its scorecard names the build's creation, or, from a scorer
246 * before that stamp, the build has recorded its benchmark as complete. A folder initialised again
247 * keeps the old benchmark/ folder until it is scored, and that must not read as done. */
248export function recapIsCurrent(project: Record<string, unknown>, raw: Raw): boolean {
249 if (raw.completion === undefined) return false;
250 const stamp = record(record(raw.scorecard).run).buildCreatedAt;
251 if (typeof stamp === 'string') return stamp === text(project.createdAt);
252 return text(record(record(project.phases).benchmark).status) === 'complete';
253}
254
255/** The completion message as written, cut with a note if it ever outgrows what Markdown draws. */
256export function recapOf(completion: string): string {
257 const text = completion.trim();
258 return text.length <= RECAP_LIMIT
259 ? text
260 : `${text.slice(0, RECAP_LIMIT)}\n\n(cut here: the whole message is in benchmark/completion.md)`;
261}
262
263/** Finished: the scorer has written this build's recap. The runner reports done once the Figma
264 * build is over, while verification and scoring still have to run. */
265export function isFinished(summary: Summary): boolean {
266 return summary.hasRecap;
267}
268
269/** The run is stopped for the person: building or checking, and the runner is gone. */
270export function isDown(summary: Summary): boolean {
271 if (!summary.found || isFinished(summary)) return false;
272 if (summary.blocker) return true;
273 const runner = summary.runner;
274 return runner !== null && (runner.state === 'building' || runner.state === 'preflight') && !runner.connected;
275}
276
277/** Before the build starts, or between builds, the runner is not needed: idle, not missing. */
278export function isIdle(runner: Runner): boolean {
279 return runner.state === 'waiting' && !runner.connected;
280}
281
282export function runnerLine(runner: Runner): string {
283 if (runner.state === 'done') return 'Figma build finished';
284 if (!runner.serverAlive) return 'runner server not responding';
285 if (runner.connected) return 'runner connected';
286 if (isIdle(runner)) return 'runner idle until the build';
287 if (runner.state === 'connecting') return 'waiting for the runner to start';
288 const minutes = Math.max(1, Math.round((runner.lastSeenMs ?? 0) / 60_000));
289 return `runner not seen for ${minutes}m`;
290}
291
292export function stepsLine(runner: Runner): string | null {
293 if ((runner.state === 'building' || runner.state === 'done') && runner.stepsTotal) {
294 const kind = runner.state === 'building' && runner.stepKind ? `, ${runner.stepKind}` : '';
295 return `steps ${runner.stepsDone ?? 0}/${runner.stepsTotal}${kind}`;
296 }
297 // The server's last word to a runner that has since gone quiet ("Connected. Waiting…") is stale.
298 return isIdle(runner) || runner.state === 'connecting' ? null : runner.message;
299}
300
301/** A length of time as the pane writes it: 41m, 3h 44m; under a minute, 40s. */
302export function durationOf(seconds: number | null): string | null {
303 if (seconds === null || seconds < 0) return null;
304 if (seconds < 60) return `${Math.round(seconds)}s`;
305 // Rounded, as the recap rounds it, so the two agree.
306 const minutes = Math.round(seconds / 60);
307 return minutes < 60 ? `${minutes}m` : `${Math.floor(minutes / 60)}h ${minutes % 60}m`;
308}
309
310/** How long the run has been going, from its start. */
311export function elapsedOf(startedAt: string | null, nowMs: number): string | null {
312 const ms = msSince(startedAt, nowMs);
313 if (ms === null || ms < 0) return null;
314 return ms < 60_000 ? '0m' : durationOf(Math.floor(ms / 60_000) * 60);
315}
316
317/** The scorecard's headline figures; null when there is no scorecard to read. */
318export function scoresOf(scorecard: unknown): Scores | null {
319 const card = record(scorecard);
320 if (!('headline' in card)) return null;
321 const headline = record(card.headline);
322 const coverage = record(headline.coverage);
323 const accuracy = record(record(headline.accuracy).corrected);
324 const effort = record(headline.effort);
325 const open = record(record(record(card.sections).conformance).open);
326 return {
327 built: count(coverage.built),
328 eligible: count(coverage.eligible),
329 withinTolerance: count(accuracy.pass),
330 widths: count(accuracy.total),
331 workingSeconds: count(effort.workingSeconds),
332 buildSeconds: count(effort.buildSeconds),
333 buildSteps: count(effort.buildSteps),
334 tokens: count(effort.tokens),
335 toolCalls: count(effort.toolCalls),
336 blockers: count(open.blocker),
337 majors: count(open.major),
338 };
339}
340
341/** The first Figma file link in the recap, without the punctuation of the sentence it ends. */
342export function figmaUrlOf(completion: string): string | null {
343 return /https:\/\/www\.figma\.com\/(?:design|file)\/[^\s)>\]]+/.exec(completion)?.[0].replace(/[.,;:]+$/, '') ?? null;
344}
345
346/** 40013109 as 40.0M, 563519 as 564K. */
347export function compactOf(value: number | null): string | null {
348 if (value === null) return null;
349 if (value >= 1_000_000) return `${(value / 1_000_000).toFixed(1)}M`;
350 if (value >= 1_000) return `${Math.round(value / 1_000)}K`;
351 return String(value);
352}
353
354/** A share as a whole percent, or null with nothing to divide by. Anything short of the whole
355 * stays below 100, so 200 of 201 reads 99% and is never judged complete. */
356export function percentOf(part: number | null, whole: number | null): number | null {
357 if (part === null || !whole) return null;
358 const percent = Math.round((part / whole) * 100);
359 return part < whole && percent >= 100 ? 99 : percent;
360}
361
362/** A bar of `cells` cells, filled in proportion: the filled run and the empty run, drawn apart. */
363export function barOf(done: number, total: number, cells: number): { filled: string; empty: string } {
364 const width = Math.max(4, cells);
365 const share = total > 0 ? Math.min(1, Math.max(0, done / total)) : 0;
366 const filled = Math.round(share * width);
367 return { filled: '━'.repeat(filled), empty: '━'.repeat(width - filled) };
368}
369
370/** Where the run stands, in one word, and the color it is drawn in. */
371export type Tone = { label: string; color: string };
372
373export function toneOf(summary: Summary): Tone {
374 const stages = stagesOf(summary);
375 if (stages.some((stage) => stage.state === 'failed')) return { label: 'Failed', color: 'red' };
376 if (needsYouOf(summary)) return { label: 'Needs you', color: 'yellow' };
377 if (isFinished(summary)) {
378 return verdictOf(summary.findings)
379 ? { label: 'Done · needs review', color: 'yellow' }
380 : { label: 'Done', color: 'green' };
381 }
382 const active = stages.find((stage) => stage.state === 'active' || stage.state === 'stopped');
383 return { label: active?.doing ?? 'Starting', color: 'cyan' };
384}
385
386/** What the person has to do, from one place, so the header, the card, the stage and the toast
387 * always agree; null when the run needs nothing from them. */
388export function needsYouOf(summary: Summary): { message: string; canResume: boolean } | null {
389 // A failure outranks every request: the Failed card says what went wrong, and resuming would
390 // only run into it again.
391 if (hasFailed(summary)) return null;
392 if (summary.blocker) return { message: summary.blocker, canResume: true };
393 if (summary.waiting) return { message: summary.waiting, canResume: false };
394 const check = (summary.preflight?.checks ?? []).find((c) => c.status === 'needs-you');
395 if (check) return { message: `${check.label}: ${check.message ?? 'needs your attention'}`, canResume: false };
396 if (isDown(summary)) {
397 return {
398 message: 'The Figma runner has stopped. Reopen it in Figma desktop, then press Resume run.',
399 canResume: true,
400 };
401 }
402 return null;
403}
404
405/** Something failed: a phase run here, the runner while the run is unfinished, or a preflight check. */
406export function hasFailed(summary: Summary): boolean {
407 return (
408 summary.phases.some((phase) => !phase.reused && phase.status === 'failed') ||
409 (summary.runner?.state === 'failed' && !isFinished(summary)) ||
410 (summary.preflight?.checks ?? []).some((check) => check.status === 'failed')
411 );
412}
413
414/** The first stage that failed, and what the card says about it; null when nothing failed. */
415export function failureOf(summary: Summary): { stage: string; text: string } | null {
416 const stage = stagesOf(summary).find((s) => s.state === 'failed');
417 if (!stage) return null;
418 const check =
419 stage.id === 'preflight' ? (summary.preflight?.checks ?? []).find((c) => c.status === 'failed') : undefined;
420 const phase = stage.phases.find((p) => !p.reused && p.status === 'failed');
421 const message = check
422 ? check.message
423 ? `${check.label}: ${check.message}`
424 : check.label
425 : ((phase && summary.phaseErrors[phase.name]) ??
426 (stage.id === 'build' && summary.runner?.state === 'failed' ? summary.runner.message : null));
427 return {
428 stage: stage.label,
429 text: message
430 ? `${stage.label} stopped with an error: ${message}`
431 : `${stage.label} stopped with an error. Ask Claude in the conversation what went wrong.`,
432 };
433}
434
435/** What the verdict card says when verification left blocking or major problems open; null otherwise. */
436export function verdictOf(findings: Findings | null): string | null {
437 if (!findings || findings.blocker + findings.major === 0) return null;
438 const { blocker, major } = findings;
439 const parts = [blocker ? `${blocker} blocking` : null, major ? `${major} major` : null].filter(Boolean);
440 const one = blocker + major === 1;
441 return `${parts.join(' and ')} problem${one ? ' is' : 's are'} still open, so the library does not yet meet the design-lab standard.`;
442}
443
444/** A phase name as a person reads it: figma-build as Figma build. */
445export function phaseLabel(name: string): string {
446 const words = name.replace(/[-_]/g, ' ');
447 return words.charAt(0).toUpperCase() + words.slice(1);
448}
449
450export const PHASE_DONE = DONE;
451
452/** The status line, or undefined to clear it once the run is over or gone. */
453export function statusOf(summary: Summary, nowMs: number): string | undefined {
454 if (!summary.found) return undefined;
455 if (isFinished(summary)) return undefined;
456 // The engine shows the plugin's name before it: `design-lab: steps 112/158 · runner connected · 41m`.
457 const parts: string[] = [];
458 const runner = summary.runner;
459 const checks = summary.preflight?.checks;
460 if (runner?.state === 'building' && runner.stepsTotal)
461 parts.push(`steps ${runner.stepsDone ?? 0}/${runner.stepsTotal}`);
462 else if (checks && summary.preflight?.status !== 'complete') {
463 parts.push(`preflight ${checks.filter((check) => check.status === 'done').length}/${checks.length}`);
464 } else if (summary.current) parts.push(summary.current);
465 if (runner) parts.push(runnerLine(runner));
466 const time = elapsedOf(summary.startedAt, nowMs);
467 if (time) parts.push(time);
468 return parts.length > 0 ? parts.join(' · ') : undefined;
469}
470
471/** The whole summary as plain text: the command's answer where nothing draws. */
472export function plainOf(summary: Summary): string {
473 if (!summary.found) return `No design-lab run in ${summary.workspace}: it has no project.json.`;
474 const lines = [`design-lab · ${summary.siteLabel ?? summary.workspace}`, ''];
475 const checks = summary.preflight?.checks;
476 if (checks && checks.length > 0) {
477 lines.push(' Preflight');
478 for (const check of checks) {
479 const message = checkMessage(check);
480 lines.push(` ${CHECK_MARKS[check.status] ?? '·'} ${check.label}${message ? `: ${message}` : ''}`);
481 }
482 lines.push('');
483 }
484 for (const phase of summary.phases) {
485 const mark = DONE.has(phase.status)
486 ? '✓'
487 : phase.status === 'failed'
488 ? '✗'
489 : phase.name === summary.current
490 ? '▸'
491 : phase.status === 'stopped'
492 ? '!'
493 : '·';
494 lines.push(` ${mark} ${phase.name}`);
495 }
496 if (summary.runner) {
497 const steps = stepsLine(summary.runner);
498 lines.push('');
499 if (steps) lines.push(` ${steps}`);
500 if (!summary.hasRecap) lines.push(` ${runnerLine(summary.runner)}`);
501 }
502 const failure = failureOf(summary);
503 if (failure) lines.push('', ` Failed: ${failure.text}`);
504 const need = needsYouOf(summary);
505 if (need) lines.push('', ` Needs you: ${need.message}`);
506 if (summary.hasRecap) lines.push('', ` Recap: ${summary.workspace}/benchmark/completion.md`);
507 return lines.join('\n');
508}
509
510/**
511 * What the resume button does after asking to fill the prompt box: nothing more once filled;
512 * send the request where the session has no prompt box at all; otherwise (a dialog holds the
513 * keys, or the reason is unknown) leave the person's draft alone and say why.
514 */
515export function afterFill(filled: { isFilled: boolean; refusal?: string } | undefined): 'done' | 'submit' | 'explain' {
516 if (filled?.isFilled) return 'done';
517 return filled?.refusal === 'no_composer' ? 'submit' : 'explain';
518}
519
520export const RESUME_PROMPT = (workspace: string) =>
521 `The design-lab runner is open again in Figma desktop. Resume the design-lab run in ${workspace} from where it stopped.`;
522
523/** What the phases recorded about the run: counts from inventory, plan and components, the file
524 * from connect (or, failing that, the recap). */
525export function factsOf(project: Record<string, unknown>, completion: string | undefined): Facts {
526 const phases = record(project.phases);
527 const detail = (name: string) => record(record(phases[name]).detail);
528 return {
529 found: count(detail('inventory').components),
530 toBuild: count(detail('plan').build),
531 built: count(detail('components').built),
532 expected: count(detail('components').expected),
533 figmaUrl: text(detail('connect').fileUrl) ?? (completion ? figmaUrlOf(completion) : null),
534 };
535}
536
537/** The open findings by severity, or null before this run's verification has written its report.
538 * A folder initialised again keeps the last run's report, so one written before this run began is
539 * not this run's. */
540export function findingsOf(report: unknown, createdAt: string | null): Findings | null {
541 const r = record(report);
542 if (!Array.isArray(r.open) || !Array.isArray(r.passed)) return null;
543 const written = Date.parse(text(r.generatedAt) ?? '');
544 if (Number.isNaN(written) || written < Date.parse(createdAt ?? '')) return null;
545 const by = (severity: string) => (r.open as unknown[]).filter((f) => record(f).severity === severity).length;
546 return {
547 blocker: by('blocker'),
548 major: by('major'),
549 minor: by('minor'),
550 passed: r.passed.length,
551 waived: Array.isArray(r.waived) ? r.waived.length : 0,
552 };
553}
554
555/** The five stages a run moves through, and the phases each holds. A phase not named here
556 * belongs to Build, so nothing the run records goes unshown. */
557export const STAGES = [
558 { id: 'preflight', label: 'Preflight', doing: 'Preflight', phases: ['preflight'] },
559 {
560 id: 'discovery',
561 label: 'Discovery',
562 doing: 'Discovering',
563 phases: ['discovery', 'inventory', 'usage', 'capture', 'tokens', 'plan'],
564 },
565 { id: 'build', label: 'Build', doing: 'Building', phases: ['connect', 'foundation', 'components', 'index'] },
566 { id: 'verify', label: 'Verify', doing: 'Verifying', phases: ['verify'] },
567 { id: 'report', label: 'Report', doing: 'Reporting', phases: ['benchmark'] },
568] as const;
569
570export type StageId = (typeof STAGES)[number]['id'];
571// flagged: finished with open blocking or major problems; failed: stopped with an error.
572export type StageState = 'done' | 'flagged' | 'active' | 'stopped' | 'failed' | 'pending' | 'reused';
573export type Stage = {
574 id: StageId;
575 label: string;
576 doing: string;
577 state: StageState;
578 phases: Phase[];
579 note: string | null;
580 noteColor?: string;
581};
582
583// Every known phase in the order the run takes them. A phase not named here sits just after index,
584// so it stays with Build and comes before verify.
585const FLOW: string[] = STAGES.flatMap((stage) => [...stage.phases]);
586const BUILD_PHASES = new Set<string>(STAGES.find((stage) => stage.id === 'build')!.phases);
587
588function flowRank(name: string): number {
589 const at = FLOW.indexOf(name);
590 return at >= 0 ? at : FLOW.indexOf('index') + 0.5;
591}
592
593function stageOf(name: string): StageId {
594 return STAGES.find((stage) => (stage.phases as readonly string[]).includes(name))?.id ?? 'build';
595}
596
597const plural = (n: number, one: string, many: string) => `${n} ${n === 1 ? one : many}`;
598
599function stageNote(
600 id: StageId,
601 state: StageState,
602 summary: Summary,
603 phases: Phase[],
604): { note: string | null; noteColor?: string } {
605 const f = summary.facts;
606 if (state === 'reused') return { note: 'reused from an earlier run' };
607 if (state === 'pending') return { note: null };
608 if (id === 'preflight') {
609 const checks = summary.preflight?.checks;
610 if (!checks) return { note: state === 'done' ? 'passed' : null };
611 const done = checks.filter((check) => check.status === 'done').length;
612 return { note: state === 'done' ? `${checks.length} checks passed` : `${done} of ${checks.length} checks` };
613 }
614 if (id === 'discovery') {
615 if (state !== 'done')
616 return { note: `${phases.filter((phase) => DONE.has(phase.status)).length} of ${phases.length} steps` };
617 const parts = [
618 f.found !== null ? `${f.found} found` : null,
619 f.toBuild !== null ? `${f.toBuild} planned` : null,
620 ].filter(Boolean);
621 return { note: parts.length ? parts.join(' · ') : 'complete' };
622 }
623 if (id === 'build') {
624 const runner = summary.runner;
625 if (state !== 'done' && runner?.state === 'building' && runner.stepsTotal)
626 return { note: `${percentOf(runner.stepsDone ?? 0, runner.stepsTotal)}%` };
627 if (f.built !== null && f.expected !== null) {
628 return {
629 note: `${f.built} of ${f.expected} planned built`,
630 noteColor: state === 'done' && f.built < f.expected ? 'yellow' : undefined,
631 };
632 }
633 return { note: state === 'done' ? 'complete' : null };
634 }
635 if (id === 'verify') {
636 const k = summary.findings;
637 if (!k || (state !== 'done' && state !== 'flagged')) return { note: null };
638 // Kept short so the row fits a narrow pane; the verdict card above the figures has the rest.
639 if (state === 'flagged') {
640 const parts = [
641 k.blocker ? plural(k.blocker, 'blocker', 'blockers') : null,
642 k.major ? `${k.major} major` : null,
643 ].filter(Boolean);
644 return { note: `${parts.join(' · ')} open`, noteColor: 'yellow' };
645 }
646 if (k.minor > 0) return { note: `${k.minor} minor open`, noteColor: 'yellow' };
647 // Checks that do not apply to this site are not problems, so the note leaves them out.
648 return {
649 note: k.waived === 0 ? `all ${k.passed} checks pass` : `${k.passed} passed · ${k.waived} waived`,
650 noteColor: 'green',
651 };
652 }
653 return { note: state === 'done' ? 'report ready' : null };
654}
655
656/** Each stage with where it stands, first match winning: copied from an earlier run, failed, done
657 * (flagged when verification left problems open), stopped for the person, under way, not started.
658 * A stage is under way when it holds the run's current phase or any of its phases is running, so
659 * more than one can be. */
660export function stagesOf(summary: Summary): Stage[] {
661 const finished = isFinished(summary);
662 const need = needsYouOf(summary);
663 const checks = summary.preflight?.checks ?? [];
664 const k = summary.findings;
665 return STAGES.map((def) => {
666 // summary.phases is already in flow order.
667 const phases = summary.phases.filter((phase) => stageOf(phase.name) === def.id);
668 const own = phases.filter((phase) => !phase.reused);
669 const running = own.some((phase) => phase.status === 'running');
670 const holdsCurrent = phases.some((phase) => phase.name === summary.current);
671 const failed =
672 own.some((phase) => phase.status === 'failed') ||
673 (def.id === 'build' && summary.runner?.state === 'failed' && !finished) ||
674 (def.id === 'preflight' && checks.some((check) => check.status === 'failed'));
675 // Once the recap is written the run is over: a stage whose phase record never caught up
676 // (verify left 'running', say) is still done, and Verify is judged by its findings.
677 const done = (phases.length > 0 && phases.every((phase) => DONE.has(phase.status))) || finished;
678 let state: StageState;
679 if (
680 phases.length > 0 &&
681 phases.every((phase) => phase.reused) &&
682 !phases.some((phase) => phase.status === 'running')
683 )
684 state = 'reused';
685 else if (failed) state = 'failed';
686 else if (done) state = def.id === 'verify' && k && k.blocker + k.major > 0 ? 'flagged' : 'done';
687 else if ((holdsCurrent || running) && !finished) {
688 const stopped =
689 (need !== null && holdsCurrent) ||
690 (def.id === 'preflight' && checks.some((check) => check.status === 'needs-you')) ||
691 phases.some((phase) => phase.status === 'stopped');
692 state = stopped ? 'stopped' : 'active';
693 } else state = 'pending';
694 return {
695 id: def.id,
696 label: def.label,
697 doing: def.doing,
698 state,
699 phases,
700 ...stageNote(def.id, state, summary, phases),
701 };
702 });
703}
704hooks/mod/locate.ts 19 lines1// Path arithmetic for finding the person's runs; the lookups that read files live in register.tsx,
2// since a mod's engine calls stay in the file that registers its hooks.
3
4export const MARKERS = ['plans', 'analysis-reports', 'design'];
5
6/** Every folder from `path` up to the root, nearest first. */
7export function ancestors(path: string): string[] {
8 const parts = path.replace(/\/+$/, '').split('/').filter(Boolean);
9 return parts.map((_, i) => `/${parts.slice(0, parts.length - i).join('/')}`);
10}
11
12export function base(path: string): string {
13 return path.split('/').filter(Boolean).pop() ?? path;
14}
15
16export function parent(path: string): string {
17 return path.replace(/\/[^/]+\/?$/, '') || '/';
18}
19src/generated/progress.ts 15 lines1// Generated from schemas/progress.schema.json. Do not edit.
2
3export interface Progress {
4 state: "waiting" | "preflight" | "building" | "done" | "failed";
5 stepsDone: number | null;
6 stepsTotal: number | null;
7 step: string | null;
8 stepKind: string | null;
9 message: string | null;
10 inflight: boolean;
11 lastSeen: string | null;
12 at: string;
13 serverPid: number;
14}
15src/generated/runner-step.ts 102 lines1// Generated from schemas/runner-step.schema.json. Do not edit.
2
3export type RunnerStep =
4 | {
5 step?: string;
6 done?: number;
7 total?: number;
8 kind: "done";
9 buildId?: string;
10 generation?: string;
11 stepToken?: string;
12 }
13 | {
14 step: string;
15 done?: number;
16 total?: number;
17 kind: "wait";
18 retryMs: number;
19 message: string;
20 buildId?: string;
21 generation?: string;
22 stepToken?: string;
23 }
24 | {
25 step: string;
26 done?: number;
27 total?: number;
28 kind: "check";
29 code: string;
30 out?: string;
31 buildId?: string;
32 generation?: string;
33 stepToken?: string;
34 }
35 | {
36 step: string;
37 done?: number;
38 total?: number;
39 kind: "dump";
40 code: string;
41 out?: string;
42 buildId?: string;
43 generation?: string;
44 stepToken?: string;
45 }
46 | ((
47 | {
48 code: string;
49 }
50 | {
51 payload: string;
52 }
53 ) & {
54 step: string;
55 done?: number;
56 total?: number;
57 kind: "use_figma";
58 code?: string;
59 payload?: string;
60 characters?: number;
61 buildId?: string;
62 generation?: string;
63 stepToken?: string;
64 })
65 | {
66 step: string;
67 done?: number;
68 total?: number;
69 kind: "upload";
70 nodeIds: string[];
71 scaleMode: "FIT" | "FILL";
72 files?: {
73 file: string;
74 contentType: string;
75 }[];
76 buildId?: string;
77 generation?: string;
78 stepToken?: string;
79 }
80 | {
81 step: string;
82 done?: number;
83 total?: number;
84 kind: "screenshot";
85 nodeId: string;
86 out: string;
87 maxDimension?: number;
88 buildId?: string;
89 generation?: string;
90 stepToken?: string;
91 }
92 | {
93 step: string;
94 done?: number;
95 total?: number;
96 kind: "skip";
97 reason: string;
98 buildId?: string;
99 generation?: string;
100 stepToken?: string;
101 };
102types/index.d.ts 92 lines1// Generated from src/protocol.ts by scripts/generate-mod-contract.ts. Do not edit.
2export type ProgressState = "waiting" | "preflight" | "building" | "done" | "failed";
3export type StepKind = "wait" | "done" | "check" | "dump" | "use_figma" | "upload" | "screenshot" | "skip";
4
5// reused: copied from an earlier run (design-lab:figma-build), not run again here
6export type Phase = { name: string; status: string; reused: boolean };
7
8// What the phases recorded that the pane says beside each stage; each null until recorded.
9export type Facts = {
10 found: number | null;
11 toBuild: number | null;
12 built: number | null;
13 expected: number | null;
14 figmaUrl: string | null;
15};
16
17// The verification report's open findings by severity, the checks it passed, and those it waived.
18export type Findings = { blocker: number; major: number; minor: number; passed: number; waived: number };
19
20export type Runner = {
21 state: ProgressState | 'connecting';
22 stepsDone: number | null;
23 stepsTotal: number | null;
24 stepKind: StepKind | null;
25 message: string | null;
26 serverAlive: boolean;
27 connected: boolean;
28 lastSeenMs: number | null;
29};
30
31export type Check = {
32 id: string;
33 label: string;
34 status: string;
35 message: string | null;
36 dependsOn: string[];
37};
38
39// The headline figures of a finished run, from its scorecard: each null when the scorer left it out.
40export type Scores = {
41 built: number | null;
42 eligible: number | null;
43 withinTolerance: number | null;
44 widths: number | null;
45 workingSeconds: number | null;
46 buildSeconds: number | null;
47 buildSteps: number | null;
48 tokens: number | null;
49 toolCalls: number | null;
50 blockers: number | null;
51 majors: number | null;
52};
53
54export type Summary = {
55 workspace: string;
56 found: boolean;
57 siteLabel: string | null;
58 phases: Phase[];
59 current: string | null;
60 // checks: the list preflight last wrote, or null when it has none newer than the recorded phase
61 preflight: { status: string; at: string | null; checks: Check[] | null } | null;
62 runner: Runner | null;
63 blocker: string | null;
64 // the run is waiting for the person to do something it will notice by itself (start the runner)
65 waiting: string | null;
66 log: string[];
67 hasRecap: boolean;
68 recap: string | null;
69 scores: Scores | null;
70 facts: Facts;
71 findings: Findings | null;
72 startedAt: string | null;
73 // the last message the phase log holds for each failed phase, by phase name
74 phaseErrors: Record<string, string>;
75 // the run's artifact files that exist, run-relative (read only once the recap is written)
76 present: string[];
77};
78
79declare module 'claude-code' {
80 interface PluginState {
81 'design-lab': {
82 run: string | null
83 summary: Summary | null
84 alarmed: boolean
85 follow: string | null
86 skip: string | null
87 // the run whose full completion message is shown under its figures, or null when folded
88 recapOpen: string | null
89 }
90 }
91}
92