SLOPSHOPPER

design-lab

Build and maintain a verified Figma component library from real code with resumable, schema-validated workflows.

newpanerowstoaststatusprompt
v0.24.1UNLICENSEDupdated 2026-10-08cosmicdreams/claude-plugins/design-lab
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · design-lab
│ ┃ design-lab ✕ › fix the failing auth test and add an audit log call │ ┃ │ ┃ design-lab ⏺ Read(src/auth.ts) │ ┃ ⎿ Read 6 lines │ ┃ No design-lab run is being watched. ⏺ Update(src/auth.ts) │ ⎿ Added 2 lines, removed 1 line │ ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · design-lab
design-lab No design-lab run is being watched.
README

design-lab

Build and maintain a Figma component library from a codebase.

Requires Node 24. Run node scripts/lab_setup.ts install playwright once to install the pinned Node dependencies and Chromium into the shared cache. See TYPESCRIPT.md for plain-copy setup, checks and isolated comparison harnesses.

Extract, model, render are separate on purpose. Writing source straight into Figma gives a one-shot script that cannot re-run, cannot diff against the source later, and cannot feed anything but Figma. components.json and tokens.json are the contract; Figma is one renderer.

Three independent plug points

Component source, token source and usage source vary separately. A Single Directory Component site has no tokens in configuration at all — they live in stylesheets. Conflating the axes forces one site down another's path.

Verified against

SiteComponentsConfig pathTokens
Drupal authoring site69 Drupal authoring bundlesconfig/default97 planned variables from authored Sass
Site Studio site A146 Site Studio + 6 customconfig/packages (declared in settings)129 custom style entities
Site Studio site B101 Site Studio + 3 customconfig/packages (declared in settings)176 custom style entities
Site Studio site C168 Site Studio + 4 customconfig/sitestudio (declared in settings)166 custom style entities
Paragraphs site43 Paragraph typesconfig/default113 base tokens via Sass source map
Paragraphs site, compiled-CSS branch43 Paragraph typesconfig/default94 authored custom properties

The Paragraphs site also has 13 custom Single Directory Components, but only 6 are invoked by a paragraph template - they are a partial rendering layer, not the component source. It was recorded as a 13-component Single Directory Component site until 2026-08-31; that profile came from a bug, not the site. See references/strategies/README.md.

Start here

Run design-lab:init once on a new machine. It settles everything about you and the machine, asking before it installs or changes anything: where runs live (next to each project as PROJECT/design/<date>, or in ~/.design/<project>/<date>; runs are personal and never committed), one shared Playwright and its browser for capture, the baseline packages, the Figma runner, and the Claude Code setting that would otherwise make runs ask you to approve commands. lab_setup.ts check reports the same without changing anything. Everything about one site belongs to preflight, at the start of each run.

Use design-lab:run for a complete library. It creates a run folder outside the repository (where, design-lab:init decided) with its project.json, records the repository commit and every strategy decision, validates artifacts before rendering, and can resume from the first incomplete phase. Use a narrower skill only when the request names a single phase.

Use design-lab:figma-build to build an earlier run's capture and plan into a new, empty Figma file: a new run folder beside the others, with its own pane, verification and report, in minutes rather than the hours a fresh capture takes. The earlier run is left as it was.

Both open the design-lab pane beside the conversation as they start (below); nobody has to ask for it.

Builds write into Figma through the design-lab runner, a Figma development plugin that must be imported once per machine by hand — Figma offers no command-line install. To get the steps with this machine's paths filled in, ask:

Using design-lab's references/relay.md, give me concise steps to install the design-lab runner plugin in Figma desktop, including the absolute path to its manifest on this machine.

SkillDoes
design-lab:initonce per machine: where runs live, capture tools, the Figma runner, Claude Code settings
design-lab:runend-to-end, resumable workflow and completion gate
design-lab:figma-buildan earlier run's capture and plan, built into a new empty Figma file as a new run, verified and scored
design-lab:detectwhich strategies apply
design-lab:inventorycomponents + fields + slots + source defects -> components.json
design-lab:usageverified anonymous example addresses + placement counts + tiers
design-lab:capturemeasures and photographs each component on a running site, per breakpoint
design-lab:tokenscolour, spacing, type per breakpoint, each with its code name -> tokens.json
design-lab:planreviewable build proposal with variant arithmetic and hard refusals
design-lab:figma-foundationvariable collections, modes, scopes, code syntax. Once per file
design-lab:figma-componentone named atomic component transaction — variants, properties, bindings, documentation card, build record
design-lab:figma-indexthe Getting Started page: inventory, linked index, coverage, known gaps. Refresh after every component
design-lab:verifychecks the whole file against the base expectations; every gap ends as a fix or a recorded waiver
design-lab:evaluatethe last step of every run: scores it into scorecard.json, a self-contained HTML report (coverage, accuracy against the live site, time and tokens, repeatability) and the fixed completion message

scripts/workflow.ts is the deterministic front door. Its init, identity, detect, select, preflight, connect, extract, usage, plan, variables, approve, target, register, record, validate, status, and watch commands write atomically and keep artifact hashes in the project manifest. The schemas in schemas/ are the machine-readable contracts; references/library-standard.md is the canonical product definition.

Component extractors cover Site Studio, SDCs, Paragraphs, and combined Drupal authoring vocabularies (block_content + Paragraphs). Token extractors cover Site Studio styles, theme-loaded CSS custom properties, Sass source maps, and source-authored Sass. Combined Drupal extraction also writes render-evidence.json, a bounded map from each authoring bundle to its existing Twig, SDC, stylesheet, root-class, and referenced-field evidence. That evidence includes deterministic styleFacts parsed from the component's own Sass: root and nested-part declarations stay separate, retain token/literal provenance, and give the model the visual facts it needs without asking it to rediscover every stylesheet rule.

Drupal database usage is deterministic too: extract_drupal_usage.ts reads the running DDEV project, preserves placements and structural references as separate measures, writes a validated usage.json, and merges tiers into components.json. Planning hard-stops when detection found usage evidence but it was neither measured nor explicitly waived as degraded. Waivers are human decisions: the workflow requires both a named decider and a reason, and final verification refuses any unresolved prerequisite phase. Registering one component receipt also cannot complete the component phase; valid non-failing receipts must cover the approved plan.

Capture: scaffold_configs.ts writes a config per component and names the ones a human must finish; capture_all.ts measures the box model and typography per breakpoint and takes element-scoped screenshots. Playwright is not vendored: the pinned packages and Chromium come from the shared cache that lab_setup.ts install fills, whatever the current working directory. DESIGN_LAB_BROWSER_EXECUTABLE selects an installed browser when a Playwright-managed Chromium download is unavailable. See skills/capture/SKILL.md.

Token sources are ranked by evidence. A substantial theme-loaded custom-property layer states runtime intent and wins. Source-authored Sass is next, then a recovered Sass source map. Weak or unloaded CSS never outranks the theme sources merely because similarly named files exist.

figma-component owns one component transaction on purpose. run may process many transactions in one model session, but each becomes durable only after its build record is validated. Every recorded assertion must explicitly pass; skipped fidelity work and an empty assertion object remain incomplete. A partial run therefore resumes safely instead of pretending the library is done.

verify is the one that runs last and the one that should have existed first. Every other skill reports on its own step, so a library can pass all of them and still be half a library — which is exactly what happened on the Paragraphs site: four empty Foundations pages, 36 of 43 components missing, no documentation links anywhere, and sixteen variables whose Dev Mode names existed nowhere in the codebase. Nothing was looking at the whole.

Planned: drift.

Fonts

Every site's font trouble has been its own, so fonts are a step of each run, not of setup. When the build is ready to write, workflow.ts connect asks Figma desktop which fonts it can draw with and writes the run's font plan, fonts.json. For each text stack the site renders, it works out the family the visitor really sees: a family declared with no source is skipped, as the browser skips it; icon fonts and generic families are set aside. Each family is then available to Figma, or missing with a stand-in chosen by genre (the same design under its other macOS name when Figma desktop leaves a system font out of its list, such as Courier New for Courier, which needs nothing installed; a metric-compatible clone where one exists, such as Arimo for Arial) and the route to the real font: the Adobe Fonts kit to activate, the open-licence files to install, or the client's desktop files or the foundry's trial for a commercial family, never the site's web font files. The run never stops for a font. The build draws each weight in the face the site really serves (a font-weight: 500 rule serving a Semibold file is drawn Semibold), matches style names generously, accepts trial and web family names, and uses a variable font's weight axis when no named style fits. Verify and the completion message name each stand-in as a decision, not as a failure. workflow.ts fonts --project <run> writes the plan again from the font list recorded at the last connect (or, before the first, without one), and workflow.ts report fonts --project <run> shows it. The step writes fonts.json, figma/available-fonts.json (Figma's list) and, for a site using Adobe Fonts, fonts-kits.json (the kit's families, read once) in the run folder.

Watching a run

design-lab:run and design-lab:figma-build open a pane beside the transcript as they start, and it stays open for the whole run. Typed as slash commands, the pane opens at any window width; asked for in words, Claude Code places a pane it was not asked for only in a wide terminal (144 columns). If the project's newest run is finished, the pane waits for the new run's folder rather than showing the old recap; if it is unfinished, that is the run being resumed and it shows at once. /design-lab:watch [run folder] opens the same pane by hand, after it was closed or for a run started elsewhere. The pane shows the run as five stages (Preflight, Discovery, Build, Verify, Report), one row each with its result, and opens the stage under way beneath its row: the preflight checklist, ticked off as preflight proves each check; the build's steps done of total and whether the runner is connected. A word in the header says where the run stands (for example Building, Needs you, Failed, Done or Done · needs review), and a card above the stages says what the person has to do, or what went wrong. It also keeps a status line such as design-lab: steps 112/158 · runner connected · 41m in that session. With no folder it shows this project's newest run, found by the convention design-lab:init chose from the folder the session is in, and moves to each newer run there as it starts; before the first run it opens and waits for one. Outside any project it falls back to the run workflow.ts init, preflight or connect last recorded in ~/.design-lab/active-run.json. Figma is first touched when the build is ready to write: workflow.ts connect asks the person to open the target file and start the runner, and the pane shows that as the one thing that needs them, with no button, until the runner connects. When the runner later stops asking for steps, the pane says what to do in Figma desktop and offers one button, Resume run, which puts the resume request in the prompt box. When the benchmark has written completion.md, a toast says the run is done, the status line clears, and the pane leads with the verdict (when verification left blocking or major problems open), the run's coverage, accuracy, time and tokens, and links to the Figma library, the benchmark report and the run's files; its Recap starts folded and shows the full completion message when Show is pressed. /design-lab:recap [run folder] shows any finished run's completion message again, with no Claude turn.

The pane is a Claude Code mod (hooks/mod/), which needs a Claude Code that loads mods (2.1.286 does; 2.1.284 does not) and draws in the terminal and the desktop app's Code tab. The desktop app runs its own copy of Claude Code, updated separately from the app; the pane draws there as a sidebar once that copy loads mods, and until then the command answers with text. It only reads the run folder and the pointer (preflight-checks.json, which workflow.ts preflight rewrites as each check starts and settles, is in the run folder); it writes nothing and never reads the runner token. Everything works without it: where the mod is not loaded, or nothing draws (claude -p, the VS Code chat panel), the same command prints the same summary from workflow.ts watch. Mod tests run with node design-lab/scripts/test-mod.ts, which invokes claude plugin test on an isolated mod-only test copy.

References

  • references/prior-art.md — read first. Look for an existing Figma file and existing tooling before extracting anything
  • references/model.md — the universal model and the provenance rule
  • references/variant-policy.md — the decision that makes or breaks the library
  • references/findability.md — how anyone finds a component in a 146-component file
  • references/tokens-and-variables.md — code syntax, and why the Figma name is not the token
  • references/defaults.md — which variant goes first, and the evidence for it
  • references/verification.md — assert numbers, do not eyeball 146 components
  • references/benchmark.md — run checklist and fixed opening prompt; every run ends with design-lab:evaluate
  • references/completion-message.md — the fixed reply design-lab:run ends with
  • references/build-records.md — the idempotency and resume contract
  • references/strategies/README.md — per-strategy mapping and counting traps

Evaluation replays

The corpus and scoreboard use explicit per-person locations from ~/.claude/design-lab.json, or the file named by DESIGN_LAB_CONFIG:

{
  "corpus": "/path/to/corpus",
  "scoreboard": {
    "ledger": "/path/to/ledger.jsonl",
    "dashboard": "/path/to/dashboard.html"
  }
}

All three keys are required. The commands never create or edit this configuration. scoreboard.ts record appends to the ledger and redraws the dashboard from the whole ledger; --open shows it.

node scripts/corpus.ts freeze --run /path/to/finished-run --label site-a
node scripts/corpus.ts list
node scripts/tier1.ts --all --out /tmp/property-results.json
node scripts/tier1.ts --run /path/to/run --out /tmp/property-results.json
node scripts/tier2.ts --site site-a --file-key SCRATCH_FILE_KEY
node scripts/scoreboard.ts record --run /path/to/evaluated-run --tier 2 --site site-a
node scripts/scoreboard.ts rows

Freezing copies the run and writes a manifest with artifact hashes and producer identity. An existing label is refused; use a new label to record a corpus refresh. Older manifests may have no producer commit; that absence stays explicit. Both component-id and older machine-name measurement files are supported.

Tier 1 rebuilds the same trees as figma_build.ts init, in a temporary directory, and compares their resolved breakpoint properties with the saved measurements. Its output is JSON; one summary per site goes to standard error. Geometry uses tree layout arithmetic, not font shaping or Figma rendering. Geometry needing font shaping says unmeasured. Older captures have styled inline descendants rather than exact character ranges, so text run counts identify distinct measured inline styles and flag flattening; their basis is recorded beside each result. Captured states absent from the default responsive tree are reported as unmeasured. This comparison does not change the builder or verify's gate.

Tier 2 requires Figma open with the runner. It creates a new workspace under the site's replays/ directory, clears the designated scratch file, rebuilds with frozen images, waits for the build and verification dumps, assembles measurements, verifies, and scores. The runner stays open between evaluations; for --all, have it open in each scratch file. The image step has no site or public fallback requests: uncached sources remain failures. For --all, add a scratchFileKey to each site's corpus.json; file keys are checked before any replay starts. --timeout limits the wait per site. Failed evaluations keep the workspace and its evidence for inspection. Verification findings do not prevent writing the scorecard. New build receipts and the master-matches-capture gate use the corrected comparison metric; historical scores retain their original metric for comparison. A report is produced even on failed verification, but the manifest records failed quality until the shared completion gate passes.

The benchmark and ledger share run_metrics.ts; neither imports the other. Missing metrics remain unmeasured or null. Recording an evaluation is a separate explicit step, so replay does not silently append a ledger row.

Variable collections default to one shared <Brand> Core, with slash groups for domains and width modes shared by invariant and responsive values. Additional collections require an independent mode axis or documented publishing/ownership boundary; inventory size alone never causes a split.

Source 7 files
hooks/mod/register.tsx 719 lines
1// design-lab's pane: where a run is and when it is done, from the files the run writes. It reads
2// only run folders (and, to find them, the person's design-lab settings, the project's runs folder
3// and the active-run pointer), writes nothing to disk, and asks nothing of the person except, when
4// the runner has stopped, a button that puts the resume request in the prompt.
5
6import { atom, read, update } from 'claude-code'
7import type { CommandRunInput, EngineInterface, Register, Timer } from 'claude-code'
8
9import type { Summary as RunSummary } from '../../src/protocol'
10type Summary = RunSummary<'mod'>
11import {
12  afterFill, barOf, CHECK_COLORS, CHECK_MARKS, checkMessage, compactOf, durationOf, elapsedOf, failureOf, needsYouOf, parseJson,
13  percentOf, PHASE_DONE, phaseLabel, plainOf, RESUME_PROMPT, isIdle, runnerLine, stagesOf, statusOf, stepsLine,
14  summaryOf, toneOf, verdictOf,
15} from './model'
16import { ancestors, base, MARKERS, parent } from './locate'
17
18const PANE = 'design-lab'
19const COMMAND = 'design-lab:watch'
20const RECAP_COMMAND = 'design-lab:recap'
21// The skills that start or resume a run: each opens the pane by itself, so nobody has to know
22// about design-lab:watch to see where a run is.
23export const RUN_SKILLS = ['design-lab:run', 'design-lab:figma-build'] as const
24import { POLL_MS } from '../../src/protocol.ts'
25export { POLL_MS } from '../../src/protocol.ts'
26// The runner log is tailed only while it is small enough to read whole every poll.
27const LOG_READ_LIMIT = 1024 * 1024
28
29const runAtom = atom({ plugin: 'design-lab', key: 'run' } as const, null)
30const summaryAtom = atom({ plugin: 'design-lab', key: 'summary' } as const, null)
31const alarmedAtom = atom({ plugin: 'design-lab', key: 'alarmed' } as const, false)
32// The runs folder the pane follows when the person named no run: it moves to each newer run there.
33const followAtom = atom({ plugin: 'design-lab', key: 'follow' } as const, null)
34// A finished run the pane passes over while a run skill is starting the next one, so the pane
35// never shows the last run's recap as if it were the new run.
36const skipAtom = atom({ plugin: 'design-lab', key: 'skip' } as const, null)
37// Whether a finished run's full completion message is open under its figures.
38// Named by run, so opening one run's recap never opens the next run's.
39const recapOpenAtom = atom({ plugin: 'design-lab', key: 'recapOpen' } as const, null)
40
41// The files read under a run folder, and nothing else there.
42export const RUN_FILES = {
43  project: 'project.json',
44  phaseLog: 'phase-log.jsonl',
45  progress: 'figma/progress.json',
46  runnerLog: 'figma/runner.log',
47  completion: 'benchmark/completion.md',
48  scorecard: 'benchmark/scorecard.json',
49  preflightChecks: 'preflight-checks.json',
50  verifyReport: 'verify-report.json',
51} as const
52
53// The files a finished run links to, when they exist: a rebuild has no plan or components of its
54// own unless it copied them.
55export const ARTIFACT_FILES = {
56  report: 'benchmark/report.html',
57  verifyReport: 'verify-report.json',
58  plan: 'plan.json',
59  components: 'components.json',
60  scorecard: 'benchmark/scorecard.json',
61} as const
62
63// The bar's fill by the color its Text would take, and its empty track.
64const BAR_COLORS: Record<string, string> = { cyan: '#4fa8d6', green: '#3fb36b', yellow: '#d6b44f', red: '#d65f5f' }
65const BAR_TRACK = '#8888884d'
66
67/** A rounded progress bar, wide and short so it scales to the pane's width. */
68function barSvg(share: number, fill: string): string {
69  const filled = share > 0 ? `<rect width="${Math.max(12, share)}" height="12" rx="6" fill="${fill}"/>` : ''
70  return `<svg xmlns="http://www.w3.org/2000/svg" width="1000" height="12" viewBox="0 0 1000 12">`
71    + `<rect width="1000" height="12" rx="6" fill="${BAR_TRACK}"/>${filled}</svg>`
72}
73
74let timer: Timer | undefined
75// The last JSON that parsed, so a file caught half-written keeps the last good reading.
76const lastGood = new Map<string, unknown>()
77
78async function readText($: EngineInterface, path: string): Promise<string | undefined> {
79  try {
80    return await $.fs.read(path)
81  } catch {
82    return undefined
83  }
84}
85
86async function readJson($: EngineInterface, path: string): Promise<unknown> {
87  const value = parseJson(await readText($, path))
88  if (value !== undefined) lastGood.set(path, value)
89  return value ?? lastGood.get(path)
90}
91
92async function readLog($: EngineInterface, path: string): Promise<string | undefined> {
93  try {
94    const stat = await $.fs.stat(path)
95    return stat.size <= LOG_READ_LIMIT ? await $.fs.read(path) : undefined
96  } catch {
97    return undefined
98  }
99}
100
101export async function summarise($: EngineInterface, run: string): Promise<Summary> {
102  const at = (file: string) => `${run}/${file}`
103  const raw = {
104    project: await readJson($, at(RUN_FILES.project)),
105    phaseLog: await readText($, at(RUN_FILES.phaseLog)),
106    progress: await readJson($, at(RUN_FILES.progress)),
107    preflightChecks: await readJson($, at(RUN_FILES.preflightChecks)),
108    verifyReport: await readJson($, at(RUN_FILES.verifyReport)),
109    runnerLog: await readLog($, at(RUN_FILES.runnerLog)),
110    ...((await $.fs.exists(at(RUN_FILES.completion)))
111      ? { completion: await readText($, at(RUN_FILES.completion)), scorecard: await readJson($, at(RUN_FILES.scorecard)),
112        present: await presentOf($, run) }
113      : {}),
114  }
115  return summaryOf(run, raw, await $.clock.now())
116}
117
118/** Which of the run's artifact files exist, so the pane links only to those. */
119async function presentOf($: EngineInterface, run: string): Promise<string[]> {
120  const found: string[] = []
121  for (const path of Object.values(ARTIFACT_FILES)) if (await exists($, `${run}/${path}`)) found.push(path)
122  return found
123}
124
125let refreshing = false
126
127async function refresh($: EngineInterface): Promise<void> {
128  // A slow file system must not stack refreshes: skip a tick while the last one is still reading.
129  if (refreshing) return
130  refreshing = true
131  try {
132    await refreshNow($)
133  } finally {
134    refreshing = false
135  }
136}
137
138// Nothing awaits a refresh the timer or an event starts, so a failure there would be an unhandled
139// rejection. It is told once, as a toast, and again only after a refresh worked in between.
140let refreshFailed = false
141function refreshInBackground($: EngineInterface): void {
142  void refresh($).then(
143    () => { refreshFailed = false },
144    () => { if (!refreshFailed) { refreshFailed = true; reportFailure($, 'refreshing the pane') } },
145  )
146}
147
148async function refreshNow($: EngineInterface): Promise<void> {
149  const follow = await read($, followAtom)
150  if (follow) {
151    const newest = await newestRun($, follow)
152    if (newest && newest !== (await read($, skipAtom)) && newest !== (await read($, runAtom))) {
153      await update($, runAtom, () => newest)
154      await update($, skipAtom, () => null)
155      await update($, alarmedAtom, () => false)
156    }
157  }
158  const run = await read($, runAtom)
159  if (!run) return
160  const summary = await summarise($, run)
161  const before = await read($, summaryAtom)
162  if (JSON.stringify(before) !== JSON.stringify(summary)) await update($, summaryAtom, () => summary)
163  $.ui.status(statusOf(summary, await $.clock.now()))
164  // Done: say so once, the moment the recap appears.
165  if (summary.hasRecap && before && before.found && !before.hasRecap) {
166    // Say what the header says: a run that left problems open is done but needs review.
167    const review = verdictOf(summary.findings) ? ', with verification problems to review' : ''
168    $.ui.toast(`design-lab: ${summary.siteLabel ?? 'the run'} is done${review}. The recap is in the design-lab pane.`, { timeoutMs: 10_000 })
169  }
170  // The watchdog: once per transition, never again until the run needs nothing from the person.
171  // It says what the Needs you card says.
172  const need = needsYouOf(summary)
173  const alarmed = await read($, alarmedAtom)
174  if (need && !alarmed) {
175    $.ui.toast(`design-lab needs you: ${need.message}`, { timeoutMs: 10_000 })
176    await update($, alarmedAtom, () => true)
177  } else if (!need && alarmed) {
178    await update($, alarmedAtom, () => false)
179  }
180}
181
182function watch($: EngineInterface): void {
183  timer?.cancel()
184  timer = $.clock.every(POLL_MS, () => refreshInBackground($))
185}
186
187/** The run a command names, or with none: the newest in this project's runs folder, which the pane
188 * then follows; else the run the machine-wide pointer names. */
189async function runOf($: EngineInterface, args: string): Promise<string | { missing: string; follow?: string }> {
190  const given = args.trim()
191  if (given) return given.startsWith('/') ? given.replace(/\/+$/, '') : `${await $.session.cwd()}/${given}`
192  const folder = await runsFolder($, await $.session.cwd())
193  if (folder) {
194    const newest = await newestRun($, folder)
195    return newest ?? { missing: `No design-lab run yet in ${folder}. The pane shows the first one as soon as it starts.`, follow: folder }
196  }
197  const home = await $.env.get('HOME')
198  const pointer = parseJson(home ? await readText($, `${home}/.design-lab/active-run.json`) : undefined)
199  const workspace = typeof pointer === 'object' && pointer !== null ? (pointer as { workspace?: unknown }).workspace : undefined
200  if (typeof workspace === 'string') return workspace
201  return { missing: (await settings($)).convention
202    ? 'No design-lab run found for this folder. Give the run folder: /design-lab:watch <run folder>'
203    : 'design-lab is not set up on this machine yet: run design-lab:init once. Or give the run folder: /design-lab:watch <run folder>' }
204}
205
206// Which run to watch when the person names none: the newest in this project's runs folder, by
207// the convention design-lab:init recorded (src/lab-config.ts holds the same rule).
208
209async function exists($: EngineInterface, path: string): Promise<boolean> {
210  return $.fs.exists(path).catch(() => false)
211}
212
213async function isDir($: EngineInterface, path: string): Promise<boolean> {
214  const found = await $.fs.stat(path, { resolve: false }).catch(() => undefined)
215  return found?.kind === 'dir'
216}
217
218async function hasMarker($: EngineInterface, folder: string): Promise<boolean> {
219  for (const marker of MARKERS) if (await isDir($, `${folder}/${marker}`)) return true
220  return false
221}
222
223/** A path with every link followed, as the scripts resolve it; the path itself when it cannot be. */
224async function real($: EngineInterface, path: string): Promise<string> {
225  const found = await $.fs.stat(path, { resolve: true }).catch(() => undefined)
226  return (found?.realPath ?? path).replace(/\/+$/, '') || '/'
227}
228
229/** The person's design-lab settings, or {} before design-lab:init has run. */
230async function settings($: EngineInterface): Promise<{ convention?: string; home?: string }> {
231  const rawHome = await $.env.get('HOME')
232  const home = rawHome ? await real($, rawHome) : undefined
233  const path = (await $.env.get('DESIGN_LAB_CONFIG')) ?? (home ? `${home}/.claude/design-lab.json` : undefined)
234  const text = path ? await $.fs.read(path).catch(() => undefined) : undefined
235  const value = parseJson(typeof text === 'string' ? text : undefined) as { runs?: { convention?: unknown } } | undefined
236  const convention = typeof value?.runs?.convention === 'string' ? value.runs.convention : undefined
237  return { convention, home }
238}
239
240/** The project folder for a session in `cwd`: above worktrees/ for PROJECT/worktrees/<name>,
241 * else the nearest folder above the repository holding plans/, analysis-reports/ or design/. */
242async function projectFolder($: EngineInterface, cwd: string, home?: string): Promise<string | undefined> {
243  const stop = (folder: string) => folder === '/' || folder === home
244  const chain = ancestors(cwd)
245  let repo: string | undefined
246  for (const folder of chain) if (await exists($, `${folder}/.git`)) { repo = folder; break }
247  if (!repo) {
248    for (const folder of chain) {
249      if (stop(folder)) return undefined
250      if (await isDir($, `${folder}/worktrees`) || await hasMarker($, folder)) return folder
251    }
252    return undefined
253  }
254  if (base(parent(repo)) === 'worktrees') return parent(parent(repo))
255  for (const folder of ancestors(parent(repo))) {
256    if (stop(folder)) return undefined
257    if (await hasMarker($, folder)) return folder
258  }
259  return undefined
260}
261
262/** Where this project's runs live, or undefined when design-lab:init has not chosen. */
263async function runsFolder($: EngineInterface, sessionCwd: string): Promise<string | undefined> {
264  const { convention, home } = await settings($)
265  const cwd = await real($, sessionCwd)
266  const project = await projectFolder($, cwd, home)
267  if (convention === 'project') return project ? `${project}/design` : undefined
268  if (convention === 'home' && home) {
269    let repo: string | undefined
270    for (const folder of ancestors(cwd)) if (await exists($, `${folder}/.git`)) { repo = folder; break }
271    return `${home}/.design/${base(project ?? repo ?? cwd)}`
272  }
273  return undefined
274}
275
276// When each run began, read once per run folder: a run's start never changes, so a poll lists the
277// runs folder and reads only the project.json of runs it has not seen.
278const startedAt = new Map<string, string>()
279
280/** The newest run in a runs folder: latest start (project.json createdAt), then folder name. */
281async function newestRun($: EngineInterface, folder: string): Promise<string | undefined> {
282  const entries = await $.fs.list(folder).catch(() => [])
283  let best: { created: string; name: string; path: string } | undefined
284  for (const entry of entries) {
285    if (entry.kind !== 'dir') continue
286    const path = `${folder}/${entry.name}`
287    let created = startedAt.get(path)
288    if (created === undefined) {
289      const text = await $.fs.read(`${path}/project.json`).catch(() => undefined)
290      const value = (parseJson(typeof text === 'string' ? text : undefined) as { createdAt?: unknown } | undefined)?.createdAt
291      if (typeof value !== 'string') continue
292      startedAt.set(path, value)
293      created = value
294    }
295    if (!best || created > best.created || (created === best.created && entry.name > best.name)) {
296      best = { created, name: entry.name, path }
297    }
298  }
299  return best?.path
300}
301
302/** A run skill is starting: follow this project's runs folder and open the pane beside the
303 * conversation. A newest run with no recap yet is the one being resumed, so it shows at once; a finished
304 * one is passed over until the new run's folder appears. Opened from the person's own slash
305 * command, Claude Code places the pane at any width; from the Skill tool, only in a wide window. */
306async function followRunSkill($: EngineInterface): Promise<void> {
307  const folder = await runsFolder($, await $.session.cwd())
308  // Not set up yet: the skill sends the person to design-lab:init first.
309  if (!folder) return
310  const newest = await newestRun($, folder)
311  const summary = newest ? await summarise($, newest) : undefined
312  const resuming = newest && summary?.found && !summary.hasRecap ? newest : null
313  await update($, followAtom, () => folder)
314  await update($, skipAtom, () => resuming ? null : newest ?? null)
315  await update($, runAtom, () => resuming)
316  await update($, summaryAtom, () => resuming ? summary! : null)
317  await update($, alarmedAtom, () => false)
318  watch($)
319  if (resuming) refreshInBackground($)
320  if ((await $.session.surfaces()).length > 0) await $.ui.open({ id: PANE, title: 'design-lab' }).catch(() => undefined)
321}
322
323async function resume($: EngineInterface, run: string): Promise<void> {
324  const text = RESUME_PROMPT(run)
325  const step = afterFill(await $.prompt.fill({ text, mode: 'replace' }).catch(() => undefined))
326  if (step === 'submit') await $.prompt.submit({ text }).catch(() => undefined)
327  if (step === 'explain') $.ui.toast('design-lab: close the open dialog, then press resume again.')
328}
329
330/** A hook whose registration catches a failure says so once, as a toast; the engine logs the error itself. */
331function reportFailure($: EngineInterface, what: string): void {
332  $.ui.toast(`design-lab: ${what} failed; claude --debug has the error`, { timeoutMs: 10_000 })
333}
334
335/** A gating command hook must answer, so a failure becomes the command's answer rather than a hang. */
336function answerFailure($: EngineInterface, command: string): { text: string } {
337  reportFailure($, command)
338  return { text: `design-lab could not answer ${command}; claude --debug has the error.` }
339}
340
341async function watchCommand($: EngineInterface, e: CommandRunInput): Promise<{ text: string }> {
342  const run = await runOf($, e.args)
343  // Named, a run is watched as named; found by convention, the pane follows the runs folder.
344  const follow = e.args.trim() ? null : typeof run === 'string'
345    ? (await runsFolder($, await $.session.cwd())) ?? null : run.follow ?? null
346  await update($, followAtom, () => follow)
347  await update($, skipAtom, () => null)
348  if (typeof run !== 'string') {
349    if (!follow) return { text: run.missing }
350    await update($, runAtom, () => null)
351    await update($, summaryAtom, () => null)
352    watch($)
353    if ((await $.session.surfaces()).length > 0) await $.ui.open({ id: PANE, title: 'design-lab', focus: true })
354    return { text: run.missing }
355  }
356  await update($, runAtom, () => run)
357  await update($, alarmedAtom, () => false)
358  await refresh($)
359  watch($)
360  const summary = (await read($, summaryAtom)) ?? (await summarise($, run))
361  if ((await $.session.surfaces()).length === 0) return { text: plainOf(summary) }
362  // The person asked for it: bring it to the front, over any other pane already open.
363  const opened = await $.ui.open({ id: PANE, title: 'design-lab', focus: true })
364  return { text: opened.isPlaced ? `Watching ${summary.siteLabel ?? run}.` : plainOf(summary) }
365}
366
367async function recapCommand($: EngineInterface, e: CommandRunInput): Promise<{ text: string }> {
368  const run = await runOf($, e.args)
369  if (typeof run !== 'string') return { text: run.missing }
370  const summary = await summarise($, run)
371  if (!summary.found) return { text: plainOf(summary) }
372  return { text: summary.recap ?? `${summary.siteLabel ?? run} has no recap yet: the run has not finished its benchmark.` }
373}
374
375export const register: Register = on => {
376  on('session.start', async ($, e, next) => {
377    // After a reload the run is still in state; pick the watch back up.
378    if ((await read($, runAtom)) || (await read($, followAtom))) {
379      watch($)
380      refreshInBackground($)
381    }
382    return next(e)
383  })
384
385  on('session.end', async (_$, e, next) => {
386    timer?.cancel()
387    timer = undefined
388    return next(e)
389  })
390
391  // A gating hook: its registration's .catch answers in its place when the work throws.
392  on('command.run', { command: COMMAND }, ($, e) => watchCommand($, e))
393    .catch(($) => answerFailure($, COMMAND))
394
395  // Typed as a slash command: the person asked, so the pane opens at any width. A failure here is
396  // reported, and the command still runs through next.
397  for (const command of RUN_SKILLS) {
398    on('command.run', { command }, async ($, e, next) => {
399      await followRunSkill($)
400      return next(e)
401    }).catch(($, e, next) => {
402      reportFailure($, command)
403      return next(e)
404    })
405  }
406
407  // Called through the Skill tool (the person asked in their own words): the same, unasked. A typed
408  // command raises this too, after command.run; following again then changes nothing.
409  on('skill.prompt', async ($, e, next) => {
410    if ((RUN_SKILLS as readonly string[]).includes(e.skill)) {
411      await followRunSkill($).catch(() => reportFailure($, e.skill))
412    }
413    return next(e)
414  })
415
416  on('command.run', { command: RECAP_COMMAND }, ($, e) => recapCommand($, e))
417    .catch(($) => answerFailure($, RECAP_COMMAND))
418
419  // The recap's output row, drawn as the Markdown it is, so its links are links.
420  on('ui.render', { component: 'CommandOutput', props: { command: RECAP_COMMAND } }, async ($, e, next) => {
421    if (!e.props.text) return next(e)
422    const { Markdown } = $.ui.resolve(e)
423    return <Markdown text={e.props.text} />
424  })
425
426  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
427    const elements = $.ui.resolve(e)
428    const { Box, Text, Button, Markdown } = elements
429    // Drawn surfaces (desktop, editor, phone) draw the bar as a vector; the terminal as cells.
430    const Svg = e.surface !== 'terminal' && 'Svg' in elements ? elements.Svg : undefined
431    const summary = await read($, summaryAtom)
432    const run = await read($, runAtom)
433    // The pane's ✕ sits on its first row: start one row lower so it never covers text.
434    const follow = await read($, followAtom)
435    if (!summary || !run) {
436      const skip = await read($, skipAtom)
437      return <Box marginTop={1} flexDirection="column" gap={1}>
438        <Text bold>design-lab</Text>
439        <Text dimColor>{follow
440          ? skip
441            ? 'Waiting for the new design-lab run to start. It appears here as soon as its folder is made.'
442            : `No design-lab run yet in ${follow}. It appears here as soon as one starts.`
443          : 'No design-lab run is being watched.'}</Text>
444      </Box>
445    }
446    if (!summary.found) return <Box marginTop={1}><Text>{plainOf(summary)}</Text></Box>
447    const now = await $.clock.now()
448    const columns = e.props.bodyColumns
449    // The card around the bar takes margin 2, border 2 and padding 2; the rest is slack.
450    const width = Math.max(4, columns - 10)
451    const tone = toneOf(summary)
452    const stages = stagesOf(summary)
453    const need = needsYouOf(summary)
454    const failure = failureOf(summary)
455    const runner = summary.runner
456    const preflight = summary.preflight
457    const scores = summary.scores
458    const finished = summary.hasRecap
459    const verdict = finished ? verdictOf(summary.findings) : null
460    const recapOpen = (await read($, recapOpenAtom)) === run
461    const time = finished ? durationOf(scores?.workingSeconds ?? null) : elapsedOf(summary.startedAt, now)
462    const name = run.split('/').pop()
463    // Finished with figures, the Time tile holds the time, so the subtitle does not repeat it.
464    const subtitle = finished && scores ? [name] : [name, time && `${time} ${finished ? 'working time' : 'elapsed'}`]
465    // Two tiles side by side need about 52 columns; narrower, they stack.
466    const narrowTiles = columns < 52
467    // Narrower still, a stage's steps go one per line.
468    const narrowSteps = columns < 40
469
470    // One figure in a quietly bordered tile: only the value carries color, so the tiles never
471    // compete with a card that asks for attention.
472    const tile = (key: string, label: string, value: string | null, details: (string | null)[], color: string | undefined) => (
473      <Box key={`tile-${key}`} flexDirection="column" flexGrow={1} flexShrink={1} width={narrowTiles ? '100%' : '50%'}
474        borderStyle="round" borderDimColor paddingX={1}>
475        <Text dimColor>{label}</Text>
476        <Text bold color={color}>{value ?? '–'}</Text>
477        {details.filter(Boolean).map((detail, i) => <Text key={`${key}-${i}`} dimColor wrap="wrap">{detail}</Text>)}
478      </Box>
479    )
480    // Coverage and accuracy are judgments: green only when nothing is missing, yellow otherwise.
481    const judged = (share: number | null) => share === null ? undefined : share >= 100 ? 'green' : 'yellow'
482    const progress = (done: number, total: number, color: string) => {
483      if (Svg) {
484        const share = Math.round(Math.min(1, Math.max(0, total > 0 ? done / total : 0)) * 1000)
485        return <Svg alt={`${done} of ${total} steps`} source={barSvg(share, BAR_COLORS[color] ?? BAR_COLORS.cyan!)} />
486      }
487      const bar = barOf(done, total, width)
488      return <Text wrap="truncate-end"><Text color={color}>{bar.filled}</Text><Text dimColor>{bar.empty}</Text></Text>
489    }
490    // A stage's phases as ticks: done, failed, under way (or waiting on the person), still to come.
491    const steps = (phases: { name: string; status: string }[], stopped: boolean) => {
492      const ticks = phases.map(phase => {
493        const done = PHASE_DONE.has(phase.status)
494        const failed = phase.status === 'failed'
495        const current = phase.name === summary.current || phase.status === 'running'
496        const color = done ? 'green' : failed ? 'red' : current ? (stopped ? 'yellow' : 'cyan') : undefined
497        const mark = done ? '✓' : failed ? '✗' : current ? (stopped ? '!' : '▸') : '○'
498        return { key: phase.name, color, current, quiet: !done && !failed && !current, text: `${mark} ${phaseLabel(phase.name)}` }
499      })
500      if (narrowSteps) {
501        return <Box flexDirection="column">{ticks.map(t =>
502          <Text key={t.key} color={t.color} dimColor={t.quiet} bold={t.current} wrap="truncate-end">{t.text}</Text>)}</Box>
503      }
504      return <Text wrap="wrap">{ticks.map((t, i) =>
505        <Text key={t.key} color={t.color} dimColor={t.quiet} bold={t.current}>{i > 0 ? '  ' : ''}{t.text}</Text>)}</Text>
506    }
507
508    // What the open stage shows beneath its row: only the detail that stage needs.
509    const detail = (id: string, state: string, phases: { name: string; status: string }[]) => {
510      if (id === 'preflight') {
511        return preflight?.checks
512          ? preflight.checks.map(check => {
513            const message = checkMessage(check)
514            return (
515              <Box key={`check-${check.id}`} flexDirection="column">
516                <Text dimColor={check.status === 'waiting'} color={CHECK_COLORS[check.status]}>
517                  {CHECK_MARKS[check.status] ?? '·'} {check.label}
518                </Text>
519                {message && <Text dimColor={check.status === 'checking'}>  {message}</Text>}
520              </Box>
521            )
522          })
523          : <Text dimColor>Checking the site, the tools and the Figma file.</Text>
524      }
525      if (id === 'build') {
526        const stopped = state === 'stopped'
527        const building = runner && runner.state === 'building' && runner.stepsTotal
528        // Stopped, the Needs you card says what to do; the step count would only repeat the bar.
529        const line = !runner || stopped || state === 'failed' ? null
530          : runner.stepsTotal && (runner.state === 'building' || runner.state === 'done')
531            ? `${runner.stepsDone ?? 0} of ${runner.stepsTotal} steps` : stepsLine(runner)
532        return (
533          <Box flexDirection="column" gap={1}>
534            {steps(phases, stopped)}
535            {(building || line) && runner && (
536              <Box flexDirection="column">
537                {building && progress(runner.stepsDone ?? 0, runner.stepsTotal!, stopped ? 'yellow' : state === 'failed' ? 'red' : 'cyan')}
538                {line && <Text dimColor wrap="truncate-end">{line}</Text>}
539              </Box>
540            )}
541            {runner && (
542              <Text wrap="truncate-end">
543                <Text color={runner.connected || runner.state === 'done' ? 'green' : isIdle(runner) ? 'gray' : 'yellow'}>● </Text>
544                <Text dimColor={isIdle(runner)}>{runnerLine(runner)}</Text>
545              </Text>
546            )}
547          </Box>
548        )
549      }
550      if (id === 'verify') return <Text dimColor>Checking the file against the design-lab standard.</Text>
551      if (id === 'report') return <Text dimColor>Scoring the run and writing the report.</Text>
552      return steps(phases, state === 'stopped')
553    }
554
555    const MARK: Record<string, string> = { done: '✓', flagged: '!', active: '▸', stopped: '!', failed: '✗', pending: '○', reused: '↺' }
556    const COLOR: Record<string, string | undefined> = {
557      done: 'green', flagged: 'yellow', active: 'cyan', stopped: 'yellow', failed: 'red', pending: undefined, reused: undefined,
558    }
559    // One card opens: a failure first, then a stop for the person, then the first stage under way.
560    const opened = stages.find(stage => stage.state === 'failed') ?? stages.find(stage => stage.state === 'stopped')
561      ?? stages.find(stage => stage.state === 'active')
562
563    return (
564      <Box flexDirection="column" marginTop={1} gap={1}>
565        {/* Header: what the run is, and where it stands in one colored word. */}
566        <Box flexDirection="column">
567          <Box flexDirection="row" justifyContent="space-between" alignItems="center">
568            <Text bold wrap="truncate-end">{summary.siteLabel ?? run}</Text>
569            <Text bold inverse color={tone.color}>{`\u00a0${tone.label}\u00a0`}</Text>
570          </Box>
571          <Text dimColor wrap="truncate-end">{subtitle.filter(Boolean).join(' · ')}</Text>
572        </Box>
573
574        {/* What went wrong, above everything else. */}
575        {failure && (
576          <Box key="failed" flexDirection="column" borderStyle="round" borderColor="red" paddingX={1}>
577            <Text bold color="red">Failed</Text>
578            <Text>{failure.text}</Text>
579          </Box>
580        )}
581
582        {/* The one thing the person has to do. */}
583        {need && (
584          <Box key="needs-you" flexDirection="column" borderStyle="round" borderColor="yellow" paddingX={1}>
585            <Text bold color="yellow">Needs you</Text>
586            <Text>{need.message}</Text>
587            {need.canResume && <Text dimColor>When the runner is open again, press Resume run.</Text>}
588            {summary.facts.figmaUrl && <Markdown text={`[Open the Figma file ↗](${summary.facts.figmaUrl})`} />}
589            {need.canResume && (
590              <Box marginTop={1}>
591                <Button key="resume" label="Resume run" onPress={() => void resume($, run)} />
592              </Box>
593            )}
594          </Box>
595        )}
596
597        {/* The five stages, in order: one row each, the open one with its detail beneath. */}
598        <Box flexDirection="column">
599          {stages.map((stage, i) => {
600            const lit = stage.state === 'active' || stage.state === 'stopped' || stage.state === 'failed'
601            const color = COLOR[stage.state]
602            return (
603              <Box key={`stage-${stage.id}`} flexDirection="column">
604                <Box flexDirection="row" justifyContent="space-between" gap={1}>
605                  <Box flexShrink={0}>
606                    <Text bold={lit} dimColor={stage.state === 'pending' || stage.state === 'reused'}
607                      color={lit || stage.state === 'flagged' ? color : undefined}>
608                      <Text color={color}>{MARK[stage.state]}</Text> {i + 1}  {stage.label}
609                    </Text>
610                  </Box>
611                  {stage.note ? <Text color={stage.noteColor} dimColor={!stage.noteColor && !lit}
612                    wrap={stage.id === 'verify' ? 'wrap' : 'truncate-end'}>{stage.note}</Text> : null}
613                </Box>
614                {opened === stage && (
615                  // Stopped, the card stays quiet: the Needs you card is the only yellow box.
616                  <Box key={`card-${stage.id}`} flexDirection="column" marginLeft={2} marginY={1} paddingX={1} borderStyle="round"
617                    borderColor={stage.state === 'stopped' ? undefined : color} borderDimColor>
618                    {detail(stage.id, stage.state, stage.phases)}
619                  </Box>
620                )}
621              </Box>
622            )
623          })}
624        </Box>
625
626        {/* Finished: the verdict, the report's figures, then what the run made and where it is. */}
627        {finished && (
628          <Box flexDirection="column" gap={1}>
629            {verdict && (
630              <Box key="verdict" flexDirection="column" borderStyle="round" borderColor="yellow" paddingX={1}>
631                <Text bold color="yellow">Verification found problems</Text>
632                <Text>{verdict}</Text>
633                <Markdown text={`[Open the verification findings](${fileUrl(run, ARTIFACT_FILES.verifyReport)})`} />
634              </Box>
635            )}
636            {scores && (() => {
637              const coverage = percentOf(scores.built, scores.eligible)
638              const accuracy = percentOf(scores.withinTolerance, scores.widths)
639              const missing = scores.built !== null && scores.eligible !== null && scores.eligible > scores.built
640                ? `${scores.eligible - scores.built} not built` : null
641              return (
642                <Box flexDirection="column" gap={1}>
643                  <Box flexDirection={narrowTiles ? 'column' : 'row'} gap={1}>
644                    {tile('coverage', 'Coverage', coverage !== null ? `${coverage}%` : null,
645                      [scores.built !== null && scores.eligible !== null ? `${scores.built} of ${scores.eligible} buildable` : null, missing],
646                      judged(coverage))}
647                    {tile('accuracy', 'Accuracy', accuracy !== null ? `${accuracy}%` : null,
648                      [scores.withinTolerance !== null && scores.widths !== null ? `${scores.withinTolerance} of ${scores.widths} widths within tolerance` : null],
649                      judged(accuracy))}
650                  </Box>
651                  <Box flexDirection={narrowTiles ? 'column' : 'row'} gap={1}>
652                    {tile('time', 'Time', durationOf(scores.workingSeconds),
653                      [scores.buildSeconds !== null ? `Figma build ${durationOf(scores.buildSeconds)}` : null], 'magenta')}
654                    {tile('tokens', 'Tokens', compactOf(scores.tokens),
655                      [scores.toolCalls !== null ? `${scores.toolCalls} tool calls` : null], 'blue')}
656                  </Box>
657                </Box>
658              )
659            })()}
660            {(() => {
661              const links = artifactsOf(run, summary.facts.figmaUrl, new Set(summary.present))
662              return (links.main || links.files) && (
663                <Box flexDirection="column" gap={1}>
664                  {links.main && (
665                    <Box flexDirection="column">
666                      <Text bold dimColor>Results</Text>
667                      <Markdown text={links.main} />
668                    </Box>
669                  )}
670                  {links.files && (
671                    <Box flexDirection="column">
672                      <Text bold dimColor>Run files</Text>
673                      <Markdown text={links.files} />
674                    </Box>
675                  )}
676                </Box>
677              )
678            })()}
679            {summary.recap && (
680              <Box flexDirection="column">
681                <Box flexDirection="row" justifyContent="space-between" alignItems="center">
682                  <Text bold dimColor>Recap</Text>
683                  <Button key="recap" label={recapOpen ? 'Hide' : 'Show'} onPress={() => void update($, recapOpenAtom, open => open === run ? null : run)} />
684                </Box>
685                {recapOpen && <Markdown text={summary.recap} />}
686              </Box>
687            )}
688          </Box>
689        )}
690      </Box>
691    )
692  })
693}
694
695/** A file in the run folder as a link: each path segment encoded, so a space, #, ? or bracket in a
696 * folder name cannot end or break the link. */
697export function fileUrl(run: string, path: string): string {
698  const encode = (segment: string) => encodeURIComponent(segment).replace(/[!'()*]/g, c => `%${c.charCodeAt(0).toString(16).toUpperCase()}`)
699  return `file://${`${run}/${path}`.split('/').map(encode).join('/')}`
700}
701
702/** What a finished run made, as links: the two that matter (the Figma file and the benchmark
703 * report), then the run's own files. Only files that exist are listed. */
704export function artifactsOf(run: string, figmaUrl: string | null, present: Set<string>): { main: string | null; files: string | null } {
705  const main = [
706    figmaUrl ? `**[Open the Figma library ↗](${figmaUrl})**` : null,
707    present.has(ARTIFACT_FILES.report) ? `**[Open the benchmark report](${fileUrl(run, ARTIFACT_FILES.report)})**` : null,
708  ].filter(Boolean)
709  const files = ([
710    ['Verification findings', ARTIFACT_FILES.verifyReport],
711    ['Build plan', ARTIFACT_FILES.plan],
712    ['Components', ARTIFACT_FILES.components],
713    ['Scorecard', ARTIFACT_FILES.scorecard],
714  ] as const).filter(([, path]) => present.has(path)).map(([label, path]) => `[${label}](${fileUrl(run, path)})`)
715  // No bullets: Markdown indents them unevenly. The two results stand one per line (a hard break
716  // is two trailing spaces); the run files share one line.
717  return { main: main.length ? main.join('  \n') : null, files: files.length ? files.join(' · ') : null }
718}
719
src/protocol.ts 210 lines
1/** Self-contained values: safe in the Node server, Figma stripping and the no-Node mod host.
2 * The only imports are erased types generated from our JSON schemas. */
3import type { Progress } from './generated/progress.ts';
4import type { RunnerStep } from './generated/runner-step.ts';
5export const PORT = 8765;
6export const WAIT_MS = 5000;
7export const RETRY_MS = WAIT_MS;
8export const HEARTBEAT_SECONDS = 10;
9export const SERVER_FRESH_MS = 3 * HEARTBEAT_SECONDS * 1000;
10export const RUNNER_ABSENT_MS = 120_000;
11export const POLL_MS = WAIT_MS;
12export type ProgressState = Progress['state'];
13export type StepKind = RunnerStep['kind'];
14export const PROGRESS_STATES = {
15  waiting: 'waiting',
16  preflight: 'preflight',
17  building: 'building',
18  done: 'done',
19  failed: 'failed',
20} as const satisfies { [State in ProgressState]: State };
21export const STEP_KINDS = {
22  wait: 'wait',
23  done: 'done',
24  check: 'check',
25  dump: 'dump',
26  use_figma: 'use_figma',
27  upload: 'upload',
28  screenshot: 'screenshot',
29  skip: 'skip',
30} as const satisfies { [Kind in StepKind]: Kind };
31export const isProgressState = (value: string): value is ProgressState =>
32  Object.values(PROGRESS_STATES).some((state) => state === value);
33export const isStepKind = (value: string): value is StepKind =>
34  Object.values(STEP_KINDS).some((kind) => kind === value);
35
36/** Finite workflow vocabulary; dynamic artifact paths remain generated typed maps. */
37export const PHASE_NAMES = [
38  'init',
39  'discovery',
40  'inventory',
41  'usage',
42  'capture',
43  'tokens',
44  'plan',
45  'variables',
46  'preflight',
47  'connect',
48  'foundation',
49  'components',
50  'index',
51  'verify',
52  'benchmark',
53] as const;
54export type PhaseName = (typeof PHASE_NAMES)[number];
55export const REGISTRABLE_PHASES = [
56  'usage',
57  'capture',
58  'foundation',
59  'components',
60  'index',
61  'verify',
62] as const satisfies readonly PhaseName[];
63export type RegistrablePhase = (typeof REGISTRABLE_PHASES)[number];
64export const RECORDABLE_PHASES = [...REGISTRABLE_PHASES, 'benchmark'] as const satisfies readonly PhaseName[];
65export type RecordablePhase = (typeof RECORDABLE_PHASES)[number];
66export const PHASE_STATUSES = [
67  'pending',
68  'running',
69  'complete',
70  'failed',
71  'waived',
72  'awaiting-approval',
73  'approved',
74  'waiting',
75  'stopped',
76  'invalidated',
77] as const;
78export type PhaseStatus = (typeof PHASE_STATUSES)[number];
79export type CheckStatus = 'done' | 'checking' | 'needs-you' | 'failed' | 'waiting';
80export interface ChecklistDocument {
81  pass: string;
82  at: string;
83  ready: boolean | null;
84  goAheadAt?: string | null;
85  checks: (Omit<Check, 'status'> & { status: CheckStatus; at: string })[];
86}
87export interface Handshake {
88  ok: boolean;
89  at?: string;
90  failure?: string;
91  runnerConnected?: boolean;
92  fileKey?: string | null;
93  fileName?: string | null;
94  fileKeyMatches?: boolean;
95  empty?: boolean;
96  onlyPreflightCover?: boolean;
97  writable?: boolean;
98  pluginData?: boolean;
99  connectionOnly?: boolean;
100  coverPageId?: string;
101  coverId?: string;
102  font?: string;
103  fontLoaded?: boolean;
104  fonts?: Record<string, string[]> | null;
105  server?: { pid: number | null; started: boolean };
106  install?: { folder: string; manifest: string; version: string; firstInstall: boolean; updated: boolean };
107  instructions?: string[];
108  outdated?: boolean;
109  runnerVersion?: string;
110}
111
112// DESIGN_LAB_MOD_CONTRACT_BEGIN
113// reused: copied from an earlier run (design-lab:figma-build), not run again here
114export type Phase = { name: string; status: string; reused: boolean };
115
116// What the phases recorded that the pane says beside each stage; each null until recorded.
117export type Facts = {
118  found: number | null;
119  toBuild: number | null;
120  built: number | null;
121  expected: number | null;
122  figmaUrl: string | null;
123};
124
125// The verification report's open findings by severity, the checks it passed, and those it waived.
126export type Findings = { blocker: number; major: number; minor: number; passed: number; waived: number };
127
128export type Runner = {
129  state: ProgressState | 'connecting';
130  stepsDone: number | null;
131  stepsTotal: number | null;
132  stepKind: StepKind | null;
133  message: string | null;
134  serverAlive: boolean;
135  connected: boolean;
136  lastSeenMs: number | null;
137};
138
139export type Check = {
140  id: string;
141  label: string;
142  status: string;
143  message: string | null;
144  dependsOn: string[];
145};
146
147// The headline figures of a finished run, from its scorecard: each null when the scorer left it out.
148export type Scores = {
149  built: number | null;
150  eligible: number | null;
151  withinTolerance: number | null;
152  widths: number | null;
153  workingSeconds: number | null;
154  buildSeconds: number | null;
155  buildSteps: number | null;
156  tokens: number | null;
157  toolCalls: number | null;
158  blockers: number | null;
159  majors: number | null;
160};
161
162type ModSummary = {
163  workspace: string;
164  found: boolean;
165  siteLabel: string | null;
166  phases: Phase[];
167  current: string | null;
168  // checks: the list preflight last wrote, or null when it has none newer than the recorded phase
169  preflight: { status: string; at: string | null; checks: Check[] | null } | null;
170  runner: Runner | null;
171  blocker: string | null;
172  // the run is waiting for the person to do something it will notice by itself (start the runner)
173  waiting: string | null;
174  log: string[];
175  hasRecap: boolean;
176  recap: string | null;
177  scores: Scores | null;
178  facts: Facts;
179  findings: Findings | null;
180  startedAt: string | null;
181  // the last message the phase log holds for each failed phase, by phase name
182  phaseErrors: Record<string, string>;
183  // the run's artifact files that exist, run-relative (read only once the recap is written)
184  present: string[];
185};
186
187// DESIGN_LAB_MOD_CONTRACT_END
188
189/** Both readers describe one run, but the CLI keeps its historical file-path/seconds
190 * projection; the pane adds presentation facts and recap text. The surface parameter
191 * makes these differences explicit without weakening either contract. */
192type CliRunner = Omit<Runner, 'lastSeenMs'> & { lastSeenSeconds: number | null };
193type CliSummary =
194  | { found: false; workspace: string; recap?: never }
195  | {
196      found: true;
197      workspace: string;
198      siteLabel?: string;
199      phases: Omit<Phase, 'reused'>[];
200      nextPhase: string | null;
201      preflightChecks: Check[] | null;
202      runner: CliRunner | null;
203      blocker: string | null;
204      waiting: string | null;
205      recap: string | null;
206      startedAt: string | null;
207      elapsedSeconds: number | null;
208    };
209export type Summary<Surface extends 'cli' | 'mod' = 'cli'> = Surface extends 'mod' ? ModSummary : CliSummary;
210
hooks/mod/model.ts 704 lines
1// What a design-lab run looks like from the files it writes, with no engine calls: the same
2// reading `workflow.ts watch` does, so the pane and the text fallback agree.
3
4import type { Check, Facts, Findings, Phase, Runner, Scores, Summary as RunSummary } from '../../src/protocol';
5import type { Summary as ManifestSummary } from '../../types';
6type Summary = RunSummary<'mod'>;
7type Assignable<Expected, Actual extends Expected> = Actual;
8// Both directions enforce the generated host contract without runtime dependencies.
9export type HostSummaryMatchesProtocol = Assignable<Summary, ManifestSummary>;
10export type ProtocolSummaryMatchesHost = Assignable<ManifestSummary, Summary>;
11
12// protocol.ts has no runtime dependencies, so these values are safe in the mod host.
13import { SERVER_FRESH_MS, RUNNER_ABSENT_MS, isProgressState, isStepKind } from '../../src/protocol.ts';
14export { SERVER_FRESH_MS, RUNNER_ABSENT_MS } from '../../src/protocol.ts';
15export const LOG_LINES = 8;
16// The Markdown element draws at most 10,000 characters.
17export const RECAP_LIMIT = 9_500;
18
19const DONE = new Set(['complete', 'approved', 'waived']);
20
21/** The files the mod reads, as text; a missing or unreadable file is undefined. */
22export type Raw = {
23  project?: unknown;
24  phaseLog?: string;
25  progress?: unknown;
26  runnerLog?: string;
27  completion?: string;
28  scorecard?: unknown;
29  preflightChecks?: unknown;
30  verifyReport?: unknown;
31  // the artifact files found in the run folder, run-relative
32  present?: string[];
33};
34
35export function parseJson(text: string | undefined): unknown {
36  if (text === undefined) return undefined;
37  try {
38    return JSON.parse(text);
39  } catch {
40    return undefined;
41  }
42}
43
44function record(value: unknown): Record<string, unknown> {
45  return typeof value === 'object' && value !== null ? (value as Record<string, unknown>) : {};
46}
47
48function text(value: unknown): string | null {
49  return typeof value === 'string' ? value : null;
50}
51
52function count(value: unknown): number | null {
53  return typeof value === 'number' ? value : null;
54}
55
56function msSince(stamp: string | null, nowMs: number): number | null {
57  if (!stamp) return null;
58  const at = Date.parse(stamp);
59  return Number.isNaN(at) ? null : nowMs - at;
60}
61
62/** The phase log's lines that parse; a half-written last line is skipped. */
63export function entriesOf(log: string | undefined): Record<string, unknown>[] {
64  if (!log) return [];
65  return log.split('\n').flatMap((line) => {
66    const value = parseJson(line.trim() || undefined);
67    return typeof value === 'object' && value !== null ? [value] : [];
68  });
69}
70
71export function tailOf(log: string | undefined, lines = LOG_LINES): string[] {
72  if (!log) return [];
73  return log
74    .split('\n')
75    .filter((line) => line.trim() !== '')
76    .slice(-lines);
77}
78
79export function runnerOf(progress: unknown, nowMs: number): Runner | null {
80  const p = record(progress);
81  if (!('state' in p)) return null;
82  const serverAge = msSince(text(p.at), nowMs);
83  const serverAlive = serverAge !== null && serverAge <= SERVER_FRESH_MS;
84  const seenAge = msSince(text(p.lastSeen), nowMs);
85  const asked = seenAge !== null && seenAge <= RUNNER_ABSENT_MS;
86  const connected = serverAlive && (asked || p.inflight === true);
87  return {
88    state: typeof p.state === 'string' && isProgressState(p.state) ? p.state : 'waiting',
89    stepsDone: count(p.stepsDone),
90    stepsTotal: count(p.stepsTotal),
91    stepKind: typeof p.stepKind === 'string' && isStepKind(p.stepKind) ? p.stepKind : null,
92    message: text(p.message),
93    serverAlive,
94    connected,
95    lastSeenMs: seenAge,
96  };
97}
98
99export function summaryOf(workspace: string, raw: Raw, nowMs: number): Summary {
100  const project = record(raw.project);
101  if (raw.project === undefined) {
102    return {
103      workspace,
104      found: false,
105      siteLabel: null,
106      phases: [],
107      current: null,
108      preflight: null,
109      runner: null,
110      blocker: null,
111      waiting: null,
112      log: [],
113      hasRecap: false,
114      recap: null,
115      scores: null,
116      facts: { found: null, toBuild: null, built: null, expected: null, figmaUrl: null },
117      findings: null,
118      startedAt: null,
119      phaseErrors: {},
120      present: [],
121    };
122  }
123  const recorded: Phase[] = Object.entries(record(project.phases)).map(([name, value]) => ({
124    name,
125    status: text(record(value).status) ?? 'pending',
126    reused: Boolean(record(value).from),
127  }));
128  const finished = recapIsCurrent(project, raw);
129  // Preflight records its phase only once it passes, so until then a run that has not reached the
130  // build is still in preflight. A rebuild copies preflight from its source and never runs it here.
131  const preflighting =
132    !recorded.some((phase) => phase.name === 'preflight') &&
133    !recorded.some((phase) => phase.reused) &&
134    !finished &&
135    !recorded.some((phase) => BUILD_PHASES.has(phase.name) && phase.status !== 'pending');
136  // In the order the run takes them, not the order project.json happens to hold them.
137  const phases = [...recorded, ...(preflighting ? [{ name: 'preflight', status: 'running', reused: false }] : [])].sort(
138    (a, b) => flowRank(a.name) - flowRank(b.name),
139  );
140  // A phase copied from an earlier run never takes the current phase, whatever status it copied.
141  const own = phases.filter((phase) => !phase.reused);
142  const running = own.find((phase) => phase.status === 'running');
143  const due = own.find((phase) => !DONE.has(phase.status));
144  const runner = runnerOf(raw.progress, nowMs);
145  const entries = entriesOf(raw.phaseLog);
146  const last = entries[entries.length - 1];
147  const phaseErrors: Record<string, string> = {};
148  for (const phase of phases) {
149    if (phase.status !== 'failed') continue;
150    const said = entries.filter((entry) => entry.phase === phase.name && typeof entry.message === 'string').pop();
151    if (said) phaseErrors[phase.name] = said.message as string;
152  }
153  // The open blocker: the newest entry stopped the run for the person, and the runner has not
154  // come back since.
155  const blocker = last && last.status === 'stopped' && !runner?.connected ? text(last.message) : null;
156  // Waiting on the person for something the run notices by itself (the runner starting at the
157  // build's connection): what to do, with nothing to press.
158  const waiting = last && last.status === 'waiting' && !runner?.connected ? text(last.message) : null;
159  // While the build waits for the person to start the runner, the runner is awaited, not idle.
160  const shown = waiting && runner && runner.state === 'waiting' ? { ...runner, state: 'connecting' as const } : runner;
161  const preflight = record(record(project.phases).preflight);
162  const preflightAt = preflight.from ? null : text(preflight.updatedAt);
163  const checks = checksOf(raw.preflightChecks, preflight);
164  return {
165    workspace,
166    found: true,
167    siteLabel: text(record(project.run).siteLabel),
168    phases,
169    current: (running ?? due)?.name ?? null,
170    preflight:
171      preflight.status || checks ? { status: text(preflight.status) ?? 'running', at: preflightAt, checks } : null,
172    runner: shown,
173    blocker,
174    waiting,
175    log: tailOf(raw.runnerLog),
176    hasRecap: finished,
177    recap: finished ? recapOf(raw.completion!) : null,
178    scores: finished ? scoresOf(raw.scorecard) : null,
179    facts: factsOf(project, finished ? raw.completion : undefined),
180    findings: findingsOf(raw.verifyReport, text(project.createdAt)),
181    startedAt: preflightAt ?? text(project.createdAt),
182    phaseErrors,
183    present: raw.present ?? [],
184  };
185}
186
187export const CHECK_MARKS: Record<string, string> = {
188  done: '✓',
189  checking: '▸',
190  'needs-you': '!',
191  failed: '✗',
192  waiting: '·',
193};
194export const CHECK_COLORS: Record<string, string | undefined> = {
195  done: 'green',
196  checking: 'cyan',
197  'needs-you': 'yellow',
198  failed: 'red',
199};
200
201/** The checklist preflight last wrote, or null when there is none or it is older than the recorded
202 * preflight phase (a run that passed preflight before the checklist existed). */
203export function checksOf(document: unknown, phase: Record<string, unknown>): Check[] | null {
204  const d = record(document);
205  if (!Array.isArray(d.checks)) return null;
206  const written = Date.parse(text(d.at) ?? '');
207  const passed = Date.parse(phase.from ? '' : (text(phase.updatedAt) ?? ''));
208  if (phase.status === 'complete' && written < passed) return null;
209  return d.checks.flatMap((value) => {
210    const c = record(value);
211    const id = text(c.id);
212    return id === null
213      ? []
214      : [
215          {
216            id,
217            label: text(c.label) ?? id,
218            status: text(c.status) ?? 'waiting',
219            message: text(c.message),
220            dependsOn: Array.isArray(c.dependsOn) ? c.dependsOn.filter((x): x is string => typeof x === 'string') : [],
221          },
222        ];
223  });
224}
225
226/** Preflight passed and every check is done: the group can fold to one line. */
227export function preflightPassed(summary: Summary): boolean {
228  const p = summary.preflight;
229  return p?.status === 'complete' && (p.checks ?? []).every((check) => check.status === 'done');
230}
231
232/** What a check says beside its label: only while it is working or needs something. */
233export function checkMessage(check: Check): string | null {
234  return ['checking', 'needs-you', 'failed'].includes(check.status) ? check.message : null;
235}
236
237/** HH:MM on the person's clock. */
238export function clockOf(stamp: string | null): string | null {
239  if (!stamp) return null;
240  const at = new Date(stamp);
241  if (Number.isNaN(at.getTime())) return null;
242  return `${String(at.getHours()).padStart(2, '0')}:${String(at.getMinutes()).padStart(2, '0')}`;
243}
244
245/** The recap belongs to this build: its scorecard names the build's creation, or, from a scorer
246 * before that stamp, the build has recorded its benchmark as complete. A folder initialised again
247 * keeps the old benchmark/ folder until it is scored, and that must not read as done. */
248export function recapIsCurrent(project: Record<string, unknown>, raw: Raw): boolean {
249  if (raw.completion === undefined) return false;
250  const stamp = record(record(raw.scorecard).run).buildCreatedAt;
251  if (typeof stamp === 'string') return stamp === text(project.createdAt);
252  return text(record(record(project.phases).benchmark).status) === 'complete';
253}
254
255/** The completion message as written, cut with a note if it ever outgrows what Markdown draws. */
256export function recapOf(completion: string): string {
257  const text = completion.trim();
258  return text.length <= RECAP_LIMIT
259    ? text
260    : `${text.slice(0, RECAP_LIMIT)}\n\n(cut here: the whole message is in benchmark/completion.md)`;
261}
262
263/** Finished: the scorer has written this build's recap. The runner reports done once the Figma
264 * build is over, while verification and scoring still have to run. */
265export function isFinished(summary: Summary): boolean {
266  return summary.hasRecap;
267}
268
269/** The run is stopped for the person: building or checking, and the runner is gone. */
270export function isDown(summary: Summary): boolean {
271  if (!summary.found || isFinished(summary)) return false;
272  if (summary.blocker) return true;
273  const runner = summary.runner;
274  return runner !== null && (runner.state === 'building' || runner.state === 'preflight') && !runner.connected;
275}
276
277/** Before the build starts, or between builds, the runner is not needed: idle, not missing. */
278export function isIdle(runner: Runner): boolean {
279  return runner.state === 'waiting' && !runner.connected;
280}
281
282export function runnerLine(runner: Runner): string {
283  if (runner.state === 'done') return 'Figma build finished';
284  if (!runner.serverAlive) return 'runner server not responding';
285  if (runner.connected) return 'runner connected';
286  if (isIdle(runner)) return 'runner idle until the build';
287  if (runner.state === 'connecting') return 'waiting for the runner to start';
288  const minutes = Math.max(1, Math.round((runner.lastSeenMs ?? 0) / 60_000));
289  return `runner not seen for ${minutes}m`;
290}
291
292export function stepsLine(runner: Runner): string | null {
293  if ((runner.state === 'building' || runner.state === 'done') && runner.stepsTotal) {
294    const kind = runner.state === 'building' && runner.stepKind ? `, ${runner.stepKind}` : '';
295    return `steps ${runner.stepsDone ?? 0}/${runner.stepsTotal}${kind}`;
296  }
297  // The server's last word to a runner that has since gone quiet ("Connected. Waiting…") is stale.
298  return isIdle(runner) || runner.state === 'connecting' ? null : runner.message;
299}
300
301/** A length of time as the pane writes it: 41m, 3h 44m; under a minute, 40s. */
302export function durationOf(seconds: number | null): string | null {
303  if (seconds === null || seconds < 0) return null;
304  if (seconds < 60) return `${Math.round(seconds)}s`;
305  // Rounded, as the recap rounds it, so the two agree.
306  const minutes = Math.round(seconds / 60);
307  return minutes < 60 ? `${minutes}m` : `${Math.floor(minutes / 60)}h ${minutes % 60}m`;
308}
309
310/** How long the run has been going, from its start. */
311export function elapsedOf(startedAt: string | null, nowMs: number): string | null {
312  const ms = msSince(startedAt, nowMs);
313  if (ms === null || ms < 0) return null;
314  return ms < 60_000 ? '0m' : durationOf(Math.floor(ms / 60_000) * 60);
315}
316
317/** The scorecard's headline figures; null when there is no scorecard to read. */
318export function scoresOf(scorecard: unknown): Scores | null {
319  const card = record(scorecard);
320  if (!('headline' in card)) return null;
321  const headline = record(card.headline);
322  const coverage = record(headline.coverage);
323  const accuracy = record(record(headline.accuracy).corrected);
324  const effort = record(headline.effort);
325  const open = record(record(record(card.sections).conformance).open);
326  return {
327    built: count(coverage.built),
328    eligible: count(coverage.eligible),
329    withinTolerance: count(accuracy.pass),
330    widths: count(accuracy.total),
331    workingSeconds: count(effort.workingSeconds),
332    buildSeconds: count(effort.buildSeconds),
333    buildSteps: count(effort.buildSteps),
334    tokens: count(effort.tokens),
335    toolCalls: count(effort.toolCalls),
336    blockers: count(open.blocker),
337    majors: count(open.major),
338  };
339}
340
341/** The first Figma file link in the recap, without the punctuation of the sentence it ends. */
342export function figmaUrlOf(completion: string): string | null {
343  return /https:\/\/www\.figma\.com\/(?:design|file)\/[^\s)>\]]+/.exec(completion)?.[0].replace(/[.,;:]+$/, '') ?? null;
344}
345
346/** 40013109 as 40.0M, 563519 as 564K. */
347export function compactOf(value: number | null): string | null {
348  if (value === null) return null;
349  if (value >= 1_000_000) return `${(value / 1_000_000).toFixed(1)}M`;
350  if (value >= 1_000) return `${Math.round(value / 1_000)}K`;
351  return String(value);
352}
353
354/** A share as a whole percent, or null with nothing to divide by. Anything short of the whole
355 * stays below 100, so 200 of 201 reads 99% and is never judged complete. */
356export function percentOf(part: number | null, whole: number | null): number | null {
357  if (part === null || !whole) return null;
358  const percent = Math.round((part / whole) * 100);
359  return part < whole && percent >= 100 ? 99 : percent;
360}
361
362/** A bar of `cells` cells, filled in proportion: the filled run and the empty run, drawn apart. */
363export function barOf(done: number, total: number, cells: number): { filled: string; empty: string } {
364  const width = Math.max(4, cells);
365  const share = total > 0 ? Math.min(1, Math.max(0, done / total)) : 0;
366  const filled = Math.round(share * width);
367  return { filled: '━'.repeat(filled), empty: '━'.repeat(width - filled) };
368}
369
370/** Where the run stands, in one word, and the color it is drawn in. */
371export type Tone = { label: string; color: string };
372
373export function toneOf(summary: Summary): Tone {
374  const stages = stagesOf(summary);
375  if (stages.some((stage) => stage.state === 'failed')) return { label: 'Failed', color: 'red' };
376  if (needsYouOf(summary)) return { label: 'Needs you', color: 'yellow' };
377  if (isFinished(summary)) {
378    return verdictOf(summary.findings)
379      ? { label: 'Done · needs review', color: 'yellow' }
380      : { label: 'Done', color: 'green' };
381  }
382  const active = stages.find((stage) => stage.state === 'active' || stage.state === 'stopped');
383  return { label: active?.doing ?? 'Starting', color: 'cyan' };
384}
385
386/** What the person has to do, from one place, so the header, the card, the stage and the toast
387 * always agree; null when the run needs nothing from them. */
388export function needsYouOf(summary: Summary): { message: string; canResume: boolean } | null {
389  // A failure outranks every request: the Failed card says what went wrong, and resuming would
390  // only run into it again.
391  if (hasFailed(summary)) return null;
392  if (summary.blocker) return { message: summary.blocker, canResume: true };
393  if (summary.waiting) return { message: summary.waiting, canResume: false };
394  const check = (summary.preflight?.checks ?? []).find((c) => c.status === 'needs-you');
395  if (check) return { message: `${check.label}: ${check.message ?? 'needs your attention'}`, canResume: false };
396  if (isDown(summary)) {
397    return {
398      message: 'The Figma runner has stopped. Reopen it in Figma desktop, then press Resume run.',
399      canResume: true,
400    };
401  }
402  return null;
403}
404
405/** Something failed: a phase run here, the runner while the run is unfinished, or a preflight check. */
406export function hasFailed(summary: Summary): boolean {
407  return (
408    summary.phases.some((phase) => !phase.reused && phase.status === 'failed') ||
409    (summary.runner?.state === 'failed' && !isFinished(summary)) ||
410    (summary.preflight?.checks ?? []).some((check) => check.status === 'failed')
411  );
412}
413
414/** The first stage that failed, and what the card says about it; null when nothing failed. */
415export function failureOf(summary: Summary): { stage: string; text: string } | null {
416  const stage = stagesOf(summary).find((s) => s.state === 'failed');
417  if (!stage) return null;
418  const check =
419    stage.id === 'preflight' ? (summary.preflight?.checks ?? []).find((c) => c.status === 'failed') : undefined;
420  const phase = stage.phases.find((p) => !p.reused && p.status === 'failed');
421  const message = check
422    ? check.message
423      ? `${check.label}: ${check.message}`
424      : check.label
425    : ((phase && summary.phaseErrors[phase.name]) ??
426      (stage.id === 'build' && summary.runner?.state === 'failed' ? summary.runner.message : null));
427  return {
428    stage: stage.label,
429    text: message
430      ? `${stage.label} stopped with an error: ${message}`
431      : `${stage.label} stopped with an error. Ask Claude in the conversation what went wrong.`,
432  };
433}
434
435/** What the verdict card says when verification left blocking or major problems open; null otherwise. */
436export function verdictOf(findings: Findings | null): string | null {
437  if (!findings || findings.blocker + findings.major === 0) return null;
438  const { blocker, major } = findings;
439  const parts = [blocker ? `${blocker} blocking` : null, major ? `${major} major` : null].filter(Boolean);
440  const one = blocker + major === 1;
441  return `${parts.join(' and ')} problem${one ? ' is' : 's are'} still open, so the library does not yet meet the design-lab standard.`;
442}
443
444/** A phase name as a person reads it: figma-build as Figma build. */
445export function phaseLabel(name: string): string {
446  const words = name.replace(/[-_]/g, ' ');
447  return words.charAt(0).toUpperCase() + words.slice(1);
448}
449
450export const PHASE_DONE = DONE;
451
452/** The status line, or undefined to clear it once the run is over or gone. */
453export function statusOf(summary: Summary, nowMs: number): string | undefined {
454  if (!summary.found) return undefined;
455  if (isFinished(summary)) return undefined;
456  // The engine shows the plugin's name before it: `design-lab: steps 112/158 · runner connected · 41m`.
457  const parts: string[] = [];
458  const runner = summary.runner;
459  const checks = summary.preflight?.checks;
460  if (runner?.state === 'building' && runner.stepsTotal)
461    parts.push(`steps ${runner.stepsDone ?? 0}/${runner.stepsTotal}`);
462  else if (checks && summary.preflight?.status !== 'complete') {
463    parts.push(`preflight ${checks.filter((check) => check.status === 'done').length}/${checks.length}`);
464  } else if (summary.current) parts.push(summary.current);
465  if (runner) parts.push(runnerLine(runner));
466  const time = elapsedOf(summary.startedAt, nowMs);
467  if (time) parts.push(time);
468  return parts.length > 0 ? parts.join(' · ') : undefined;
469}
470
471/** The whole summary as plain text: the command's answer where nothing draws. */
472export function plainOf(summary: Summary): string {
473  if (!summary.found) return `No design-lab run in ${summary.workspace}: it has no project.json.`;
474  const lines = [`design-lab · ${summary.siteLabel ?? summary.workspace}`, ''];
475  const checks = summary.preflight?.checks;
476  if (checks && checks.length > 0) {
477    lines.push('  Preflight');
478    for (const check of checks) {
479      const message = checkMessage(check);
480      lines.push(`    ${CHECK_MARKS[check.status] ?? '·'} ${check.label}${message ? `: ${message}` : ''}`);
481    }
482    lines.push('');
483  }
484  for (const phase of summary.phases) {
485    const mark = DONE.has(phase.status)
486      ? '✓'
487      : phase.status === 'failed'
488        ? '✗'
489        : phase.name === summary.current
490          ? '▸'
491          : phase.status === 'stopped'
492            ? '!'
493            : '·';
494    lines.push(`  ${mark} ${phase.name}`);
495  }
496  if (summary.runner) {
497    const steps = stepsLine(summary.runner);
498    lines.push('');
499    if (steps) lines.push(`  ${steps}`);
500    if (!summary.hasRecap) lines.push(`  ${runnerLine(summary.runner)}`);
501  }
502  const failure = failureOf(summary);
503  if (failure) lines.push('', `  Failed: ${failure.text}`);
504  const need = needsYouOf(summary);
505  if (need) lines.push('', `  Needs you: ${need.message}`);
506  if (summary.hasRecap) lines.push('', `  Recap: ${summary.workspace}/benchmark/completion.md`);
507  return lines.join('\n');
508}
509
510/**
511 * What the resume button does after asking to fill the prompt box: nothing more once filled;
512 * send the request where the session has no prompt box at all; otherwise (a dialog holds the
513 * keys, or the reason is unknown) leave the person's draft alone and say why.
514 */
515export function afterFill(filled: { isFilled: boolean; refusal?: string } | undefined): 'done' | 'submit' | 'explain' {
516  if (filled?.isFilled) return 'done';
517  return filled?.refusal === 'no_composer' ? 'submit' : 'explain';
518}
519
520export const RESUME_PROMPT = (workspace: string) =>
521  `The design-lab runner is open again in Figma desktop. Resume the design-lab run in ${workspace} from where it stopped.`;
522
523/** What the phases recorded about the run: counts from inventory, plan and components, the file
524 * from connect (or, failing that, the recap). */
525export function factsOf(project: Record<string, unknown>, completion: string | undefined): Facts {
526  const phases = record(project.phases);
527  const detail = (name: string) => record(record(phases[name]).detail);
528  return {
529    found: count(detail('inventory').components),
530    toBuild: count(detail('plan').build),
531    built: count(detail('components').built),
532    expected: count(detail('components').expected),
533    figmaUrl: text(detail('connect').fileUrl) ?? (completion ? figmaUrlOf(completion) : null),
534  };
535}
536
537/** The open findings by severity, or null before this run's verification has written its report.
538 * A folder initialised again keeps the last run's report, so one written before this run began is
539 * not this run's. */
540export function findingsOf(report: unknown, createdAt: string | null): Findings | null {
541  const r = record(report);
542  if (!Array.isArray(r.open) || !Array.isArray(r.passed)) return null;
543  const written = Date.parse(text(r.generatedAt) ?? '');
544  if (Number.isNaN(written) || written < Date.parse(createdAt ?? '')) return null;
545  const by = (severity: string) => (r.open as unknown[]).filter((f) => record(f).severity === severity).length;
546  return {
547    blocker: by('blocker'),
548    major: by('major'),
549    minor: by('minor'),
550    passed: r.passed.length,
551    waived: Array.isArray(r.waived) ? r.waived.length : 0,
552  };
553}
554
555/** The five stages a run moves through, and the phases each holds. A phase not named here
556 * belongs to Build, so nothing the run records goes unshown. */
557export const STAGES = [
558  { id: 'preflight', label: 'Preflight', doing: 'Preflight', phases: ['preflight'] },
559  {
560    id: 'discovery',
561    label: 'Discovery',
562    doing: 'Discovering',
563    phases: ['discovery', 'inventory', 'usage', 'capture', 'tokens', 'plan'],
564  },
565  { id: 'build', label: 'Build', doing: 'Building', phases: ['connect', 'foundation', 'components', 'index'] },
566  { id: 'verify', label: 'Verify', doing: 'Verifying', phases: ['verify'] },
567  { id: 'report', label: 'Report', doing: 'Reporting', phases: ['benchmark'] },
568] as const;
569
570export type StageId = (typeof STAGES)[number]['id'];
571// flagged: finished with open blocking or major problems; failed: stopped with an error.
572export type StageState = 'done' | 'flagged' | 'active' | 'stopped' | 'failed' | 'pending' | 'reused';
573export type Stage = {
574  id: StageId;
575  label: string;
576  doing: string;
577  state: StageState;
578  phases: Phase[];
579  note: string | null;
580  noteColor?: string;
581};
582
583// Every known phase in the order the run takes them. A phase not named here sits just after index,
584// so it stays with Build and comes before verify.
585const FLOW: string[] = STAGES.flatMap((stage) => [...stage.phases]);
586const BUILD_PHASES = new Set<string>(STAGES.find((stage) => stage.id === 'build')!.phases);
587
588function flowRank(name: string): number {
589  const at = FLOW.indexOf(name);
590  return at >= 0 ? at : FLOW.indexOf('index') + 0.5;
591}
592
593function stageOf(name: string): StageId {
594  return STAGES.find((stage) => (stage.phases as readonly string[]).includes(name))?.id ?? 'build';
595}
596
597const plural = (n: number, one: string, many: string) => `${n} ${n === 1 ? one : many}`;
598
599function stageNote(
600  id: StageId,
601  state: StageState,
602  summary: Summary,
603  phases: Phase[],
604): { note: string | null; noteColor?: string } {
605  const f = summary.facts;
606  if (state === 'reused') return { note: 'reused from an earlier run' };
607  if (state === 'pending') return { note: null };
608  if (id === 'preflight') {
609    const checks = summary.preflight?.checks;
610    if (!checks) return { note: state === 'done' ? 'passed' : null };
611    const done = checks.filter((check) => check.status === 'done').length;
612    return { note: state === 'done' ? `${checks.length} checks passed` : `${done} of ${checks.length} checks` };
613  }
614  if (id === 'discovery') {
615    if (state !== 'done')
616      return { note: `${phases.filter((phase) => DONE.has(phase.status)).length} of ${phases.length} steps` };
617    const parts = [
618      f.found !== null ? `${f.found} found` : null,
619      f.toBuild !== null ? `${f.toBuild} planned` : null,
620    ].filter(Boolean);
621    return { note: parts.length ? parts.join(' · ') : 'complete' };
622  }
623  if (id === 'build') {
624    const runner = summary.runner;
625    if (state !== 'done' && runner?.state === 'building' && runner.stepsTotal)
626      return { note: `${percentOf(runner.stepsDone ?? 0, runner.stepsTotal)}%` };
627    if (f.built !== null && f.expected !== null) {
628      return {
629        note: `${f.built} of ${f.expected} planned built`,
630        noteColor: state === 'done' && f.built < f.expected ? 'yellow' : undefined,
631      };
632    }
633    return { note: state === 'done' ? 'complete' : null };
634  }
635  if (id === 'verify') {
636    const k = summary.findings;
637    if (!k || (state !== 'done' && state !== 'flagged')) return { note: null };
638    // Kept short so the row fits a narrow pane; the verdict card above the figures has the rest.
639    if (state === 'flagged') {
640      const parts = [
641        k.blocker ? plural(k.blocker, 'blocker', 'blockers') : null,
642        k.major ? `${k.major} major` : null,
643      ].filter(Boolean);
644      return { note: `${parts.join(' · ')} open`, noteColor: 'yellow' };
645    }
646    if (k.minor > 0) return { note: `${k.minor} minor open`, noteColor: 'yellow' };
647    // Checks that do not apply to this site are not problems, so the note leaves them out.
648    return {
649      note: k.waived === 0 ? `all ${k.passed} checks pass` : `${k.passed} passed · ${k.waived} waived`,
650      noteColor: 'green',
651    };
652  }
653  return { note: state === 'done' ? 'report ready' : null };
654}
655
656/** Each stage with where it stands, first match winning: copied from an earlier run, failed, done
657 * (flagged when verification left problems open), stopped for the person, under way, not started.
658 * A stage is under way when it holds the run's current phase or any of its phases is running, so
659 * more than one can be. */
660export function stagesOf(summary: Summary): Stage[] {
661  const finished = isFinished(summary);
662  const need = needsYouOf(summary);
663  const checks = summary.preflight?.checks ?? [];
664  const k = summary.findings;
665  return STAGES.map((def) => {
666    // summary.phases is already in flow order.
667    const phases = summary.phases.filter((phase) => stageOf(phase.name) === def.id);
668    const own = phases.filter((phase) => !phase.reused);
669    const running = own.some((phase) => phase.status === 'running');
670    const holdsCurrent = phases.some((phase) => phase.name === summary.current);
671    const failed =
672      own.some((phase) => phase.status === 'failed') ||
673      (def.id === 'build' && summary.runner?.state === 'failed' && !finished) ||
674      (def.id === 'preflight' && checks.some((check) => check.status === 'failed'));
675    // Once the recap is written the run is over: a stage whose phase record never caught up
676    // (verify left 'running', say) is still done, and Verify is judged by its findings.
677    const done = (phases.length > 0 && phases.every((phase) => DONE.has(phase.status))) || finished;
678    let state: StageState;
679    if (
680      phases.length > 0 &&
681      phases.every((phase) => phase.reused) &&
682      !phases.some((phase) => phase.status === 'running')
683    )
684      state = 'reused';
685    else if (failed) state = 'failed';
686    else if (done) state = def.id === 'verify' && k && k.blocker + k.major > 0 ? 'flagged' : 'done';
687    else if ((holdsCurrent || running) && !finished) {
688      const stopped =
689        (need !== null && holdsCurrent) ||
690        (def.id === 'preflight' && checks.some((check) => check.status === 'needs-you')) ||
691        phases.some((phase) => phase.status === 'stopped');
692      state = stopped ? 'stopped' : 'active';
693    } else state = 'pending';
694    return {
695      id: def.id,
696      label: def.label,
697      doing: def.doing,
698      state,
699      phases,
700      ...stageNote(def.id, state, summary, phases),
701    };
702  });
703}
704
hooks/mod/locate.ts 19 lines
1// Path arithmetic for finding the person's runs; the lookups that read files live in register.tsx,
2// since a mod's engine calls stay in the file that registers its hooks.
3
4export const MARKERS = ['plans', 'analysis-reports', 'design'];
5
6/** Every folder from `path` up to the root, nearest first. */
7export function ancestors(path: string): string[] {
8  const parts = path.replace(/\/+$/, '').split('/').filter(Boolean);
9  return parts.map((_, i) => `/${parts.slice(0, parts.length - i).join('/')}`);
10}
11
12export function base(path: string): string {
13  return path.split('/').filter(Boolean).pop() ?? path;
14}
15
16export function parent(path: string): string {
17  return path.replace(/\/[^/]+\/?$/, '') || '/';
18}
19
src/generated/progress.ts 15 lines
1// Generated from schemas/progress.schema.json. Do not edit.
2
3export interface Progress {
4  state: "waiting" | "preflight" | "building" | "done" | "failed";
5  stepsDone: number | null;
6  stepsTotal: number | null;
7  step: string | null;
8  stepKind: string | null;
9  message: string | null;
10  inflight: boolean;
11  lastSeen: string | null;
12  at: string;
13  serverPid: number;
14}
15
src/generated/runner-step.ts 102 lines
1// Generated from schemas/runner-step.schema.json. Do not edit.
2
3export type RunnerStep =
4  | {
5      step?: string;
6      done?: number;
7      total?: number;
8      kind: "done";
9      buildId?: string;
10      generation?: string;
11      stepToken?: string;
12    }
13  | {
14      step: string;
15      done?: number;
16      total?: number;
17      kind: "wait";
18      retryMs: number;
19      message: string;
20      buildId?: string;
21      generation?: string;
22      stepToken?: string;
23    }
24  | {
25      step: string;
26      done?: number;
27      total?: number;
28      kind: "check";
29      code: string;
30      out?: string;
31      buildId?: string;
32      generation?: string;
33      stepToken?: string;
34    }
35  | {
36      step: string;
37      done?: number;
38      total?: number;
39      kind: "dump";
40      code: string;
41      out?: string;
42      buildId?: string;
43      generation?: string;
44      stepToken?: string;
45    }
46  | ((
47      | {
48          code: string;
49        }
50      | {
51          payload: string;
52        }
53    ) & {
54      step: string;
55      done?: number;
56      total?: number;
57      kind: "use_figma";
58      code?: string;
59      payload?: string;
60      characters?: number;
61      buildId?: string;
62      generation?: string;
63      stepToken?: string;
64    })
65  | {
66      step: string;
67      done?: number;
68      total?: number;
69      kind: "upload";
70      nodeIds: string[];
71      scaleMode: "FIT" | "FILL";
72      files?: {
73        file: string;
74        contentType: string;
75      }[];
76      buildId?: string;
77      generation?: string;
78      stepToken?: string;
79    }
80  | {
81      step: string;
82      done?: number;
83      total?: number;
84      kind: "screenshot";
85      nodeId: string;
86      out: string;
87      maxDimension?: number;
88      buildId?: string;
89      generation?: string;
90      stepToken?: string;
91    }
92  | {
93      step: string;
94      done?: number;
95      total?: number;
96      kind: "skip";
97      reason: string;
98      buildId?: string;
99      generation?: string;
100      stepToken?: string;
101    };
102
types/index.d.ts 92 lines
1// Generated from src/protocol.ts by scripts/generate-mod-contract.ts. Do not edit.
2export type ProgressState = "waiting" | "preflight" | "building" | "done" | "failed";
3export type StepKind = "wait" | "done" | "check" | "dump" | "use_figma" | "upload" | "screenshot" | "skip";
4
5// reused: copied from an earlier run (design-lab:figma-build), not run again here
6export type Phase = { name: string; status: string; reused: boolean };
7
8// What the phases recorded that the pane says beside each stage; each null until recorded.
9export type Facts = {
10  found: number | null;
11  toBuild: number | null;
12  built: number | null;
13  expected: number | null;
14  figmaUrl: string | null;
15};
16
17// The verification report's open findings by severity, the checks it passed, and those it waived.
18export type Findings = { blocker: number; major: number; minor: number; passed: number; waived: number };
19
20export type Runner = {
21  state: ProgressState | 'connecting';
22  stepsDone: number | null;
23  stepsTotal: number | null;
24  stepKind: StepKind | null;
25  message: string | null;
26  serverAlive: boolean;
27  connected: boolean;
28  lastSeenMs: number | null;
29};
30
31export type Check = {
32  id: string;
33  label: string;
34  status: string;
35  message: string | null;
36  dependsOn: string[];
37};
38
39// The headline figures of a finished run, from its scorecard: each null when the scorer left it out.
40export type Scores = {
41  built: number | null;
42  eligible: number | null;
43  withinTolerance: number | null;
44  widths: number | null;
45  workingSeconds: number | null;
46  buildSeconds: number | null;
47  buildSteps: number | null;
48  tokens: number | null;
49  toolCalls: number | null;
50  blockers: number | null;
51  majors: number | null;
52};
53
54export type Summary = {
55  workspace: string;
56  found: boolean;
57  siteLabel: string | null;
58  phases: Phase[];
59  current: string | null;
60  // checks: the list preflight last wrote, or null when it has none newer than the recorded phase
61  preflight: { status: string; at: string | null; checks: Check[] | null } | null;
62  runner: Runner | null;
63  blocker: string | null;
64  // the run is waiting for the person to do something it will notice by itself (start the runner)
65  waiting: string | null;
66  log: string[];
67  hasRecap: boolean;
68  recap: string | null;
69  scores: Scores | null;
70  facts: Facts;
71  findings: Findings | null;
72  startedAt: string | null;
73  // the last message the phase log holds for each failed phase, by phase name
74  phaseErrors: Record<string, string>;
75  // the run's artifact files that exist, run-relative (read only once the recap is written)
76  present: string[];
77};
78
79declare module 'claude-code' {
80  interface PluginState {
81    'design-lab': {
82      run: string | null
83      summary: Summary | null
84      alarmed: boolean
85      follow: string | null
86      skip: string | null
87      // the run whose full completion message is shown under its figures, or null when folded
88      recapOpen: string | null
89    }
90  }
91}
92