SLOPSHOPPER

test-user

A Haiku test user that clicks through your locally running app and reports UI/UX issues, with a live pane of its tasks, findings and failures

newpaneguardcommandtoaststatus
v0.2.0no licenseupdated 2026-10-09WeaponizedLego/claude-code-mods/test-user
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · test-user
│ ┃ Test user ✕ › fix the failing auth test and add an audit log call │ ┃ ╭──────────────────────────────────────────╮ │ ┃ │ No test run yet │ ⏺ Read(src/auth.ts) │ ┃ │ Ask Claude to test the app, or run │ ⎿ Read 6 lines │ ┃ │ /test-user run [what to test]. │ ⏺ Update(src/auth.ts) │ ┃ ╰──────────────────────────────────────────╯ ⎿ Added 2 lines, removed 1 line │ ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /test-user │ ⎿ test-user: No test run yet: ask Claude to test the app, or run / │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · Test user
╭──────────────────────────────────────────────────────────╮ │ No test run yet │ │ Ask Claude to test the app, or run /test-user run [what │ │ to test]. │ ╰──────────────────────────────────────────────────────────╯
README

claude-code-mods

My Claude Code mods, one folder per mod. Needs Claude Code 2.1.287 or newer.

ModWhat it does
usage-bandA band above the prompt showing context fill, tokens, cost and rate limits: stat tiles in the desktop app, a coloured row in the terminal
plan-progressReads a phased plan into a live progress tree: phases, steps, elapsed time and an estimate of what is left
test-userA Haiku test user that clicks through your locally running app and reports UI/UX issues: its task list, findings and failures in a live pane

All three mods share one look: dark violet cards with a lilac accent. In the desktop app's Code tab they draw as images (cards, tiles, gradients); in the terminal they use the same colours as text.

usage-band in the desktop app

Install

At the prompt of a terminal Claude Code session:

/plugin install usage-band --marketplace WeaponizedLego/claude-code-mods
/plugin install plan-progress --marketplace WeaponizedLego/claude-code-mods

Answer y to add the marketplace, then pick the user scope. Or from a shell:

claude plugin marketplace add WeaponizedLego/claude-code-mods
claude plugin install usage-band@claude-code-mods
claude plugin install plan-progress@claude-code-mods

Update later with claude plugin marketplace update claude-code-mods, then claude plugin update usage-band@claude-code-mods.

plan-progress

plan-progress in the desktop app

When a plan with several phases is approved in plan mode, the mod reads its phases and steps (## Phase 1: ... headings, other work headings, or a nested list) into a tree and opens it in a pane. In the terminal it looks like this:

Move auth to sessions                         20m in · ~40m left
▰▰▰▰▰▰▱▱▱▱▱▱▱▱▱▱▱▱ 33%  2/6 steps · phase 2/3

✓ 1 Setup                                                 20m
  ├ ✓ Add the sessions table                              10m
  └ ✓ Write the session store                             10m
● 2 Switch over
  ├ ● Replace token checks in the middleware
  ├ ○ Update the login endpoint
  └ ○ Remove the old token helpers
○ 3 Verify
  └ ○ Run the auth test suite

The status line carries the short form (▰▰▱▱ 2/6 · phase 2/3: Switch over · 20m in · ~40m left). Steps move when Claude calls the mod's plan_mark tool, or by themselves when Claude's tasks or todos are named like a step. Without plan mode, ask Claude to lay the work out in phases and it calls plan_set. The estimate is the pace so far times the steps left.

/plan-progress opens the tree (it also works mid-turn); /plan-progress clear stops tracking.

test-user

test-user in the desktop app

A test user for the app you are running locally. Ask Claude to "test the app" (or a feature, a page, a flow) and it hands the job to test-user:tester, a subagent on Haiku that drives the app in the browser the way a person would: it works out what to test, finds the app (the URL you gave, .claude/launch.json, package.json), plans a handful of user tasks, works through them and reports each problem it sees. It never edits code, stays on localhost and only uses test data.

When you just say "test the app", Claude points it at the feature most recently built in the session; with nothing built, the tester looks at git diff and git log for a recent feature, and failing that tests the whole application.

The pane shows the run as it goes: status, tasks passed out of total, a count per severity (blocker, major, minor, polish), the task list with a note on each failure, a card per finding (worst first, with where and how to reproduce), a red card when it could not test at all, and earlier runs folded at the bottom. The status line carries the short form (test-user ● 3/6 tasks · 2 findings · 4m · Pay with the test card) and a toast says how it ended.

The session gets the full record, not just the tester's closing message: it rides back on the Agent result when the tester ran in the foreground, and on the next prompt (or the background task's notification) otherwise; test_report reads it any time. A tester whose turn ends early and is resumed carries on with the same run.

Claude gets the same reporting tools (test_plan, test_task, test_finding, test_finish), so a test it runs by hand shows in the pane too, and it can amend the latest tester run with what it re-checked. If the tester stops without a verdict, the record decides it; a "passed" never hides a failed task.

/test-user opens the pane; /test-user run [what to test] starts a run; /test-user clear forgets the runs.

Adding a mod

Put it in its own folder (<mod>/.claude-plugin/plugin.json, <mod>/hooks/...) and add an entry to .claude-plugin/marketplace.json. Check it with claude plugin validate ..

Source 5 files
hooks/register.tsx 545 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { Run, RunStatus, Severity, TaskStatus } from '../types'
5import { C, SEVERITY_COLOR, STATUS_COLOR, TASK_COLOR, TASK_GLYPH, emptySvg, runSvg } from './look'
6import { AGENT_DESCRIPTION, TESTER_PROMPT } from './prompt'
7import {
8  MAX_RUNS,
9  SEVERITIES,
10  STATUS_LABEL,
11  addFinding,
12  advance,
13  bySeverity,
14  clean,
15  cleanDetail,
16  duration,
17  finish,
18  findingsLine,
19  historyLine,
20  newRun,
21  plan,
22  plural,
23  reopen,
24  report,
25  setTask,
26  stats,
27  statusLine,
28  verdictOf,
29} from './run'
30
31const PANE = 'test-user'
32const TITLE = 'Test user'
33const AGENT = 'test-user:tester'
34const PLAN = 'mcp__test-user__test_plan'
35const TASK = 'mcp__test-user__test_task'
36const FINDING = 'mcp__test-user__test_finding'
37const FINISH = 'mcp__test-user__test_finish'
38const REPORT = 'mcp__test-user__test_report'
39const TICK_MS = 30_000
40// Roughly one cell of the desktop's code font, in CSS pixels.
41const CELL_PX = 8
42const MAIN = 'main'
43
44const runsAtom = atom({ plugin: 'test-user', key: 'runs' } as const, [])
45const nowAtom = atom({ plugin: 'test-user', key: 'now' } as const, 0)
46
47type PlanInput = { tasks?: unknown[]; target?: string; scope?: 'feature' | 'app' }
48type TaskInput = { task?: number; status?: TaskStatus; note?: string }
49type FindingInput = { severity?: Severity; title?: string; detail?: string; where?: string; task?: number }
50type FinishInput = { verdict?: 'passed' | 'issues' | 'blocked'; summary?: string; reason?: string }
51
52// The loop a call came from: a tester subagent's id, or the main conversation.
53const loopOf = (e: { agentId?: string }) => e.agentId ?? MAIN
54
55async function save($: EngineInterface, fn: (runs: Run[], now: number) => Run[]): Promise<Run[]> {
56  const now = await $.clock.now()
57  const runs = await update($, runsAtom, r => fn(r, now).slice(0, MAX_RUNS))
58  await update($, nowAtom, () => now)
59  $.ui.status(runs[0] ? statusLine(runs[0], now) : undefined)
60  return runs
61}
62
63const same = (a: Run, b: Run) => a.agentId === b.agentId && a.startedAt === b.startedAt
64
65// The run a loop's updates go to: a tester's own latest run, ended or not (a
66// tester whose turn ended early and was resumed carries on with it); for the
67// main conversation, its own run in progress, else the latest run of all, so
68// Claude can amend a tester's run with what it re-checked.
69function targetOf(runs: Run[], loop: string): number {
70  if (loop !== MAIN) return runs.findIndex(r => r.agentId === loop)
71  const own = runs.findIndex(r => r.agentId === MAIN && r.status === 'running')
72  return own >= 0 ? own : runs.length ? 0 : -1
73}
74
75/** Applies `fn` to the run this loop updates; null when there is none. */
76async function changeRun($: EngineInterface, loop: string, fn: (run: Run, now: number) => Run): Promise<Run | null> {
77  let changed: Run | null = null
78  await save($, (runs, now) => {
79    // `update` may run this again on a version miss: start each pass afresh.
80    changed = null
81    const i = targetOf(runs, loop)
82    return runs.map((r, j) => {
83      if (j !== i) return r
84      if (r.agentId === loop) return (changed = advance(fn(reopen(r), now), now))
85      // Claude amending a tester's run: an ended run stays ended, its verdict redone.
86      const amended = fn(r, now)
87      return (changed = r.status === 'running' ? advance(amended, now) : { ...amended, status: verdictOf(amended), isDelivered: false })
88    })
89  })
90  return changed
91}
92
93async function startRun($: EngineInterface, loop: string, tester: string, brief: string, toolUseId?: string): Promise<Run> {
94  const runs = await save($, (runs, now) => {
95    // A loop has one run going at a time; an older one it left open is closed.
96    const rest = runs.map(r => (r.agentId === loop && r.status === 'running' ? finish(r, verdictOf(r), now, 'turn') : r))
97    return [{ ...newRun(loop, tester, brief, now), toolUseId }, ...rest]
98  })
99  await openPane($, false)
100  return runs[0]!
101}
102
103async function markDelivered($: EngineInterface, delivered: Run[]) {
104  await save($, runs => runs.map(r => (delivered.some(d => same(d, r)) ? { ...r, isDelivered: true } : r)))
105}
106
107const deliveredText = (run: Run, now: number) =>
108  `test-user results (the full record from the test-user pane, including updates the tester's own final message may leave out):\n\n${report(run, now)}`
109
110// A pane that cannot open (refused, or no surface to draw on) never costs the run.
111async function openPane($: EngineInterface, isAsked: boolean) {
112  const opened = await $.ui.open({ id: PANE, title: TITLE }).catch(() => ({ isPlaced: false }))
113  if (!opened.isPlaced && !isAsked) $.ui.toast('test-user: run /test-user to watch the test run')
114}
115
116const outcomeToast = (run: Run) => {
117  const n = run.findings.length
118  if (run.status === 'passed') return `test-user: all ${run.tasks.length} tasks passed`
119  if (run.status === 'issues') return `test-user: done — ${plural(n, 'finding')}, ${run.tasks.filter(t => t.status === 'failed').length} failed`
120  return `test-user: ${STATUS_LABEL[run.status].toLowerCase()}${run.failure ? ` — ${run.failure}` : ''}`
121}
122
123const firstLine = (t: string) => t.split(/\r?\n/).find(l => l.trim()) ?? ''
124
125export const register: Register = on => {
126  on('session.start', async ($, e, next) => {
127    const result = await next(e)
128
129    await $.agent.register({
130      name: 'tester',
131      description: AGENT_DESCRIPTION,
132      prompt: TESTER_PROMPT,
133      model: 'haiku',
134      disallowedTools: ['Edit', 'Write', 'NotebookEdit', 'Agent', 'Workflow', 'Artifact', REPORT],
135      maxTurns: 120,
136    })
137
138    await $.tool.register({
139      name: 'test_plan',
140      description:
141        'Start (or replace) the task list of a UI/UX test run, shown live to the developer. Call once you know what you are testing. ' +
142        'Used by the test-user:tester agent; Claude may also use it to record a test it runs by hand.',
143      inputSchema: {
144        type: 'object',
145        properties: {
146          tasks: { type: 'array', minItems: 1, maxItems: 15, items: { type: 'string' }, description: 'Things a user does, in order' },
147          target: { type: 'string', description: 'Short name of what is tested, e.g. "Checkout · localhost:5173"' },
148          scope: { type: 'string', enum: ['feature', 'app'], description: 'A recent feature, or the whole application' },
149        },
150        required: ['tasks'],
151      },
152      isDeferred: false,
153    })
154    await $.tool.register({
155      name: 'test_task',
156      description: 'Mark a task of the current test run active, passed, failed or skipped. Tasks count from 1. The next task becomes active by itself.',
157      inputSchema: {
158        type: 'object',
159        properties: {
160          task: { type: 'integer', minimum: 1 },
161          status: { type: 'string', enum: ['active', 'passed', 'failed', 'skipped'] },
162          note: { type: 'string', description: 'One line: what happened' },
163        },
164        required: ['task', 'status'],
165      },
166      isDeferred: false,
167    })
168    await $.tool.register({
169      name: 'test_finding',
170      description: 'Report one UI/UX problem seen during the current test run. One call per problem.',
171      inputSchema: {
172        type: 'object',
173        properties: {
174          severity: { type: 'string', enum: SEVERITIES, description: 'blocker: cannot complete; major: works badly; minor: friction or visual defect; polish: nit' },
175          title: { type: 'string', description: 'The problem in a few words' },
176          detail: { type: 'string', description: 'What you did, what happened, what you expected' },
177          where: { type: 'string', description: 'The screen, URL or element' },
178          task: { type: 'integer', minimum: 1, description: 'The task it came up in' },
179        },
180        required: ['severity', 'title'],
181      },
182      isDeferred: false,
183    })
184    await $.tool.register({
185      name: 'test_finish',
186      description: 'End the current test run with a verdict. "blocked" when the app could not be tested at all (give the reason).',
187      inputSchema: {
188        type: 'object',
189        properties: {
190          verdict: { type: 'string', enum: ['passed', 'issues', 'blocked'] },
191          summary: { type: 'string', description: 'Two to four sentences for the developer' },
192          reason: { type: 'string', description: 'Why it was blocked' },
193        },
194        required: ['verdict', 'summary'],
195      },
196      isDeferred: false,
197    })
198    await $.tool.register({
199      name: 'test_report',
200      description:
201        'Read the latest test-user run (or the one in progress) as the developer sees it in the test-user pane: its tasks, findings and why it failed. ' +
202        'Use it to answer questions about a test run or to check what the tester recorded.',
203      inputSchema: { type: 'object', properties: {} },
204      isDeferred: false,
205    })
206    await $.command.register({
207      name: 'test-user',
208      description: 'Show the test user pane (run [what]: start a test run; clear: forget the runs)',
209      argumentHint: '[run [what to test] | clear]',
210      immediate: true,
211    })
212
213    $.clock.every(TICK_MS, () => {
214      void (async () => {
215        const runs = await read($, runsAtom)
216        if (runs[0]?.status !== 'running') return
217        const t = await $.clock.now()
218        await update($, nowAtom, () => t)
219        $.ui.status(statusLine(runs[0], t))
220      })()
221    })
222
223    // After a reload the runs are still in the session's state.
224    const runs = await read($, runsAtom)
225    if (runs[0]) $.ui.status(statusLine(runs[0], await $.clock.now()))
226    return result
227  })
228
229  // A tester subagent starting is a run starting.
230  on('agent.spawn', async ($, e, next) => {
231    const started = await next(e)
232    if (e.subagentType !== AGENT || !('agentId' in started) || !started.agentId) return started
233    await startRun($, started.agentId, 'Haiku test user', e.description || firstLine(e.prompt), e.tool_use_id)
234    return started
235  })
236
237  // Its turn ending ends the run, whatever the tester remembered to say; a
238  // tester resumed after that carries on with the same run (targetOf).
239  on('turn.complete', async ($, e, next) => {
240    const result = await next(e)
241    const loop = e.agentId
242    if (!loop) return result
243    const runs = await read($, runsAtom)
244    if (!runs.some(r => r.agentId === loop && r.status === 'running')) return result
245
246    const ended = await save($, (runs, now) =>
247      runs.map(r => {
248        if (r.agentId !== loop || r.status !== 'running') return r
249        if (e.reason === 'aborted') return finish(r, 'failed', now, 'turn', undefined, 'The run was interrupted.')
250        if (e.reason !== 'answer') {
251          return finish(r, 'failed', now, 'turn', undefined, `The tester stopped: ${e.reason === 'refusal' ? 'the model refused' : 'an API error'}.`)
252        }
253        return finish(r, verdictOf(r), now, 'turn', r.summary ?? (clean(e.answer, 300) || undefined))
254      }),
255    )
256    const run = ended.find(r => r.agentId === loop)
257    if (run) $.ui.toast(outcomeToast(run))
258    return result
259  })
260
261  // The session gets the full record: a foreground tester's on its Agent result...
262  on('tool.call', { tool: 'Agent' }, async ($, e, next) => {
263    const ran = await next(e)
264    if ((e as unknown as { subagent_type?: string }).subagent_type !== AGENT || ran.deny !== undefined) return ran
265    const run = (await read($, runsAtom)).find(r => r.toolUseId === e.tool_use_id)
266    if (!run) return ran
267    const now = await $.clock.now()
268    if (run.status === 'running') {
269      const note = `test-user: this run is still going (${statusLine(run, now)}). Its full record reaches you when it ends; ${REPORT} reads it at any time.`
270      return { ...ran, context: [...(ran.context ?? []), note] }
271    }
272    await markDelivered($, [run])
273    return { ...ran, context: [...(ran.context ?? []), deliveredText(run, now)] }
274  }).catch(($, e, next) => next(e))
275
276  // ...and any run that ended unseen (a background tester, a resumed one) rides
277  // on the next prompt, a background task's notification included.
278  on('prompt.submit', async ($, e, next) => {
279    const fresh = (await read($, runsAtom)).filter(r => r.status !== 'running' && r.isDelivered === false)
280    if (fresh.length === 0) return next(e)
281    const now = await $.clock.now()
282    await markDelivered($, fresh)
283    return next({ ...e, context: [...(e.context ?? []), ...fresh.map(r => deliveredText(r, now))] })
284  }).catch(($, e, next) => next(e))
285
286  on('tool.call', { tool: PLAN }, async ($, e) => {
287    const input = e as unknown as PlanInput
288    const tasks = (input.tasks ?? []).map(t => clean(t)).filter(Boolean)
289    if (tasks.length === 0) return { deny: 'test_plan needs at least one task.' }
290    const loop = loopOf(e)
291    const runs = await read($, runsAtom)
292    // A tester carries on with its run if its turn ended it; anything else is a new run.
293    const own = runs.find(r => r.agentId === loop)
294    const carriesOn = own && (own.status === 'running' || (loop !== MAIN && own.endedBy === 'turn'))
295    if (!carriesOn) await startRun($, loop, loop === MAIN ? 'Claude' : 'Haiku test user', input.target ?? 'Test run')
296    const run = (await changeRun($, loop, (r, now) => ({
297      ...plan(r, tasks, now),
298      target: input.target ? clean(input.target, 80) : r.target,
299      scope: input.scope ?? r.scope,
300    })))!
301    const at = run.tasks.findIndex(t => t.status === 'active')
302    const list = run.tasks.map((t, i) => `${i + 1}. ${t.title}${t.status === 'pending' || t.status === 'active' ? '' : ` (${t.status})`}`)
303    return {
304      result: `Task list shown to the developer:\n${list.join('\n')}\n\n${at >= 0 ? `Task ${at + 1} is active. ` : ''}Mark each with test_task; report problems with test_finding.`,
305    }
306  })
307
308  on('tool.call', { tool: TASK }, async ($, e) => {
309    const input = e as unknown as TaskInput
310    const loop = loopOf(e)
311    const status = input.status ?? 'passed'
312    let error = ''
313    const run = await changeRun($, loop, (r, now) => {
314      const i = Number(input.task) - 1
315      if (!r.tasks[i]) {
316        error = `There is no task ${input.task}; the run has ${r.tasks.length}.`
317        return r
318      }
319      return setTask(r, i, status, input.note ? clean(input.note, 160) : undefined, now)
320    })
321    if (!run) return { deny: 'No test run here yet. Call test_plan first.' }
322    if (error) return { deny: error }
323    const s = stats(run, await $.clock.now())
324    const cur = s.current != null ? ` Now on task ${s.current + 1}: ${run.tasks[s.current]!.title}.` : s.closed === s.total ? ' All tasks done: call test_finish.' : ''
325    return { result: `Task ${input.task} ${status}. ${s.closed}/${s.total} done, ${plural(run.findings.length, 'finding')}.${cur}` }
326  })
327
328  on('tool.call', { tool: FINDING }, async ($, e) => {
329    const input = e as unknown as FindingInput
330    const title = clean(input.title)
331    if (!title) return { deny: 'A finding needs a title.' }
332    const severity: Severity = SEVERITIES.includes(input.severity as Severity) ? (input.severity as Severity) : 'minor'
333    const run = await changeRun($, loopOf(e), (r, now) => {
334      const active = r.tasks.findIndex(t => t.status === 'active')
335      const task = input.task != null && r.tasks[Number(input.task) - 1] ? Number(input.task) - 1 : active >= 0 ? active : undefined
336      return addFinding(
337        r,
338        { severity, title, detail: input.detail ? cleanDetail(input.detail) : undefined, where: input.where ? clean(input.where, 80) : undefined, task },
339        now,
340      )
341    })
342    if (!run) return { deny: 'No test run here yet. Call test_plan first.' }
343    return { result: `Recorded (${severity}). ${findingsLine(stats(run, 0))} so far.` }
344  })
345
346  on('tool.call', { tool: FINISH }, async ($, e) => {
347    const input = e as unknown as FinishInput
348    const loop = loopOf(e)
349    const summary = input.summary ? cleanDetail(input.summary) : undefined
350    const end = (r: Run, now: number): Run => {
351      // The record wins over a verdict it contradicts: failed tasks or findings are issues.
352      const verdict: RunStatus = input.verdict === 'blocked' ? 'blocked' : verdictOf(r) === 'issues' || input.verdict === 'issues' ? 'issues' : 'passed'
353      return finish(r, verdict, now, 'tester', summary, verdict === 'blocked' ? clean(input.reason ?? input.summary ?? 'No reason given', 300) : undefined)
354    }
355    // Blocked before it planned anything: still a run the developer should see.
356    if (targetOf(await read($, runsAtom), loop) < 0) {
357      if (input.verdict !== 'blocked') return { deny: 'No test run here yet. Call test_plan first.' }
358      await startRun($, loop, loop === MAIN ? 'Claude' : 'Haiku test user', 'Test run')
359    }
360    const done = (await changeRun($, loop, end))!
361    $.ui.toast(outcomeToast(done))
362    return { result: `Run ended: ${STATUS_LABEL[done.status]}. Now give your final report.` }
363  })
364
365  on('tool.call', { tool: REPORT }, async ($, e) => {
366    const runs = await read($, runsAtom)
367    if (!runs[0]) return { result: 'No test run yet. Spawn the test-user:tester agent to run one.' }
368    const now = await $.clock.now()
369    if (runs[0].status !== 'running') await markDelivered($, [runs[0]])
370    const earlier = runs.slice(1).map(r => `- ${historyLine(r, now)}`)
371    return { result: report(runs[0], now) + (earlier.length ? `\n\nEarlier runs:\n${earlier.join('\n')}` : '') }
372  })
373
374  on('command.run', { command: 'test-user' }, async ($, e) => {
375    const args = e.args.trim()
376    if (args === 'clear') {
377      await save($, () => [])
378      await $.ui.close({ id: PANE }).catch(() => {})
379      return { text: 'Cleared the test runs.' }
380    }
381    await openPane($, true)
382    const run = /^run\b\s*(.*)$/is.exec(args)
383    if (run) {
384      const what = run[1]!.trim()
385      void $.prompt.submit({
386        text: what
387          ? `Use the ${AGENT} agent to UI/UX test: ${what}`
388          : `Use the ${AGENT} agent to UI/UX test the application. Focus on the feature most recently developed in this session; if nothing was built in this session, test the entire application.`,
389      })
390      return { text: `Starting a test run${what ? `: ${what}` : ''}.` }
391    }
392    const runs = await read($, runsAtom)
393    if (!runs[0]) return { text: 'No test run yet: ask Claude to test the app, or run /test-user run [what to test].' }
394    return { text: statusLine(runs[0], await $.clock.now()) }
395  })
396
397  on('session.end', async ($, e, next) => {
398    if (e.reason === 'clear') await save($, () => [])
399    return next(e)
400  })
401
402  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
403    const runs = await read($, runsAtom)
404    await read($, nowAtom)
405    const now = await $.clock.now()
406
407    // Where the surface draws images, the pane is a stack of cards.
408    if (e.surface !== 'terminal') {
409      const { Box, Svg } = $.ui.resolve(e)
410      const width = e.props.bodyColumns * CELL_PX
411      if (!runs[0]) {
412        const empty = emptySvg(width)
413        return <Svg source={empty.source} alt="No test run yet" width={Math.max(300, width)} height={empty.height} />
414      }
415      const drawn = runSvg(runs, now, width)
416      return (
417        <Box flexDirection="column" key="test-user">
418          <Svg source={drawn.source} alt={drawn.alt} width={Math.max(300, width)} height={drawn.height} />
419        </Box>
420      )
421    }
422
423    const { Box, Text } = $.ui.resolve(e)
424    const run = runs[0]
425    if (!run) {
426      return (
427        <Box flexDirection="column" borderStyle="round" borderColor={C.stroke} paddingX={1}>
428          <Text bold color={C.text}>
429            No test run yet
430          </Text>
431          <Text color={C.muted} wrap="wrap">
432            Ask Claude to test the app, or run /test-user run [what to test].
433          </Text>
434        </Box>
435      )
436    }
437
438    const s = stats(run, now)
439    const tone = STATUS_COLOR[run.status]
440    return (
441      <Box flexDirection="column" key="test-user">
442        <Box flexDirection="row" justifyContent="space-between">
443          <Text bold color={C.text} wrap="truncate">
444            {run.target ?? run.brief}
445          </Text>
446          <Text color={tone} bold>{` ${STATUS_LABEL[run.status]}`}</Text>
447        </Box>
448        <Box flexDirection="row">
449          <Text bold color={C.text}>{`${s.passed}/${s.total || '–'}`}</Text>
450          <Text color={C.muted}>{` passed · ${s.failed} failed · `}</Text>
451          <Text color={s.counts.blocker ? C.hot : s.counts.major ? C.warn : C.accent}>{plural(run.findings.length, 'finding')}</Text>
452          <Text color={C.muted}>{` · ${duration(s.elapsedMs)} · ${run.tester}`}</Text>
453        </Box>
454
455        {run.failure && (
456          <Box flexDirection="column" borderStyle="round" borderColor={C.hot} paddingX={1} key="failure">
457            <Text bold color={C.hot}>
458              {run.status === 'blocked' ? 'Could not test' : 'The run failed'}
459            </Text>
460            <Text color={C.soft} wrap="wrap">
461              {run.failure}
462            </Text>
463          </Box>
464        )}
465
466        <Text> </Text>
467        <Box flexDirection="column" borderStyle="round" borderColor={run.status === 'running' ? C.deep : C.stroke} paddingX={1} key="tasks">
468          <Box flexDirection="row" justifyContent="space-between">
469            <Text bold color={C.text}>
470              Tasks
471            </Text>
472            <Text color={C.muted}>{`${s.closed}/${s.total}`}</Text>
473          </Box>
474          {run.tasks.length === 0 && <Text color={C.muted}>Working out what to test…</Text>}
475          {run.tasks.flatMap((t, i) => {
476            const took = t.startedAt != null && t.status !== 'skipped' ? duration((t.endedAt ?? now) - t.startedAt) : ''
477            const row = (
478              <Box flexDirection="row" justifyContent="space-between" key={`task-${i}`}>
479                <Box flexDirection="row" flexShrink={1}>
480                  <Text color={TASK_COLOR[t.status]}>{`${TASK_GLYPH[t.status]} `}</Text>
481                  <Text
482                    color={t.status === 'active' || t.status === 'failed' ? C.text : t.status === 'passed' ? C.soft : C.muted}
483                    bold={t.status === 'active'}
484                    strikethrough={t.status === 'skipped'}
485                    wrap="truncate"
486                  >
487                    {t.title}
488                  </Text>
489                </Box>
490                <Text color={t.status === 'active' ? C.accent : C.muted}>{took ? ` ${took}` : ''}</Text>
491              </Box>
492            )
493            if (!t.note || (t.status !== 'failed' && t.status !== 'active')) return [row]
494            return [
495              row,
496              <Text color={t.status === 'failed' ? C.hot : C.muted} wrap="truncate" key={`note-${i}`}>
497                {`  ${t.note}`}
498              </Text>,
499            ]
500          })}
501        </Box>
502
503        {bySeverity(run.findings).map((f, i) => (
504          <Box flexDirection="column" borderStyle="round" borderColor={SEVERITY_COLOR[f.severity]} paddingX={1} key={`finding-${i}`}>
505            <Box flexDirection="row" justifyContent="space-between">
506              <Text bold color={C.text} wrap="truncate">
507                {f.title}
508              </Text>
509              <Text color={SEVERITY_COLOR[f.severity]} bold>{` ${f.severity.toUpperCase()}`}</Text>
510            </Box>
511            {(f.where || f.task != null) && (
512              <Text color={C.muted} wrap="truncate">
513                {[f.task != null && run.tasks[f.task] ? `task ${f.task + 1}` : '', f.where ?? ''].filter(Boolean).join(' · ')}
514              </Text>
515            )}
516            {f.detail && (
517              <Text color={C.soft} wrap="wrap">
518                {f.detail}
519              </Text>
520            )}
521          </Box>
522        ))}
523
524        {run.summary && (
525          <Box flexDirection="column" paddingX={1} key="summary">
526            <Text bold color={C.text}>
527              Summary
528            </Text>
529            <Text color={C.soft} wrap="wrap">
530              {run.summary}
531            </Text>
532          </Box>
533        )}
534
535        {runs.slice(1).map((old, i) => (
536          <Box flexDirection="row" key={`old-${i}`}>
537            <Text color={STATUS_COLOR[old.status]}>{'● '}</Text>
538            <Text color={C.muted} wrap="truncate">{`Earlier: ${historyLine(old, now)}`}</Text>
539          </Box>
540        ))}
541      </Box>
542    )
543  })
544}
545
hooks/look.ts 243 lines
1import type { Run, RunStatus, Severity, Task, TaskStatus } from '../types'
2import { STATUS_LABEL, SEVERITIES, bySeverity, duration, historyLine, plural, stats } from './run'
3
4// The palette the other mods share: dark violet cards with a lilac accent,
5// orange and red for what needs a look. An SVG is drawn as an image, so it
6// takes real colours, not theme keys; the terminal uses the same.
7export const C = {
8  card: '#23222b',
9  raised: '#2a2933',
10  stroke: '#363541',
11  text: '#ececf1',
12  soft: '#c9c7d3',
13  muted: '#8e8c9a',
14  dim: '#5f5d6b',
15  accent: '#a98bff',
16  deep: '#7d68c9',
17  glow: '#c4b2ff',
18  track: '#3b3650',
19  warn: '#ff8a4c',
20  hot: '#ff5c7a',
21} as const
22
23export const SEVERITY_COLOR: Record<Severity, string> = { blocker: C.hot, major: C.warn, minor: C.accent, polish: C.muted }
24export const STATUS_COLOR: Record<RunStatus, string> = { running: C.glow, passed: C.accent, issues: C.warn, blocked: C.hot, failed: C.hot }
25export const TASK_COLOR: Record<TaskStatus, string> = { pending: C.dim, active: C.glow, passed: C.accent, failed: C.hot, skipped: C.dim }
26export const TASK_GLYPH: Record<TaskStatus, string> = { pending: '○', active: '●', passed: '✓', failed: '✗', skipped: '–' }
27
28const FONT = `font-family="Inter, 'Segoe UI', system-ui, -apple-system, sans-serif"`
29
30export const esc = (t: string) => t.replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/>/g, '&gt;').replace(/"/g, '&quot;')
31
32// Text in an image cannot wrap or truncate itself: cut it to an estimated width.
33const chars = (px: number, size: number) => Math.max(3, Math.floor(px / (size * 0.52)))
34const fit = (t: string, px: number, size: number) => {
35  const max = chars(px, size)
36  return t.length > max ? `${t.slice(0, max - 1)}…` : t
37}
38function wrap(t: string, px: number, size: number, maxLines: number): string[] {
39  const max = chars(px, size)
40  const lines: string[] = []
41  let line = ''
42  for (const word of t.split(' ')) {
43    if (!line) line = word
44    else if (line.length + 1 + word.length <= max) line += ` ${word}`
45    else {
46      lines.push(line)
47      line = word
48    }
49  }
50  if (line) lines.push(line)
51  if (lines.length > maxLines) {
52    const kept = lines.slice(0, maxLines)
53    kept[maxLines - 1] = fit(`${kept[maxLines - 1]} ${lines[maxLines]}`, px - size, size).replace(/…?$/, '…')
54    return kept.map(l => fit(l, px, size))
55  }
56  return lines.map(l => fit(l, px, size))
57}
58
59const text = (x: number, y: number, size: number, fill: string, body: string, extra = '') =>
60  `<text x="${x}" y="${y}" font-size="${size}" fill="${fill}" ${extra}>${body}</text>`
61
62function taskDot(t: Task, cx: number, cy: number, r: number): string {
63  if (t.status === 'passed') {
64    return `<circle cx="${cx}" cy="${cy}" r="${r}" fill="${C.accent}"/><path d="M${cx - r * 0.45} ${cy}l${r * 0.32} ${r * 0.34} ${r * 0.6}-${r * 0.68}" fill="none" stroke="${C.card}" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"/>`
65  }
66  if (t.status === 'failed') {
67    const d = r * 0.42
68    return `<circle cx="${cx}" cy="${cy}" r="${r}" fill="${C.hot}"/><path d="M${cx - d} ${cy - d}l${2 * d} ${2 * d}M${cx + d} ${cy - d}l-${2 * d} ${2 * d}" stroke="${C.card}" stroke-width="1.5" stroke-linecap="round"/>`
69  }
70  if (t.status === 'active') return `<circle cx="${cx}" cy="${cy}" r="${r}" fill="${C.deep}" stroke="${C.glow}" stroke-width="1.2"/>`
71  if (t.status === 'skipped') return `<circle cx="${cx}" cy="${cy}" r="${r - 1}" fill="${C.dim}"/>`
72  return `<circle cx="${cx}" cy="${cy}" r="${r - 0.6}" fill="none" stroke="${C.dim}" stroke-width="1.2"/>`
73}
74
75// A person with a check: the test user.
76const avatar = (cx: number, cy: number, color: string) =>
77  `<circle cx="${cx}" cy="${cy}" r="14" fill="${C.text}"/>` +
78  `<circle cx="${cx}" cy="${cy - 4}" r="3.6" fill="${C.card}"/>` +
79  `<path d="M${cx - 7} ${cy + 8}a7 6 0 0 1 14 0z" fill="${C.card}"/>` +
80  `<circle cx="${cx + 10}" cy="${cy + 9}" r="5" fill="${color}" stroke="${C.card}" stroke-width="1.5"/>`
81
82const GAP = 10
83const HEADER_H = 160
84const TASK_H = 24
85
86/** The whole pane as one SVG: a summary card, the tasks, then a card per finding. */
87export function runSvg(runs: Run[], now: number, width: number): { source: string; height: number; alt: string } {
88  const W = Math.max(300, Math.round(width))
89  const run = runs[0]!
90  const s = stats(run, now)
91  const parts: string[] = []
92  const tone = STATUS_COLOR[run.status]
93
94  // ---- the summary card
95  parts.push(`<rect x="0.5" y="0.5" width="${W - 1}" height="${HEADER_H - 1}" rx="16" fill="${C.card}" stroke="${C.stroke}"/>`)
96  parts.push(avatar(32, 32, tone))
97  const chip = STATUS_LABEL[run.status]
98  const chipW = Math.round(chip.length * 6.4 + 24)
99  parts.push(`<rect x="${W - 18 - chipW}" y="19" width="${chipW}" height="26" rx="13" fill="${C.raised}" stroke="${run.status === 'running' ? C.stroke : tone}"/>`)
100  parts.push(text(W - 18 - chipW / 2, 36, 11.5, run.status === 'running' ? C.soft : tone, esc(chip), 'text-anchor="middle"'))
101  parts.push(text(56, 30, 14.5, C.text, esc(fit(run.target ?? run.brief, W - 56 - chipW - 30, 14.5)), 'font-weight="600"'))
102  parts.push(text(56, 46, 11, C.muted, esc(fit(`${run.scope === 'app' ? 'Whole app' : run.scope === 'feature' ? 'Feature' : 'Scope pending'} · by ${run.tester}`, W - 56 - chipW - 30, 11))))
103
104  parts.push(
105    `<text x="18" y="96" font-weight="500" letter-spacing="-0.5"><tspan font-size="34" fill="${C.text}">${s.passed}</tspan><tspan font-size="18" fill="${C.muted}">/${s.total || '–'}</tspan></text>`,
106  )
107  const lead = run.findings.length ? plural(run.findings.length, 'finding') : run.status === 'running' ? 'nothing found yet' : 'nothing found'
108  const leadColor = s.counts.blocker ? C.hot : s.counts.major ? C.warn : C.accent
109  const rest = ` · ${s.failed ? `${s.failed} failed · ` : ''}${duration(s.elapsedMs)}${run.status === 'running' ? ' in' : ''}`
110  parts.push(`<text x="18" y="116" font-size="12"><tspan fill="${leadColor}" font-weight="600">${esc(lead)}</tspan><tspan fill="${C.muted}">${esc(rest)}</tspan></text>`)
111
112  // Severity counters, right of the big number, where there is room for them.
113  let sx = W - 18
114  for (const sev of W >= 480 ? [...SEVERITIES].reverse() : []) {
115    const n = s.counts[sev]
116    const label = `${n} ${sev}`
117    const w = Math.round(label.length * 6 + 22)
118    sx -= w
119    parts.push(`<rect x="${sx}" y="76" width="${w}" height="22" rx="11" fill="${n ? C.raised : 'none'}" stroke="${n ? SEVERITY_COLOR[sev] : C.stroke}"/>`)
120    parts.push(`<circle cx="${sx + 11}" cy="87" r="3" fill="${n ? SEVERITY_COLOR[sev] : C.dim}"/>`)
121    parts.push(text(sx + 18, 91, 10.5, n ? C.soft : C.dim, esc(label)))
122    sx -= 6
123  }
124
125  // One cell per task, as in plan-progress: filled passed, red failed, glowing current, hatched to come.
126  const cellsY = 130
127  if (run.tasks.length) {
128    const avail = W - 36
129    const cell = Math.max(4, Math.min(56, (avail - (run.tasks.length - 1) * 5) / run.tasks.length))
130    const gap = run.tasks.length > 1 ? Math.min(5, (avail - cell * run.tasks.length) / (run.tasks.length - 1)) : 0
131    run.tasks.forEach((t, i) => {
132      const x = 18 + i * (cell + gap)
133      const fill = t.status === 'passed' ? C.accent : t.status === 'failed' ? C.hot : t.status === 'active' ? C.deep : t.status === 'skipped' ? C.dim : 'url(#hatch)'
134      const stroke = t.status === 'active' ? ` stroke="${C.glow}" stroke-width="1.2"` : ''
135      parts.push(`<rect x="${x.toFixed(1)}" y="${cellsY}" width="${cell.toFixed(1)}" height="18" rx="${Math.min(5, cell / 3).toFixed(1)}" fill="${fill}"${stroke}/>`)
136    })
137  } else {
138    parts.push(`<rect x="18" y="${cellsY}" width="${W - 36}" height="18" rx="5" fill="url(#hatch)"/>`)
139  }
140
141  let y = HEADER_H + GAP
142
143  // ---- why it could not test
144  if (run.failure) {
145    const lines = wrap(run.failure, W - 36, 12, 4)
146    const h = 40 + lines.length * 17
147    parts.push(`<rect x="0.5" y="${y + 0.5}" width="${W - 1}" height="${h - 1}" rx="14" fill="url(#hotwash)" stroke="${C.hot}"/>`)
148    parts.push(text(18, y + 25, 13.5, C.text, run.status === 'blocked' ? 'Could not test' : 'The run failed', 'font-weight="600"'))
149    lines.forEach((l, i) => parts.push(text(18, y + 46 + i * 17, 12, C.soft, esc(l))))
150    y += h + GAP
151  }
152
153  // ---- the summary, once given
154  if (run.summary) {
155    const lines = wrap(run.summary, W - 36, 12, 5)
156    const h = 40 + lines.length * 17
157    parts.push(`<rect x="0.5" y="${y + 0.5}" width="${W - 1}" height="${h - 1}" rx="14" fill="${C.raised}" stroke="${C.stroke}"/>`)
158    parts.push(text(18, y + 25, 13.5, C.text, 'Summary', 'font-weight="600"'))
159    lines.forEach((l, i) => parts.push(text(18, y + 46 + i * 17, 12, C.soft, esc(l))))
160    y += h + GAP
161  }
162
163  // ---- the task list
164  if (run.tasks.length) {
165    const rows = run.tasks.map(t => (t.note && (t.status === 'failed' || t.status === 'active') ? 2 : 1))
166    const h = 42 + rows.reduce((a, b) => a + b, 0) * TASK_H - 4 - rows.filter(r => r === 2).length * 4
167    const isNow = run.status === 'running'
168    parts.push(
169      `<rect x="0.5" y="${y + 0.5}" width="${W - 1}" height="${h - 1}" rx="14" fill="${isNow ? 'url(#now)' : C.raised}" stroke="${isNow ? C.deep : C.stroke}"/>`,
170    )
171    parts.push(text(16, y + 25, 13.5, C.text, 'Tasks', 'font-weight="600"'))
172    parts.push(text(W - 18, y + 25, 11, C.muted, esc(`${s.closed} of ${s.total} done`), 'text-anchor="end"'))
173    let ty = y + 50
174    run.tasks.forEach((t, i) => {
175      parts.push(taskDot(t, 24, ty - 4, 6))
176      const color = t.status === 'active' ? C.text : t.status === 'passed' ? C.soft : t.status === 'failed' ? C.text : C.muted
177      const weight = t.status === 'active' ? ' font-weight="600"' : ''
178      const deco = t.status === 'skipped' ? ' text-decoration="line-through"' : ''
179      const took = t.startedAt != null && t.status !== 'skipped' ? duration((t.endedAt ?? now) - t.startedAt) : ''
180      parts.push(text(38, ty, 12, color, esc(fit(t.title, W - 100, 12)), `${weight}${deco}`))
181      if (took) parts.push(text(W - 18, ty, 11, t.status === 'active' ? C.accent : C.muted, took, 'text-anchor="end"'))
182      if (rows[i] === 2) {
183        ty += TASK_H - 6
184        parts.push(text(38, ty, 11, t.status === 'failed' ? C.hot : C.muted, esc(fit(t.note!, W - 60, 11))))
185        ty += TASK_H + 2
186      } else ty += TASK_H
187    })
188    y += h + GAP
189  }
190
191  // ---- the findings, worst first
192  for (const f of bySeverity(run.findings)) {
193    const color = SEVERITY_COLOR[f.severity]
194    const detail = f.detail ? wrap(f.detail, W - 44, 11.5, 3) : []
195    const h = 50 + (f.where ? 0 : -2) + detail.length * 16 + (detail.length ? 4 : 0)
196    parts.push(`<rect x="0.5" y="${y + 0.5}" width="${W - 1}" height="${h - 1}" rx="14" fill="${C.raised}" stroke="${C.stroke}"/>`)
197    parts.push(`<rect x="0.5" y="${y + 12}" width="3.5" height="${h - 24}" rx="1.75" fill="${color}"/>`)
198    const tag = f.severity.toUpperCase()
199    const tagW = Math.round(tag.length * 6.6 + 16)
200    parts.push(`<rect x="${W - 16 - tagW}" y="${y + 12}" width="${tagW}" height="18" rx="9" fill="none" stroke="${color}"/>`)
201    parts.push(text(W - 16 - tagW / 2, y + 24.5, 9.5, color, tag, 'text-anchor="middle" font-weight="600" letter-spacing="0.6"'))
202    parts.push(text(18, y + 25, 13, C.text, esc(fit(f.title, W - 40 - tagW, 13)), 'font-weight="600"'))
203    const meta = [f.task != null && run.tasks[f.task] ? `task ${f.task + 1}` : '', f.where ?? ''].filter(Boolean).join(' · ')
204    parts.push(text(18, y + 42, 11, C.muted, esc(fit(meta || 'general', W - 40, 11))))
205    detail.forEach((l, i) => parts.push(text(18, y + 62 + i * 16, 11.5, C.soft, esc(l))))
206    y += h + GAP
207  }
208
209  // ---- earlier runs, folded
210  for (const old of runs.slice(1)) {
211    parts.push(`<rect x="0.5" y="${y + 0.5}" width="${W - 1}" height="32" rx="12" fill="none" stroke="${C.stroke}" stroke-dasharray="4 4"/>`)
212    parts.push(`<circle cx="18" cy="${y + 16.5}" r="4" fill="${STATUS_COLOR[old.status]}"/>`)
213    parts.push(text(30, y + 21, 11.5, C.muted, esc(fit(`Earlier: ${historyLine(old, now)}`, W - 50, 11.5))))
214    y += 32 + GAP
215  }
216
217  const height = Math.round(y - GAP)
218  const defs =
219    `<defs>` +
220    `<pattern id="hatch" width="5" height="5" patternUnits="userSpaceOnUse" patternTransform="rotate(45)"><rect width="5" height="5" fill="${C.track}"/><rect width="2" height="5" fill="#4a4463"/></pattern>` +
221    `<linearGradient id="now" x1="0" y1="0" x2="1" y2="1"><stop offset="0" stop-color="#3d3260"/><stop offset="1" stop-color="${C.raised}"/></linearGradient>` +
222    `<linearGradient id="hotwash" x1="0" y1="0" x2="1" y2="1"><stop offset="0" stop-color="#4a2433"/><stop offset="1" stop-color="${C.raised}"/></linearGradient>` +
223    `</defs>`
224  const source = `<svg xmlns="http://www.w3.org/2000/svg" width="${W}" height="${height}" viewBox="0 0 ${W} ${height}" ${FONT}>${defs}${parts.join('')}</svg>`
225
226  const alt =
227    `${run.target ?? run.brief}: ${STATUS_LABEL[run.status]}, ${s.passed} of ${s.total} tasks passed, ${s.failed} failed, ` +
228    `${plural(run.findings.length, 'finding')}${run.failure ? `. ${run.failure}` : ''}`
229  return { source, height, alt }
230}
231
232export function emptySvg(width: number): { source: string; height: number } {
233  const W = Math.max(300, Math.round(width))
234  const H = 92
235  const source =
236    `<svg xmlns="http://www.w3.org/2000/svg" width="${W}" height="${H}" viewBox="0 0 ${W} ${H}" ${FONT}>` +
237    `<rect x="0.5" y="0.5" width="${W - 1}" height="${H - 1}" rx="16" fill="${C.card}" stroke="${C.stroke}" stroke-dasharray="4 4"/>` +
238    text(20, 38, 14, C.text, 'No test run yet', 'font-weight="600"') +
239    text(20, 60, 11.5, C.muted, esc(fit('Ask Claude to test the app, or run /test-user run [what to test].', W - 40, 11.5))) +
240    `</svg>`
241  return { source, height: H }
242}
243
hooks/prompt.ts 48 lines
1// What the tester is told: the agent type's system prompt and its listing line.
2
3export const AGENT_DESCRIPTION =
4  'A Haiku-powered test user: drives a locally running app in the browser the way a real person would and reports UI/UX problems, ' +
5  'with its task list and findings shown live in the test-user pane. Use it when the user asks to test, QA, try out or click through the app. ' +
6  'In the prompt, say what to test, the local URL if known, and what changed. When the user just says "test the app", ' +
7  'make the focus the feature most recently developed in this session (what it does, where it lives, the files touched); ' +
8  'if nothing was built in this session, ask it to test the entire application. It never edits code. ' +
9  'Its full record (every task, finding and failure, as the developer sees it in the pane) comes back with its result, or with the next prompt when it ran in the background; mcp__test-user__test_report reads it any time.'
10
11export const TESTER_PROMPT = `You are a test user. You try a locally running application the way a real person would, and you report what works, what breaks, and what is confusing. You are testing UI and UX, not reading code for its own sake. You never edit, write or delete project files.
12
13# Your tools
14- Browser tools: mcp__Claude_Browser__* (preview_start, navigate, read_page, find, computer, form_input, get_page_text, read_console_messages, read_network_requests, resize_window) — or mcp__claude-in-chrome__* if those are the ones you have. Prefer read_page / get_page_text / find to read the page; take a screenshot when you judge layout, spacing, contrast or anything visual.
15- Read, Grep, Glob and read-only Bash (git status, git diff, git log, cat) to work out what to test and how to reach the app.
16- Reporting tools, which the developer watches live in a pane. Use them as you go, not only at the end:
17  - mcp__test-user__test_plan — your task list (call it once you know what to test; call again to change it)
18  - mcp__test-user__test_task — mark a task active, passed, failed or skipped, with a short note
19  - mcp__test-user__test_finding — one problem you saw (one call per problem)
20  - mcp__test-user__test_finish — your verdict and summary, last thing before your final answer
21
22# How to run a test
231. Decide the scope.
24   - If the brief names a feature, flow or page, test that (scope "feature"), plus a 30-second smoke check that the app's home screen still loads.
25   - If the brief only says to test "the app" / "the application", or gives no focus: run git status, git diff --stat and git log -5 --stat. If they show a recent feature (changed UI files, a recent commit message), test that feature (scope "feature"). If there is no git history or nothing recent, test the whole application's main flows (scope "app").
262. Find the app. Use the URL in the brief. Otherwise look for it in .claude/launch.json, package.json scripts, vite/next/webpack config or the README, and try the likely localhost port. If it is not running and .claude/launch.json has a configuration, start it with preview_start. If you still cannot reach it, call test_finish with verdict "blocked" and the reason, then stop.
273. Call test_plan with 3–10 concrete tasks phrased as things a user does ("Sign up with a new account", "Add an item to the cart and change its quantity", "Open settings on a narrow window"). Set target to a short name of what you are testing, e.g. "Checkout flow · localhost:5173".
284. Work through the tasks in order. For each: mark it active, do it, look carefully, report each problem with test_finding, then mark it passed or failed (failed = the user could not complete it or it behaved wrongly) with a one-line note.
295. Call test_finish, then give your final answer: a short report — what you tested, the verdict, and the findings worst first, each with where it happened and how to reproduce it.
30
31# What to look for
32- Does it work: actions complete, data saves and shows up, navigation goes where it says, no dead buttons, no errors in the console (read_console_messages) or failing requests.
33- Feedback: loading states, success and error messages, disabled states, what happens on double-click or a slow response.
34- Forms: validation messages, required fields, bad input (empty, too long, wrong format), keyboard (Tab order, Enter to submit, Esc to close).
35- Clarity: labels and copy a newcomer understands, obvious next step, consistent naming.
36- Layout: overlap, cut-off text, misalignment, scroll traps, a narrow viewport (resize_window mobile) when the page is meant to work there.
37- Empty, first-run and edge states: no data, one item, many items.
38- Accessibility basics: buttons and inputs have names in read_page, focus is visible, contrast is readable.
39
40Severities: blocker = a user cannot complete the task; major = it works but badly or loses data / misleads; minor = noticeable friction or a visual defect; polish = small copy or alignment nits.
41
42# Ground rules
43- Stay on the local app (localhost, 127.0.0.1, *.localhost, *.test). Do not visit outside sites beyond what the app itself loads.
44- Use test data only: values you invent, or seed/fixture/example-config values from the project. Never enter real credentials, payment or personal data. If a flow needs a real account you do not have, mark that task skipped and say why.
45- Do not delete data you did not create. Do not change system or browser settings.
46- Be efficient: one look per screen is usually enough; do not re-screenshot what read_page already told you.
47- Report what you saw, not guesses. If you are unsure whether something is a bug, say so in the detail.`
48
hooks/run.ts 182 lines
1import type { Finding, Run, RunStatus, Severity, Task, TaskStatus } from '../types'
2
3export const MAX_TASKS = 15
4export const MAX_FINDINGS = 40
5export const MAX_RUNS = 5
6const MAX_TITLE = 90
7const MAX_DETAIL = 400
8
9export const SEVERITIES: Severity[] = ['blocker', 'major', 'minor', 'polish']
10
11export function clean(text: unknown, max = MAX_TITLE): string {
12  const out = String(text ?? '')
13    .replace(/\s+/g, ' ')
14    .trim()
15  return out.length > max ? `${out.slice(0, max - 1)}…` : out
16}
17
18export const cleanDetail = (text: unknown) => clean(text, MAX_DETAIL)
19
20export function newRun(agentId: string, tester: string, brief: string, now: number): Run {
21  return { agentId, tester, brief: clean(brief, 160) || 'Test the application', status: 'running', startedAt: now, tasks: [], findings: [] }
22}
23
24const isClosed = (t: Task) => t.status === 'passed' || t.status === 'failed' || t.status === 'skipped'
25
26/** A new task list; tasks already closed under the same title keep their outcome. */
27export function plan(run: Run, titles: string[], now: number): Run {
28  const before = new Map(run.tasks.map(t => [t.title.toLowerCase(), t]))
29  const tasks = titles
30    .map(t => clean(t))
31    .filter(Boolean)
32    .slice(0, MAX_TASKS)
33    .map(title => before.get(title.toLowerCase()) ?? { title, status: 'pending' as const })
34  return advance({ ...run, tasks }, now)
35}
36
37/** With nothing in progress, the first pending task becomes the current one. */
38export function advance(run: Run, now: number): Run {
39  if (run.status !== 'running' || run.tasks.some(t => t.status === 'active')) return run
40  const i = run.tasks.findIndex(t => t.status === 'pending')
41  return i < 0 ? run : setTask(run, i, 'active', undefined, now)
42}
43
44export function setTask(run: Run, index: number, status: TaskStatus, note: string | undefined, now: number): Run {
45  const tasks = run.tasks.map((t, i): Task => {
46    if (i === index) {
47      // Marked by hand, it is no longer a task the run never reached.
48      if (t.isAutoSkipped) t = { title: t.title, status: t.status, note: t.note, startedAt: t.startedAt, endedAt: t.endedAt }
49      if (status === 'pending') return { title: t.title, status }
50      if (status === 'active') return { ...t, status, note: note ?? t.note, startedAt: t.startedAt ?? now, endedAt: undefined }
51      return { ...t, status, note: note ?? t.note, startedAt: t.startedAt ?? now, endedAt: now }
52    }
53    // Starting one task puts any other in progress back in the queue.
54    if (status === 'active' && t.status === 'active') return { ...t, status: 'pending' }
55    return t
56  })
57  return { ...run, tasks }
58}
59
60export function addFinding(run: Run, f: Omit<Finding, 'at'>, now: number): Run {
61  if (run.findings.length >= MAX_FINDINGS) return run
62  return { ...run, findings: [...run.findings, { ...f, at: now }] }
63}
64
65const RANK: Record<Severity, number> = { blocker: 0, major: 1, minor: 2, polish: 3 }
66export const bySeverity = (fs: Finding[]) => [...fs].sort((a, b) => RANK[a.severity] - RANK[b.severity] || a.at - b.at)
67
68/** The outcome the run's own record supports, for a tester that did not say. */
69export function verdictOf(run: Run): RunStatus {
70  if (run.tasks.some(t => t.status === 'failed') || run.findings.length > 0) return 'issues'
71  if (run.tasks.length === 0) return 'blocked'
72  return 'passed'
73}
74
75export function finish(run: Run, status: RunStatus, now: number, endedBy: 'tester' | 'turn', summary?: string, failure?: string): Run {
76  // Tasks left open when the run ends were never tried.
77  const tasks = run.tasks.map((t): Task =>
78    isClosed(t) ? t : { ...t, status: 'skipped', isAutoSkipped: true, endedAt: t.startedAt != null ? now : undefined },
79  )
80  return { ...run, tasks, status, endedAt: now, endedBy, isDelivered: false, summary: summary ?? run.summary, failure: failure ?? run.failure }
81}
82
83/** A run that ended takes updates again: its auto-skipped tasks go back in the queue. */
84export function reopen(run: Run): Run {
85  if (run.status === 'running') return run
86  const tasks = run.tasks.map((t): Task => (t.isAutoSkipped ? { title: t.title, status: 'pending', note: t.note } : t))
87  return { ...run, tasks, status: 'running', endedAt: undefined, endedBy: undefined, failure: undefined, isDelivered: undefined }
88}
89
90export type Stats = {
91  total: number
92  closed: number
93  passed: number
94  failed: number
95  counts: Record<Severity, number>
96  elapsedMs: number
97  current: number | null
98}
99
100export function stats(run: Run, now: number): Stats {
101  const counts = { blocker: 0, major: 0, minor: 0, polish: 0 }
102  for (const f of run.findings) counts[f.severity]++
103  const current = run.tasks.findIndex(t => t.status === 'active')
104  return {
105    total: run.tasks.length,
106    closed: run.tasks.filter(isClosed).length,
107    passed: run.tasks.filter(t => t.status === 'passed').length,
108    failed: run.tasks.filter(t => t.status === 'failed').length,
109    counts,
110    elapsedMs: Math.max(0, (run.endedAt ?? now) - run.startedAt),
111    current: current < 0 ? null : current,
112  }
113}
114
115export function duration(ms: number): string {
116  const min = Math.round(ms / 60_000)
117  if (min < 1) return `${Math.max(0, Math.round(ms / 1000))}s`
118  if (min < 60) return `${min}m`
119  const h = Math.floor(min / 60)
120  return `${h}h${String(min % 60).padStart(2, '0')}m`
121}
122
123export const plural = (n: number, word: string) => `${n} ${word}${n === 1 ? '' : 's'}`
124
125export const STATUS_LABEL: Record<RunStatus, string> = {
126  running: 'Testing',
127  passed: 'Passed',
128  issues: 'Issues found',
129  blocked: 'Blocked',
130  failed: 'Failed',
131}
132
133export function findingsLine(s: Stats): string {
134  const n = SEVERITIES.reduce((sum, k) => sum + s.counts[k], 0)
135  if (n === 0) return 'no findings'
136  const parts = SEVERITIES.filter(k => s.counts[k] > 0).map(k => `${s.counts[k]} ${k}`)
137  return `${plural(n, 'finding')} (${parts.join(', ')})`
138}
139
140/** One line for the status bar. */
141export function statusLine(run: Run, now: number): string {
142  const s = stats(run, now)
143  const n = run.findings.length
144  if (run.status === 'running') {
145    const at = s.current != null ? ` · ${run.tasks[s.current]!.title}` : ''
146    return `test-user ● ${s.closed}/${s.total || '?'} tasks · ${plural(n, 'finding')} · ${duration(s.elapsedMs)}${at}`
147  }
148  if (run.status === 'passed') return `test-user ✓ ${s.passed}/${s.total} passed · ${duration(s.elapsedMs)}`
149  if (run.status === 'issues') return `test-user ! ${s.failed} failed · ${plural(n, 'finding')} · ${duration(s.elapsedMs)}`
150  return `test-user × ${STATUS_LABEL[run.status].toLowerCase()}${run.failure ? `: ${run.failure}` : ''}`
151}
152
153const GLYPH: Record<TaskStatus, string> = { pending: '[ ]', active: '[>]', passed: '[x]', failed: '[!]', skipped: '[-]' }
154
155/** The run as plain text, numbered as the tools take it, for the model. */
156export function report(run: Run, now: number): string {
157  const s = stats(run, now)
158  const lines = [
159    `${STATUS_LABEL[run.status]}: ${run.target ?? run.brief} (${run.scope === 'app' ? 'whole app' : run.scope === 'feature' ? 'feature' : 'scope not set'}, by ${run.tester}, ${duration(s.elapsedMs)})`,
160  ]
161  if (run.failure) lines.push(`Failure: ${run.failure}`)
162  if (run.summary) lines.push(`Summary: ${run.summary}`)
163  lines.push('', `Tasks (${s.passed} passed, ${s.failed} failed, ${s.total} total):`)
164  if (run.tasks.length === 0) lines.push('  (none planned)')
165  run.tasks.forEach((t, i) =>
166    lines.push(`  ${i + 1}. ${GLYPH[t.status]} ${t.title}${t.isAutoSkipped ? ' (not reached)' : ''}${t.note ? ` — ${t.note}` : ''}`),
167  )
168  lines.push('', `Findings: ${findingsLine(s)}`)
169  for (const f of bySeverity(run.findings)) {
170    const task = f.task != null && run.tasks[f.task] ? ` [task ${f.task + 1}]` : ''
171    lines.push(`  - ${f.severity.toUpperCase()}${task}: ${f.title}${f.where ? ` @ ${f.where}` : ''}${f.detail ? `\n    ${f.detail}` : ''}`)
172  }
173  return lines.join('\n')
174}
175
176/** The earlier runs, one line each. */
177export function historyLine(run: Run, now: number): string {
178  const s = stats(run, now)
179  const what = run.status === 'passed' ? `${s.passed}/${s.total} passed` : run.status === 'running' ? 'running' : run.findings.length ? plural(run.findings.length, 'finding') : STATUS_LABEL[run.status].toLowerCase()
180  return `${run.target ?? run.brief} · ${what}`
181}
182
types/index.d.ts 64 lines
1export type TaskStatus = 'pending' | 'active' | 'passed' | 'failed' | 'skipped'
2
3export type Task = {
4  title: string
5  status: TaskStatus
6  note?: string
7  startedAt?: number
8  endedAt?: number
9  // Skipped only because the run ended with it open; a reopened run takes it up again.
10  isAutoSkipped?: true
11}
12
13export type Severity = 'blocker' | 'major' | 'minor' | 'polish'
14
15export type Finding = {
16  severity: Severity
17  title: string
18  detail?: string
19  // The screen or URL it was seen on.
20  where?: string
21  // The task it came up in, from 0.
22  task?: number
23  at: number
24}
25
26// running: still going. passed: every task passed, nothing found. issues: done,
27// with failed tasks or findings. blocked: the tester could not test (app not
28// reachable, no way in). failed: the run itself died (error, interrupt).
29export type RunStatus = 'running' | 'passed' | 'issues' | 'blocked' | 'failed'
30
31export type Run = {
32  // The tester's agent id, or 'main' when Claude records a run itself.
33  agentId: string
34  tester: string
35  brief: string
36  target?: string
37  scope?: 'feature' | 'app'
38  status: RunStatus
39  startedAt: number
40  endedAt?: number
41  tasks: Task[]
42  findings: Finding[]
43  summary?: string
44  // Why the run was blocked or failed.
45  failure?: string
46  // The Agent call that started it, so its result can carry the report.
47  toolUseId?: string
48  // Who ended it: the tester through test_finish, or its turn ending without one.
49  endedBy?: 'tester' | 'turn'
50  // Whether the main conversation has been handed the finished report.
51  isDelivered?: boolean
52}
53
54declare module 'claude-code' {
55  interface PluginState {
56    'test-user': {
57      // Newest first; the pane shows the first.
58      runs: Run[]
59      // Bumped by the ticker so elapsed times redraw while a run goes.
60      now: number
61    }
62  }
63}
64