A Haiku test user that clicks through your locally running app and reports UI/UX issues, with a live pane of its tasks, findings and failures

My Claude Code mods, one folder per mod. Needs Claude Code 2.1.287 or newer.
| Mod | What it does |
|---|---|
| usage-band | A band above the prompt showing context fill, tokens, cost and rate limits: stat tiles in the desktop app, a coloured row in the terminal |
| plan-progress | Reads a phased plan into a live progress tree: phases, steps, elapsed time and an estimate of what is left |
| test-user | A Haiku test user that clicks through your locally running app and reports UI/UX issues: its task list, findings and failures in a live pane |
All three mods share one look: dark violet cards with a lilac accent. In the desktop app's Code tab they draw as images (cards, tiles, gradients); in the terminal they use the same colours as text.

At the prompt of a terminal Claude Code session:
/plugin install usage-band --marketplace WeaponizedLego/claude-code-mods
/plugin install plan-progress --marketplace WeaponizedLego/claude-code-mods
Answer y to add the marketplace, then pick the user scope. Or from a shell:
claude plugin marketplace add WeaponizedLego/claude-code-mods
claude plugin install usage-band@claude-code-mods
claude plugin install plan-progress@claude-code-mods
Update later with claude plugin marketplace update claude-code-mods, then claude plugin update usage-band@claude-code-mods.

When a plan with several phases is approved in plan mode, the mod reads its phases and steps (## Phase 1: ... headings, other work headings, or a nested list) into a tree and opens it in a pane. In the terminal it looks like this:
Move auth to sessions 20m in · ~40m left
▰▰▰▰▰▰▱▱▱▱▱▱▱▱▱▱▱▱ 33% 2/6 steps · phase 2/3
✓ 1 Setup 20m
├ ✓ Add the sessions table 10m
└ ✓ Write the session store 10m
● 2 Switch over
├ ● Replace token checks in the middleware
├ ○ Update the login endpoint
└ ○ Remove the old token helpers
○ 3 Verify
└ ○ Run the auth test suite
The status line carries the short form (▰▰▱▱ 2/6 · phase 2/3: Switch over · 20m in · ~40m left). Steps move when Claude calls the mod's plan_mark tool, or by themselves when Claude's tasks or todos are named like a step. Without plan mode, ask Claude to lay the work out in phases and it calls plan_set. The estimate is the pace so far times the steps left.
/plan-progress opens the tree (it also works mid-turn); /plan-progress clear stops tracking.

A test user for the app you are running locally. Ask Claude to "test the app" (or a feature, a page, a flow) and it hands the job to test-user:tester, a subagent on Haiku that drives the app in the browser the way a person would: it works out what to test, finds the app (the URL you gave, .claude/launch.json, package.json), plans a handful of user tasks, works through them and reports each problem it sees. It never edits code, stays on localhost and only uses test data.
When you just say "test the app", Claude points it at the feature most recently built in the session; with nothing built, the tester looks at git diff and git log for a recent feature, and failing that tests the whole application.
The pane shows the run as it goes: status, tasks passed out of total, a count per severity (blocker, major, minor, polish), the task list with a note on each failure, a card per finding (worst first, with where and how to reproduce), a red card when it could not test at all, and earlier runs folded at the bottom. The status line carries the short form (test-user ● 3/6 tasks · 2 findings · 4m · Pay with the test card) and a toast says how it ended.
The session gets the full record, not just the tester's closing message: it rides back on the Agent result when the tester ran in the foreground, and on the next prompt (or the background task's notification) otherwise; test_report reads it any time. A tester whose turn ends early and is resumed carries on with the same run.
Claude gets the same reporting tools (test_plan, test_task, test_finding, test_finish), so a test it runs by hand shows in the pane too, and it can amend the latest tester run with what it re-checked. If the tester stops without a verdict, the record decides it; a "passed" never hides a failed task.
/test-user opens the pane; /test-user run [what to test] starts a run; /test-user clear forgets the runs.
Put it in its own folder (<mod>/.claude-plugin/plugin.json, <mod>/hooks/...) and add an entry to .claude-plugin/marketplace.json. Check it with claude plugin validate ..
hooks/register.tsx 545 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { Run, RunStatus, Severity, TaskStatus } from '../types'
5import { C, SEVERITY_COLOR, STATUS_COLOR, TASK_COLOR, TASK_GLYPH, emptySvg, runSvg } from './look'
6import { AGENT_DESCRIPTION, TESTER_PROMPT } from './prompt'
7import {
8 MAX_RUNS,
9 SEVERITIES,
10 STATUS_LABEL,
11 addFinding,
12 advance,
13 bySeverity,
14 clean,
15 cleanDetail,
16 duration,
17 finish,
18 findingsLine,
19 historyLine,
20 newRun,
21 plan,
22 plural,
23 reopen,
24 report,
25 setTask,
26 stats,
27 statusLine,
28 verdictOf,
29} from './run'
30
31const PANE = 'test-user'
32const TITLE = 'Test user'
33const AGENT = 'test-user:tester'
34const PLAN = 'mcp__test-user__test_plan'
35const TASK = 'mcp__test-user__test_task'
36const FINDING = 'mcp__test-user__test_finding'
37const FINISH = 'mcp__test-user__test_finish'
38const REPORT = 'mcp__test-user__test_report'
39const TICK_MS = 30_000
40// Roughly one cell of the desktop's code font, in CSS pixels.
41const CELL_PX = 8
42const MAIN = 'main'
43
44const runsAtom = atom({ plugin: 'test-user', key: 'runs' } as const, [])
45const nowAtom = atom({ plugin: 'test-user', key: 'now' } as const, 0)
46
47type PlanInput = { tasks?: unknown[]; target?: string; scope?: 'feature' | 'app' }
48type TaskInput = { task?: number; status?: TaskStatus; note?: string }
49type FindingInput = { severity?: Severity; title?: string; detail?: string; where?: string; task?: number }
50type FinishInput = { verdict?: 'passed' | 'issues' | 'blocked'; summary?: string; reason?: string }
51
52// The loop a call came from: a tester subagent's id, or the main conversation.
53const loopOf = (e: { agentId?: string }) => e.agentId ?? MAIN
54
55async function save($: EngineInterface, fn: (runs: Run[], now: number) => Run[]): Promise<Run[]> {
56 const now = await $.clock.now()
57 const runs = await update($, runsAtom, r => fn(r, now).slice(0, MAX_RUNS))
58 await update($, nowAtom, () => now)
59 $.ui.status(runs[0] ? statusLine(runs[0], now) : undefined)
60 return runs
61}
62
63const same = (a: Run, b: Run) => a.agentId === b.agentId && a.startedAt === b.startedAt
64
65// The run a loop's updates go to: a tester's own latest run, ended or not (a
66// tester whose turn ended early and was resumed carries on with it); for the
67// main conversation, its own run in progress, else the latest run of all, so
68// Claude can amend a tester's run with what it re-checked.
69function targetOf(runs: Run[], loop: string): number {
70 if (loop !== MAIN) return runs.findIndex(r => r.agentId === loop)
71 const own = runs.findIndex(r => r.agentId === MAIN && r.status === 'running')
72 return own >= 0 ? own : runs.length ? 0 : -1
73}
74
75/** Applies `fn` to the run this loop updates; null when there is none. */
76async function changeRun($: EngineInterface, loop: string, fn: (run: Run, now: number) => Run): Promise<Run | null> {
77 let changed: Run | null = null
78 await save($, (runs, now) => {
79 // `update` may run this again on a version miss: start each pass afresh.
80 changed = null
81 const i = targetOf(runs, loop)
82 return runs.map((r, j) => {
83 if (j !== i) return r
84 if (r.agentId === loop) return (changed = advance(fn(reopen(r), now), now))
85 // Claude amending a tester's run: an ended run stays ended, its verdict redone.
86 const amended = fn(r, now)
87 return (changed = r.status === 'running' ? advance(amended, now) : { ...amended, status: verdictOf(amended), isDelivered: false })
88 })
89 })
90 return changed
91}
92
93async function startRun($: EngineInterface, loop: string, tester: string, brief: string, toolUseId?: string): Promise<Run> {
94 const runs = await save($, (runs, now) => {
95 // A loop has one run going at a time; an older one it left open is closed.
96 const rest = runs.map(r => (r.agentId === loop && r.status === 'running' ? finish(r, verdictOf(r), now, 'turn') : r))
97 return [{ ...newRun(loop, tester, brief, now), toolUseId }, ...rest]
98 })
99 await openPane($, false)
100 return runs[0]!
101}
102
103async function markDelivered($: EngineInterface, delivered: Run[]) {
104 await save($, runs => runs.map(r => (delivered.some(d => same(d, r)) ? { ...r, isDelivered: true } : r)))
105}
106
107const deliveredText = (run: Run, now: number) =>
108 `test-user results (the full record from the test-user pane, including updates the tester's own final message may leave out):\n\n${report(run, now)}`
109
110// A pane that cannot open (refused, or no surface to draw on) never costs the run.
111async function openPane($: EngineInterface, isAsked: boolean) {
112 const opened = await $.ui.open({ id: PANE, title: TITLE }).catch(() => ({ isPlaced: false }))
113 if (!opened.isPlaced && !isAsked) $.ui.toast('test-user: run /test-user to watch the test run')
114}
115
116const outcomeToast = (run: Run) => {
117 const n = run.findings.length
118 if (run.status === 'passed') return `test-user: all ${run.tasks.length} tasks passed`
119 if (run.status === 'issues') return `test-user: done — ${plural(n, 'finding')}, ${run.tasks.filter(t => t.status === 'failed').length} failed`
120 return `test-user: ${STATUS_LABEL[run.status].toLowerCase()}${run.failure ? ` — ${run.failure}` : ''}`
121}
122
123const firstLine = (t: string) => t.split(/\r?\n/).find(l => l.trim()) ?? ''
124
125export const register: Register = on => {
126 on('session.start', async ($, e, next) => {
127 const result = await next(e)
128
129 await $.agent.register({
130 name: 'tester',
131 description: AGENT_DESCRIPTION,
132 prompt: TESTER_PROMPT,
133 model: 'haiku',
134 disallowedTools: ['Edit', 'Write', 'NotebookEdit', 'Agent', 'Workflow', 'Artifact', REPORT],
135 maxTurns: 120,
136 })
137
138 await $.tool.register({
139 name: 'test_plan',
140 description:
141 'Start (or replace) the task list of a UI/UX test run, shown live to the developer. Call once you know what you are testing. ' +
142 'Used by the test-user:tester agent; Claude may also use it to record a test it runs by hand.',
143 inputSchema: {
144 type: 'object',
145 properties: {
146 tasks: { type: 'array', minItems: 1, maxItems: 15, items: { type: 'string' }, description: 'Things a user does, in order' },
147 target: { type: 'string', description: 'Short name of what is tested, e.g. "Checkout · localhost:5173"' },
148 scope: { type: 'string', enum: ['feature', 'app'], description: 'A recent feature, or the whole application' },
149 },
150 required: ['tasks'],
151 },
152 isDeferred: false,
153 })
154 await $.tool.register({
155 name: 'test_task',
156 description: 'Mark a task of the current test run active, passed, failed or skipped. Tasks count from 1. The next task becomes active by itself.',
157 inputSchema: {
158 type: 'object',
159 properties: {
160 task: { type: 'integer', minimum: 1 },
161 status: { type: 'string', enum: ['active', 'passed', 'failed', 'skipped'] },
162 note: { type: 'string', description: 'One line: what happened' },
163 },
164 required: ['task', 'status'],
165 },
166 isDeferred: false,
167 })
168 await $.tool.register({
169 name: 'test_finding',
170 description: 'Report one UI/UX problem seen during the current test run. One call per problem.',
171 inputSchema: {
172 type: 'object',
173 properties: {
174 severity: { type: 'string', enum: SEVERITIES, description: 'blocker: cannot complete; major: works badly; minor: friction or visual defect; polish: nit' },
175 title: { type: 'string', description: 'The problem in a few words' },
176 detail: { type: 'string', description: 'What you did, what happened, what you expected' },
177 where: { type: 'string', description: 'The screen, URL or element' },
178 task: { type: 'integer', minimum: 1, description: 'The task it came up in' },
179 },
180 required: ['severity', 'title'],
181 },
182 isDeferred: false,
183 })
184 await $.tool.register({
185 name: 'test_finish',
186 description: 'End the current test run with a verdict. "blocked" when the app could not be tested at all (give the reason).',
187 inputSchema: {
188 type: 'object',
189 properties: {
190 verdict: { type: 'string', enum: ['passed', 'issues', 'blocked'] },
191 summary: { type: 'string', description: 'Two to four sentences for the developer' },
192 reason: { type: 'string', description: 'Why it was blocked' },
193 },
194 required: ['verdict', 'summary'],
195 },
196 isDeferred: false,
197 })
198 await $.tool.register({
199 name: 'test_report',
200 description:
201 'Read the latest test-user run (or the one in progress) as the developer sees it in the test-user pane: its tasks, findings and why it failed. ' +
202 'Use it to answer questions about a test run or to check what the tester recorded.',
203 inputSchema: { type: 'object', properties: {} },
204 isDeferred: false,
205 })
206 await $.command.register({
207 name: 'test-user',
208 description: 'Show the test user pane (run [what]: start a test run; clear: forget the runs)',
209 argumentHint: '[run [what to test] | clear]',
210 immediate: true,
211 })
212
213 $.clock.every(TICK_MS, () => {
214 void (async () => {
215 const runs = await read($, runsAtom)
216 if (runs[0]?.status !== 'running') return
217 const t = await $.clock.now()
218 await update($, nowAtom, () => t)
219 $.ui.status(statusLine(runs[0], t))
220 })()
221 })
222
223 // After a reload the runs are still in the session's state.
224 const runs = await read($, runsAtom)
225 if (runs[0]) $.ui.status(statusLine(runs[0], await $.clock.now()))
226 return result
227 })
228
229 // A tester subagent starting is a run starting.
230 on('agent.spawn', async ($, e, next) => {
231 const started = await next(e)
232 if (e.subagentType !== AGENT || !('agentId' in started) || !started.agentId) return started
233 await startRun($, started.agentId, 'Haiku test user', e.description || firstLine(e.prompt), e.tool_use_id)
234 return started
235 })
236
237 // Its turn ending ends the run, whatever the tester remembered to say; a
238 // tester resumed after that carries on with the same run (targetOf).
239 on('turn.complete', async ($, e, next) => {
240 const result = await next(e)
241 const loop = e.agentId
242 if (!loop) return result
243 const runs = await read($, runsAtom)
244 if (!runs.some(r => r.agentId === loop && r.status === 'running')) return result
245
246 const ended = await save($, (runs, now) =>
247 runs.map(r => {
248 if (r.agentId !== loop || r.status !== 'running') return r
249 if (e.reason === 'aborted') return finish(r, 'failed', now, 'turn', undefined, 'The run was interrupted.')
250 if (e.reason !== 'answer') {
251 return finish(r, 'failed', now, 'turn', undefined, `The tester stopped: ${e.reason === 'refusal' ? 'the model refused' : 'an API error'}.`)
252 }
253 return finish(r, verdictOf(r), now, 'turn', r.summary ?? (clean(e.answer, 300) || undefined))
254 }),
255 )
256 const run = ended.find(r => r.agentId === loop)
257 if (run) $.ui.toast(outcomeToast(run))
258 return result
259 })
260
261 // The session gets the full record: a foreground tester's on its Agent result...
262 on('tool.call', { tool: 'Agent' }, async ($, e, next) => {
263 const ran = await next(e)
264 if ((e as unknown as { subagent_type?: string }).subagent_type !== AGENT || ran.deny !== undefined) return ran
265 const run = (await read($, runsAtom)).find(r => r.toolUseId === e.tool_use_id)
266 if (!run) return ran
267 const now = await $.clock.now()
268 if (run.status === 'running') {
269 const note = `test-user: this run is still going (${statusLine(run, now)}). Its full record reaches you when it ends; ${REPORT} reads it at any time.`
270 return { ...ran, context: [...(ran.context ?? []), note] }
271 }
272 await markDelivered($, [run])
273 return { ...ran, context: [...(ran.context ?? []), deliveredText(run, now)] }
274 }).catch(($, e, next) => next(e))
275
276 // ...and any run that ended unseen (a background tester, a resumed one) rides
277 // on the next prompt, a background task's notification included.
278 on('prompt.submit', async ($, e, next) => {
279 const fresh = (await read($, runsAtom)).filter(r => r.status !== 'running' && r.isDelivered === false)
280 if (fresh.length === 0) return next(e)
281 const now = await $.clock.now()
282 await markDelivered($, fresh)
283 return next({ ...e, context: [...(e.context ?? []), ...fresh.map(r => deliveredText(r, now))] })
284 }).catch(($, e, next) => next(e))
285
286 on('tool.call', { tool: PLAN }, async ($, e) => {
287 const input = e as unknown as PlanInput
288 const tasks = (input.tasks ?? []).map(t => clean(t)).filter(Boolean)
289 if (tasks.length === 0) return { deny: 'test_plan needs at least one task.' }
290 const loop = loopOf(e)
291 const runs = await read($, runsAtom)
292 // A tester carries on with its run if its turn ended it; anything else is a new run.
293 const own = runs.find(r => r.agentId === loop)
294 const carriesOn = own && (own.status === 'running' || (loop !== MAIN && own.endedBy === 'turn'))
295 if (!carriesOn) await startRun($, loop, loop === MAIN ? 'Claude' : 'Haiku test user', input.target ?? 'Test run')
296 const run = (await changeRun($, loop, (r, now) => ({
297 ...plan(r, tasks, now),
298 target: input.target ? clean(input.target, 80) : r.target,
299 scope: input.scope ?? r.scope,
300 })))!
301 const at = run.tasks.findIndex(t => t.status === 'active')
302 const list = run.tasks.map((t, i) => `${i + 1}. ${t.title}${t.status === 'pending' || t.status === 'active' ? '' : ` (${t.status})`}`)
303 return {
304 result: `Task list shown to the developer:\n${list.join('\n')}\n\n${at >= 0 ? `Task ${at + 1} is active. ` : ''}Mark each with test_task; report problems with test_finding.`,
305 }
306 })
307
308 on('tool.call', { tool: TASK }, async ($, e) => {
309 const input = e as unknown as TaskInput
310 const loop = loopOf(e)
311 const status = input.status ?? 'passed'
312 let error = ''
313 const run = await changeRun($, loop, (r, now) => {
314 const i = Number(input.task) - 1
315 if (!r.tasks[i]) {
316 error = `There is no task ${input.task}; the run has ${r.tasks.length}.`
317 return r
318 }
319 return setTask(r, i, status, input.note ? clean(input.note, 160) : undefined, now)
320 })
321 if (!run) return { deny: 'No test run here yet. Call test_plan first.' }
322 if (error) return { deny: error }
323 const s = stats(run, await $.clock.now())
324 const cur = s.current != null ? ` Now on task ${s.current + 1}: ${run.tasks[s.current]!.title}.` : s.closed === s.total ? ' All tasks done: call test_finish.' : ''
325 return { result: `Task ${input.task} ${status}. ${s.closed}/${s.total} done, ${plural(run.findings.length, 'finding')}.${cur}` }
326 })
327
328 on('tool.call', { tool: FINDING }, async ($, e) => {
329 const input = e as unknown as FindingInput
330 const title = clean(input.title)
331 if (!title) return { deny: 'A finding needs a title.' }
332 const severity: Severity = SEVERITIES.includes(input.severity as Severity) ? (input.severity as Severity) : 'minor'
333 const run = await changeRun($, loopOf(e), (r, now) => {
334 const active = r.tasks.findIndex(t => t.status === 'active')
335 const task = input.task != null && r.tasks[Number(input.task) - 1] ? Number(input.task) - 1 : active >= 0 ? active : undefined
336 return addFinding(
337 r,
338 { severity, title, detail: input.detail ? cleanDetail(input.detail) : undefined, where: input.where ? clean(input.where, 80) : undefined, task },
339 now,
340 )
341 })
342 if (!run) return { deny: 'No test run here yet. Call test_plan first.' }
343 return { result: `Recorded (${severity}). ${findingsLine(stats(run, 0))} so far.` }
344 })
345
346 on('tool.call', { tool: FINISH }, async ($, e) => {
347 const input = e as unknown as FinishInput
348 const loop = loopOf(e)
349 const summary = input.summary ? cleanDetail(input.summary) : undefined
350 const end = (r: Run, now: number): Run => {
351 // The record wins over a verdict it contradicts: failed tasks or findings are issues.
352 const verdict: RunStatus = input.verdict === 'blocked' ? 'blocked' : verdictOf(r) === 'issues' || input.verdict === 'issues' ? 'issues' : 'passed'
353 return finish(r, verdict, now, 'tester', summary, verdict === 'blocked' ? clean(input.reason ?? input.summary ?? 'No reason given', 300) : undefined)
354 }
355 // Blocked before it planned anything: still a run the developer should see.
356 if (targetOf(await read($, runsAtom), loop) < 0) {
357 if (input.verdict !== 'blocked') return { deny: 'No test run here yet. Call test_plan first.' }
358 await startRun($, loop, loop === MAIN ? 'Claude' : 'Haiku test user', 'Test run')
359 }
360 const done = (await changeRun($, loop, end))!
361 $.ui.toast(outcomeToast(done))
362 return { result: `Run ended: ${STATUS_LABEL[done.status]}. Now give your final report.` }
363 })
364
365 on('tool.call', { tool: REPORT }, async ($, e) => {
366 const runs = await read($, runsAtom)
367 if (!runs[0]) return { result: 'No test run yet. Spawn the test-user:tester agent to run one.' }
368 const now = await $.clock.now()
369 if (runs[0].status !== 'running') await markDelivered($, [runs[0]])
370 const earlier = runs.slice(1).map(r => `- ${historyLine(r, now)}`)
371 return { result: report(runs[0], now) + (earlier.length ? `\n\nEarlier runs:\n${earlier.join('\n')}` : '') }
372 })
373
374 on('command.run', { command: 'test-user' }, async ($, e) => {
375 const args = e.args.trim()
376 if (args === 'clear') {
377 await save($, () => [])
378 await $.ui.close({ id: PANE }).catch(() => {})
379 return { text: 'Cleared the test runs.' }
380 }
381 await openPane($, true)
382 const run = /^run\b\s*(.*)$/is.exec(args)
383 if (run) {
384 const what = run[1]!.trim()
385 void $.prompt.submit({
386 text: what
387 ? `Use the ${AGENT} agent to UI/UX test: ${what}`
388 : `Use the ${AGENT} agent to UI/UX test the application. Focus on the feature most recently developed in this session; if nothing was built in this session, test the entire application.`,
389 })
390 return { text: `Starting a test run${what ? `: ${what}` : ''}.` }
391 }
392 const runs = await read($, runsAtom)
393 if (!runs[0]) return { text: 'No test run yet: ask Claude to test the app, or run /test-user run [what to test].' }
394 return { text: statusLine(runs[0], await $.clock.now()) }
395 })
396
397 on('session.end', async ($, e, next) => {
398 if (e.reason === 'clear') await save($, () => [])
399 return next(e)
400 })
401
402 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
403 const runs = await read($, runsAtom)
404 await read($, nowAtom)
405 const now = await $.clock.now()
406
407 // Where the surface draws images, the pane is a stack of cards.
408 if (e.surface !== 'terminal') {
409 const { Box, Svg } = $.ui.resolve(e)
410 const width = e.props.bodyColumns * CELL_PX
411 if (!runs[0]) {
412 const empty = emptySvg(width)
413 return <Svg source={empty.source} alt="No test run yet" width={Math.max(300, width)} height={empty.height} />
414 }
415 const drawn = runSvg(runs, now, width)
416 return (
417 <Box flexDirection="column" key="test-user">
418 <Svg source={drawn.source} alt={drawn.alt} width={Math.max(300, width)} height={drawn.height} />
419 </Box>
420 )
421 }
422
423 const { Box, Text } = $.ui.resolve(e)
424 const run = runs[0]
425 if (!run) {
426 return (
427 <Box flexDirection="column" borderStyle="round" borderColor={C.stroke} paddingX={1}>
428 <Text bold color={C.text}>
429 No test run yet
430 </Text>
431 <Text color={C.muted} wrap="wrap">
432 Ask Claude to test the app, or run /test-user run [what to test].
433 </Text>
434 </Box>
435 )
436 }
437
438 const s = stats(run, now)
439 const tone = STATUS_COLOR[run.status]
440 return (
441 <Box flexDirection="column" key="test-user">
442 <Box flexDirection="row" justifyContent="space-between">
443 <Text bold color={C.text} wrap="truncate">
444 {run.target ?? run.brief}
445 </Text>
446 <Text color={tone} bold>{` ${STATUS_LABEL[run.status]}`}</Text>
447 </Box>
448 <Box flexDirection="row">
449 <Text bold color={C.text}>{`${s.passed}/${s.total || '–'}`}</Text>
450 <Text color={C.muted}>{` passed · ${s.failed} failed · `}</Text>
451 <Text color={s.counts.blocker ? C.hot : s.counts.major ? C.warn : C.accent}>{plural(run.findings.length, 'finding')}</Text>
452 <Text color={C.muted}>{` · ${duration(s.elapsedMs)} · ${run.tester}`}</Text>
453 </Box>
454
455 {run.failure && (
456 <Box flexDirection="column" borderStyle="round" borderColor={C.hot} paddingX={1} key="failure">
457 <Text bold color={C.hot}>
458 {run.status === 'blocked' ? 'Could not test' : 'The run failed'}
459 </Text>
460 <Text color={C.soft} wrap="wrap">
461 {run.failure}
462 </Text>
463 </Box>
464 )}
465
466 <Text> </Text>
467 <Box flexDirection="column" borderStyle="round" borderColor={run.status === 'running' ? C.deep : C.stroke} paddingX={1} key="tasks">
468 <Box flexDirection="row" justifyContent="space-between">
469 <Text bold color={C.text}>
470 Tasks
471 </Text>
472 <Text color={C.muted}>{`${s.closed}/${s.total}`}</Text>
473 </Box>
474 {run.tasks.length === 0 && <Text color={C.muted}>Working out what to test…</Text>}
475 {run.tasks.flatMap((t, i) => {
476 const took = t.startedAt != null && t.status !== 'skipped' ? duration((t.endedAt ?? now) - t.startedAt) : ''
477 const row = (
478 <Box flexDirection="row" justifyContent="space-between" key={`task-${i}`}>
479 <Box flexDirection="row" flexShrink={1}>
480 <Text color={TASK_COLOR[t.status]}>{`${TASK_GLYPH[t.status]} `}</Text>
481 <Text
482 color={t.status === 'active' || t.status === 'failed' ? C.text : t.status === 'passed' ? C.soft : C.muted}
483 bold={t.status === 'active'}
484 strikethrough={t.status === 'skipped'}
485 wrap="truncate"
486 >
487 {t.title}
488 </Text>
489 </Box>
490 <Text color={t.status === 'active' ? C.accent : C.muted}>{took ? ` ${took}` : ''}</Text>
491 </Box>
492 )
493 if (!t.note || (t.status !== 'failed' && t.status !== 'active')) return [row]
494 return [
495 row,
496 <Text color={t.status === 'failed' ? C.hot : C.muted} wrap="truncate" key={`note-${i}`}>
497 {` ${t.note}`}
498 </Text>,
499 ]
500 })}
501 </Box>
502
503 {bySeverity(run.findings).map((f, i) => (
504 <Box flexDirection="column" borderStyle="round" borderColor={SEVERITY_COLOR[f.severity]} paddingX={1} key={`finding-${i}`}>
505 <Box flexDirection="row" justifyContent="space-between">
506 <Text bold color={C.text} wrap="truncate">
507 {f.title}
508 </Text>
509 <Text color={SEVERITY_COLOR[f.severity]} bold>{` ${f.severity.toUpperCase()}`}</Text>
510 </Box>
511 {(f.where || f.task != null) && (
512 <Text color={C.muted} wrap="truncate">
513 {[f.task != null && run.tasks[f.task] ? `task ${f.task + 1}` : '', f.where ?? ''].filter(Boolean).join(' · ')}
514 </Text>
515 )}
516 {f.detail && (
517 <Text color={C.soft} wrap="wrap">
518 {f.detail}
519 </Text>
520 )}
521 </Box>
522 ))}
523
524 {run.summary && (
525 <Box flexDirection="column" paddingX={1} key="summary">
526 <Text bold color={C.text}>
527 Summary
528 </Text>
529 <Text color={C.soft} wrap="wrap">
530 {run.summary}
531 </Text>
532 </Box>
533 )}
534
535 {runs.slice(1).map((old, i) => (
536 <Box flexDirection="row" key={`old-${i}`}>
537 <Text color={STATUS_COLOR[old.status]}>{'● '}</Text>
538 <Text color={C.muted} wrap="truncate">{`Earlier: ${historyLine(old, now)}`}</Text>
539 </Box>
540 ))}
541 </Box>
542 )
543 })
544}
545hooks/look.ts 243 lines1import type { Run, RunStatus, Severity, Task, TaskStatus } from '../types'
2import { STATUS_LABEL, SEVERITIES, bySeverity, duration, historyLine, plural, stats } from './run'
3
4// The palette the other mods share: dark violet cards with a lilac accent,
5// orange and red for what needs a look. An SVG is drawn as an image, so it
6// takes real colours, not theme keys; the terminal uses the same.
7export const C = {
8 card: '#23222b',
9 raised: '#2a2933',
10 stroke: '#363541',
11 text: '#ececf1',
12 soft: '#c9c7d3',
13 muted: '#8e8c9a',
14 dim: '#5f5d6b',
15 accent: '#a98bff',
16 deep: '#7d68c9',
17 glow: '#c4b2ff',
18 track: '#3b3650',
19 warn: '#ff8a4c',
20 hot: '#ff5c7a',
21} as const
22
23export const SEVERITY_COLOR: Record<Severity, string> = { blocker: C.hot, major: C.warn, minor: C.accent, polish: C.muted }
24export const STATUS_COLOR: Record<RunStatus, string> = { running: C.glow, passed: C.accent, issues: C.warn, blocked: C.hot, failed: C.hot }
25export const TASK_COLOR: Record<TaskStatus, string> = { pending: C.dim, active: C.glow, passed: C.accent, failed: C.hot, skipped: C.dim }
26export const TASK_GLYPH: Record<TaskStatus, string> = { pending: '○', active: '●', passed: '✓', failed: '✗', skipped: '–' }
27
28const FONT = `font-family="Inter, 'Segoe UI', system-ui, -apple-system, sans-serif"`
29
30export const esc = (t: string) => t.replace(/&/g, '&').replace(/</g, '<').replace(/>/g, '>').replace(/"/g, '"')
31
32// Text in an image cannot wrap or truncate itself: cut it to an estimated width.
33const chars = (px: number, size: number) => Math.max(3, Math.floor(px / (size * 0.52)))
34const fit = (t: string, px: number, size: number) => {
35 const max = chars(px, size)
36 return t.length > max ? `${t.slice(0, max - 1)}…` : t
37}
38function wrap(t: string, px: number, size: number, maxLines: number): string[] {
39 const max = chars(px, size)
40 const lines: string[] = []
41 let line = ''
42 for (const word of t.split(' ')) {
43 if (!line) line = word
44 else if (line.length + 1 + word.length <= max) line += ` ${word}`
45 else {
46 lines.push(line)
47 line = word
48 }
49 }
50 if (line) lines.push(line)
51 if (lines.length > maxLines) {
52 const kept = lines.slice(0, maxLines)
53 kept[maxLines - 1] = fit(`${kept[maxLines - 1]} ${lines[maxLines]}`, px - size, size).replace(/…?$/, '…')
54 return kept.map(l => fit(l, px, size))
55 }
56 return lines.map(l => fit(l, px, size))
57}
58
59const text = (x: number, y: number, size: number, fill: string, body: string, extra = '') =>
60 `<text x="${x}" y="${y}" font-size="${size}" fill="${fill}" ${extra}>${body}</text>`
61
62function taskDot(t: Task, cx: number, cy: number, r: number): string {
63 if (t.status === 'passed') {
64 return `<circle cx="${cx}" cy="${cy}" r="${r}" fill="${C.accent}"/><path d="M${cx - r * 0.45} ${cy}l${r * 0.32} ${r * 0.34} ${r * 0.6}-${r * 0.68}" fill="none" stroke="${C.card}" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"/>`
65 }
66 if (t.status === 'failed') {
67 const d = r * 0.42
68 return `<circle cx="${cx}" cy="${cy}" r="${r}" fill="${C.hot}"/><path d="M${cx - d} ${cy - d}l${2 * d} ${2 * d}M${cx + d} ${cy - d}l-${2 * d} ${2 * d}" stroke="${C.card}" stroke-width="1.5" stroke-linecap="round"/>`
69 }
70 if (t.status === 'active') return `<circle cx="${cx}" cy="${cy}" r="${r}" fill="${C.deep}" stroke="${C.glow}" stroke-width="1.2"/>`
71 if (t.status === 'skipped') return `<circle cx="${cx}" cy="${cy}" r="${r - 1}" fill="${C.dim}"/>`
72 return `<circle cx="${cx}" cy="${cy}" r="${r - 0.6}" fill="none" stroke="${C.dim}" stroke-width="1.2"/>`
73}
74
75// A person with a check: the test user.
76const avatar = (cx: number, cy: number, color: string) =>
77 `<circle cx="${cx}" cy="${cy}" r="14" fill="${C.text}"/>` +
78 `<circle cx="${cx}" cy="${cy - 4}" r="3.6" fill="${C.card}"/>` +
79 `<path d="M${cx - 7} ${cy + 8}a7 6 0 0 1 14 0z" fill="${C.card}"/>` +
80 `<circle cx="${cx + 10}" cy="${cy + 9}" r="5" fill="${color}" stroke="${C.card}" stroke-width="1.5"/>`
81
82const GAP = 10
83const HEADER_H = 160
84const TASK_H = 24
85
86/** The whole pane as one SVG: a summary card, the tasks, then a card per finding. */
87export function runSvg(runs: Run[], now: number, width: number): { source: string; height: number; alt: string } {
88 const W = Math.max(300, Math.round(width))
89 const run = runs[0]!
90 const s = stats(run, now)
91 const parts: string[] = []
92 const tone = STATUS_COLOR[run.status]
93
94 // ---- the summary card
95 parts.push(`<rect x="0.5" y="0.5" width="${W - 1}" height="${HEADER_H - 1}" rx="16" fill="${C.card}" stroke="${C.stroke}"/>`)
96 parts.push(avatar(32, 32, tone))
97 const chip = STATUS_LABEL[run.status]
98 const chipW = Math.round(chip.length * 6.4 + 24)
99 parts.push(`<rect x="${W - 18 - chipW}" y="19" width="${chipW}" height="26" rx="13" fill="${C.raised}" stroke="${run.status === 'running' ? C.stroke : tone}"/>`)
100 parts.push(text(W - 18 - chipW / 2, 36, 11.5, run.status === 'running' ? C.soft : tone, esc(chip), 'text-anchor="middle"'))
101 parts.push(text(56, 30, 14.5, C.text, esc(fit(run.target ?? run.brief, W - 56 - chipW - 30, 14.5)), 'font-weight="600"'))
102 parts.push(text(56, 46, 11, C.muted, esc(fit(`${run.scope === 'app' ? 'Whole app' : run.scope === 'feature' ? 'Feature' : 'Scope pending'} · by ${run.tester}`, W - 56 - chipW - 30, 11))))
103
104 parts.push(
105 `<text x="18" y="96" font-weight="500" letter-spacing="-0.5"><tspan font-size="34" fill="${C.text}">${s.passed}</tspan><tspan font-size="18" fill="${C.muted}">/${s.total || '–'}</tspan></text>`,
106 )
107 const lead = run.findings.length ? plural(run.findings.length, 'finding') : run.status === 'running' ? 'nothing found yet' : 'nothing found'
108 const leadColor = s.counts.blocker ? C.hot : s.counts.major ? C.warn : C.accent
109 const rest = ` · ${s.failed ? `${s.failed} failed · ` : ''}${duration(s.elapsedMs)}${run.status === 'running' ? ' in' : ''}`
110 parts.push(`<text x="18" y="116" font-size="12"><tspan fill="${leadColor}" font-weight="600">${esc(lead)}</tspan><tspan fill="${C.muted}">${esc(rest)}</tspan></text>`)
111
112 // Severity counters, right of the big number, where there is room for them.
113 let sx = W - 18
114 for (const sev of W >= 480 ? [...SEVERITIES].reverse() : []) {
115 const n = s.counts[sev]
116 const label = `${n} ${sev}`
117 const w = Math.round(label.length * 6 + 22)
118 sx -= w
119 parts.push(`<rect x="${sx}" y="76" width="${w}" height="22" rx="11" fill="${n ? C.raised : 'none'}" stroke="${n ? SEVERITY_COLOR[sev] : C.stroke}"/>`)
120 parts.push(`<circle cx="${sx + 11}" cy="87" r="3" fill="${n ? SEVERITY_COLOR[sev] : C.dim}"/>`)
121 parts.push(text(sx + 18, 91, 10.5, n ? C.soft : C.dim, esc(label)))
122 sx -= 6
123 }
124
125 // One cell per task, as in plan-progress: filled passed, red failed, glowing current, hatched to come.
126 const cellsY = 130
127 if (run.tasks.length) {
128 const avail = W - 36
129 const cell = Math.max(4, Math.min(56, (avail - (run.tasks.length - 1) * 5) / run.tasks.length))
130 const gap = run.tasks.length > 1 ? Math.min(5, (avail - cell * run.tasks.length) / (run.tasks.length - 1)) : 0
131 run.tasks.forEach((t, i) => {
132 const x = 18 + i * (cell + gap)
133 const fill = t.status === 'passed' ? C.accent : t.status === 'failed' ? C.hot : t.status === 'active' ? C.deep : t.status === 'skipped' ? C.dim : 'url(#hatch)'
134 const stroke = t.status === 'active' ? ` stroke="${C.glow}" stroke-width="1.2"` : ''
135 parts.push(`<rect x="${x.toFixed(1)}" y="${cellsY}" width="${cell.toFixed(1)}" height="18" rx="${Math.min(5, cell / 3).toFixed(1)}" fill="${fill}"${stroke}/>`)
136 })
137 } else {
138 parts.push(`<rect x="18" y="${cellsY}" width="${W - 36}" height="18" rx="5" fill="url(#hatch)"/>`)
139 }
140
141 let y = HEADER_H + GAP
142
143 // ---- why it could not test
144 if (run.failure) {
145 const lines = wrap(run.failure, W - 36, 12, 4)
146 const h = 40 + lines.length * 17
147 parts.push(`<rect x="0.5" y="${y + 0.5}" width="${W - 1}" height="${h - 1}" rx="14" fill="url(#hotwash)" stroke="${C.hot}"/>`)
148 parts.push(text(18, y + 25, 13.5, C.text, run.status === 'blocked' ? 'Could not test' : 'The run failed', 'font-weight="600"'))
149 lines.forEach((l, i) => parts.push(text(18, y + 46 + i * 17, 12, C.soft, esc(l))))
150 y += h + GAP
151 }
152
153 // ---- the summary, once given
154 if (run.summary) {
155 const lines = wrap(run.summary, W - 36, 12, 5)
156 const h = 40 + lines.length * 17
157 parts.push(`<rect x="0.5" y="${y + 0.5}" width="${W - 1}" height="${h - 1}" rx="14" fill="${C.raised}" stroke="${C.stroke}"/>`)
158 parts.push(text(18, y + 25, 13.5, C.text, 'Summary', 'font-weight="600"'))
159 lines.forEach((l, i) => parts.push(text(18, y + 46 + i * 17, 12, C.soft, esc(l))))
160 y += h + GAP
161 }
162
163 // ---- the task list
164 if (run.tasks.length) {
165 const rows = run.tasks.map(t => (t.note && (t.status === 'failed' || t.status === 'active') ? 2 : 1))
166 const h = 42 + rows.reduce((a, b) => a + b, 0) * TASK_H - 4 - rows.filter(r => r === 2).length * 4
167 const isNow = run.status === 'running'
168 parts.push(
169 `<rect x="0.5" y="${y + 0.5}" width="${W - 1}" height="${h - 1}" rx="14" fill="${isNow ? 'url(#now)' : C.raised}" stroke="${isNow ? C.deep : C.stroke}"/>`,
170 )
171 parts.push(text(16, y + 25, 13.5, C.text, 'Tasks', 'font-weight="600"'))
172 parts.push(text(W - 18, y + 25, 11, C.muted, esc(`${s.closed} of ${s.total} done`), 'text-anchor="end"'))
173 let ty = y + 50
174 run.tasks.forEach((t, i) => {
175 parts.push(taskDot(t, 24, ty - 4, 6))
176 const color = t.status === 'active' ? C.text : t.status === 'passed' ? C.soft : t.status === 'failed' ? C.text : C.muted
177 const weight = t.status === 'active' ? ' font-weight="600"' : ''
178 const deco = t.status === 'skipped' ? ' text-decoration="line-through"' : ''
179 const took = t.startedAt != null && t.status !== 'skipped' ? duration((t.endedAt ?? now) - t.startedAt) : ''
180 parts.push(text(38, ty, 12, color, esc(fit(t.title, W - 100, 12)), `${weight}${deco}`))
181 if (took) parts.push(text(W - 18, ty, 11, t.status === 'active' ? C.accent : C.muted, took, 'text-anchor="end"'))
182 if (rows[i] === 2) {
183 ty += TASK_H - 6
184 parts.push(text(38, ty, 11, t.status === 'failed' ? C.hot : C.muted, esc(fit(t.note!, W - 60, 11))))
185 ty += TASK_H + 2
186 } else ty += TASK_H
187 })
188 y += h + GAP
189 }
190
191 // ---- the findings, worst first
192 for (const f of bySeverity(run.findings)) {
193 const color = SEVERITY_COLOR[f.severity]
194 const detail = f.detail ? wrap(f.detail, W - 44, 11.5, 3) : []
195 const h = 50 + (f.where ? 0 : -2) + detail.length * 16 + (detail.length ? 4 : 0)
196 parts.push(`<rect x="0.5" y="${y + 0.5}" width="${W - 1}" height="${h - 1}" rx="14" fill="${C.raised}" stroke="${C.stroke}"/>`)
197 parts.push(`<rect x="0.5" y="${y + 12}" width="3.5" height="${h - 24}" rx="1.75" fill="${color}"/>`)
198 const tag = f.severity.toUpperCase()
199 const tagW = Math.round(tag.length * 6.6 + 16)
200 parts.push(`<rect x="${W - 16 - tagW}" y="${y + 12}" width="${tagW}" height="18" rx="9" fill="none" stroke="${color}"/>`)
201 parts.push(text(W - 16 - tagW / 2, y + 24.5, 9.5, color, tag, 'text-anchor="middle" font-weight="600" letter-spacing="0.6"'))
202 parts.push(text(18, y + 25, 13, C.text, esc(fit(f.title, W - 40 - tagW, 13)), 'font-weight="600"'))
203 const meta = [f.task != null && run.tasks[f.task] ? `task ${f.task + 1}` : '', f.where ?? ''].filter(Boolean).join(' · ')
204 parts.push(text(18, y + 42, 11, C.muted, esc(fit(meta || 'general', W - 40, 11))))
205 detail.forEach((l, i) => parts.push(text(18, y + 62 + i * 16, 11.5, C.soft, esc(l))))
206 y += h + GAP
207 }
208
209 // ---- earlier runs, folded
210 for (const old of runs.slice(1)) {
211 parts.push(`<rect x="0.5" y="${y + 0.5}" width="${W - 1}" height="32" rx="12" fill="none" stroke="${C.stroke}" stroke-dasharray="4 4"/>`)
212 parts.push(`<circle cx="18" cy="${y + 16.5}" r="4" fill="${STATUS_COLOR[old.status]}"/>`)
213 parts.push(text(30, y + 21, 11.5, C.muted, esc(fit(`Earlier: ${historyLine(old, now)}`, W - 50, 11.5))))
214 y += 32 + GAP
215 }
216
217 const height = Math.round(y - GAP)
218 const defs =
219 `<defs>` +
220 `<pattern id="hatch" width="5" height="5" patternUnits="userSpaceOnUse" patternTransform="rotate(45)"><rect width="5" height="5" fill="${C.track}"/><rect width="2" height="5" fill="#4a4463"/></pattern>` +
221 `<linearGradient id="now" x1="0" y1="0" x2="1" y2="1"><stop offset="0" stop-color="#3d3260"/><stop offset="1" stop-color="${C.raised}"/></linearGradient>` +
222 `<linearGradient id="hotwash" x1="0" y1="0" x2="1" y2="1"><stop offset="0" stop-color="#4a2433"/><stop offset="1" stop-color="${C.raised}"/></linearGradient>` +
223 `</defs>`
224 const source = `<svg xmlns="http://www.w3.org/2000/svg" width="${W}" height="${height}" viewBox="0 0 ${W} ${height}" ${FONT}>${defs}${parts.join('')}</svg>`
225
226 const alt =
227 `${run.target ?? run.brief}: ${STATUS_LABEL[run.status]}, ${s.passed} of ${s.total} tasks passed, ${s.failed} failed, ` +
228 `${plural(run.findings.length, 'finding')}${run.failure ? `. ${run.failure}` : ''}`
229 return { source, height, alt }
230}
231
232export function emptySvg(width: number): { source: string; height: number } {
233 const W = Math.max(300, Math.round(width))
234 const H = 92
235 const source =
236 `<svg xmlns="http://www.w3.org/2000/svg" width="${W}" height="${H}" viewBox="0 0 ${W} ${H}" ${FONT}>` +
237 `<rect x="0.5" y="0.5" width="${W - 1}" height="${H - 1}" rx="16" fill="${C.card}" stroke="${C.stroke}" stroke-dasharray="4 4"/>` +
238 text(20, 38, 14, C.text, 'No test run yet', 'font-weight="600"') +
239 text(20, 60, 11.5, C.muted, esc(fit('Ask Claude to test the app, or run /test-user run [what to test].', W - 40, 11.5))) +
240 `</svg>`
241 return { source, height: H }
242}
243hooks/prompt.ts 48 lines1// What the tester is told: the agent type's system prompt and its listing line.
2
3export const AGENT_DESCRIPTION =
4 'A Haiku-powered test user: drives a locally running app in the browser the way a real person would and reports UI/UX problems, ' +
5 'with its task list and findings shown live in the test-user pane. Use it when the user asks to test, QA, try out or click through the app. ' +
6 'In the prompt, say what to test, the local URL if known, and what changed. When the user just says "test the app", ' +
7 'make the focus the feature most recently developed in this session (what it does, where it lives, the files touched); ' +
8 'if nothing was built in this session, ask it to test the entire application. It never edits code. ' +
9 'Its full record (every task, finding and failure, as the developer sees it in the pane) comes back with its result, or with the next prompt when it ran in the background; mcp__test-user__test_report reads it any time.'
10
11export const TESTER_PROMPT = `You are a test user. You try a locally running application the way a real person would, and you report what works, what breaks, and what is confusing. You are testing UI and UX, not reading code for its own sake. You never edit, write or delete project files.
12
13# Your tools
14- Browser tools: mcp__Claude_Browser__* (preview_start, navigate, read_page, find, computer, form_input, get_page_text, read_console_messages, read_network_requests, resize_window) — or mcp__claude-in-chrome__* if those are the ones you have. Prefer read_page / get_page_text / find to read the page; take a screenshot when you judge layout, spacing, contrast or anything visual.
15- Read, Grep, Glob and read-only Bash (git status, git diff, git log, cat) to work out what to test and how to reach the app.
16- Reporting tools, which the developer watches live in a pane. Use them as you go, not only at the end:
17 - mcp__test-user__test_plan — your task list (call it once you know what to test; call again to change it)
18 - mcp__test-user__test_task — mark a task active, passed, failed or skipped, with a short note
19 - mcp__test-user__test_finding — one problem you saw (one call per problem)
20 - mcp__test-user__test_finish — your verdict and summary, last thing before your final answer
21
22# How to run a test
231. Decide the scope.
24 - If the brief names a feature, flow or page, test that (scope "feature"), plus a 30-second smoke check that the app's home screen still loads.
25 - If the brief only says to test "the app" / "the application", or gives no focus: run git status, git diff --stat and git log -5 --stat. If they show a recent feature (changed UI files, a recent commit message), test that feature (scope "feature"). If there is no git history or nothing recent, test the whole application's main flows (scope "app").
262. Find the app. Use the URL in the brief. Otherwise look for it in .claude/launch.json, package.json scripts, vite/next/webpack config or the README, and try the likely localhost port. If it is not running and .claude/launch.json has a configuration, start it with preview_start. If you still cannot reach it, call test_finish with verdict "blocked" and the reason, then stop.
273. Call test_plan with 3–10 concrete tasks phrased as things a user does ("Sign up with a new account", "Add an item to the cart and change its quantity", "Open settings on a narrow window"). Set target to a short name of what you are testing, e.g. "Checkout flow · localhost:5173".
284. Work through the tasks in order. For each: mark it active, do it, look carefully, report each problem with test_finding, then mark it passed or failed (failed = the user could not complete it or it behaved wrongly) with a one-line note.
295. Call test_finish, then give your final answer: a short report — what you tested, the verdict, and the findings worst first, each with where it happened and how to reproduce it.
30
31# What to look for
32- Does it work: actions complete, data saves and shows up, navigation goes where it says, no dead buttons, no errors in the console (read_console_messages) or failing requests.
33- Feedback: loading states, success and error messages, disabled states, what happens on double-click or a slow response.
34- Forms: validation messages, required fields, bad input (empty, too long, wrong format), keyboard (Tab order, Enter to submit, Esc to close).
35- Clarity: labels and copy a newcomer understands, obvious next step, consistent naming.
36- Layout: overlap, cut-off text, misalignment, scroll traps, a narrow viewport (resize_window mobile) when the page is meant to work there.
37- Empty, first-run and edge states: no data, one item, many items.
38- Accessibility basics: buttons and inputs have names in read_page, focus is visible, contrast is readable.
39
40Severities: blocker = a user cannot complete the task; major = it works but badly or loses data / misleads; minor = noticeable friction or a visual defect; polish = small copy or alignment nits.
41
42# Ground rules
43- Stay on the local app (localhost, 127.0.0.1, *.localhost, *.test). Do not visit outside sites beyond what the app itself loads.
44- Use test data only: values you invent, or seed/fixture/example-config values from the project. Never enter real credentials, payment or personal data. If a flow needs a real account you do not have, mark that task skipped and say why.
45- Do not delete data you did not create. Do not change system or browser settings.
46- Be efficient: one look per screen is usually enough; do not re-screenshot what read_page already told you.
47- Report what you saw, not guesses. If you are unsure whether something is a bug, say so in the detail.`
48hooks/run.ts 182 lines1import type { Finding, Run, RunStatus, Severity, Task, TaskStatus } from '../types'
2
3export const MAX_TASKS = 15
4export const MAX_FINDINGS = 40
5export const MAX_RUNS = 5
6const MAX_TITLE = 90
7const MAX_DETAIL = 400
8
9export const SEVERITIES: Severity[] = ['blocker', 'major', 'minor', 'polish']
10
11export function clean(text: unknown, max = MAX_TITLE): string {
12 const out = String(text ?? '')
13 .replace(/\s+/g, ' ')
14 .trim()
15 return out.length > max ? `${out.slice(0, max - 1)}…` : out
16}
17
18export const cleanDetail = (text: unknown) => clean(text, MAX_DETAIL)
19
20export function newRun(agentId: string, tester: string, brief: string, now: number): Run {
21 return { agentId, tester, brief: clean(brief, 160) || 'Test the application', status: 'running', startedAt: now, tasks: [], findings: [] }
22}
23
24const isClosed = (t: Task) => t.status === 'passed' || t.status === 'failed' || t.status === 'skipped'
25
26/** A new task list; tasks already closed under the same title keep their outcome. */
27export function plan(run: Run, titles: string[], now: number): Run {
28 const before = new Map(run.tasks.map(t => [t.title.toLowerCase(), t]))
29 const tasks = titles
30 .map(t => clean(t))
31 .filter(Boolean)
32 .slice(0, MAX_TASKS)
33 .map(title => before.get(title.toLowerCase()) ?? { title, status: 'pending' as const })
34 return advance({ ...run, tasks }, now)
35}
36
37/** With nothing in progress, the first pending task becomes the current one. */
38export function advance(run: Run, now: number): Run {
39 if (run.status !== 'running' || run.tasks.some(t => t.status === 'active')) return run
40 const i = run.tasks.findIndex(t => t.status === 'pending')
41 return i < 0 ? run : setTask(run, i, 'active', undefined, now)
42}
43
44export function setTask(run: Run, index: number, status: TaskStatus, note: string | undefined, now: number): Run {
45 const tasks = run.tasks.map((t, i): Task => {
46 if (i === index) {
47 // Marked by hand, it is no longer a task the run never reached.
48 if (t.isAutoSkipped) t = { title: t.title, status: t.status, note: t.note, startedAt: t.startedAt, endedAt: t.endedAt }
49 if (status === 'pending') return { title: t.title, status }
50 if (status === 'active') return { ...t, status, note: note ?? t.note, startedAt: t.startedAt ?? now, endedAt: undefined }
51 return { ...t, status, note: note ?? t.note, startedAt: t.startedAt ?? now, endedAt: now }
52 }
53 // Starting one task puts any other in progress back in the queue.
54 if (status === 'active' && t.status === 'active') return { ...t, status: 'pending' }
55 return t
56 })
57 return { ...run, tasks }
58}
59
60export function addFinding(run: Run, f: Omit<Finding, 'at'>, now: number): Run {
61 if (run.findings.length >= MAX_FINDINGS) return run
62 return { ...run, findings: [...run.findings, { ...f, at: now }] }
63}
64
65const RANK: Record<Severity, number> = { blocker: 0, major: 1, minor: 2, polish: 3 }
66export const bySeverity = (fs: Finding[]) => [...fs].sort((a, b) => RANK[a.severity] - RANK[b.severity] || a.at - b.at)
67
68/** The outcome the run's own record supports, for a tester that did not say. */
69export function verdictOf(run: Run): RunStatus {
70 if (run.tasks.some(t => t.status === 'failed') || run.findings.length > 0) return 'issues'
71 if (run.tasks.length === 0) return 'blocked'
72 return 'passed'
73}
74
75export function finish(run: Run, status: RunStatus, now: number, endedBy: 'tester' | 'turn', summary?: string, failure?: string): Run {
76 // Tasks left open when the run ends were never tried.
77 const tasks = run.tasks.map((t): Task =>
78 isClosed(t) ? t : { ...t, status: 'skipped', isAutoSkipped: true, endedAt: t.startedAt != null ? now : undefined },
79 )
80 return { ...run, tasks, status, endedAt: now, endedBy, isDelivered: false, summary: summary ?? run.summary, failure: failure ?? run.failure }
81}
82
83/** A run that ended takes updates again: its auto-skipped tasks go back in the queue. */
84export function reopen(run: Run): Run {
85 if (run.status === 'running') return run
86 const tasks = run.tasks.map((t): Task => (t.isAutoSkipped ? { title: t.title, status: 'pending', note: t.note } : t))
87 return { ...run, tasks, status: 'running', endedAt: undefined, endedBy: undefined, failure: undefined, isDelivered: undefined }
88}
89
90export type Stats = {
91 total: number
92 closed: number
93 passed: number
94 failed: number
95 counts: Record<Severity, number>
96 elapsedMs: number
97 current: number | null
98}
99
100export function stats(run: Run, now: number): Stats {
101 const counts = { blocker: 0, major: 0, minor: 0, polish: 0 }
102 for (const f of run.findings) counts[f.severity]++
103 const current = run.tasks.findIndex(t => t.status === 'active')
104 return {
105 total: run.tasks.length,
106 closed: run.tasks.filter(isClosed).length,
107 passed: run.tasks.filter(t => t.status === 'passed').length,
108 failed: run.tasks.filter(t => t.status === 'failed').length,
109 counts,
110 elapsedMs: Math.max(0, (run.endedAt ?? now) - run.startedAt),
111 current: current < 0 ? null : current,
112 }
113}
114
115export function duration(ms: number): string {
116 const min = Math.round(ms / 60_000)
117 if (min < 1) return `${Math.max(0, Math.round(ms / 1000))}s`
118 if (min < 60) return `${min}m`
119 const h = Math.floor(min / 60)
120 return `${h}h${String(min % 60).padStart(2, '0')}m`
121}
122
123export const plural = (n: number, word: string) => `${n} ${word}${n === 1 ? '' : 's'}`
124
125export const STATUS_LABEL: Record<RunStatus, string> = {
126 running: 'Testing',
127 passed: 'Passed',
128 issues: 'Issues found',
129 blocked: 'Blocked',
130 failed: 'Failed',
131}
132
133export function findingsLine(s: Stats): string {
134 const n = SEVERITIES.reduce((sum, k) => sum + s.counts[k], 0)
135 if (n === 0) return 'no findings'
136 const parts = SEVERITIES.filter(k => s.counts[k] > 0).map(k => `${s.counts[k]} ${k}`)
137 return `${plural(n, 'finding')} (${parts.join(', ')})`
138}
139
140/** One line for the status bar. */
141export function statusLine(run: Run, now: number): string {
142 const s = stats(run, now)
143 const n = run.findings.length
144 if (run.status === 'running') {
145 const at = s.current != null ? ` · ${run.tasks[s.current]!.title}` : ''
146 return `test-user ● ${s.closed}/${s.total || '?'} tasks · ${plural(n, 'finding')} · ${duration(s.elapsedMs)}${at}`
147 }
148 if (run.status === 'passed') return `test-user ✓ ${s.passed}/${s.total} passed · ${duration(s.elapsedMs)}`
149 if (run.status === 'issues') return `test-user ! ${s.failed} failed · ${plural(n, 'finding')} · ${duration(s.elapsedMs)}`
150 return `test-user × ${STATUS_LABEL[run.status].toLowerCase()}${run.failure ? `: ${run.failure}` : ''}`
151}
152
153const GLYPH: Record<TaskStatus, string> = { pending: '[ ]', active: '[>]', passed: '[x]', failed: '[!]', skipped: '[-]' }
154
155/** The run as plain text, numbered as the tools take it, for the model. */
156export function report(run: Run, now: number): string {
157 const s = stats(run, now)
158 const lines = [
159 `${STATUS_LABEL[run.status]}: ${run.target ?? run.brief} (${run.scope === 'app' ? 'whole app' : run.scope === 'feature' ? 'feature' : 'scope not set'}, by ${run.tester}, ${duration(s.elapsedMs)})`,
160 ]
161 if (run.failure) lines.push(`Failure: ${run.failure}`)
162 if (run.summary) lines.push(`Summary: ${run.summary}`)
163 lines.push('', `Tasks (${s.passed} passed, ${s.failed} failed, ${s.total} total):`)
164 if (run.tasks.length === 0) lines.push(' (none planned)')
165 run.tasks.forEach((t, i) =>
166 lines.push(` ${i + 1}. ${GLYPH[t.status]} ${t.title}${t.isAutoSkipped ? ' (not reached)' : ''}${t.note ? ` — ${t.note}` : ''}`),
167 )
168 lines.push('', `Findings: ${findingsLine(s)}`)
169 for (const f of bySeverity(run.findings)) {
170 const task = f.task != null && run.tasks[f.task] ? ` [task ${f.task + 1}]` : ''
171 lines.push(` - ${f.severity.toUpperCase()}${task}: ${f.title}${f.where ? ` @ ${f.where}` : ''}${f.detail ? `\n ${f.detail}` : ''}`)
172 }
173 return lines.join('\n')
174}
175
176/** The earlier runs, one line each. */
177export function historyLine(run: Run, now: number): string {
178 const s = stats(run, now)
179 const what = run.status === 'passed' ? `${s.passed}/${s.total} passed` : run.status === 'running' ? 'running' : run.findings.length ? plural(run.findings.length, 'finding') : STATUS_LABEL[run.status].toLowerCase()
180 return `${run.target ?? run.brief} · ${what}`
181}
182types/index.d.ts 64 lines1export type TaskStatus = 'pending' | 'active' | 'passed' | 'failed' | 'skipped'
2
3export type Task = {
4 title: string
5 status: TaskStatus
6 note?: string
7 startedAt?: number
8 endedAt?: number
9 // Skipped only because the run ended with it open; a reopened run takes it up again.
10 isAutoSkipped?: true
11}
12
13export type Severity = 'blocker' | 'major' | 'minor' | 'polish'
14
15export type Finding = {
16 severity: Severity
17 title: string
18 detail?: string
19 // The screen or URL it was seen on.
20 where?: string
21 // The task it came up in, from 0.
22 task?: number
23 at: number
24}
25
26// running: still going. passed: every task passed, nothing found. issues: done,
27// with failed tasks or findings. blocked: the tester could not test (app not
28// reachable, no way in). failed: the run itself died (error, interrupt).
29export type RunStatus = 'running' | 'passed' | 'issues' | 'blocked' | 'failed'
30
31export type Run = {
32 // The tester's agent id, or 'main' when Claude records a run itself.
33 agentId: string
34 tester: string
35 brief: string
36 target?: string
37 scope?: 'feature' | 'app'
38 status: RunStatus
39 startedAt: number
40 endedAt?: number
41 tasks: Task[]
42 findings: Finding[]
43 summary?: string
44 // Why the run was blocked or failed.
45 failure?: string
46 // The Agent call that started it, so its result can carry the report.
47 toolUseId?: string
48 // Who ended it: the tester through test_finish, or its turn ending without one.
49 endedBy?: 'tester' | 'turn'
50 // Whether the main conversation has been handed the finished report.
51 isDelivered?: boolean
52}
53
54declare module 'claude-code' {
55 interface PluginState {
56 'test-user': {
57 // Newest first; the pane shows the first.
58 runs: Run[]
59 // Bumped by the ticker so elapsed times redraw while a run goes.
60 now: number
61 }
62 }
63}
64