SLOPSHOPPER

test-hud

Test runs at a glance: passing over total in the status line, a sparkline of failures across runs, the failing tests, and a toast when the suite turns green

newpaneguardcommandtoaststatus
★ 2v0.3.0MITupdated 2026-10-08ice-lfernandes/claude-code-mods/test-hud
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · test-hud
│ ┃ Tests ✕ › fix the failing auth test and add an audit log call │ ┃ Test runs in Bash this session. Press a │ ┃ failing test to ask for a fix. ⏺ Read(src/auth.ts) │ ┃ ⎿ Read 6 lines │ ┃ ✗ 3/4 passing · 1 failed ⏺ Update(src/auth.ts) │ ┃ bun · #1 · 0s · bun test ↻ run again ⎿ Added 2 lines, removed 1 line │ ┃ ⏺ Bash(bun test) │ ┃ █ failures over the last 1 bun run (most ⎿ 3 pass, 1 fail │ ┃ 1) │ ┃ ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ Failing │ ┃ The output named no failing tests in a ✻ Worked for 42s · done 4:20 PM │ ┃ shape this mod reads. │ ┃ › /test-hud │ ┃ Runs press one to see it │ ┃ ▸ #1 bun 3/4 0s bun test │ ┃ │ ┃ /test-hud clear · demo · help [ Close ] │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ test-hud: ✗ tests 3/4

Draws

Pane · Tests
Test runs in Bash this session. Press a failing test to ask for a fix. ✗ 3/4 passing · 1 failed bun · #1 · 0s · bun test ↻ run again █ failures over the last 1 bun run (most 1) Failing The output named no failing tests in a shape this mod reads. Runs press one to see it ▸ #1 bun 3/4 0s bun test /test-hud clear · demo · help [ Close ]
README

test-hud

Test runs at a glance: how many pass, which ones fail, and whether the count is going down. The serious cousin of boss-fight: same idea, no boss.

test-hud: the /test-hud pane, the status line and the green toast

  • Reads every test run in Bash, by the lead or a subagent: vitest, jest, pytest, Maven (surefire and failsafe), Gradle, cargo, go, bun, mocha, rspec, claude plugin test, and npm, pnpm or yarn test scripts. The command must start the runner (./mvnw test, cd api && npx vitest run); a command that only mentions one (grep jest package.json) does not count. Maven runs with -DskipTests and Gradle runs with -x test do not count.
  • Status line: ✗ tests 41/43 ▃▅█▅▂. That is the last run, passing over total (skipped left out), and a sparkline of failures over the runner's last 8 runs. ▁ is a green run; ▂ to █ scale to the most failures in the line.
  • Toast when the suite turns green after red runs, once: Tests green: 43/43 (vitest) after 4 red runs in 12m 05s. With the regressionToast option, also a toast when a runner turns red right after a green run.
  • /test-hud opens the pane: the last run with its failing tests by name, the tests fixed since the run before, sparklines of failures and duration, and the last 10 runs. The command runs mid-turn. The name is not /tests, which a user's own skill or command often takes.
  • A test that was not failing in the run before is marked new. A test that failed, then passed, then failed again in the command's last 10 runs is marked flaky?.
  • Press a run in the list to see it in the pane; ↩ latest run goes back.
  • With two runners or more in the session, tabs pick which runner the pane shows.
  • Every press waits in the prompt. A failing test puts "Investigate and fix the failure in ..." there, with the whole command that ran. run again asks for the same command, so the run goes through Bash and the pane reads it. The footer's clear puts /test-hud clear there too. Nothing runs until you press Enter.
  • Commands: /test-hud clear drops the runs, /test-hud demo adds six fake runs that go from red to green, /test-hud help lists them.

Options

In /config:

OptionDefaultWhat it does
languageautopt-BR or en for the pane, the status line and the toasts. auto follows LANG: Portuguese for pt_*, English otherwise
regressionToastoffA toast when a runner turns red right after a green run
keepHistoryoffKeep the last runs per project, so the next session in the same project starts with them. Demo runs are not kept

Counts come from the runner's own summary line. Failing test names come from the runner's failure lines, at most 20 per run. For mocha the mod reads counts only, no names. Gradle prints counts only when a test fails, and Maven prints none with -q: a successful run of either shows as green with no counts (✓ tests pass). A run that ends with no summary (a compile error, a missing script) is not recorded.

Install

/plugin install test-hud --marketplace ice-lfernandes/claude-code-mods

Or for one session: claude --plugin-dir ./test-hud

What it reaches

ModNetworkRuns processesFilesCalls a modelSends data anywhere
test-hudNoNoReads Bash's saved copy of an output too long to show wholeNoNo

It reads the Bash calls through the tool.call event, after they run. The hook only observes: it never changes a command or its result, and if it fails, the result stands. When an output is too long for Claude Code to show whole, the summary at its end is cut off. The mod then reads the full output from the file Claude Code saved under tool-results/. It reads no other file. It keeps the last 30 runs in the session's plugin state. Only with keepHistory does it keep them across sessions, in the plugin's store, under the project root. It reads LANG for the auto language.

The test runner parser follows boss-fight from OneWave-AI/claude-code-mods (MIT).

Source 5 files
hooks/register.tsx 368 lines
1// test-hud: test runs at a glance, the serious cousin of boss-fight.
2//
3//   runs     every Bash command that starts a known runner (vitest, jest, pytest, maven, gradle,
4//            cargo, go, bun, mocha, rspec, an npm test script...) is read for its summary:
5//            failed, passed, skipped, and the failing tests by name.
6//   status   "✗ tests 41/43 ▃▅█▅▂": the last run and a sparkline of failures over the runner's
7//            last 8 runs. ▁ is a green run.
8//   toasts   when a runner turns green after red runs, once, with how many red runs it took;
9//            with the `regressionToast` option, also when it turns red after a green run.
10//   /test-hud  opens the pane: the last run, its failing tests (new and flaky ones marked), the
11//            tests it fixed, sparklines of failures and duration, and the recent runs. A failing
12//            test is a button that asks for a fix in the prompt; `run again` puts the command
13//            there. Nothing runs on a press. A recent run is a button that shows it; with two
14//            runners or more, tabs pick one. /test-hud clear drops the runs; /test-hud demo
15//            seeds a red-to-green streak.
16//   history  with the `keepHistory` option, the runs are kept per project root in the plugin's
17//            store, so the next session starts with them.
18//
19// Reads one file: Bash's saved copy of an output too long to show whole, so the summary at its
20// end is not lost. Runs no process, calls no model.
21
22import { atom, read, update } from 'claude-code'
23import type { EngineInterface, Register, ToolCallResult } from 'claude-code'
24
25import type { Parsed, Run } from '../types'
26import { commandLine, durationSpark, fixed, flaky, fresh, greenText, parseRun, quietPass, record, redText, runnerOf, runnersOf, score, spark, statusText, trail, wholeCommand } from './hud'
27import type { Lang, Verb } from './ui'
28import { clip, elapsed, fillArgs, langOf, linesOf, verbRow } from './ui'
29import { COMMAND, WORDS } from './words'
30
31const PANE = 'test-hud'
32const PANE_SPARK = 24
33const PANE_RUNS = 10
34
35const runs = atom({ plugin: 'test-hud', key: 'runs' } as const, [] as Run[])
36const selected = atom({ plugin: 'test-hud', key: 'selected' } as const, null as number | null)
37const tab = atom({ plugin: 'test-hud', key: 'tab' } as const, null as string | null)
38
39let shownStatus: string | undefined
40// Set by register from the options, and by session.start from the system's LANG.
41let lang: Lang = 'en'
42let regressionToast = false
43let keepHistory = false
44let root = ''
45
46const storeKey = () => `runs:${root}`
47
48/** Keeps the runs for the next session in this project, demo runs left out; with the option only. */
49const save = async ($: EngineInterface, list: readonly Run[]) => {
50  if (keepHistory) await $.store.set(storeKey(), list.filter(r => !r.isDemo))
51}
52
53const isRun = (x: unknown): x is Run =>
54  typeof x === 'object' && x !== null && typeof (x as Run).n === 'number' && typeof (x as Run).runner === 'string' && Array.isArray((x as Run).failures)
55
56const show = ($: EngineInterface, list: readonly Run[]) => {
57  const text = statusText(list, lang)
58  if (text === shownStatus) return
59  shownStatus = text
60  $.ui.status(text)
61}
62
63type BashRecord = { stdout?: string; stderr?: string; interrupted?: boolean; backgroundTaskId?: string; persistedOutputPath?: string }
64
65// The engine's own note when an output is too long to show whole.
66const SAVED = /Full output saved to: (\S+\/tool-results\/\S+)/
67
68/** The whole output: the engine's saved copy when Bash cut it short, else what the model read. */
69const outputOf = async ($: EngineInterface, ran: ToolCallResult<'Bash'>) => {
70  const rec = (ran.isError ? undefined : ran.result) as BashRecord | undefined
71  const saved = rec?.persistedOutputPath ?? SAVED.exec(ran.text ?? '')?.[1]
72  if (saved) {
73    const full = await $.fs.read(saved).catch(() => null)
74    if (typeof full === 'string') return full
75  }
76  return rec?.stdout !== undefined ? `${rec.stdout}\n${rec.stderr ?? ''}` : (ran.text ?? '')
77}
78
79const add = async ($: EngineInterface, parsed: Parsed, meta: Omit<Run, keyof Parsed | 'n'>) => {
80  let run: Run | undefined
81  let list: Run[] = []
82  await update($, runs, l => {
83    run = { ...parsed, ...meta, n: (l.at(-1)?.n ?? 0) + 1 }
84    list = record(l, run)
85    return list
86  })
87  show($, list)
88  if (!run) return
89  if (!run.isDemo) await save($, list)
90  const green = greenText(list, run, lang)
91  if (green) $.ui.toast(`${green} /${COMMAND}`, { timeoutMs: 8000 })
92  const red = regressionToast ? redText(list, run, lang) : null
93  if (red) $.ui.toast(`${red} /${COMMAND}`, { timeoutMs: 8000 })
94}
95
96const DEMO_NAMES = [
97  'src/parse.test.ts > parse > handles empty input',
98  'src/parse.test.ts > parse > keeps trailing commas',
99  'src/auth.test.ts > session > refreshes an expired token',
100  'src/auth.test.ts > session > rejects a revoked token',
101  'src/cart.test.ts > totals > rounds half up',
102  'src/cart.test.ts > totals > applies the coupon once',
103  'src/api.test.ts > routes > returns 404 for unknown ids',
104]
105
106const DEMO: readonly { failing: number[]; ms: number }[] = [
107  { failing: [0, 1, 2, 3, 4, 5, 6], ms: 14_000 },
108  { failing: [0, 1, 2, 3, 6], ms: 12_000 },
109  { failing: [0, 1, 2, 3, 6, 4], ms: 13_000 },
110  { failing: [2, 3, 6], ms: 11_000 },
111  { failing: [6], ms: 12_000 },
112  { failing: [], ms: 12_000 },
113]
114
115const runDemo = async ($: EngineInterface) => {
116  const now = await $.clock.now()
117  await update($, runs, l => l.filter(r => !r.isDemo))
118  let at = now - 9 * 60_000
119  for (const d of DEMO) {
120    at += 90_000
121    const failures = d.failing.map(i => DEMO_NAMES[i]!)
122    await add($, { failed: failures.length, passed: 43 - failures.length, skipped: 1, failures }, { runner: 'vitest', command: 'npm test', fullCommand: 'npm test', endedAt: at, durationMs: d.ms, isDemo: true })
123  }
124}
125
126const open = ($: EngineInterface) => $.ui.open({ id: PANE, title: WORDS[lang].pane }).catch(() => null)
127
128/** /test-hud and its arguments: what the command answers, and what the pane's verbs run. */
129const runCommand = async ($: EngineInterface, args: string): Promise<{ text?: string }> => {
130  const w = WORDS[lang]
131  switch (args.trim().toLowerCase()) {
132    case 'clear':
133      await update($, runs, () => [])
134      await update($, selected, () => null)
135      await update($, tab, () => null)
136      show($, [])
137      await save($, [])
138      return { text: w.cleared }
139    case 'demo':
140      await runDemo($)
141      await open($)
142      return { text: w.demoAdded }
143    case 'help':
144      return { text: w.help }
145    case '': {
146      const opened = await open($)
147      if (opened?.isPlaced) return {}
148      const list = await read($, runs)
149      const r = list.at(-1)
150      if (!r) return { text: w.noRuns }
151      return { text: `${w.lastRun(r.runner, score(r, lang), r.failed)} ${spark(trail(list))}${r.failures.length ? ` ${w.failingList(r.failures.slice(0, 5).join('; '))}` : ''}` }
152    }
153    default:
154      return { text: w.help }
155  }
156}
157
158/** Puts a text in the prompt for the person to send; runs nothing. */
159const fill = async ($: EngineInterface, text: string) => {
160  await $.prompt.fill(fillArgs(text))
161}
162
163/** The pane's verbs: clear drops the person's runs, so it waits in the prompt. */
164const VERBS: readonly Verb[] = [{ verb: 'clear', fill: `/${COMMAND} clear` }, { verb: 'demo' }, { verb: 'help' }]
165
166/** A verb pressed in the pane: fills the prompt, or runs and writes its answer to the transcript. */
167const pressVerb = async ($: EngineInterface, v: Verb) => {
168  try {
169    if (v.fill) return await fill($, v.fill)
170    const { text } = await runCommand($, v.verb)
171    for (const line of linesOf(text)) $.ui.log(line)
172  } catch {
173    $.ui.toast(WORDS[lang].failedToRun(`/${COMMAND} ${v.verb}`))
174  }
175}
176
177export const register: Register = (on, options) => {
178  lang = langOf(options.language)
179  regressionToast = options.regressionToast === true
180  keepHistory = options.keepHistory === true
181
182  on('session.start', async ($, e, next) => {
183    const result = await next(e)
184    lang = langOf(options.language, await $.env.get('LANG').catch(() => undefined))
185    await $.command.register({
186      name: COMMAND,
187      description: WORDS[lang].description,
188      argumentHint: '[clear|demo|help]',
189      immediate: true,
190    })
191    if (keepHistory) {
192      root = await $.session.root()
193      const stored = await $.store.get(storeKey()).catch(() => undefined)
194      const kept = Array.isArray(stored) ? stored.filter(isRun) : []
195      if (kept.length) {
196        await update($, runs, () => kept)
197        show($, kept)
198      }
199    }
200    return result
201  })
202
203  on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
204    const runner = runnerOf(e.command)
205    if (!runner || e.run_in_background) return next(e)
206    const start = await $.clock.now()
207    const ran = await next(e)
208    if (ran.deny !== undefined) return ran
209    const rec = (ran.isError ? undefined : ran.result) as BashRecord | undefined
210    if (rec?.backgroundTaskId || rec?.interrupted) return ran
211    const parsed = parseRun(runner, await outputOf($, ran)) ?? (ran.isError ? null : quietPass(runner))
212    if (!parsed) return ran
213    const end = await $.clock.now()
214    await add($, parsed, {
215      runner,
216      command: commandLine(e.command),
217      fullCommand: wholeCommand(e.command),
218      endedAt: end,
219      durationMs: end - start,
220      ...(e.agentId ? { agentId: e.agentId } : {}),
221    })
222    return ran
223  }).catch(($, e, next) => next(e)) // an observer: fail open, the call's result stands
224
225  on('command.run', { command: COMMAND }, ($, e) => runCommand($, e.args))
226
227  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
228    const { Box, Text, Button } = $.ui.resolve(e)
229    const w = WORDS[lang]
230    const list = await read($, runs)
231    const width = Math.max(40, (e.props.bodyColumns || e.viewport?.columns || 80) - 2)
232    const footer = (
233      <Box flexDirection="row" flexWrap="wrap" gap={2}>
234        {verbRow({ Box, Text, Button }, COMMAND, VERBS, v => pressVerb($, v))}
235        <Button key="close" role="dismiss" label={w.close} onPress={() => $.ui.close({ id: PANE })} />
236      </Box>
237    )
238
239    if (list.length === 0) {
240      return (
241        <Box flexDirection="column" paddingX={1} gap={1}>
242          <Box flexDirection="column">
243            <Text dimColor>{w.empty}</Text>
244            <Text dimColor>{w.emptyNext}</Text>
245          </Box>
246          {footer}
247        </Box>
248      )
249    }
250
251    // The runner on show: the tab picked, else the latest run's. The run on show: the one
252    // picked in the list, else that runner's latest.
253    const runners = runnersOf(list)
254    const picked = await read($, tab)
255    const runner = picked !== null && runners.includes(picked) ? picked : list.at(-1)!.runner
256    const mine = list.filter(x => x.runner === runner)
257    const chosen = await read($, selected)
258    const r = mine.find(x => x.n === chosen) ?? mine.at(-1)!
259    const isEarlier = r !== mine.at(-1)
260
261    const isRed = r.failed > 0
262    const counts = [
263      r.passed === null ? (isRed ? '' : w.passedNoCounts) : w.passing(r.passed, r.passed + r.failed),
264      isRed ? w.failed(r.failed) : '',
265      r.skipped ? w.skipped(r.skipped) : '',
266    ].filter(Boolean)
267    const line = mine.filter(x => x.n <= r.n).slice(-PANE_SPARK)
268    const most = Math.max(...line.map(x => x.failed))
269    const ms = line.map(x => x.durationMs)
270    const isNew = fresh(list, r)
271    const isFlaky = flaky(list, r)
272    const gone = fixed(list, r)
273    const recent = mine.slice(-PANE_RUNS).reverse()
274    const tone = isRed ? 'error' : 'success'
275    const pickTab = async (name: string) => {
276      await update($, tab, () => name)
277      await update($, selected, () => null)
278    }
279
280    return (
281      <Box flexDirection="column" paddingX={1} gap={1}>
282        <Text dimColor>{w.hint}</Text>
283
284        {runners.length > 1 && (
285          <Box flexDirection="row" flexWrap="wrap" gap={2}>
286            {runners.map(name => (
287              <Button key={`tab:${name}`} plain dimColor={name !== runner} label={name === runner ? `▸ ${name}` : name} onPress={() => pickTab(name)} />
288            ))}
289          </Box>
290        )}
291
292        <Box flexDirection="column">
293          <Text color={tone} bold>{`${isRed ? '✗' : '✓'} ${counts.join(' · ')}`}</Text>
294          <Box flexDirection="row" flexWrap="wrap" gap={2}>
295            <Text dimColor>{clip(`${r.runner} · #${r.n} · ${elapsed(r.durationMs)}${r.agentId ? ` · ${w.subagent}` : ''} · ${r.command}`, width - w.runAgain.length - 2)}</Text>
296            <Button key="rerun" plain label={w.runAgain} onPress={() => fill($, w.rerun(r.fullCommand))} />
297          </Box>
298          {isEarlier && (
299            <Box flexDirection="row" gap={2}>
300              <Text color="warning">{w.earlier(r.n)}</Text>
301              <Button key="latest" plain label={w.latest} onPress={() => update($, selected, () => null)} />
302            </Box>
303          )}
304        </Box>
305
306        <Box flexDirection="column">
307          <Text>
308            <Text color={tone}>{spark(line)}</Text>
309            <Text dimColor>{`  ${w.sparkNote(line.length, r.runner, most)}`}</Text>
310          </Text>
311          {line.length > 1 && (
312            <Text>
313              <Text>{durationSpark(line)}</Text>
314              <Text dimColor>{`  ${w.durationNote(line.length, elapsed(Math.min(...ms)), elapsed(Math.max(...ms)))}`}</Text>
315            </Text>
316          )}
317        </Box>
318
319        {isRed && (
320          <Box flexDirection="column">
321            <Text bold>{w.failing(r.failures.length, r.failed)}</Text>
322            {r.failures.length === 0 && <Text dimColor>{`  ${w.noNames}`}</Text>}
323            {r.failures.map(f => (
324              <Box key={`fail:${f}`} flexDirection="row">
325                <Text color="error">{'  ✗ '}</Text>
326                <Button key={`fix:${f}`} plain label={clip(f, width - 22)} onPress={() => fill($, w.askFix(f, r.fullCommand))} />
327                {isNew.has(f) && !isFlaky.has(f) && <Text color="warning">{`  ${w.isNew}`}</Text>}
328                {/* flaky says more than new: it failed here before. */}
329                {isFlaky.has(f) && <Text color="warning">{`  ${w.flaky}`}</Text>}
330              </Box>
331            ))}
332          </Box>
333        )}
334
335        {gone.length > 0 && (
336          <Box flexDirection="column">
337            <Text bold>{w.fixed(gone.length)}</Text>
338            {gone.map(f => (
339              <Text key={`fixed:${f}`}>
340                <Text color="success">{'  ✓ '}</Text>
341                <Text>{clip(f, width - 6)}</Text>
342              </Text>
343            ))}
344          </Box>
345        )}
346
347        <Box flexDirection="column">
348          <Text>
349            <Text bold>{w.runs}</Text>
350            <Text dimColor>{`  ${w.pickRun}`}</Text>
351          </Text>
352          {recent.map(x => (
353            <Box key={`run:${x.n}`} flexDirection="row">
354              <Text color={x === r ? 'claude' : undefined}>{x === r ? '▸ ' : '  '}</Text>
355              <Button key={`run:${x.n}`} plain dimColor={x !== r} label={`#${String(x.n).padEnd(4)}`} onPress={() => update($, selected, () => x.n)} />
356              <Text>{x.runner.padEnd(8)}</Text>
357              <Text color={x.failed ? 'error' : 'success'}>{score(x, lang).padEnd(8)}</Text>
358              <Text dimColor>{`${elapsed(x.durationMs).padStart(6)}  ${clip(x.command, Math.max(10, width - 34))}${x.agentId ? ` (${w.subagent})` : ''}`}</Text>
359            </Box>
360          ))}
361        </Box>
362
363        {footer}
364      </Box>
365    )
366  })
367}
368
hooks/hud.ts 280 lines
1// Pure test-run bookkeeping: no engine, so the tests drive it directly.
2// The summary shapes follow boss-fight's parser (OneWave-AI/claude-code-mods, MIT), with Maven
3// (surefire and failsafe), Gradle, skipped counts and failing test names added.
4
5import type { Parsed, Run } from '../types'
6import type { Lang } from './ui'
7import { clip, elapsed } from './ui'
8import { WORDS } from './words'
9
10const MAX_RUNS = 30
11const MAX_NAMES = 20
12/** How far back a test that fails, passes and fails again counts as flaky. */
13export const FLAKY_RUNS = 10
14/** The most of a command kept whole, for the prompt to run it again. */
15const MAX_COMMAND = 2000
16export const SPARK_RUNS = 8
17
18// Where a command starts: line start or after ; & | ( , past env assignments, wrappers and a path.
19const LEAD = String.raw`(?:^|[;&|(])\s*(?:\w+=\S*\s+)*(?:(?:npx|bunx|pnpm(?:\s+exec)?|yarn|uv\s+run|poetry\s+run|pipenv\s+run|hatch\s+run|bundle\s+exec|time|nice|timeout\s+\S+)\s+)*(?:[\w.~-]*\/)*`
20const at = (body: string) => new RegExp(LEAD + body, 'm')
21
22const RUNNERS: readonly [string, RegExp][] = [
23  ['maven', at(String.raw`mvnw?(?:\.cmd)?\b(?=[^;&|\n]*\b(?:test|verify|install|package|integration-test|surefire:test)\b)(?![^;&|\n]*-D(?:skipTests|maven\.test\.skip)\b)`)],
24  ['gradle', at(String.raw`gradlew?(?:\.bat)?\b(?=[^;&|\n]*[\s:](?:\w*[tT]est|check|build)\b)(?![^;&|\n]*(?:-x|--exclude-task)\s+:?test\b)`)],
25  ['vitest', at(String.raw`vitest\b`)],
26  ['jest', at(String.raw`jest\b`)],
27  ['pytest', at(String.raw`(?:pytest|py\.test|python3?\s+-m\s+pytest)\b`)],
28  ['bun', at(String.raw`bun\s+test\b`)],
29  ['cargo', at(String.raw`cargo\s+(?:test|nextest)\b`)],
30  ['go', at(String.raw`go\s+test\b`)],
31  ['claude', at(String.raw`claude\s+plugin\s+test\b`)],
32  ['mocha', at(String.raw`mocha\b`)],
33  ['rspec', at(String.raw`rspec\b`)],
34  ['npm', at(String.raw`(?:npm|pnpm|yarn|bun)\s+(?:run\s+)?(?:test|t)(?::[\w:-]+)?(?=[\s;&|)]|$)`)],
35]
36
37/** Which runner a Bash command starts, or null when it runs no tests. */
38export const runnerOf = (command: string): string | null => {
39  for (const [name, re] of RUNNERS) if (re.test(command)) return name
40  return null
41}
42
43const ANSI = /\u001b\[[0-9;?]*[A-Za-z]/g
44
45const num = (re: RegExp, s: string) => Number(re.exec(s)?.[1] ?? 0)
46
47/** The last match's number: summaries come last, so it wins over per-file lines. */
48const last = (text: string, re: RegExp) => {
49  let out: number | null = null
50  for (const m of text.matchAll(re)) out = Number(m[1])
51  return out
52}
53
54/** Failing tests by name, in the shapes the common runners print. Deduplicated, at most 20. */
55export const failingNames = (text: string): string[] => {
56  const shapes: RegExp[] = [
57    /^\s*● (.+?)\s*$/gm, // jest
58    /^\s*FAIL\s+(.+ > .+?)\s*$/gm, // vitest
59    /^(?:FAILED|ERROR) (\S+::\S+)/gm, // pytest
60    /^\[ERROR\]\s{2,}([\w$]+(?:[.>][\w$[\]]+)+)(?=[:\s]|$)/gm, // maven
61    /^(\S.*? > .+?) FAILED\s*$/gm, // gradle
62    /^test (\S+) \.\.\. FAILED\s*$/gm, // cargo
63    /^\s*--- FAIL: (\S+)/gm, // go
64    /^\(fail\) (.+?)(?: \[[\d.]+m?s\])?\s*$/gm, // bun, claude plugin test
65    /^rspec \S+ # (.+?)\s*$/gm, // rspec
66  ]
67  const seen = new Set<string>()
68  for (const re of shapes) {
69    for (const m of text.matchAll(re)) {
70      const name = m[1]!.trim()
71      if (name === 'Console' || seen.has(name)) continue
72      seen.add(name)
73      if (seen.size === MAX_NAMES) return [...seen].map(n => clip(n, 100))
74    }
75  }
76  return [...seen].map(n => clip(n, 100))
77}
78
79/**
80 * Counts from a runner's output, tolerant of the common summary shapes. Null when the output
81 * carries no summary at all.
82 */
83export const parseRun = (runner: string, raw: string): Parsed | null => {
84  const text = raw.replace(ANSI, '').replace(/\r/g, '')
85  const counts = summary(runner, text)
86  return counts && { ...counts, failures: counts.failed ? failingNames(text) : [] }
87}
88
89const summary = (runner: string, text: string): Omit<Parsed, 'failures'> | null => {
90  // maven: one "Tests run: 43, Failures: 1, Errors: 1, Skipped: 2" per module, after the
91  // per-class lines (which go on with ", Time elapsed"): add the modules up.
92  const maven = [...text.matchAll(/^(?:\[(?:INFO|WARNING|ERROR)\]\s+)?Tests run:\s*(\d+),\s*Failures:\s*(\d+),\s*Errors:\s*(\d+),\s*Skipped:\s*(\d+)\s*$/gm)]
93  if (maven.length) {
94    const [run, failures, errors, skipped] = [1, 2, 3, 4].map(i => maven.reduce((n, m) => n + Number(m[i]), 0)) as [number, number, number, number]
95    return { failed: failures + errors, passed: run - failures - errors - skipped, skipped }
96  }
97
98  // gradle prints counts only when something failed: "43 tests completed, 2 failed, 1 skipped".
99  const gradle = [...text.matchAll(/^\s*(\d+) tests? completed(?:, (\d+) failed)?(?:, (\d+) skipped)?/gm)]
100  if (gradle.length) {
101    const [done, failed, skipped] = [1, 2, 3].map(i => gradle.reduce((n, m) => n + Number(m[i] ?? 0), 0)) as [number, number, number]
102    return { failed, passed: done - failed - skipped, skipped }
103  }
104
105  // jest: "Tests:       2 failed, 1 skipped, 5 passed, 8 total"
106  const jest = /^\s*Tests:\s+(.*\d+ total)/m.exec(text)?.[1]
107  if (jest) return { failed: num(/(\d+) failed/, jest), passed: num(/(\d+) passed/, jest), skipped: num(/(\d+) skipped/, jest) + num(/(\d+) todo/, jest) }
108
109  // vitest: "Tests  2 failed | 5 passed | 1 skipped (8)"
110  const vitest = /^\s*Tests\s+(.*?(?:failed|passed).*?)\s*\(\d+\)\s*$/m.exec(text)?.[1]
111  if (vitest) return { failed: num(/(\d+) failed/, vitest), passed: num(/(\d+) passed/, vitest), skipped: num(/(\d+) skipped/, vitest) + num(/(\d+) todo/, vitest) }
112
113  // cargo: "test result: FAILED. 3 passed; 2 failed; 1 ignored;" once per crate: add them.
114  const cargo = [...text.matchAll(/test result: \w+\.\s+(\d+) passed;\s+(\d+) failed;\s+(\d+) ignored/g)]
115  if (cargo.length) {
116    const [passed, failed, skipped] = [1, 2, 3].map(i => cargo.reduce((n, m) => n + Number(m[i]), 0)) as [number, number, number]
117    return { failed, passed, skipped }
118  }
119
120  // bun, claude plugin test: " 12 pass\n 1 fail"
121  const pass = last(text, /^\s*(\d+) pass\s*$/gm)
122  const fail = last(text, /^\s*(\d+) fail\s*$/gm)
123  if (pass !== null || fail !== null) return { failed: fail ?? 0, passed: pass ?? 0, skipped: last(text, /^\s*(\d+) skip\s*$/gm) ?? 0 }
124
125  // pytest: "=== 2 failed, 5 passed, 1 skipped in 0.12s ===", or the same bare with -q.
126  const py = /^[=\s]*(\d+ (?:failed|passed|errors?|skipped|xfailed|xpassed|deselected|warnings?)\b.*?) in [\d.]+s\b/m.exec(text)?.[1]
127  if (py) return { failed: num(/(\d+) failed/, py) + num(/(\d+) errors?/, py), passed: num(/(\d+) passed/, py), skipped: num(/(\d+) skipped/, py) }
128  if (/^=+ no tests ran\b/m.test(text)) return { failed: 0, passed: 0, skipped: 0 }
129
130  // mocha: "5 passing", "2 failing", "1 pending"
131  const passing = last(text, /^\s*(\d+) passing\b/gm)
132  const failing = last(text, /^\s*(\d+) failing\b/gm)
133  if (passing !== null || failing !== null) return { failed: failing ?? 0, passed: passing ?? 0, skipped: last(text, /^\s*(\d+) pending\b/gm) ?? 0 }
134
135  // rspec: "7 examples, 2 failures, 1 pending"
136  const rspec = /(\d+) examples?, (\d+) failures?(?:, (\d+) pending)?/.exec(text)
137  if (rspec) {
138    const [all, failed, skipped] = [rspec[1], rspec[2], rspec[3]].map(n => Number(n ?? 0)) as [number, number, number]
139    return { failed, passed: all - failed - skipped, skipped }
140  }
141
142  // go: "--- FAIL: TestX" per failing top-level test; "--- PASS" only with -v.
143  if (runner === 'go') {
144    const fails = (text.match(/^--- FAIL:/gm) ?? []).length
145    const passes = (text.match(/^--- PASS:/gm) ?? []).length
146    const failed = fails || (/^FAIL\b/m.test(text) ? 1 : 0)
147    if (failed || passes) return { failed, passed: passes || null, skipped: (text.match(/^--- SKIP:/gm) ?? []).length }
148    if (/^ok\s/m.test(text)) return { failed: 0, passed: null, skipped: 0 }
149  }
150
151  return null
152}
153
154/**
155 * A run with no summary still says something when the runner stays quiet on success: Gradle
156 * always, Maven with -q. Exit 0 then reads as green with no counts.
157 */
158export const quietPass = (runner: string): Parsed | null =>
159  runner === 'gradle' || runner === 'maven' ? { failed: 0, passed: null, skipped: 0, failures: [] } : null
160
161export const record = (list: readonly Run[], run: Run): Run[] => {
162  const next = [...list, run]
163  return next.length > MAX_RUNS ? next.slice(next.length - MAX_RUNS) : next
164}
165
166/** The run of the same runner before `run`. */
167export const previous = (list: readonly Run[], run: Run): Run | undefined =>
168  list.filter(r => r.runner === run.runner && r.n < run.n).at(-1)
169
170/**
171 * Tests the runner's run before `run` failed and `run` does not: all of them when `run` is green.
172 * None when either run failed without names this mod could read.
173 */
174export const fixed = (list: readonly Run[], run: Run): string[] => {
175  const prev = previous(list, run)
176  if (!prev || prev.failures.length === 0) return []
177  if (run.failed > 0 && run.failures.length === 0) return []
178  return prev.failures.filter(f => !run.failures.includes(f))
179}
180
181/** Failing tests of `run` that were not failing in the runner's run before it. */
182export const fresh = (list: readonly Run[], run: Run): Set<string> => {
183  const prev = previous(list, run)
184  // An earlier run that failed without names we could read says nothing about which are new.
185  if (!prev || (prev.failed > 0 && prev.failures.length === 0)) return new Set()
186  return new Set(run.failures.filter(f => !prev.failures.includes(f)))
187}
188
189/** Runs of the same command as `run`, oldest first, up to `run`. */
190export const sameCommand = (list: readonly Run[], run: Run) =>
191  list.filter(r => r.runner === run.runner && r.fullCommand === run.fullCommand && r.n <= run.n)
192
193/** A run whose names say which tests failed: none failed, or the names are all there. */
194const isReadable = (r: Run) => r.failed === 0 || (r.failures.length > 0 && r.failures.length < MAX_NAMES)
195
196/**
197 * Failing tests of `run` that look flaky: over the command's last FLAKY_RUNS runs, the test
198 * failed, then passed, then failed again. Runs whose names were cut or unreadable say nothing.
199 */
200export const flaky = (list: readonly Run[], run: Run, size = FLAKY_RUNS): Set<string> => {
201  const runs = sameCommand(list, run).filter(isReadable).slice(-size)
202  const out = new Set<string>()
203  for (const name of run.failures) {
204    // 0: before its first failure, 1: failed, waiting for a pass, 2: passed, waiting for a fail.
205    let stage = 0
206    for (const r of runs) {
207      const failed = r.failures.includes(name)
208      if (stage === 0 && failed) stage = 1
209      else if (stage === 1 && !failed) stage = 2
210      else if (stage === 2 && failed) {
211        out.add(name)
212        break
213      }
214    }
215  }
216  return out
217}
218
219/** The runners of the runs, in the order they first ran. */
220export const runnersOf = (list: readonly Run[]) => [...new Set(list.map(r => r.runner))]
221
222const LEVELS = '▂▃▄▅▆▇█'
223const BARS = '▁▂▃▄▅▆▇█'
224
225/** Failures per run, oldest first: ▁ is a green run, ▂ to █ scale to the most failures. */
226export const spark = (list: readonly Run[]) => {
227  const max = Math.max(0, ...list.map(r => r.failed))
228  return list.map(r => (r.failed === 0 ? '▁' : LEVELS[Math.min(6, Math.ceil((r.failed / max) * 7) - 1)])).join('')
229}
230
231/** Durations per run, oldest first: ▁ the shortest, █ the longest; ▄ when all take as long. */
232export const durationSpark = (list: readonly Run[]) => {
233  const ms = list.map(r => r.durationMs)
234  const lo = Math.min(...ms)
235  const hi = Math.max(...ms)
236  return ms.map(d => (hi === lo ? '▄' : BARS[Math.round(((d - lo) / (hi - lo)) * 7)])).join('')
237}
238
239/** The last runs of a runner, the latest run's by default, for the sparklines. */
240export const trail = (list: readonly Run[], size = SPARK_RUNS, runner = list.at(-1)?.runner) =>
241  runner ? list.filter(r => r.runner === runner).slice(-size) : []
242
243/** "41/43", "2 failed", "pass". */
244export const score = (r: Run, lang: Lang = 'en') =>
245  r.passed === null ? (r.failed ? WORDS[lang].failedScore(r.failed) : WORDS[lang].pass) : `${r.passed}/${r.passed + r.failed}`
246
247/** "✗ tests 41/43 ▃▅█▅▂", or undefined before any run. */
248export const statusText = (list: readonly Run[], lang: Lang = 'en') => {
249  const r = list.at(-1)
250  if (!r) return undefined
251  const line = spark(trail(list))
252  return `${r.failed ? '✗' : '✓'} ${WORDS[lang].tests} ${score(r, lang)}${line.length > 1 ? ` ${line}` : ''}`
253}
254
255/** The toast for a run that turned its runner green after red runs, else null. */
256export const greenText = (list: readonly Run[], run: Run, lang: Lang = 'en'): string | null => {
257  if (run.failed > 0) return null
258  const same = list.filter(r => r.runner === run.runner && r.n < run.n)
259  let reds = 0
260  while (reds < same.length && same[same.length - 1 - reds]!.failed > 0) reds++
261  if (reds === 0) return null
262  const first = same[same.length - reds]!
263  const span = elapsed(run.endedAt - (first.endedAt - first.durationMs))
264  return WORDS[lang].green(score(run, lang), run.runner, reds, span)
265}
266
267/** The toast for a run that turned its runner red right after a green run, else null. */
268export const redText = (list: readonly Run[], run: Run, lang: Lang = 'en'): string | null => {
269  if (run.failed === 0) return null
270  const prev = previous(list, run)
271  if (!prev || prev.failed > 0) return null
272  return WORDS[lang].red(score(run, lang), run.runner, run.failed)
273}
274
275/** The command's first line, clipped, for the pane. */
276export const commandLine = (command: string) => clip(command.split('\n')[0]!.trim(), 60)
277
278/** The whole command, for the prompt to run it again; cut only when it is very long. */
279export const wholeCommand = (command: string) => clip(command.trim(), MAX_COMMAND)
280
hooks/ui.tsx 114 lines
1// Shared UI helpers, after launchpad's patterns. The same file in every mod that has one: a mod
2// installs alone and cannot import another's code, so scripts/check-shared.sh keeps the copies
3// equal. Change one, copy it to the others.
4//
5// Nothing here takes `$`: the engine follows `$` only into functions of the file that uses it,
6// never across an import. A call on `$` stays in register.tsx; this file gives it its arguments.
7//
8//   language and icons  the `language` and `icons` options, else the system's LANG and terminal.
9//   prompt              the arguments of $.prompt.fill: a text with its first `[blank]` marked.
10//   lists               the window of a long list a pane shows, for ui.scroll.
11//   verbs               a row of a command's arguments, one press each.
12//   numbers             tokens, elapsed time, clipped text and short model names.
13
14import type { Elements, PromptFillArgs } from 'claude-code'
15
16export type Lang = 'pt-BR' | 'en'
17export type IconStyle = 'emoji' | 'symbol'
18
19/** The language: the `language` option when it names one, else Portuguese for a pt LANG, else English. */
20export const langOf = (option: unknown, systemLang?: string | null): Lang =>
21  option === 'en' || option === 'pt-BR' ? option : /^pt([_.@-]|$)/i.test(systemLang ?? '') ? 'pt-BR' : 'en'
22
23/**
24 * The icon style: the `icons` option when it names one; on `auto` (or none), symbols in a
25 * JetBrains IDE's terminal (TERMINAL_EMULATOR=JetBrains-JediTerm), which gives many emoji one
26 * column where Claude Code counts two, and emoji everywhere else.
27 */
28export const styleOf = (option: unknown, terminal?: string | null): IconStyle =>
29  option === 'emoji' || option === 'symbol' ? option : /^JetBrains/i.test(terminal ?? '') ? 'symbol' : 'emoji'
30
31/** An icon in the style: its emoji, or the one-cell symbol that stands in for it. */
32export const glyph = (style: IconStyle, icon: { emoji: string; symbol: string }) => icon[style]
33
34const BLANK = /\[[^\]\n]+\]/
35
36/** The first `[blank]` in a text, as offsets, so the prompt can mark what to replace. */
37export const blankIn = (text: string): { start: number; end: number } | null => {
38  const m = BLANK.exec(text)
39  return m ? { start: m.index, end: m.index + m[0].length } : null
40}
41
42/** What `$.prompt.fill` takes to put a text in the prompt, its `[blank]` marked to replace. */
43export const fillArgs = (text: string): PromptFillArgs => {
44  const blank = blankIn(text)
45  return blank ? { text, decorations: [{ ...blank, bold: true, underline: true }] } : { text }
46}
47
48/** The rows of a list of `total` a pane shows from `offset`, kept inside the list. */
49export const windowOf = (total: number, offset: number, rows: number): { start: number; end: number } => {
50  const size = Math.max(1, rows)
51  const start = Math.max(0, Math.min(offset, total - size))
52  return { start, end: Math.min(total, start + size) }
53}
54
55/**
56 * One of a command's arguments in a verb row. `fill` is the text the prompt waits with, for a
57 * verb that takes an argument or undoes something: a stray click then loses nothing. A verb with
58 * no `fill` runs, and its answer goes to the transcript a line at a time (`linesOf`).
59 */
60export type Verb = { verb: string; label?: string; fill?: string }
61
62/** A command's answer as the lines `$.ui.log` writes, one row each; blank lines dropped. */
63export const linesOf = (text: string | undefined) => (text ?? '').split('\n').filter(line => line.trim() !== '')
64
65/** `/command verb · verb · verb`, each verb a plain button; `lead` goes dim before the command. */
66export function verbRow(
67  ui: Pick<Elements[keyof Elements], 'Box' | 'Text' | 'Button'>,
68  command: string,
69  verbs: readonly Verb[],
70  onPress: (v: Verb) => void,
71  lead = '',
72) {
73  const { Box, Text, Button } = ui
74  return (
75    <Box flexDirection="row" flexWrap="wrap">
76      {lead !== '' && <Text dimColor>{`${lead} · `}</Text>}
77      <Text dimColor>{`/${command} `}</Text>
78      {verbs.map((v, i) => (
79        <Box key={`verbrow:${v.verb}`} flexDirection="row">
80          {i > 0 && <Text dimColor> · </Text>}
81          <Button key={`verb:${v.verb}`} plain dimColor label={v.label ?? v.verb} onPress={() => onPress(v)} />
82        </Box>
83      ))}
84    </Box>
85  )
86}
87
88/** 950, 1.2k, 46k, 1.2M. */
89export const tokens = (n: number) => {
90  if (n < 1000) return String(n)
91  if (n < 1_000_000) return `${(n / 1000).toFixed(n < 10_000 ? 1 : 0)}k`
92  return `${(n / 1_000_000).toFixed(1)}M`
93}
94
95/** 42s, 6m 05s, 1h 02m. */
96export const elapsed = (ms: number) => {
97  const s = Math.max(0, Math.round(ms / 1000))
98  if (s < 60) return `${s}s`
99  const m = Math.floor(s / 60)
100  if (m < 60) return `${m}m ${String(s % 60).padStart(2, '0')}s`
101  return `${Math.floor(m / 60)}h ${String(m % 60).padStart(2, '0')}m`
102}
103
104/** The text cut to `n` characters, an ellipsis last when it was longer. */
105export const clip = (s: string, n: number) => (s.length > n ? `${s.slice(0, n - 1)}…` : s)
106
107/** claude-haiku-4-5-20251001 -> haiku 4.5 */
108export const shortModel = (m?: string) => {
109  if (!m) return ''
110  const hit = /(opus|sonnet|haiku|fable)[-\s]?(\d+(?:[-.]\d+)?)?/i.exec(m)
111  if (!hit) return m.length > 14 ? `${m.slice(0, 13)}…` : m
112  return `${hit[1]!.toLowerCase()}${hit[2] ? ` ${hit[2].replace('-', '.')}` : ''}`
113}
114
hooks/words.ts 142 lines
1// What test-hud says, in Portuguese and English. The language comes from the `language` option,
2// else the system's LANG (ui.tsx's langOf).
3
4import type { Lang } from './ui'
5
6export const COMMAND = 'test-hud'
7
8type Words = {
9  description: string
10  pane: string
11  hint: string
12  passing: (passed: number, total: number) => string
13  passedNoCounts: string
14  failed: (n: number) => string
15  skipped: (n: number) => string
16  subagent: string
17  sparkNote: (runs: number, runner: string, most: number) => string
18  failing: (shown: number, of: number) => string
19  noNames: string
20  isNew: string
21  /** The badge on a test that failed, passed and failed again in the command's last runs. */
22  flaky: string
23  durationNote: (runs: number, shortest: string, longest: string) => string
24  /** The button back to the runner's latest run, and the note while an earlier one shows. */
25  latest: string
26  earlier: (n: number) => string
27  pickRun: string
28  fixed: (n: number) => string
29  runs: string
30  runAgain: string
31  close: string
32  empty: string
33  emptyNext: string
34  /** What a failing test's button puts in the prompt. */
35  askFix: (test: string, command: string) => string
36  /** What `run again` puts in the prompt: a request, so the run goes through Bash and is read. */
37  rerun: (command: string) => string
38  /** The status line's word: `✗ tests 41/43`. */
39  tests: string
40  /** A score with no counts: `2 failed`, `pass`. */
41  failedScore: (n: number) => string
42  pass: string
43  green: (score: string, runner: string, reds: number, span: string) => string
44  red: (score: string, runner: string, failing: number) => string
45  noRuns: string
46  lastRun: (runner: string, score: string, failed: number) => string
47  failingList: (names: string) => string
48  cleared: string
49  demoAdded: string
50  help: string
51  failedToRun: (what: string) => string
52}
53
54export const WORDS: Record<Lang, Words> = {
55  'pt-BR': {
56    description: 'Testes: /test-hud abre o painel; clear, demo, help',
57    pane: 'Testes',
58    hint: 'Execuções de teste no Bash desta sessão. Clique numa falha para pedir a correção.',
59    passing: (p, t) => `${p}/${t} passando`,
60    passedNoCounts: 'passou, sem contagem impressa',
61    failed: n => `${n} ${n === 1 ? 'falhou' : 'falharam'}`,
62    skipped: n => `${n} ${n === 1 ? 'ignorado' : 'ignorados'}`,
63    subagent: 'subagente',
64    sparkNote: (runs, runner, most) => `falhas ${runs === 1 ? 'na última execução' : `nas últimas ${runs} execuções`} do ${runner}${most ? ` (máx. ${most})` : ''}`,
65    failing: (shown, of) => `Falhando${shown ? ` (${shown}${shown < of ? ` de ${of}` : ''})` : ''}`,
66    noNames: 'A saída não nomeou os testes que falharam num formato que este mod lê.',
67    isNew: 'nova',
68    flaky: 'instável?',
69    durationNote: (runs, lo, hi) => `duração ${runs === 1 ? 'da última execução' : `das últimas ${runs} execuções`} (${lo === hi ? lo : `de ${lo} a ${hi}`})`,
70    latest: '↩ última execução',
71    earlier: n => `mostrando a execução #${n}`,
72    pickRun: 'clique numa para ver',
73    fixed: n => `Corrigidos desde a execução anterior (${n})`,
74    runs: 'Execuções',
75    runAgain: '↻ rodar de novo',
76    close: 'Fechar',
77    empty: 'Nenhuma execução de teste nesta sessão. Comandos de teste rodados no Bash aparecem aqui.',
78    emptyNext: '/test-hud demo mostra como fica.',
79    askFix: (test, command) => `Investigue e corrija a falha em ${test}. Comando: ${command}`,
80    rerun: command => `Rode os testes de novo: ${command}`,
81    tests: 'testes',
82    failedScore: n => `${n} ${n === 1 ? 'falhou' : 'falharam'}`,
83    pass: 'ok',
84    green: (score, runner, reds, span) => `Testes verdes: ${score} (${runner}) depois de ${reds} ${reds === 1 ? 'execução vermelha' : 'execuções vermelhas'} em ${span}.`,
85    red: (score, runner, failing) => `Testes ficaram vermelhos: ${score} (${runner}), ${failing} ${failing === 1 ? 'falha' : 'falhas'}.`,
86    noRuns: 'Nenhuma execução de teste nesta sessão.',
87    lastRun: (runner, score, failed) => `Última execução: ${runner} ${score}${failed ? `, ${failed} ${failed === 1 ? 'falhou' : 'falharam'}` : ''}.`,
88    failingList: names => `Falhando: ${names}.`,
89    cleared: 'Execuções de teste apagadas.',
90    demoAdded: 'Seis execuções de demonstração, do vermelho ao verde. /test-hud clear tira.',
91    help: [
92      '/test-hud          abre o painel',
93      '/test-hud clear    apaga as execuções',
94      '/test-hud demo     mostra execuções de exemplo',
95    ].join('\n'),
96    failedToRun: what => `Não deu para rodar: ${what}`,
97  },
98  en: {
99    description: 'Test runs: /test-hud opens the pane; clear, demo, help',
100    pane: 'Tests',
101    hint: 'Test runs in Bash this session. Press a failing test to ask for a fix.',
102    passing: (p, t) => `${p}/${t} passing`,
103    passedNoCounts: 'passed, no counts printed',
104    failed: n => `${n} failed`,
105    skipped: n => `${n} skipped`,
106    subagent: 'subagent',
107    sparkNote: (runs, runner, most) => `failures over the last ${runs} ${runner} run${runs === 1 ? '' : 's'}${most ? ` (most ${most})` : ''}`,
108    failing: (shown, of) => `Failing${shown ? ` (${shown}${shown < of ? ` of ${of}` : ''})` : ''}`,
109    noNames: 'The output named no failing tests in a shape this mod reads.',
110    isNew: 'new',
111    flaky: 'flaky?',
112    durationNote: (runs, lo, hi) => `duration of the last ${runs} run${runs === 1 ? '' : 's'} (${lo === hi ? lo : `${lo} to ${hi}`})`,
113    latest: '↩ latest run',
114    earlier: n => `showing run #${n}`,
115    pickRun: 'press one to see it',
116    fixed: n => `Fixed since the previous run (${n})`,
117    runs: 'Runs',
118    runAgain: '↻ run again',
119    close: 'Close',
120    empty: 'No test runs yet this session. Test commands run in Bash show here.',
121    emptyNext: '/test-hud demo shows what this looks like.',
122    askFix: (test, command) => `Investigate and fix the failure in ${test}. Command: ${command}`,
123    rerun: command => `Run the tests again: ${command}`,
124    tests: 'tests',
125    failedScore: n => `${n} failed`,
126    pass: 'pass',
127    green: (score, runner, reds, span) => `Tests green: ${score} (${runner}) after ${reds} red run${reds === 1 ? '' : 's'} in ${span}.`,
128    red: (score, runner, failing) => `Tests turned red: ${score} (${runner}), ${failing} failing.`,
129    noRuns: 'No test runs yet this session.',
130    lastRun: (runner, score, failed) => `Last run: ${runner} ${score}${failed ? `, ${failed} failed` : ''}.`,
131    failingList: names => `Failing: ${names}.`,
132    cleared: 'Test runs cleared.',
133    demoAdded: 'Six demo runs added, red to green. /test-hud clear removes them.',
134    help: [
135      '/test-hud          open the pane',
136      '/test-hud clear    drop the runs',
137      '/test-hud demo     show sample runs',
138    ].join('\n'),
139    failedToRun: what => `Could not run: ${what}`,
140  },
141}
142
types/index.d.ts 40 lines
1/** What a test runner's output says, before it becomes a Run. */
2export type Parsed = {
3  failed: number
4  /** Null when the runner printed no counts: a quiet Gradle or Maven success. */
5  passed: number | null
6  skipped: number
7  /** Failing tests by name, as the runner prints them. At most 20. */
8  failures: string[]
9}
10
11/** One test command that ran in Bash. Times are epoch milliseconds. */
12export type Run = Parsed & {
13  /** Counts up through the session; the pane shows it as #n. */
14  n: number
15  /** `vitest`, `jest`, `pytest`, `maven`, `gradle`, `npm`... */
16  runner: string
17  /** The command's first line, clipped. */
18  command: string
19  /** The whole command, as it ran, for the prompt to run it again. */
20  fullCommand: string
21  endedAt: number
22  durationMs: number
23  /** Set when a subagent ran it. */
24  agentId?: string
25  isDemo?: boolean
26}
27
28declare module 'claude-code' {
29  interface PluginState {
30    'test-hud': {
31      /** The last runs, oldest first. */
32      runs: Run[]
33      /** The run the pane shows, picked in its list; null for the runner's latest. */
34      selected: number | null
35      /** The runner the pane shows, picked in its tabs; null for the latest run's. */
36      tab: string | null
37    }
38  }
39}
40