SLOPSHOPPER

flaky-memory

Remembers which tests failed on which code, and tells the model when a failing test has both passed and failed on the same code, so it runs the test again…

newguardcommandprocess
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · flaky-memory
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /flaky-memory ⎿ flaky-memory: on · no flaky test in the last 7 days (1 tests seen) ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

flaky-memory

A test fails, the model assumes its last change broke it and starts "fixing" code that was fine. Sometimes the test is simply flaky: on the very same code it passed an hour ago. This mod remembers which tests failed on which code. When a test fails that both passed and failed on the same code in the last 7 days, it adds a note to that Bash result, so the model runs the test again instead of changing code for a flaky test.

What it does

Before a Bash command that runs tests, the mod takes a fingerprint of the working tree: git rev-parse HEAD, git diff HEAD and the names of the untracked files, hashed with 64-bit FNV-1a. After the command, it reads which tests the output names as passed or failed and stores one run per test with that fingerprint.

A test is flaky when one fingerprint holds both a pass and a failure of it. A failure after a code change is not flaky, because the fingerprint is different.

After the Bash result of a failed run, the model reads this note:

flaky-memory: go:TestFlip failed 2 of 3 runs in the last 7 days and both passed and failed on the same code once. It may be flaky rather than broken by this change: run it again before you change code for it.

At the same moment you get one line, so you see what the model was told. It holds the finding alone, without the instruction:

flaky-memory: go:TestFlip failed 2 of 3 runs in the last 7 days and both passed and failed on the same code once

With the sidebar open, that line goes into its stream instead, with failed 2 of 3 in red and the explanation faint, and the transcript stays clean. When the window no longer holds a pass and a failure of that test on one tree, the entry goes away and a new one says so, with is no longer flaky in green:

flaky-memory: no longer flaky go:TestFlip is no longer flaky: nothing in the last 7 days has it passing and failing on the same code

/flaky-memory reset takes the entry down without a closing line, because you asked for it. Without the sidebar, the line lands in the transcript as above.

Test commands

A Bash command counts as a test command when it contains one of these: go test, pytest, python -m pytest (also python3), jest, vitest, bun test, cargo test, cargo nextest, phpunit (also vendor/bin/phpunit), npm test, pnpm test, yarn test (also with run), bun run test, deno test, rspec, make test, mvn test, gradle test (also ./gradlew test), dotnet test. Other commands pass through untouched, and no git command runs for them.

What the output must show

RunnerFailedPassed
go test--- FAIL: TestX--- PASS: TestX (with -v)
pytestFAILED path::test, ERROR path::testpath::test PASSED (-v), PASSED path::test (-rA)
jest, vitest, bun✕, ×, ✗, (fail) lines✓, √, (pass) lines
cargo testtest x ... FAILEDtest x ... ok
PHPUnit1) Class::methodnone
deno testname ... FAILEDname ... ok
dotnet testFailed Name [12 ms]Passed Name [1 ms]
rspecthe rspec path:line # name rerun listnone
Maven surefirename(Class) Time elapsed … <<< FAILURE!none
GradleClass > test FAILEDnone

A run that names no passing test still counts as a pass for the tests the same command failed in its last failing run, as long as it exits 0. That way go test ./... without -v, PHPUnit, rspec, Maven and Gradle work too: their failures are read, and their next run that exits 0 counts those tests as passed.

Only tests that failed within the window are stored, so a suite of thousands of passing tests stores nothing. Each test keeps at most 50 runs from the last 7 days.

Command

/flaky-memory the flaky tests of this repository, the most failing first /flaky-memory reset forget the runs of this repository /flaky-memory reset <test id> forget the runs of one test, for example go:TestFlip /flaky-memory on | off record test runs or not (on by default); off keeps the stored runs

The repository is its git common directory, so the worktrees of one repository share their runs.

Install

claude plugin marketplace add KilimcininKorOglu/claude-code-mods claude plugin install flaky-memory@kilimcininkoroglu-mods

Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.

After installing

  1. Restart Claude Code. The mod needs no key and no setting.
  2. It starts recording at the first test run inside a git repository.

What it can reach

Validated with claude plugin validate on Claude Code 2.1.283:

❯ ./register.ts hooks: session.start, command.run{command=flaky-memory}, tool.call{tool=Bash} ❯ ./register.ts calls: $.clock.now (via learn, runCommand), $.command.register, $.process.run (via git), $.session.cwd, $.sidebar.clear (via dropEntry), $.sidebar.set (via toPerson), $.store.delete (via forget), $.store.get (via isEnabled, loadHistory), $.store.set (via forget, learn, runCommand), $.ui.log

Reach L2: it runs git.

  1. Reads: the output of each Bash test command; the working tree through git
  2. Runs: git rev-parse, git diff HEAD and git ls-files --others, read-only, by argv, before each test command
  3. Sends: a note to the model after a failed run of a flaky test, and one line to you; nothing leaves the machine
  4. Persists: per repository, in $.store: each failed test's runs of the last 7 days (time, fingerprint, passed) and the tests each command failed last
  5. Hostile input: test output is untrusted text; it is matched against fixed line patterns, and a test name is only stored and echoed back, never run

Limits

  • Outside a git repository nothing is recorded.
  • A diff over 4 MiB gets no fingerprint, and that run is not recorded.
  • Only the names of untracked files go into the fingerprint, not their content. A change inside an untracked file does not change the fingerprint.
  • State outside the tree (a database, a cache, a file under /tmp) is not in the fingerprint. A test that depends on it can look flaky.
  • An interrupted run and a run sent to the background are not recorded, because their output is partial.
  • The runner prefixes of the ids (go:, pytest:, js:, cargo:, phpunit:, deno:, dotnet:, rspec:, maven:, gradle:) keep two runners' test names apart. The go id carries no package name, so two packages with the same test name share one id.

Development

make install # eslint, typescript-eslint, typescript make lint # complexity limit 10, the build fails above it make typecheck # needs .claude/types/ from /plugin-types make validate make test # claude plugin test

Source 4 files
hooks/register.ts 193 lines
1import type { EngineInterface, Register, ToolCallResult } from 'claude-code'
2import { fingerprintOf, type TreeState } from './fingerprint.ts'
3import { doneLines, doneLog, emptyHistory, findingsFor, isFlaky, listText, logText, noteText, readHistory, record, sectionKey, sidebarLines, type History, type Line } from './history.ts'
4import { isTestCommand, parseOutput } from './parse.ts'
5
6const ENABLED_KEY = 'enabled'
7
8const USAGE = 'expects nothing (the flaky tests), reset, reset <test id>, on or off'
9
10const CONSUMER = 'flaky-memory'
11
12/** The repository a run belongs to, and its fingerprint before the run. */
13type Tree = { project: string; fp: string }
14
15/** The tests reported to the person and not yet closed, so each closing is written once. */
16type State = { open: Set<string> }
17
18function errorText(err: unknown): string {
19  return err instanceof Error ? err.message : String(err)
20}
21
22async function isEnabled($: EngineInterface): Promise<boolean> {
23  return (await $.store.get(ENABLED_KEY)) !== false
24}
25
26/** Runs git read-only by argv; another exit code answers undefined. */
27async function git($: EngineInterface, cwd: string, args: string[]): Promise<string | undefined> {
28  const r = await $.process.run(['git', ...args], { cwd, timeoutMs: 10_000 })
29  return r.exitCode === 0 ? r.stdout : undefined
30}
31
32/** The repository's common git dir (shared by its worktrees), or undefined outside git. */
33async function projectOf($: EngineInterface, cwd: string): Promise<string | undefined> {
34  return (await git($, cwd, ['rev-parse', '--path-format=absolute', '--git-common-dir']))?.trim()
35}
36
37/** The repository and the fingerprint of its working tree, or undefined outside git or for a diff too large. */
38async function treeOf($: EngineInterface, cwd: string): Promise<Tree | undefined> {
39  const project = await projectOf($, cwd)
40  if (project === undefined) return undefined
41  const head = (await git($, cwd, ['rev-parse', '--verify', '--quiet', 'HEAD'])) ?? ''
42  const diffArgs = head === '' ? ['diff', '--no-ext-diff', '--no-color'] : ['diff', '--no-ext-diff', '--no-color', 'HEAD']
43  const state: TreeState = { head, diff: (await git($, cwd, diffArgs)) ?? '', untracked: (await git($, cwd, ['ls-files', '--others', '--exclude-standard'])) ?? '' }
44  const fp = fingerprintOf(state)
45  return fp === undefined ? undefined : { project, fp }
46}
47
48/** The project's stored history; a value of another shape is reported and started over. */
49async function loadHistory($: EngineInterface, project: string): Promise<History> {
50  const h = readHistory(await $.store.get(`runs:${project}`))
51  if (h !== undefined) return h
52  $.ui.log(`the stored runs of ${project} have an unknown shape; starting over`)
53  return emptyHistory()
54}
55
56/** A run that finished in the foreground; an interrupted or backgrounded one printed only part of its output. */
57function finished(r: ToolCallResult<'Bash'>): boolean {
58  if (r.deny !== undefined) return false
59  if (r.isError === true) return true
60  return !r.result.interrupted && r.result.backgroundTaskId === undefined
61}
62
63/**
64 * The finding the person reads: an entry in the shared sidebar's stream while it is open, else the
65 * transcript line. The model's note is another channel and carries the instruction the person does not read.
66 */
67async function toPerson($: EngineInterface, key: string, title: string, lines: Line[], line: string): Promise<void> {
68  try {
69    if (await $.sidebar.set({ consumer: CONSUMER, key: sectionKey(key), title, lines, until: 'stream' })) return
70  } catch {
71    // The sidebar mod is not installed.
72  }
73  $.ui.log(line)
74}
75
76/** Drops the sidebar entries of one test, so a closed finding leaves no warning behind. */
77async function dropEntry($: EngineInterface, id: string): Promise<void> {
78  try {
79    await $.sidebar.clear({ consumer: CONSUMER, key: sectionKey(id) })
80  } catch {
81    // The sidebar mod is not installed.
82  }
83}
84
85/** Closes each reported test the window no longer holds: its runs aged out, or they were forgotten. */
86async function closeResolved($: EngineInterface, state: State, h: History, now: number): Promise<void> {
87  for (const id of [...state.open]) {
88    if (isFlaky(h, id, now)) continue
89    state.open.delete(id)
90    await dropEntry($, id)
91    await toPerson($, id, 'no longer flaky', doneLines(id), doneLog(id))
92  }
93}
94
95/** Reports the flaky tests this run failed to the person, and answers the notes for the model. */
96async function tell($: EngineInterface, state: State, h: History, failed: readonly string[], now: number): Promise<string[]> {
97  const findings = findingsFor(h, failed, now)
98  for (const f of findings) {
99    state.open.add(f.id)
100    // The note goes to the model, the entry to the person: neither reads the other's channel.
101    await toPerson($, f.id, 'flaky test', sidebarLines(f.id, f.v), logText(f.id, f.v))
102  }
103  await closeResolved($, state, h, now)
104  return findings.map(f => noteText(f.id, f.v))
105}
106
107/** Records the run and answers the notes for its flaky failures. */
108async function learn($: EngineInterface, state: State, tree: Tree, command: string, r: ToolCallResult<'Bash'>): Promise<string[]> {
109  const outcome = parseOutput(r.text ?? '')
110  const exitedOk = r.isError !== true
111  const before = await loadHistory($, tree.project)
112  if (outcome.failed.length === 0 && outcome.passed.length === 0 && !(exitedOk && command in before.failedBy)) return []
113  const now = await $.clock.now()
114  const after = record(before, { now, fp: tree.fp, command, outcome, exitedOk })
115  await $.store.set(`runs:${tree.project}`, after)
116  return tell($, state, after, outcome.failed, now)
117}
118
119function withNotes(r: ToolCallResult<'Bash'>, notes: string[]): ToolCallResult<'Bash'> {
120  if (notes.length === 0 || r.deny !== undefined) return r
121  return { ...r, context: [...(r.context ?? []), notes.join('\n')] }
122}
123
124/** Takes the entries of the forgotten tests down; a forgotten test writes no closing line. */
125async function forgetEntries($: EngineInterface, state: State, id: string): Promise<void> {
126  for (const open of [...state.open]) {
127    if (id !== '' && open !== id) continue
128    state.open.delete(open)
129    await dropEntry($, open)
130  }
131}
132
133/** Forgets the runs of one test, or of the whole repository when `id` is empty. */
134async function forget($: EngineInterface, state: State, project: string, id: string): Promise<string> {
135  if (id === '') {
136    await $.store.delete(`runs:${project}`)
137    await forgetEntries($, state, '')
138    return 'the runs of this repository are forgotten'
139  }
140  const h = await loadHistory($, project)
141  if (h.tests[id] === undefined) return `no runs of ${id}`
142  delete h.tests[id]
143  await $.store.set(`runs:${project}`, h)
144  await forgetEntries($, state, id)
145  return `the runs of ${id} are forgotten`
146}
147
148async function runCommand($: EngineInterface, state: State, args: string): Promise<string> {
149  const [word = '', ...rest] = args.trim().split(/\s+/).filter(Boolean)
150  if (word === 'on' || word === 'off') {
151    await $.store.set(ENABLED_KEY, word === 'on')
152    return word === 'on' ? 'on: test runs are recorded' : 'off: test runs are not recorded; the stored runs stay'
153  }
154  const project = await projectOf($, await $.session.cwd())
155  if (project === undefined) return 'not in a git repository: no runs are recorded here'
156  if (word === '') return `${(await isEnabled($)) ? 'on' : 'off'} · ${listText(await loadHistory($, project), await $.clock.now())}`
157  return word === 'reset' ? forget($, state, project, rest.join(' ')) : USAGE
158}
159
160export const register: Register = on => {
161  const state: State = { open: new Set() }
162
163  on('session.start', async ($, e, next) => {
164    const r = await next(e)
165    await $.command.register({
166      name: 'flaky-memory',
167      description: 'Flaky tests of this repository: the list, reset [test id], on, off (flaky-memory)',
168      argumentHint: '[reset [test id] | on | off]',
169    })
170    return r
171  })
172
173  on('command.run', { command: 'flaky-memory' }, async ($, e) => ({ text: await runCommand($, state, String(e.args ?? '')) }))
174
175  on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
176    if (!isTestCommand(e.command) || !(await isEnabled($))) return next(e)
177    let tree: Tree | undefined
178    try {
179      tree = await treeOf($, await $.session.cwd())
180    } catch (err) {
181      $.ui.log(`this test run is not recorded: ${errorText(err)}`)
182    }
183    const r = await next(e)
184    if (tree === undefined || !finished(r)) return r
185    try {
186      return withNotes(r, await learn($, state, tree, e.command, r))
187    } catch (err) {
188      $.ui.log(`this test run is not recorded: ${errorText(err)}`)
189      return r
190    }
191  })
192}
193
hooks/fingerprint.ts 28 lines
1/** The fingerprint of a working tree: two runs with the same one ran on the same code. */
2
3/** A diff over this is not hashed, and the run is not recorded. */
4export const MAX_DIFF_CHARS = 4 * 1024 * 1024
5
6const OFFSET = 0xcbf29ce484222325n
7const PRIME = 0x100000001b3n
8const MASK = 0xffffffffffffffffn
9
10/** FNV-1a over the UTF-16 code units, 64 bits, as 16 hex digits. */
11export function fnv1a(text: string): string {
12  let hash = OFFSET
13  for (let i = 0; i < text.length; i++) {
14    hash ^= BigInt(text.charCodeAt(i))
15    hash = (hash * PRIME) & MASK
16  }
17  return hash.toString(16).padStart(16, '0')
18}
19
20/** What git says about the tree: the commit, the changes against it, and the untracked file names. */
21export type TreeState = { head: string; diff: string; untracked: string }
22
23/** The fingerprint, or undefined when the diff is too large to hash. */
24export function fingerprintOf(s: TreeState): string | undefined {
25  if (s.diff.length > MAX_DIFF_CHARS) return undefined
26  return fnv1a(`${s.head.trim()}\n${s.diff}\n${s.untracked}`)
27}
28
hooks/history.ts 161 lines
1/** Each test's recent runs, and whether one of them passed and failed on the same code. */
2import type { Outcome } from './parse.ts'
3
4export const WINDOW_MS = 7 * 24 * 60 * 60 * 1000
5export const MAX_RUNS = 50
6const MAX_COMMANDS = 50
7
8export type Run = { at: number; fp: string; ok: boolean }
9
10/**
11 * `tests`: the runs of each test. `failedBy`: the tests each command failed at
12 * its last failing run, so a later run of the same command that exits 0 counts
13 * them as passed even when its output names no passing test.
14 */
15export type History = { tests: Record<string, Run[]>; failedBy: Record<string, string[]> }
16
17export function emptyHistory(): History {
18  return { tests: {}, failedBy: {} }
19}
20
21function isRun(v: unknown): v is Run {
22  if (typeof v !== 'object' || v === null) return false
23  const r = v as Record<string, unknown>
24  return typeof r.at === 'number' && typeof r.fp === 'string' && typeof r.ok === 'boolean'
25}
26
27function isRecordOf(v: unknown, item: (x: unknown) => boolean): boolean {
28  return typeof v === 'object' && v !== null && !Array.isArray(v) && Object.values(v).every(x => Array.isArray(x) && x.every(item))
29}
30
31/** The stored history, or undefined when the value has another shape. */
32export function readHistory(value: unknown): History | undefined {
33  if (value === undefined) return emptyHistory()
34  if (typeof value !== 'object' || value === null) return undefined
35  const h = value as Record<string, unknown>
36  const ok = isRecordOf(h.tests, isRun) && isRecordOf(h.failedBy, x => typeof x === 'string')
37  return ok ? (value as History) : undefined
38}
39
40function recent(runs: readonly Run[], now: number): Run[] {
41  return runs.filter(r => now - r.at <= WINDOW_MS).slice(-MAX_RUNS)
42}
43
44/** Keeps the newest commands only, so the map does not grow without end. */
45function trimCommands(failedBy: Record<string, string[]>): Record<string, string[]> {
46  const entries = Object.entries(failedBy)
47  return Object.fromEntries(entries.slice(Math.max(0, entries.length - MAX_COMMANDS)))
48}
49
50/** One run of `command` on the tree `fp`: what it printed, and whether it exited 0. */
51export type Recorded = { now: number; fp: string; command: string; outcome: Outcome; exitedOk: boolean }
52
53/**
54 * The history after one run, older runs than the window dropped. A pass is
55 * kept only for a test that has failed in the window, so a suite of thousands
56 * of passing tests stores nothing; a flaky test shows once it fails and later
57 * passes on the same code.
58 */
59export function record(h: History, r: Recorded): History {
60  const failed = new Set(r.outcome.failed)
61  const inferred = r.exitedOk ? (h.failedBy[r.command] ?? []).filter(id => !failed.has(id)) : []
62  const tests: Record<string, Run[]> = {}
63  for (const [id, runs] of Object.entries(h.tests)) tests[id] = recent(runs, r.now)
64  const passed = new Set([...r.outcome.passed, ...inferred].filter(id => tests[id]?.some(run => !run.ok) === true))
65  for (const id of passed) tests[id] = [...(tests[id] ?? []), { at: r.now, fp: r.fp, ok: true }]
66  for (const id of failed) tests[id] = [...(tests[id] ?? []), { at: r.now, fp: r.fp, ok: false }]
67  const failedBy = { ...h.failedBy }
68  delete failedBy[r.command]
69  if (failed.size > 0) failedBy[r.command] = [...failed]
70  const kept = Object.fromEntries(Object.entries(tests).filter(([, runs]) => runs.length > 0))
71  return { tests: kept, failedBy: trimCommands(failedBy) }
72}
73
74/** A test's runs in the window, and on how many trees it both passed and failed. */
75export type Verdict = { runs: number; failures: number; sameCode: number }
76
77export function verdictOf(runs: readonly Run[], now: number): Verdict {
78  const inWindow = recent(runs, now)
79  const byTree = new Map<string, Set<boolean>>()
80  for (const r of inWindow) byTree.set(r.fp, (byTree.get(r.fp) ?? new Set()).add(r.ok))
81  const sameCode = [...byTree.values()].filter(s => s.size === 2).length
82  return { runs: inWindow.length, failures: inWindow.filter(r => !r.ok).length, sameCode }
83}
84
85function times(n: number): string {
86  return n === 1 ? 'once' : `${n} times`
87}
88
89/** The finding both channels carry: what the runs of one test say, without any instruction. */
90function findingText(id: string, v: Verdict): string {
91  return `${id} failed ${v.failures} of ${v.runs} runs in the last 7 days and both passed and failed on the same code ${times(v.sameCode)}`
92}
93
94/** What the model reads after a run in which a flaky test failed. */
95export function noteText(id: string, v: Verdict): string {
96  return `flaky-memory: ${findingText(id, v)}. It may be flaky rather than broken by this change: run it again before you change code for it.`
97}
98
99/** The transcript line: the finding alone, without the instruction the model reads. The engine adds the mod name. */
100export function logText(id: string, v: Verdict): string {
101  return findingText(id, v)
102}
103
104/** How the sidebar colours a line or a part of one. */
105type Tone = 'ok' | 'warn' | 'error' | 'dim'
106export type Part = { text: string; kind?: Tone }
107/** A sidebar line; `parts` colour pieces of it, and `text` holds the whole line for a sidebar that draws no parts. */
108export type Line = { text: string; kind?: Tone; parts?: Part[] }
109
110const part = (text: string, kind: Tone | undefined): Part => (kind === undefined ? { text } : { text, kind })
111
112/** A line made of parts, its `text` their texts joined. */
113const partsLine = (parts: Part[]): Line => ({ text: parts.map(p => p.text).join(''), parts })
114
115/** The finding as sidebar lines: the test id default, the failure count red, the explanation faint. */
116export function sidebarLines(id: string, v: Verdict): Line[] {
117  const why = ` runs in the last 7 days and both passed and failed on the same code ${times(v.sameCode)}`
118  return [partsLine([part(`${id} `, undefined), part(`failed ${v.failures} of ${v.runs}`, 'error'), part(why, 'dim')])]
119}
120
121const NO_LONGER = 'is no longer flaky'
122const NO_LONGER_WHY = ': nothing in the last 7 days has it passing and failing on the same code'
123
124/** The transcript line of a finding the window no longer holds. */
125export function doneLog(id: string): string {
126  return `${id} ${NO_LONGER}${NO_LONGER_WHY}`
127}
128
129/** The closing as sidebar lines: the test id default, `is no longer flaky` green, the explanation faint. */
130export function doneLines(id: string): Line[] {
131  return [partsLine([part(`${id} `, undefined), part(NO_LONGER, 'ok'), part(NO_LONGER_WHY, 'dim')])]
132}
133
134/** Whether the test still both passed and failed on one tree inside the window. */
135export function isFlaky(h: History, id: string, now: number): boolean {
136  return verdictOf(h.tests[id] ?? [], now).sameCode > 0
137}
138
139/** The tests this run failed that have passed and failed on the same code, with their verdicts. */
140export function findingsFor(h: History, failed: readonly string[], now: number): { id: string; v: Verdict }[] {
141  return failed.flatMap(id => {
142    const v = verdictOf(h.tests[id] ?? [], now)
143    return v.sameCode > 0 ? [{ id, v }] : []
144  })
145}
146
147/** A sidebar section key: the test id cut to what the sidebar takes, so one test keeps one key. */
148export function sectionKey(id: string): string {
149  return id.replace(/[^A-Za-z0-9._:-]+/g, '-').slice(0, 64) || 'test'
150}
151
152/** The /flaky-memory listing: every flaky test of the project, the most failing first. */
153export function listText(h: History, now: number): string {
154  const rows = Object.entries(h.tests)
155    .map(([id, runs]) => ({ id, v: verdictOf(runs, now) }))
156    .filter(r => r.v.sameCode > 0)
157    .sort((a, b) => b.v.failures / b.v.runs - a.v.failures / a.v.runs)
158  if (rows.length === 0) return `no flaky test in the last 7 days (${Object.keys(h.tests).length} tests seen)`
159  return rows.map(r => `${r.id} · failed ${r.v.failures}/${r.v.runs} · same code ${times(r.v.sameCode)}`).join('\n')
160}
161
hooks/parse.ts 74 lines
1/** Which Bash commands run tests, and which tests a run's output names as passed or failed. */
2
3/** A command that runs one of the supported test runners, directly or through npm, make or a vendor path. */
4const TEST_COMMAND =
5  /(^|[\s;&|(/])(go\s+test|pytest|python3?\s+-m\s+pytest|jest|vitest|bun\s+(run\s+)?test|deno\s+test|cargo\s+(test|nextest)|phpunit|rspec|(npm|pnpm|yarn)\s+(run\s+)?test|make\s+test|mvn\s+test|gradlew?\s+test|dotnet\s+test)(\s|$|[;&|)])/
6
7export function isTestCommand(command: string): boolean {
8  return TEST_COMMAND.test(command)
9}
10
11/** The tests one run names, each id prefixed with its runner so two runners never share an id. */
12export type Outcome = { passed: string[]; failed: string[] }
13
14type LineRule = { runner: string; pattern: RegExp; ok: boolean }
15
16/** A trailing duration that jest, vitest and bun print after a test name. */
17const DURATION = /\s*(\[[\d.]+\s?m?s\]|\(\d+(\.\d+)?\s?m?s\)|\d+(\.\d+)?m?s)\s*$/
18
19const RULES: readonly LineRule[] = [
20  { runner: 'go', pattern: /^\s*--- FAIL: (\S+)/, ok: false },
21  { runner: 'go', pattern: /^\s*--- PASS: (\S+)/, ok: true },
22  { runner: 'pytest', pattern: /^(?:FAILED|ERROR) (\S+::\S+)/, ok: false },
23  { runner: 'pytest', pattern: /^(\S+::\S+) PASSED\b/, ok: true },
24  { runner: 'pytest', pattern: /^PASSED (\S+::\S+)/, ok: true },
25  { runner: 'cargo', pattern: /^test (\S+) \.\.\. FAILED$/, ok: false },
26  { runner: 'cargo', pattern: /^test (\S+) \.\.\. ok$/, ok: true },
27  // After the cargo rules: a cargo line ends at `ok`, a deno one carries its duration after it.
28  { runner: 'deno', pattern: /^(.+?) \.\.\. FAILED\b/, ok: false },
29  { runner: 'deno', pattern: /^(.+?) \.\.\. ok\b/, ok: true },
30  { runner: 'phpunit', pattern: /^\d+\) ([\w\\]+::\w+)/, ok: false },
31  // The duration is required, so PHPUnit's `Failed asserting that ...` prose is not read as a test.
32  { runner: 'dotnet', pattern: /^\s*Failed\s+(\S+)\s+\[/, ok: false },
33  { runner: 'dotnet', pattern: /^\s*Passed\s+(\S+)\s+\[/, ok: true },
34  // The rerun list rspec prints after a failing run; a passing run names no test.
35  { runner: 'rspec', pattern: /^rspec\s+\S+ # (.+)$/, ok: false },
36  // Maven surefire and Gradle print a failure per test; neither names a passing one.
37  { runner: 'maven', pattern: /^\s*(\S+)\s+Time elapsed.*<<< (?:FAILURE|ERROR)!/, ok: false },
38  { runner: 'gradle', pattern: /^(\S+ > .+?) FAILED$/, ok: false },
39  { runner: 'js', pattern: /^\s*(?:✕|×|✗|\(fail\))\s+(.+)$/, ok: false },
40  { runner: 'js', pattern: /^\s*(?:✓|√|\(pass\))\s+(.+)$/, ok: true },
41]
42
43/** A vitest file summary (`✓ src/a.test.ts (3 tests) 5ms`) names a file, not a test. */
44const FILE_SUMMARY = /\(\d+ tests?(?: \|[^)]*)?\)/
45
46function testId(rule: LineRule, name: string): string | undefined {
47  if (rule.runner !== 'js') return `${rule.runner}:${name}`
48  if (FILE_SUMMARY.test(name)) return undefined
49  const bare = name.replace(DURATION, '').trim()
50  return bare === '' ? undefined : `js:${bare}`
51}
52
53function ruleFor(line: string): { rule: LineRule; name: string } | undefined {
54  for (const rule of RULES) {
55    const m = rule.pattern.exec(line)
56    if (m?.[1] !== undefined) return { rule, name: m[1] }
57  }
58  return undefined
59}
60
61/** Reads the passed and failed tests a run prints; a test named both ways counts as failed. */
62export function parseOutput(text: string): Outcome {
63  const passed = new Set<string>()
64  const failed = new Set<string>()
65  for (const line of text.split('\n')) {
66    const hit = ruleFor(line)
67    if (hit === undefined) continue
68    const id = testId(hit.rule, hit.name)
69    if (id !== undefined) (hit.rule.ok ? passed : failed).add(id)
70  }
71  for (const id of failed) passed.delete(id)
72  return { passed: [...passed], failed: [...failed] }
73}
74