Remembers which tests failed on which code, and tells the model when a failing test has both passed and failed on the same code, so it runs the test again…

A test fails, the model assumes its last change broke it and starts "fixing" code that was fine. Sometimes the test is simply flaky: on the very same code it passed an hour ago. This mod remembers which tests failed on which code. When a test fails that both passed and failed on the same code in the last 7 days, it adds a note to that Bash result, so the model runs the test again instead of changing code for a flaky test.
Before a Bash command that runs tests, the mod takes a fingerprint of the working tree: git rev-parse HEAD, git diff HEAD and the names of the untracked files, hashed with 64-bit FNV-1a. After the command, it reads which tests the output names as passed or failed and stores one run per test with that fingerprint.
A test is flaky when one fingerprint holds both a pass and a failure of it. A failure after a code change is not flaky, because the fingerprint is different.
After the Bash result of a failed run, the model reads this note:
flaky-memory: go:TestFlip failed 2 of 3 runs in the last 7 days and both passed and failed on the same code once. It may be flaky rather than broken by this change: run it again before you change code for it.
At the same moment you get one line, so you see what the model was told. It holds the finding alone, without the instruction:
flaky-memory: go:TestFlip failed 2 of 3 runs in the last 7 days and both passed and failed on the same code once
With the sidebar open, that line goes into its stream instead, with failed 2 of 3 in red and the explanation faint, and the transcript stays clean. When the window no longer holds a pass and a failure of that test on one tree, the entry goes away and a new one says so, with is no longer flaky in green:
flaky-memory: no longer flaky go:TestFlip is no longer flaky: nothing in the last 7 days has it passing and failing on the same code
/flaky-memory reset takes the entry down without a closing line, because you asked for it. Without the sidebar, the line lands in the transcript as above.
A Bash command counts as a test command when it contains one of these: go test, pytest, python -m pytest (also python3), jest, vitest, bun test, cargo test, cargo nextest, phpunit (also vendor/bin/phpunit), npm test, pnpm test, yarn test (also with run), bun run test, deno test, rspec, make test, mvn test, gradle test (also ./gradlew test), dotnet test. Other commands pass through untouched, and no git command runs for them.
| Runner | Failed | Passed |
|---|---|---|
| go test | --- FAIL: TestX | --- PASS: TestX (with -v) |
| pytest | FAILED path::test, ERROR path::test | path::test PASSED (-v), PASSED path::test (-rA) |
| jest, vitest, bun | ✕, ×, ✗, (fail) lines | ✓, √, (pass) lines |
| cargo test | test x ... FAILED | test x ... ok |
| PHPUnit | 1) Class::method | none |
| deno test | name ... FAILED | name ... ok |
| dotnet test | Failed Name [12 ms] | Passed Name [1 ms] |
| rspec | the rspec path:line # name rerun list | none |
| Maven surefire | name(Class) Time elapsed … <<< FAILURE! | none |
| Gradle | Class > test FAILED | none |
A run that names no passing test still counts as a pass for the tests the same command failed in its last failing run, as long as it exits 0. That way go test ./... without -v, PHPUnit, rspec, Maven and Gradle work too: their failures are read, and their next run that exits 0 counts those tests as passed.
Only tests that failed within the window are stored, so a suite of thousands of passing tests stores nothing. Each test keeps at most 50 runs from the last 7 days.
/flaky-memory the flaky tests of this repository, the most failing first /flaky-memory reset forget the runs of this repository /flaky-memory reset <test id> forget the runs of one test, for example go:TestFlip /flaky-memory on | off record test runs or not (on by default); off keeps the stored runs
The repository is its git common directory, so the worktrees of one repository share their runs.
claude plugin marketplace add KilimcininKorOglu/claude-code-mods claude plugin install flaky-memory@kilimcininkoroglu-mods
Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.
Validated with claude plugin validate on Claude Code 2.1.283:
❯ ./register.ts hooks: session.start, command.run{command=flaky-memory}, tool.call{tool=Bash} ❯ ./register.ts calls: $.clock.now (via learn, runCommand), $.command.register, $.process.run (via git), $.session.cwd, $.sidebar.clear (via dropEntry), $.sidebar.set (via toPerson), $.store.delete (via forget), $.store.get (via isEnabled, loadHistory), $.store.set (via forget, learn, runCommand), $.ui.log
Reach L2: it runs git.
/tmp) is not in the fingerprint. A test that depends on it can look flaky.go:, pytest:, js:, cargo:, phpunit:, deno:, dotnet:, rspec:, maven:, gradle:) keep two runners' test names apart. The go id carries no package name, so two packages with the same test name share one id.make install # eslint, typescript-eslint, typescript make lint # complexity limit 10, the build fails above it make typecheck # needs .claude/types/ from /plugin-types make validate make test # claude plugin test
hooks/register.ts 193 lines1import type { EngineInterface, Register, ToolCallResult } from 'claude-code'
2import { fingerprintOf, type TreeState } from './fingerprint.ts'
3import { doneLines, doneLog, emptyHistory, findingsFor, isFlaky, listText, logText, noteText, readHistory, record, sectionKey, sidebarLines, type History, type Line } from './history.ts'
4import { isTestCommand, parseOutput } from './parse.ts'
5
6const ENABLED_KEY = 'enabled'
7
8const USAGE = 'expects nothing (the flaky tests), reset, reset <test id>, on or off'
9
10const CONSUMER = 'flaky-memory'
11
12/** The repository a run belongs to, and its fingerprint before the run. */
13type Tree = { project: string; fp: string }
14
15/** The tests reported to the person and not yet closed, so each closing is written once. */
16type State = { open: Set<string> }
17
18function errorText(err: unknown): string {
19 return err instanceof Error ? err.message : String(err)
20}
21
22async function isEnabled($: EngineInterface): Promise<boolean> {
23 return (await $.store.get(ENABLED_KEY)) !== false
24}
25
26/** Runs git read-only by argv; another exit code answers undefined. */
27async function git($: EngineInterface, cwd: string, args: string[]): Promise<string | undefined> {
28 const r = await $.process.run(['git', ...args], { cwd, timeoutMs: 10_000 })
29 return r.exitCode === 0 ? r.stdout : undefined
30}
31
32/** The repository's common git dir (shared by its worktrees), or undefined outside git. */
33async function projectOf($: EngineInterface, cwd: string): Promise<string | undefined> {
34 return (await git($, cwd, ['rev-parse', '--path-format=absolute', '--git-common-dir']))?.trim()
35}
36
37/** The repository and the fingerprint of its working tree, or undefined outside git or for a diff too large. */
38async function treeOf($: EngineInterface, cwd: string): Promise<Tree | undefined> {
39 const project = await projectOf($, cwd)
40 if (project === undefined) return undefined
41 const head = (await git($, cwd, ['rev-parse', '--verify', '--quiet', 'HEAD'])) ?? ''
42 const diffArgs = head === '' ? ['diff', '--no-ext-diff', '--no-color'] : ['diff', '--no-ext-diff', '--no-color', 'HEAD']
43 const state: TreeState = { head, diff: (await git($, cwd, diffArgs)) ?? '', untracked: (await git($, cwd, ['ls-files', '--others', '--exclude-standard'])) ?? '' }
44 const fp = fingerprintOf(state)
45 return fp === undefined ? undefined : { project, fp }
46}
47
48/** The project's stored history; a value of another shape is reported and started over. */
49async function loadHistory($: EngineInterface, project: string): Promise<History> {
50 const h = readHistory(await $.store.get(`runs:${project}`))
51 if (h !== undefined) return h
52 $.ui.log(`the stored runs of ${project} have an unknown shape; starting over`)
53 return emptyHistory()
54}
55
56/** A run that finished in the foreground; an interrupted or backgrounded one printed only part of its output. */
57function finished(r: ToolCallResult<'Bash'>): boolean {
58 if (r.deny !== undefined) return false
59 if (r.isError === true) return true
60 return !r.result.interrupted && r.result.backgroundTaskId === undefined
61}
62
63/**
64 * The finding the person reads: an entry in the shared sidebar's stream while it is open, else the
65 * transcript line. The model's note is another channel and carries the instruction the person does not read.
66 */
67async function toPerson($: EngineInterface, key: string, title: string, lines: Line[], line: string): Promise<void> {
68 try {
69 if (await $.sidebar.set({ consumer: CONSUMER, key: sectionKey(key), title, lines, until: 'stream' })) return
70 } catch {
71 // The sidebar mod is not installed.
72 }
73 $.ui.log(line)
74}
75
76/** Drops the sidebar entries of one test, so a closed finding leaves no warning behind. */
77async function dropEntry($: EngineInterface, id: string): Promise<void> {
78 try {
79 await $.sidebar.clear({ consumer: CONSUMER, key: sectionKey(id) })
80 } catch {
81 // The sidebar mod is not installed.
82 }
83}
84
85/** Closes each reported test the window no longer holds: its runs aged out, or they were forgotten. */
86async function closeResolved($: EngineInterface, state: State, h: History, now: number): Promise<void> {
87 for (const id of [...state.open]) {
88 if (isFlaky(h, id, now)) continue
89 state.open.delete(id)
90 await dropEntry($, id)
91 await toPerson($, id, 'no longer flaky', doneLines(id), doneLog(id))
92 }
93}
94
95/** Reports the flaky tests this run failed to the person, and answers the notes for the model. */
96async function tell($: EngineInterface, state: State, h: History, failed: readonly string[], now: number): Promise<string[]> {
97 const findings = findingsFor(h, failed, now)
98 for (const f of findings) {
99 state.open.add(f.id)
100 // The note goes to the model, the entry to the person: neither reads the other's channel.
101 await toPerson($, f.id, 'flaky test', sidebarLines(f.id, f.v), logText(f.id, f.v))
102 }
103 await closeResolved($, state, h, now)
104 return findings.map(f => noteText(f.id, f.v))
105}
106
107/** Records the run and answers the notes for its flaky failures. */
108async function learn($: EngineInterface, state: State, tree: Tree, command: string, r: ToolCallResult<'Bash'>): Promise<string[]> {
109 const outcome = parseOutput(r.text ?? '')
110 const exitedOk = r.isError !== true
111 const before = await loadHistory($, tree.project)
112 if (outcome.failed.length === 0 && outcome.passed.length === 0 && !(exitedOk && command in before.failedBy)) return []
113 const now = await $.clock.now()
114 const after = record(before, { now, fp: tree.fp, command, outcome, exitedOk })
115 await $.store.set(`runs:${tree.project}`, after)
116 return tell($, state, after, outcome.failed, now)
117}
118
119function withNotes(r: ToolCallResult<'Bash'>, notes: string[]): ToolCallResult<'Bash'> {
120 if (notes.length === 0 || r.deny !== undefined) return r
121 return { ...r, context: [...(r.context ?? []), notes.join('\n')] }
122}
123
124/** Takes the entries of the forgotten tests down; a forgotten test writes no closing line. */
125async function forgetEntries($: EngineInterface, state: State, id: string): Promise<void> {
126 for (const open of [...state.open]) {
127 if (id !== '' && open !== id) continue
128 state.open.delete(open)
129 await dropEntry($, open)
130 }
131}
132
133/** Forgets the runs of one test, or of the whole repository when `id` is empty. */
134async function forget($: EngineInterface, state: State, project: string, id: string): Promise<string> {
135 if (id === '') {
136 await $.store.delete(`runs:${project}`)
137 await forgetEntries($, state, '')
138 return 'the runs of this repository are forgotten'
139 }
140 const h = await loadHistory($, project)
141 if (h.tests[id] === undefined) return `no runs of ${id}`
142 delete h.tests[id]
143 await $.store.set(`runs:${project}`, h)
144 await forgetEntries($, state, id)
145 return `the runs of ${id} are forgotten`
146}
147
148async function runCommand($: EngineInterface, state: State, args: string): Promise<string> {
149 const [word = '', ...rest] = args.trim().split(/\s+/).filter(Boolean)
150 if (word === 'on' || word === 'off') {
151 await $.store.set(ENABLED_KEY, word === 'on')
152 return word === 'on' ? 'on: test runs are recorded' : 'off: test runs are not recorded; the stored runs stay'
153 }
154 const project = await projectOf($, await $.session.cwd())
155 if (project === undefined) return 'not in a git repository: no runs are recorded here'
156 if (word === '') return `${(await isEnabled($)) ? 'on' : 'off'} · ${listText(await loadHistory($, project), await $.clock.now())}`
157 return word === 'reset' ? forget($, state, project, rest.join(' ')) : USAGE
158}
159
160export const register: Register = on => {
161 const state: State = { open: new Set() }
162
163 on('session.start', async ($, e, next) => {
164 const r = await next(e)
165 await $.command.register({
166 name: 'flaky-memory',
167 description: 'Flaky tests of this repository: the list, reset [test id], on, off (flaky-memory)',
168 argumentHint: '[reset [test id] | on | off]',
169 })
170 return r
171 })
172
173 on('command.run', { command: 'flaky-memory' }, async ($, e) => ({ text: await runCommand($, state, String(e.args ?? '')) }))
174
175 on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
176 if (!isTestCommand(e.command) || !(await isEnabled($))) return next(e)
177 let tree: Tree | undefined
178 try {
179 tree = await treeOf($, await $.session.cwd())
180 } catch (err) {
181 $.ui.log(`this test run is not recorded: ${errorText(err)}`)
182 }
183 const r = await next(e)
184 if (tree === undefined || !finished(r)) return r
185 try {
186 return withNotes(r, await learn($, state, tree, e.command, r))
187 } catch (err) {
188 $.ui.log(`this test run is not recorded: ${errorText(err)}`)
189 return r
190 }
191 })
192}
193hooks/fingerprint.ts 28 lines1/** The fingerprint of a working tree: two runs with the same one ran on the same code. */
2
3/** A diff over this is not hashed, and the run is not recorded. */
4export const MAX_DIFF_CHARS = 4 * 1024 * 1024
5
6const OFFSET = 0xcbf29ce484222325n
7const PRIME = 0x100000001b3n
8const MASK = 0xffffffffffffffffn
9
10/** FNV-1a over the UTF-16 code units, 64 bits, as 16 hex digits. */
11export function fnv1a(text: string): string {
12 let hash = OFFSET
13 for (let i = 0; i < text.length; i++) {
14 hash ^= BigInt(text.charCodeAt(i))
15 hash = (hash * PRIME) & MASK
16 }
17 return hash.toString(16).padStart(16, '0')
18}
19
20/** What git says about the tree: the commit, the changes against it, and the untracked file names. */
21export type TreeState = { head: string; diff: string; untracked: string }
22
23/** The fingerprint, or undefined when the diff is too large to hash. */
24export function fingerprintOf(s: TreeState): string | undefined {
25 if (s.diff.length > MAX_DIFF_CHARS) return undefined
26 return fnv1a(`${s.head.trim()}\n${s.diff}\n${s.untracked}`)
27}
28hooks/history.ts 161 lines1/** Each test's recent runs, and whether one of them passed and failed on the same code. */
2import type { Outcome } from './parse.ts'
3
4export const WINDOW_MS = 7 * 24 * 60 * 60 * 1000
5export const MAX_RUNS = 50
6const MAX_COMMANDS = 50
7
8export type Run = { at: number; fp: string; ok: boolean }
9
10/**
11 * `tests`: the runs of each test. `failedBy`: the tests each command failed at
12 * its last failing run, so a later run of the same command that exits 0 counts
13 * them as passed even when its output names no passing test.
14 */
15export type History = { tests: Record<string, Run[]>; failedBy: Record<string, string[]> }
16
17export function emptyHistory(): History {
18 return { tests: {}, failedBy: {} }
19}
20
21function isRun(v: unknown): v is Run {
22 if (typeof v !== 'object' || v === null) return false
23 const r = v as Record<string, unknown>
24 return typeof r.at === 'number' && typeof r.fp === 'string' && typeof r.ok === 'boolean'
25}
26
27function isRecordOf(v: unknown, item: (x: unknown) => boolean): boolean {
28 return typeof v === 'object' && v !== null && !Array.isArray(v) && Object.values(v).every(x => Array.isArray(x) && x.every(item))
29}
30
31/** The stored history, or undefined when the value has another shape. */
32export function readHistory(value: unknown): History | undefined {
33 if (value === undefined) return emptyHistory()
34 if (typeof value !== 'object' || value === null) return undefined
35 const h = value as Record<string, unknown>
36 const ok = isRecordOf(h.tests, isRun) && isRecordOf(h.failedBy, x => typeof x === 'string')
37 return ok ? (value as History) : undefined
38}
39
40function recent(runs: readonly Run[], now: number): Run[] {
41 return runs.filter(r => now - r.at <= WINDOW_MS).slice(-MAX_RUNS)
42}
43
44/** Keeps the newest commands only, so the map does not grow without end. */
45function trimCommands(failedBy: Record<string, string[]>): Record<string, string[]> {
46 const entries = Object.entries(failedBy)
47 return Object.fromEntries(entries.slice(Math.max(0, entries.length - MAX_COMMANDS)))
48}
49
50/** One run of `command` on the tree `fp`: what it printed, and whether it exited 0. */
51export type Recorded = { now: number; fp: string; command: string; outcome: Outcome; exitedOk: boolean }
52
53/**
54 * The history after one run, older runs than the window dropped. A pass is
55 * kept only for a test that has failed in the window, so a suite of thousands
56 * of passing tests stores nothing; a flaky test shows once it fails and later
57 * passes on the same code.
58 */
59export function record(h: History, r: Recorded): History {
60 const failed = new Set(r.outcome.failed)
61 const inferred = r.exitedOk ? (h.failedBy[r.command] ?? []).filter(id => !failed.has(id)) : []
62 const tests: Record<string, Run[]> = {}
63 for (const [id, runs] of Object.entries(h.tests)) tests[id] = recent(runs, r.now)
64 const passed = new Set([...r.outcome.passed, ...inferred].filter(id => tests[id]?.some(run => !run.ok) === true))
65 for (const id of passed) tests[id] = [...(tests[id] ?? []), { at: r.now, fp: r.fp, ok: true }]
66 for (const id of failed) tests[id] = [...(tests[id] ?? []), { at: r.now, fp: r.fp, ok: false }]
67 const failedBy = { ...h.failedBy }
68 delete failedBy[r.command]
69 if (failed.size > 0) failedBy[r.command] = [...failed]
70 const kept = Object.fromEntries(Object.entries(tests).filter(([, runs]) => runs.length > 0))
71 return { tests: kept, failedBy: trimCommands(failedBy) }
72}
73
74/** A test's runs in the window, and on how many trees it both passed and failed. */
75export type Verdict = { runs: number; failures: number; sameCode: number }
76
77export function verdictOf(runs: readonly Run[], now: number): Verdict {
78 const inWindow = recent(runs, now)
79 const byTree = new Map<string, Set<boolean>>()
80 for (const r of inWindow) byTree.set(r.fp, (byTree.get(r.fp) ?? new Set()).add(r.ok))
81 const sameCode = [...byTree.values()].filter(s => s.size === 2).length
82 return { runs: inWindow.length, failures: inWindow.filter(r => !r.ok).length, sameCode }
83}
84
85function times(n: number): string {
86 return n === 1 ? 'once' : `${n} times`
87}
88
89/** The finding both channels carry: what the runs of one test say, without any instruction. */
90function findingText(id: string, v: Verdict): string {
91 return `${id} failed ${v.failures} of ${v.runs} runs in the last 7 days and both passed and failed on the same code ${times(v.sameCode)}`
92}
93
94/** What the model reads after a run in which a flaky test failed. */
95export function noteText(id: string, v: Verdict): string {
96 return `flaky-memory: ${findingText(id, v)}. It may be flaky rather than broken by this change: run it again before you change code for it.`
97}
98
99/** The transcript line: the finding alone, without the instruction the model reads. The engine adds the mod name. */
100export function logText(id: string, v: Verdict): string {
101 return findingText(id, v)
102}
103
104/** How the sidebar colours a line or a part of one. */
105type Tone = 'ok' | 'warn' | 'error' | 'dim'
106export type Part = { text: string; kind?: Tone }
107/** A sidebar line; `parts` colour pieces of it, and `text` holds the whole line for a sidebar that draws no parts. */
108export type Line = { text: string; kind?: Tone; parts?: Part[] }
109
110const part = (text: string, kind: Tone | undefined): Part => (kind === undefined ? { text } : { text, kind })
111
112/** A line made of parts, its `text` their texts joined. */
113const partsLine = (parts: Part[]): Line => ({ text: parts.map(p => p.text).join(''), parts })
114
115/** The finding as sidebar lines: the test id default, the failure count red, the explanation faint. */
116export function sidebarLines(id: string, v: Verdict): Line[] {
117 const why = ` runs in the last 7 days and both passed and failed on the same code ${times(v.sameCode)}`
118 return [partsLine([part(`${id} `, undefined), part(`failed ${v.failures} of ${v.runs}`, 'error'), part(why, 'dim')])]
119}
120
121const NO_LONGER = 'is no longer flaky'
122const NO_LONGER_WHY = ': nothing in the last 7 days has it passing and failing on the same code'
123
124/** The transcript line of a finding the window no longer holds. */
125export function doneLog(id: string): string {
126 return `${id} ${NO_LONGER}${NO_LONGER_WHY}`
127}
128
129/** The closing as sidebar lines: the test id default, `is no longer flaky` green, the explanation faint. */
130export function doneLines(id: string): Line[] {
131 return [partsLine([part(`${id} `, undefined), part(NO_LONGER, 'ok'), part(NO_LONGER_WHY, 'dim')])]
132}
133
134/** Whether the test still both passed and failed on one tree inside the window. */
135export function isFlaky(h: History, id: string, now: number): boolean {
136 return verdictOf(h.tests[id] ?? [], now).sameCode > 0
137}
138
139/** The tests this run failed that have passed and failed on the same code, with their verdicts. */
140export function findingsFor(h: History, failed: readonly string[], now: number): { id: string; v: Verdict }[] {
141 return failed.flatMap(id => {
142 const v = verdictOf(h.tests[id] ?? [], now)
143 return v.sameCode > 0 ? [{ id, v }] : []
144 })
145}
146
147/** A sidebar section key: the test id cut to what the sidebar takes, so one test keeps one key. */
148export function sectionKey(id: string): string {
149 return id.replace(/[^A-Za-z0-9._:-]+/g, '-').slice(0, 64) || 'test'
150}
151
152/** The /flaky-memory listing: every flaky test of the project, the most failing first. */
153export function listText(h: History, now: number): string {
154 const rows = Object.entries(h.tests)
155 .map(([id, runs]) => ({ id, v: verdictOf(runs, now) }))
156 .filter(r => r.v.sameCode > 0)
157 .sort((a, b) => b.v.failures / b.v.runs - a.v.failures / a.v.runs)
158 if (rows.length === 0) return `no flaky test in the last 7 days (${Object.keys(h.tests).length} tests seen)`
159 return rows.map(r => `${r.id} · failed ${r.v.failures}/${r.v.runs} · same code ${times(r.v.sameCode)}`).join('\n')
160}
161hooks/parse.ts 74 lines1/** Which Bash commands run tests, and which tests a run's output names as passed or failed. */
2
3/** A command that runs one of the supported test runners, directly or through npm, make or a vendor path. */
4const TEST_COMMAND =
5 /(^|[\s;&|(/])(go\s+test|pytest|python3?\s+-m\s+pytest|jest|vitest|bun\s+(run\s+)?test|deno\s+test|cargo\s+(test|nextest)|phpunit|rspec|(npm|pnpm|yarn)\s+(run\s+)?test|make\s+test|mvn\s+test|gradlew?\s+test|dotnet\s+test)(\s|$|[;&|)])/
6
7export function isTestCommand(command: string): boolean {
8 return TEST_COMMAND.test(command)
9}
10
11/** The tests one run names, each id prefixed with its runner so two runners never share an id. */
12export type Outcome = { passed: string[]; failed: string[] }
13
14type LineRule = { runner: string; pattern: RegExp; ok: boolean }
15
16/** A trailing duration that jest, vitest and bun print after a test name. */
17const DURATION = /\s*(\[[\d.]+\s?m?s\]|\(\d+(\.\d+)?\s?m?s\)|\d+(\.\d+)?m?s)\s*$/
18
19const RULES: readonly LineRule[] = [
20 { runner: 'go', pattern: /^\s*--- FAIL: (\S+)/, ok: false },
21 { runner: 'go', pattern: /^\s*--- PASS: (\S+)/, ok: true },
22 { runner: 'pytest', pattern: /^(?:FAILED|ERROR) (\S+::\S+)/, ok: false },
23 { runner: 'pytest', pattern: /^(\S+::\S+) PASSED\b/, ok: true },
24 { runner: 'pytest', pattern: /^PASSED (\S+::\S+)/, ok: true },
25 { runner: 'cargo', pattern: /^test (\S+) \.\.\. FAILED$/, ok: false },
26 { runner: 'cargo', pattern: /^test (\S+) \.\.\. ok$/, ok: true },
27 // After the cargo rules: a cargo line ends at `ok`, a deno one carries its duration after it.
28 { runner: 'deno', pattern: /^(.+?) \.\.\. FAILED\b/, ok: false },
29 { runner: 'deno', pattern: /^(.+?) \.\.\. ok\b/, ok: true },
30 { runner: 'phpunit', pattern: /^\d+\) ([\w\\]+::\w+)/, ok: false },
31 // The duration is required, so PHPUnit's `Failed asserting that ...` prose is not read as a test.
32 { runner: 'dotnet', pattern: /^\s*Failed\s+(\S+)\s+\[/, ok: false },
33 { runner: 'dotnet', pattern: /^\s*Passed\s+(\S+)\s+\[/, ok: true },
34 // The rerun list rspec prints after a failing run; a passing run names no test.
35 { runner: 'rspec', pattern: /^rspec\s+\S+ # (.+)$/, ok: false },
36 // Maven surefire and Gradle print a failure per test; neither names a passing one.
37 { runner: 'maven', pattern: /^\s*(\S+)\s+Time elapsed.*<<< (?:FAILURE|ERROR)!/, ok: false },
38 { runner: 'gradle', pattern: /^(\S+ > .+?) FAILED$/, ok: false },
39 { runner: 'js', pattern: /^\s*(?:✕|×|✗|\(fail\))\s+(.+)$/, ok: false },
40 { runner: 'js', pattern: /^\s*(?:✓|√|\(pass\))\s+(.+)$/, ok: true },
41]
42
43/** A vitest file summary (`✓ src/a.test.ts (3 tests) 5ms`) names a file, not a test. */
44const FILE_SUMMARY = /\(\d+ tests?(?: \|[^)]*)?\)/
45
46function testId(rule: LineRule, name: string): string | undefined {
47 if (rule.runner !== 'js') return `${rule.runner}:${name}`
48 if (FILE_SUMMARY.test(name)) return undefined
49 const bare = name.replace(DURATION, '').trim()
50 return bare === '' ? undefined : `js:${bare}`
51}
52
53function ruleFor(line: string): { rule: LineRule; name: string } | undefined {
54 for (const rule of RULES) {
55 const m = rule.pattern.exec(line)
56 if (m?.[1] !== undefined) return { rule, name: m[1] }
57 }
58 return undefined
59}
60
61/** Reads the passed and failed tests a run prints; a test named both ways counts as failed. */
62export function parseOutput(text: string): Outcome {
63 const passed = new Set<string>()
64 const failed = new Set<string>()
65 for (const line of text.split('\n')) {
66 const hit = ruleFor(line)
67 if (hit === undefined) continue
68 const id = testId(hit.rule, hit.name)
69 if (id !== undefined) (hit.rule.ok ? passed : failed).add(id)
70 }
71 for (const id of failed) passed.delete(id)
72 return { passed: [...passed], failed: [...failed] }
73}
74