Blocks file writes that break Red-Green-Refactor on coding tasks (port of probity's enforceTdd).

Experimental. This is an experiment in enforcing test-driven development on a coding agent. It works in a kata, has not been used on a real codebase yet, and its settings, behaviour and the Claude Code API it is built on may change without notice.
A Claude Code mod that keeps the agent to Red → Green → Refactor. Before each file write on a coding task, a judge model checks the write against the TDD rules and blocks it, with the reason, when it skips a failing test or implements more than the failing test needs. The agent reads the reason and corrects course.
The TDD rules and the judging approach are ported from probity's enforceTdd rule (MIT). Probity runs as an external PreToolUse command; tdd-mod runs in-process as a Claude Code function-hooks plugin (a "mod"), so it needs no extra install, API key or process per write.
Write, Edit and NotebookEdit. The judge sees the last 10 prompts and tool calls of the session, the file as it is, and the file as the write would leave it, and answers pass or violation.TDD on or TDD off (not coding).cat > file, >>, tee, sed -i, cp, inline Python/Node writes) and tells the agent to use Write or Edit, so every change can be judged. Moving files (mv, git mv) is allowed.Install it from this repository's marketplace, in Claude Code:
/plugin marketplace add koslowskyj/tdd-mod
/plugin install tdd-mod@tdd-mod
Or run it from a clone, for one session:
git clone https://github.com/koslowskyj/tdd-mod.git ~/tdd-mod
cd your-project
claude --plugin-dir ~/tdd-mod
Then give the agent a coding task. Each blocked write shows up as a tool error starting with tdd-mod:.
To see every verdict, start with --debug and follow the debug log:
claude --plugin-dir ~/tdd-mod --debug
tail -f "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/debug/latest" | grep --line-buffered "tdd-mod: "
Lines look like tdd-mod: pass (first opinion) src/cart.ts — Green phase: …, tdd-mod: violation (second opinion) … or tdd-mod: task other: …. For every blocked write, the full judge prompt follows in numbered parts (blocked prompt … part 1/3), so the verdict can be replayed exactly.
It helps to tell the agent, in your project's CLAUDE.md, to work test-first and to change files with Write and Edit only.
Change them in /config, with /plugin configure tdd-mod@tdd-mod, or under pluginConfigs in settings.json. Installed from the marketplace the plugin is keyed tdd-mod@tdd-mod; loaded with --plugin-dir, tdd-mod (or tdd-mod@inline):
{
"pluginConfigs": {
"tdd-mod@tdd-mod": { "options": { "judgeModel": "sonnet" } }
}
}
| Setting | Default | What it does |
|---|---|---|
enabled | true | Off, the mod does nothing. |
judgeModel | haiku | The model that judges each write on a coding task. It writes no code; the session's own model does. Haiku is the default because it scored best in /tdd-eval. |
classifierModel | haiku | The model that labels each typed prompt coding or other. |
classifyTasks | true | Off, every task counts as coding and every write is judged. |
fastPath | true | Pass a write that adds exactly one test without a model call. Faster, but a new test no longer checks for a refactor left unmade. |
secondOpinion | true | Block only when a second judge call agrees. |
extensions | ts, tsx, mts, cts, js, jsx, mjs, cjs, py, go, rs, java, kt, kts, scala, rb, php, cs, fs, swift, c, cc, cpp, h, hpp, ex, exs, erl, clj, dart, lua, ipynb, vue, svelte | Comma-separated extensions of the files judged. |
ignore | /node_modules/, /dist/, /build/, /vendor/, /.git/, /.claude/ | Comma-separated globs of paths never judged. |
testPatterns | (empty: built-in patterns) | How the fast path recognises a test declaration, per extension. A JSON object from comma-separated extensions to a regex; see below. |
A change in /config reloads the mod with the new values.
The fast path counts test declarations with a regex per file extension. The built-in patterns cover it(/test( in JS/TS, def test_ in Python, func Test in Go, @Test in Java/Kotlin, #[test] in Rust, it "/def test_ in Ruby and [Fact]/[Test] in C#. testPatterns replaces or adds patterns per extension; an empty regex turns the fast path off for those extensions:
{
"pluginConfigs": {
"tdd-mod@tdd-mod": {
"options": {
"testPatterns": "{\"ts,js\": \"\\\\bscenario\\\\(\", \"ex,exs\": \"^\\\\s*test \\\"\", \"go\": \"\"}"
}
}
}
}
That is the JSON object {"ts,js": "\\bscenario\\(", "ex,exs": "^\\s*test \"", "go": ""} written as a string. The regexes run with the g and m flags, so ^ matches at each line. An invalid value is reported in the transcript when the session starts, and the built-in patterns stay in place.
/tdd-evaleval/ holds a small regression set built from kata runs: 20 labelled writes (verdicts.json) and 15 labelled prompts (tasks.json). /tdd-eval [runs] replays them through exactly the decisions a session makes, with the configured models, and reports the hit rate per case:
claude -p "/tdd-eval 3" --plugin-dir ~/tdd-mod
Results so far, 3 runs per case:
| Judge model | Writes that should pass | Writes that should be blocked | Total |
|---|---|---|---|
| haiku (default) | 51/51 | 9/9 | 60/60 |
| opus | 51/51 | 2/9 | 53/60 |
| sonnet | 51/51 | 0/9 | 51/60 |
The classifier (haiku) labelled all 45 prompts correctly. The set is small (three distinct violations), so treat these numbers as a direction, not a benchmark. Add a case whenever the judge gets a real write wrong.
claude plugin validate .claude-plugin/plugin.json # manifest and hooks, as the engine reads them
claude plugin validate .claude-plugin/marketplace.json # the marketplace entry
claude plugin test . # unit and hook tests (test/*.test.ts)
npx -p typescript@5 tsc -p .
tsconfig.json extends .claude-plugin/types/tsconfig.json, which Claude Code writes when it loads the mod. Load it once (claude --plugin-dir .) before type-checking.
GitHub Actions (.github/workflows/ci.yml) runs these checks on every push to main and on pull requests, against the pinned Claude Code version. None of them needs credentials. /tdd-eval is not part of CI: it calls the models.
Every push to main that passes the checks is released. The release job bumps the version in .claude-plugin/plugin.json from the commit messages since the last tag, commits it, tags vX.Y.Z and creates a GitHub release:
| Commits since the last release | Bump |
|---|---|
feat!: or BREAKING CHANGE | major |
feat: | minor |
| anything else | patch |
Installed copies get the new version with /plugin marketplace update tdd-mod (or Claude Code's marketplace auto-update). The official Anthropic marketplace takes third-party plugins only through its submission form, so that step is manual.
| File | What it holds |
|---|---|
.claude-plugin/plugin.json | The plugin manifest: name, version, settings |
.claude-plugin/marketplace.json | Makes this repository a marketplace that lists the plugin |
hooks/hooks.json | Points Claude Code at src/register.ts |
src/register.ts | The hooks: prompt classification, the write and Bash guards, /tdd-eval |
src/guard.ts | The two decisions: task kind, and pass/violation for a write |
src/prompt.ts | The judge's TDD rules, ported from probity |
src/tdd.ts | Prompt building, session history, applying an Edit, the one-new-test check |
src/task.ts | The task classifier prompt and the guard-override rule |
src/bash.ts | Which source files a shell command would write |
src/config.ts | Settings, globs, which files are in scope |
src/eval.ts | The /tdd-eval runner |
test/ | Unit and hook tests, run by claude plugin test . |
eval/ | The regression cases |
fastPath to get it back).The TDD rules in src/prompt.ts are from probity by Nizar Selander, MIT License.
MIT, see LICENSE.
src/register.ts 140 lines1import type { EngineInterface, Register } from 'claude-code'
2import { bashWriteTargets } from './bash.ts'
3import { isInScope, readConfig, type Config } from './config.ts'
4import { runEval } from './eval.ts'
5import { classifyTask, decide, type Complete, type Pending } from './guard.ts'
6import { isGuardInstruction, type TaskKind } from './task.ts'
7import { applyEdit, toHistory, trimHistory, type FileContent, type HistoryEvent, type Verdict } from './tdd.ts'
8
9// $.ui.log refuses a line over 4096 characters; leave room for the header.
10const LOG_CHUNK = 3800
11
12export const register: Register = (on, options) => {
13 const config = readConfig(options)
14 if (!config.enabled) return
15 const inScope = (path: string) => isInScope(path, config)
16
17 // The kind of the task at hand, from the latest prompt. Starts as coding,
18 // also after a reload, so the guard is on until a prompt says otherwise.
19 let task: TaskKind = 'coding'
20 const judging = () => task === 'coding'
21
22 on('session.start', async ($, e, next) => {
23 for (const problem of config.problems) $.ui.log(`tdd-mod: ${problem}`)
24 await $.command.register({
25 name: 'tdd-eval',
26 description: 'Replays the tdd-mod regression cases through the configured models.',
27 argumentHint: '[runs per case]',
28 })
29 return next(e)
30 })
31
32 on('command.run', { command: 'tdd-eval' }, async ($, e) => {
33 const runs = Math.max(1, Math.min(10, Number.parseInt(e.args, 10) || 3))
34 const complete: Complete = request => $.model.complete(request)
35 const root = $.plugin.root
36 const text = await runEval(complete, file => $.fs.read(`${root}/eval/${file}`), config, runs)
37 return { text }
38 })
39
40 // Only what the person types changes the task; notifications and messages
41 // from other agents arrive as prompts too, and keep it.
42 on('prompt.submit', async ($, e, next) => {
43 if (!config.classifyTasks || e.origin.kind !== 'composer' || isGuardInstruction(e.text)) return next(e)
44 task = await classifyTask(request => $.model.complete(request), config, e.text, task)
45 $.ui.status(task === 'coding' ? 'TDD on' : 'TDD off (not coding)')
46 $.ui.log(`tdd-mod: task ${task}: ${e.text.slice(0, 80).replace(/\s+/g, ' ')}`, { to: 'debug' })
47 return next(e)
48 })
49
50 on('tool.call', { tool: 'Write' }, async ($, e, next) => {
51 if (!judging() || !inScope(e.file_path)) return next(e)
52 const before = await readBefore($, e.file_path)
53 const verdict = await judgeWrite($, config, e.agentId, before, { path: e.file_path, content: e.content })
54 return verdict.kind === 'pass' ? next(e) : { deny: denial(verdict) }
55 })
56
57 on('tool.call', { tool: 'Edit' }, async ($, e, next) => {
58 if (!judging() || !inScope(e.file_path)) return next(e)
59 const before = await readBefore($, e.file_path)
60 if (before.kind !== 'present') return next(e)
61 const after = applyEdit(before.content, e.old_string, e.new_string, e.replace_all)
62 // An edit that cannot apply changes nothing; the Edit tool reports the miss.
63 if (after === undefined) return next(e)
64 const verdict = await judgeWrite($, config, e.agentId, before, { path: e.file_path, content: after })
65 return verdict.kind === 'pass' ? next(e) : { deny: denial(verdict) }
66 })
67
68 // Shell writes bypass the TDD check: send them through Write/Edit instead.
69 on('tool.call', { tool: 'Bash' }, ($, e, next) => {
70 if (!judging()) return next(e)
71 const targets = bashWriteTargets(e.command, inScope)
72 if (targets.length === 0) return next(e)
73 $.ui.log(`tdd-mod: violation (bash write) ${targets.join(', ')}`, { to: 'debug' })
74 return {
75 deny:
76 `tdd-mod: this command writes to ${targets.join(', ')}. Create and change ` +
77 'source files with the Write or Edit tool, so the TDD check can judge the change.',
78 }
79 })
80
81 on('tool.call', { tool: 'NotebookEdit' }, async ($, e, next) => {
82 if (!judging() || !inScope(e.notebook_path)) return next(e)
83 const before = await readBefore($, e.notebook_path)
84 const mode = e.edit_mode ?? 'replace'
85 const content =
86 `Notebook cell edit (${mode}) on cell ${e.cell_id ?? '(first)'}` +
87 `${e.cell_type ? `, type ${e.cell_type}` : ''}. New cell source:\n\n${e.new_source}`
88 const verdict = await judgeWrite($, config, e.agentId, before, { path: e.notebook_path, content })
89 return verdict.kind === 'pass' ? next(e) : { deny: denial(verdict) }
90 })
91}
92
93async function readBefore($: EngineInterface, path: string): Promise<FileContent> {
94 try {
95 if (!(await $.fs.exists(path))) return { kind: 'absent' }
96 return { kind: 'present', content: await $.fs.read(path) }
97 } catch {
98 return { kind: 'unknown' }
99 }
100}
101
102async function judgeWrite(
103 $: EngineInterface,
104 config: Config,
105 agentId: string | undefined,
106 before: FileContent,
107 pending: Pending,
108): Promise<Verdict> {
109 const complete: Complete = request => $.model.complete(request)
110 const decision = await decide(complete, config, () => recentHistory($, agentId), before, pending)
111 decision.opinions.forEach((verdict, i) =>
112 log($, verdict, pending, decision.fastPath ? 'fast path: one new test' : i === 0 ? 'first opinion' : 'second opinion'),
113 )
114 if (decision.verdict.kind === 'violation' && decision.prompt) logPrompt($, decision.prompt, pending)
115 return decision.verdict
116}
117
118async function recentHistory($: EngineInterface, agentId: string | undefined): Promise<HistoryEvent[]> {
119 const messages = await $.session.messages(agentId === undefined ? {} : { agentId })
120 return trimHistory(Array.isArray(messages) ? toHistory(messages) : [])
121}
122
123function log($: EngineInterface, verdict: Verdict, pending: Pending, which: string): void {
124 const reason = verdict.reason ? ` — ${verdict.reason}` : ''
125 $.ui.log(`tdd-mod: ${verdict.kind} (${which}) ${pending.path}${reason}`, { to: 'debug' })
126}
127
128/** The blocked write's prompt, in numbered parts, so it can be replayed exactly. */
129function logPrompt($: EngineInterface, prompt: string, pending: Pending): void {
130 const parts = Math.ceil(prompt.length / LOG_CHUNK)
131 for (let i = 0; i < parts; i++) {
132 const part = prompt.slice(i * LOG_CHUNK, (i + 1) * LOG_CHUNK)
133 $.ui.log(`tdd-mod: blocked prompt ${pending.path} part ${i + 1}/${parts}:\n${part}`, { to: 'debug' })
134 }
135}
136
137function denial(verdict: Verdict): string {
138 return `tdd-mod: ${verdict.reason}`
139}
140src/bash.ts 84 lines1// Finds the source files a Bash command would write, so the guard can send
2// those writes through Write/Edit where the TDD check judges them. A
3// best-effort shell reading, not a parser: it covers the ways agents write
4// files from a shell (redirects, tee, sed -i, cp, dd, inline scripts). A
5// move (mv, git mv) keeps the code as it is, so it is no write here.
6
7// `<<EOF`, `<<-EOF`, `<<'EOF'`, `<<"EOF"`, then the body up to the line that ends it.
8const HEREDOC = /<<-?\s*(['"]?)(\w+)\1[^\n]*\n[\s\S]*?\n[ \t]*\2[ \t]*(?=\n|$)/g
9// `>`, `>>`, `>|`, `2>`, `&>` followed by a target; not `>&2` and not `->`/`=>`.
10const REDIRECT = /(?:^|[^-=<>])(?:\d|&)?>[>|]?(?!&)\s*(['"]?)([^\s;|&<>'"()]+)\1/g
11// Inline interpreters: Python open(..., 'w'|'a'), write_text/write_bytes, Node writeFile(Sync).
12const SCRIPT_WRITE = /open\([^)]*['"][wax]\+?b?['"]|\.write_(?:text|bytes)\(|writeFile(?:Sync)?\(|appendFile(?:Sync)?\(/
13const PATH_LITERAL = /['"`]([^'"`\s]+\.\w+)['"`]/g
14
15/**
16 * The simple commands of `shell` as words, split on unquoted `;`, `|`, `&`
17 * and newlines, quotes removed: `sed -i 's/a;b/c/' f` stays one command.
18 */
19function commands(shell: string): string[][] {
20 const result: string[][] = []
21 let words: string[] = []
22 let word = ''
23 let quote: string | undefined
24 const endWord = () => {
25 if (word !== '') words.push(word)
26 word = ''
27 }
28 const endCommand = () => {
29 endWord()
30 if (words.length > 0) result.push(words)
31 words = []
32 }
33 for (let i = 0; i < shell.length; i++) {
34 const c = shell[i] ?? ''
35 if (quote !== undefined) {
36 if (c === quote) quote = undefined
37 else if (c === '\\' && quote === '"') word += shell[++i] ?? ''
38 else word += c
39 } else if (c === "'" || c === '"') quote = c
40 else if (c === '\\') word += shell[++i] ?? ''
41 else if (c === ';' || c === '|' || c === '&' || c === '\n') endCommand()
42 else if (c === ' ' || c === '\t') endWord()
43 else word += c
44 }
45 endCommand()
46 return result
47}
48
49/** The source files `command` would write that `inScope` accepts, deduplicated, in order of appearance. */
50export function bashWriteTargets(command: string, inScope: (path: string) => boolean): string[] {
51 const targets: string[] = []
52 const shell = command.replace(HEREDOC, (match, _q, _tag) => match.slice(0, match.indexOf('\n')))
53
54 for (const m of shell.matchAll(REDIRECT)) targets.push(m[2] ?? '')
55
56 for (const [cmd, ...args] of commands(shell)) {
57 const operands = args.filter(a => !a.startsWith('-'))
58 switch (cmd) {
59 case 'tee':
60 targets.push(...operands)
61 break
62 case 'sed':
63 case 'perl':
64 if (args.some(a => /^-[a-zA-Z]*i/.test(a) || a.startsWith('--in-place'))) targets.push(...operands)
65 break
66 case 'cp':
67 case 'install':
68 if (operands.length > 0) targets.push(operands[operands.length - 1] ?? '')
69 break
70 case 'dd':
71 targets.push(...args.filter(a => a.startsWith('of=')).map(a => a.slice(3)))
72 break
73 }
74 }
75
76 // An inline script that writes files and names a source file: its target is
77 // often a variable, so every source path it names counts.
78 if (SCRIPT_WRITE.test(command)) {
79 for (const m of command.matchAll(PATH_LITERAL)) targets.push(m[1] ?? '')
80 }
81
82 return [...new Set(targets.filter(inScope))]
83}
84src/config.ts 114 lines1// The plugin's settings (plugin.json `userConfig`) as the hooks use them.
2import type { PluginOptions } from 'claude-code'
3import { DEFAULT_TEST_PATTERNS, type TestPatterns } from './tdd.ts'
4
5export type Config = {
6 enabled: boolean
7 judgeModel: string
8 classifierModel: string
9 classifyTasks: boolean
10 fastPath: boolean
11 secondOpinion: boolean
12 extensions: ReadonlySet<string>
13 ignore: readonly RegExp[]
14 testPatterns: TestPatterns
15 /** Settings that could not be used, each with what was wrong; the mod shows them at session start. */
16 problems: readonly string[]
17}
18
19function list(value: unknown): string[] {
20 return typeof value === 'string' ? value.split(',').map(s => s.trim()).filter(Boolean) : []
21}
22
23function text(value: unknown, fallback: string): string {
24 return typeof value === 'string' && value.trim() !== '' ? value.trim() : fallback
25}
26
27function flag(value: unknown, fallback: boolean): boolean {
28 return typeof value === 'boolean' ? value : fallback
29}
30
31/** `**` any run of folders, `*` any part of one name, `?` one character. */
32export function globToRegExp(glob: string): RegExp {
33 let source = ''
34 for (let i = 0; i < glob.length; i++) {
35 const c = glob[i] ?? ''
36 if (c === '*' && glob[i + 1] === '*') {
37 const slash = glob[i + 2] === '/'
38 source += slash ? '(?:.*/)?' : '.*'
39 i += slash ? 2 : 1
40 } else if (c === '*') source += '[^/]*'
41 else if (c === '?') source += '[^/]'
42 else source += c.replace(/[.+^${}()|[\]\\]/g, '\\$&')
43 }
44 return new RegExp(`^${source}$`)
45}
46
47/**
48 * The built-in test patterns with the `testPatterns` setting applied: a JSON
49 * object from comma-separated extensions to a regex source, e.g.
50 * `{"ts,js": "\\bscenario\\(", "py": "^\\s*def should_"}`. Each entry replaces
51 * the built-in pattern for its extensions; an empty string removes it.
52 */
53export function readTestPatterns(value: unknown): { patterns: TestPatterns; problems: string[] } {
54 const patterns: Record<string, RegExp> = { ...DEFAULT_TEST_PATTERNS }
55 const problems: string[] = []
56 if (typeof value !== 'string' || value.trim() === '') return { patterns, problems }
57 let entries: unknown
58 try {
59 entries = JSON.parse(value)
60 } catch (error) {
61 problems.push(`testPatterns is not valid JSON (${(error as Error).message}); using the built-in patterns`)
62 return { patterns, problems }
63 }
64 if (typeof entries !== 'object' || entries === null || Array.isArray(entries)) {
65 problems.push('testPatterns must be a JSON object like {"ts,js": "regex"}; using the built-in patterns')
66 return { patterns, problems }
67 }
68 for (const [keys, source] of Object.entries(entries)) {
69 const extensions = list(keys).map(e => e.replace(/^\./, '').toLowerCase())
70 if (typeof source !== 'string') {
71 problems.push(`testPatterns["${keys}"] is not a string; kept the built-in pattern`)
72 continue
73 }
74 if (source === '') {
75 for (const ext of extensions) delete patterns[ext]
76 continue
77 }
78 try {
79 // g: count every declaration; m: ^ and $ match per line.
80 const pattern = new RegExp(source, 'gm')
81 for (const ext of extensions) patterns[ext] = pattern
82 } catch (error) {
83 problems.push(`testPatterns["${keys}"] is not a valid regex (${(error as Error).message}); kept the built-in pattern`)
84 }
85 }
86 return { patterns, problems }
87}
88
89/** The engine fills in the manifest's defaults; the fallbacks cover a test or a hand-edited value. */
90export function readConfig(options: PluginOptions): Config {
91 const tests = readTestPatterns(options.testPatterns)
92 return {
93 enabled: flag(options.enabled, true),
94 judgeModel: text(options.judgeModel, 'haiku'),
95 classifierModel: text(options.classifierModel, 'haiku'),
96 classifyTasks: flag(options.classifyTasks, true),
97 fastPath: flag(options.fastPath, true),
98 secondOpinion: flag(options.secondOpinion, true),
99 extensions: new Set(list(options.extensions).map(e => e.replace(/^\./, '').toLowerCase())),
100 ignore: list(options.ignore).map(globToRegExp),
101 testPatterns: tests.patterns,
102 problems: tests.problems,
103 }
104}
105
106/** Source files the guard judges: a listed extension, outside every ignored path. */
107export function isInScope(path: string, config: Config): boolean {
108 const posix = path.replace(/\\/g, '/')
109 if (config.ignore.some(glob => glob.test(posix))) return false
110 const name = posix.slice(posix.lastIndexOf('/') + 1)
111 const dot = name.lastIndexOf('.')
112 return dot > 0 && config.extensions.has(name.slice(dot + 1).toLowerCase())
113}
114src/eval.ts 84 lines1// /tdd-eval: replays the regression cases in eval/ through the same decisions
2// a session makes, with the configured models, and reports the hit rate.
3import type { Config } from './config.ts'
4import { classifyTask, decide, type Complete, type Pending } from './guard.ts'
5import { isGuardInstruction, type TaskKind } from './task.ts'
6import type { FileContent, HistoryEvent } from './tdd.ts'
7
8export type VerdictCase = {
9 name: string
10 expect: 'pass' | 'violation'
11 /** Why this is the right verdict, for whoever reads a failure. */
12 why: string
13 history: HistoryEvent[]
14 before: FileContent
15 pending: Pending
16}
17
18export type TaskCase = { text: string; previous: TaskKind; expect: TaskKind }
19
20const CONCURRENCY = 6
21
22async function pool<T>(jobs: (() => Promise<T>)[], size: number): Promise<T[]> {
23 const results: T[] = []
24 let next = 0
25 const worker = async () => {
26 while (next < jobs.length) {
27 const i = next++
28 results[i] = await (jobs[i] as () => Promise<T>)()
29 }
30 }
31 await Promise.all(Array.from({ length: Math.min(size, jobs.length) }, worker))
32 return results
33}
34
35/** Reads a file of eval/ by name; the hooks module passes `$.fs.read` in. */
36export type ReadEvalFile = (file: string) => Promise<string>
37
38function row(hits: number, runs: number, label: string): string {
39 return `${hits === runs ? '✓' : '✗'} ${hits}/${runs} ${label}`
40}
41
42export async function runEval(complete: Complete, read: ReadEvalFile, config: Config, runs: number): Promise<string> {
43 const verdictCases = JSON.parse(await read('verdicts.json')) as VerdictCase[]
44 const taskCases = JSON.parse(await read('tasks.json')) as TaskCase[]
45 const lines: string[] = []
46
47 const verdictJobs = verdictCases.flatMap(c =>
48 Array.from({ length: runs }, () => () => decide(complete, config, async () => c.history, c.before, c.pending)),
49 )
50 const decisions = await pool(verdictJobs, CONCURRENCY)
51 let verdictHits = 0
52 lines.push(
53 `Verdicts — model ${config.judgeModel}, fast path ${config.fastPath ? 'on' : 'off'}, ` +
54 `second opinion ${config.secondOpinion ? 'on' : 'off'}, ${runs} runs each`,
55 )
56 verdictCases.forEach((c, i) => {
57 const mine = decisions.slice(i * runs, (i + 1) * runs)
58 const hits = mine.filter(d => d.verdict.kind === c.expect).length
59 verdictHits += hits
60 lines.push(row(hits, runs, `expect ${c.expect.padEnd(9)} ${c.name}`))
61 const miss = mine.find(d => d.verdict.kind !== c.expect)
62 if (miss) lines.push(` got ${miss.verdict.kind}: ${miss.verdict.reason.slice(0, 160)}`)
63 })
64
65 const taskJobs = taskCases.flatMap(c =>
66 Array.from({ length: runs }, () => async () =>
67 isGuardInstruction(c.text) ? c.previous : classifyTask(complete, config, c.text, c.previous),
68 ),
69 )
70 const kinds = await pool(taskJobs, CONCURRENCY)
71 let taskHits = 0
72 lines.push('', `Tasks — model ${config.classifierModel}, ${runs} runs each`)
73 taskCases.forEach((c, i) => {
74 const hits = kinds.slice(i * runs, (i + 1) * runs).filter(k => k === c.expect).length
75 taskHits += hits
76 lines.push(row(hits, runs, `expect ${c.expect.padEnd(6)} (after ${c.previous}) ${c.text.slice(0, 70)}`))
77 })
78
79 const verdictTotal = verdictCases.length * runs
80 const taskTotal = taskCases.length * runs
81 lines.push('', `Verdicts ${verdictHits}/${verdictTotal}, tasks ${taskHits}/${taskTotal}`)
82 return lines.join('\n')
83}
84src/guard.ts 72 lines1// The guard's two decisions, shared by the session's hooks and /tdd-eval:
2// which kind of task a prompt starts, and whether a write follows TDD.
3import type { ModelCompleteRequest, ModelCompleteResult } from 'claude-code'
4import type { Config } from './config.ts'
5import { buildTaskPrompt, parseTaskKind, type TaskKind } from './task.ts'
6import { addsExactlyOneTest, buildPrompt, parseVerdict, type FileContent, type HistoryEvent, type Verdict } from './tdd.ts'
7
8const TIMEOUT_MS = 120_000
9
10export type Pending = { path: string; content: string }
11
12/** A model call; the hooks module passes `$.model.complete` in, since `$` cannot cross an import. */
13export type Complete = (request: ModelCompleteRequest) => Promise<ModelCompleteResult>
14
15/** The task kind after `text`; a failed call counts as coding, so the guard stays on. */
16export async function classifyTask(
17 complete: Complete,
18 config: Config,
19 text: string,
20 previous: TaskKind,
21): Promise<TaskKind> {
22 const reply = await complete({
23 model: config.classifierModel,
24 prompt: buildTaskPrompt(text, previous),
25 maxTokens: 10,
26 timeoutMs: TIMEOUT_MS,
27 })
28 return reply.isAnswered ? parseTaskKind(reply.text) : 'coding'
29}
30
31export type Decision = {
32 verdict: Verdict
33 /** Every verdict in order: the fast path's pass alone, or one or two model verdicts. */
34 opinions: Verdict[]
35 fastPath: boolean
36 prompt?: string
37}
38
39/**
40 * The guard's decision on one write, as a session makes it and /tdd-eval
41 * replays it: the fast path, then the verdict model, then a second opinion.
42 */
43export async function decide(
44 complete: Complete,
45 config: Config,
46 history: () => Promise<HistoryEvent[]>,
47 before: FileContent,
48 pending: Pending,
49): Promise<Decision> {
50 // Adding one test is the red step itself.
51 if (config.fastPath && addsExactlyOneTest(pending.path, before, pending.content, config.testPatterns)) {
52 const verdict: Verdict = { kind: 'pass', reason: '' }
53 return { verdict, opinions: [verdict], fastPath: true }
54 }
55 const prompt = buildPrompt(await history(), before, pending)
56 const first = await ask(complete, config, prompt)
57 if (first.kind === 'pass' || !config.secondOpinion) return { verdict: first, opinions: [first], fastPath: false, prompt }
58 // A single validator call misfires now and then: block only when a second
59 // opinion agrees. Still fail-closed: two failed calls block.
60 const second = await ask(complete, config, prompt)
61 return { verdict: second, opinions: [first, second], fastPath: false, prompt }
62}
63
64async function ask(complete: Complete, config: Config, prompt: string): Promise<Verdict> {
65 const reply = await complete({ model: config.judgeModel, prompt, timeoutMs: TIMEOUT_MS })
66 // Fail-closed: no answer from the validator counts as a violation.
67 return reply.isAnswered
68 ? parseVerdict(reply.text)
69 : { kind: 'violation', reason: `the TDD validator gave no answer (${reply.reason}); retry the write.` }
70}
71
72src/task.ts 48 lines1// Which kind of task a prompt asks for. The guard judges writes only while
2// the task is coding; for anything else (moving or renaming code, docs,
3// config, questions) nothing is blocked.
4
5export type TaskKind = 'coding' | 'other'
6
7const MAX_PROMPT_CHARS = 4000
8
9/** The classifier prompt; `previous` decides a follow-up like "continue". */
10export function buildTaskPrompt(prompt: string, previous: TaskKind): string {
11 const text = prompt.length > MAX_PROMPT_CHARS ? `${prompt.slice(0, MAX_PROMPT_CHARS)}\n[truncated]` : prompt
12 return `Classify a request a developer gave a coding agent. Answer with exactly one label:
13
14coding: adds, changes or fixes behaviour — a feature, a bug fix, a kata or exercise
15 step, implementing something, writing tests for new behaviour.
16other: anything else — moving or renaming files, classes or functions, extracting or
17 inlining code, reorganising folders, formatting, dependencies, config, docs,
18 questions, reviews, git work.
19
20When a request mixes coding with other work, answer coding.
21The previous request was labelled ${previous}. A follow-up that names no task of its
22own gets the previous label: ${previous}. Follow-ups are "continue", "go on", "retry",
23"yes" and answers to the agent's question.
24
25Request:
26<request>
27${text}
28</request>
29
30Label:`
31}
32
33// Overriding the guard ("let it through", "I'm overriding the TDD check") is
34// about the task at hand, not a new one; haiku labels it other, which would
35// switch the guard off until the next coding request.
36const GUARD_OVERRIDE = /\boverrid(?:e|es|ing)\b|\blet (?:it|this|that|the \w+) through\b|\b(?:tdd|guard)\b/i
37
38/** Whether `prompt` speaks to the guard, so the task kind stays as it is. */
39export function isGuardInstruction(prompt: string): boolean {
40 return GUARD_OVERRIDE.test(prompt)
41}
42
43/** The kind the reply names; an unclear reply counts as coding, so the guard stays on when in doubt. */
44export function parseTaskKind(reply: string): TaskKind {
45 const match = /\b(coding|other)\b/i.exec(reply)
46 return match?.[1]?.toLowerCase() === 'other' ? 'other' : 'coding'
47}
48src/tdd.ts 200 lines1// Pure logic for the TDD guard: no `$`, so the tests can drive it directly.
2// Shapes follow probity's enforceTdd, trimHistory, applyEdit and toVerdict.
3import type { SessionMessage } from 'claude-code'
4import { DEFAULT_TDD_RULES, PROCESS_INSTRUCTIONS, RESPONSE_SPEC } from './prompt.ts'
5
6export const MAX_EVENTS = 10
7export const MAX_CONTENT_CHARS = 6000
8
9export type FileContent =
10 | { kind: 'present'; content: string }
11 | { kind: 'absent' }
12 | { kind: 'unknown' }
13
14export type HistoryEvent =
15 | { kind: 'prompt'; text: string }
16 | { kind: 'tool'; tool: string; input: unknown; output: string }
17
18export type Verdict = { kind: 'pass' | 'violation'; reason: string }
19
20// Test declarations per language, by extension. A regex stand-in for
21// probity's ast-grep matchers: the module has no Node to run ast-grep.
22// The `testPatterns` setting overrides or extends them per extension.
23export type TestPatterns = Readonly<Record<string, RegExp>>
24
25const JS_TEST = /\b(?:it|test)(?:\.(?:only|skip|todo|concurrent))?\s*\(\s*['"`]/g
26export const DEFAULT_TEST_PATTERNS: TestPatterns = {
27 ts: JS_TEST, tsx: JS_TEST, mts: JS_TEST, cts: JS_TEST,
28 js: JS_TEST, jsx: JS_TEST, mjs: JS_TEST, cjs: JS_TEST,
29 py: /^[ \t]*(?:async[ \t]+)?def[ \t]+test\w*[ \t]*\(/gm,
30 go: /^func[ \t]+Test\w*[ \t]*\(/gm,
31 java: /@Test\b/g, kt: /@Test\b/g, kts: /@Test\b/g,
32 rs: /#\[(?:tokio::)?test\]/g,
33 rb: /^[ \t]*(?:it[ \t]+['"]|def[ \t]+test_)/gm,
34 cs: /\[(?:Fact|Theory|Test|TestMethod)\b/g,
35}
36
37function countTests(pattern: RegExp, text: string): number {
38 return text.match(pattern)?.length ?? 0
39}
40
41/**
42 * The fast path: a write that adds exactly one test declaration is the red
43 * step itself and passes without a model call. Undefined content before the
44 * write, or a language without a pattern, never qualifies.
45 */
46export function addsExactlyOneTest(
47 path: string,
48 before: FileContent,
49 after: string,
50 patterns: TestPatterns = DEFAULT_TEST_PATTERNS,
51): boolean {
52 if (before.kind === 'unknown') return false
53 const pattern = patterns[path.slice(path.lastIndexOf('.') + 1).toLowerCase()]
54 if (pattern === undefined) return false
55 const beforeText = before.kind === 'present' ? before.content : ''
56 return countTests(pattern, after) - countTests(pattern, beforeText) === 1
57}
58
59/**
60 * The file as the Edit leaves it, or undefined when the edit cannot apply
61 * (the Edit tool itself then reports the miss).
62 */
63export function applyEdit(
64 current: string,
65 oldString: string,
66 newString: string,
67 replaceAll = false,
68): string | undefined {
69 const occurrences = current.split(oldString).length - 1
70 if (oldString === '' || occurrences === 0) return undefined
71 if (occurrences > 1 && !replaceAll) return undefined
72 // A replacer keeps $&, $$ and friends in newString literal.
73 const replacer = () => newString
74 return replaceAll
75 ? current.replaceAll(oldString, replacer)
76 : current.replace(oldString, replacer)
77}
78
79/** Prompts and answered tool calls, oldest first, as the validator reads them. */
80export function toHistory(messages: readonly SessionMessage[]): HistoryEvent[] {
81 const events: HistoryEvent[] = []
82 for (const message of messages) {
83 if (message.role === 'user' && message.text.trim() !== '') {
84 events.push({ kind: 'prompt', text: message.text })
85 }
86 for (const use of message.toolUses) {
87 if (use.text === undefined) continue
88 events.push({ kind: 'tool', tool: use.tool, input: use.input, output: use.text })
89 }
90 }
91 return events
92}
93
94export function trimHistory(
95 events: readonly HistoryEvent[],
96 maxEvents = MAX_EVENTS,
97 maxContentChars = MAX_CONTENT_CHARS,
98): HistoryEvent[] {
99 return events.slice(-maxEvents).map(event =>
100 event.kind === 'prompt'
101 ? { ...event, text: clip(event.text, maxContentChars) }
102 : { ...event, output: clip(event.output, maxContentChars) },
103 )
104}
105
106function clip(s: string, max: number): string {
107 if (s.length <= max) return s
108 const half = Math.floor(max / 2)
109 const omitted = s.length - 2 * half
110 return `${s.slice(0, half)}\n[${omitted} more characters truncated]\n${s.slice(s.length - half)}`
111}
112
113function formatEvent(event: HistoryEvent): string {
114 if (event.kind === 'prompt') return `User: ${event.text}`
115 const input = typeof event.input === 'string' ? event.input : JSON.stringify(event.input)
116 return `${event.tool}(${input}) → ${event.output}`
117}
118
119function formatBefore(before: FileContent): string {
120 switch (before.kind) {
121 case 'present':
122 return before.content
123 case 'absent':
124 return '(file does not exist)'
125 case 'unknown':
126 return '(current file content unavailable)'
127 }
128}
129
130export function buildPrompt(
131 history: readonly HistoryEvent[],
132 before: FileContent,
133 action: { path: string; content: string },
134): string {
135 const sections = [PROCESS_INSTRUCTIONS, DEFAULT_TDD_RULES]
136 if (history.length > 0) {
137 sections.push(`## Recent session\n\n${history.map(formatEvent).join('\n')}`)
138 }
139 sections.push(`## Current file content\n\n${formatBefore(before)}`)
140 sections.push(`## Pending action\n\nFile: ${action.path}\n\n${action.content}`)
141 sections.push(RESPONSE_SPEC)
142 return sections.join('\n\n')
143}
144
145/** Fail-closed: anything but a well-formed verdict is a violation. */
146export function parseVerdict(text: string): Verdict {
147 const parsed = safeParse(text.trim()) ?? safeParse(stripFence(text)) ?? findEmbeddedObject(text)
148 if (
149 typeof parsed === 'object' && parsed !== null &&
150 'kind' in parsed && (parsed.kind === 'pass' || parsed.kind === 'violation') &&
151 'reason' in parsed && typeof parsed.reason === 'string'
152 ) {
153 return { kind: parsed.kind, reason: parsed.reason }
154 }
155 return {
156 kind: 'violation',
157 reason: `could not parse verdict from validator output: ${text.slice(0, 500)}`,
158 }
159}
160
161function safeParse(text: string): unknown {
162 try {
163 return JSON.parse(text)
164 } catch {
165 return undefined
166 }
167}
168
169function stripFence(text: string): string {
170 return text.replace(/^\s*```(?:json)?\s*/, '').replace(/\s*```\s*$/, '').trim()
171}
172
173/** The last balanced `{...}` that parses: models sometimes reason before answering. */
174function findEmbeddedObject(text: string): unknown {
175 for (let start = text.lastIndexOf('{'); start >= 0; start = text.lastIndexOf('{', start - 1)) {
176 const span = scanBalanced(text, start)
177 const parsed = span === undefined ? undefined : safeParse(span)
178 if (parsed !== undefined) return parsed
179 if (start === 0) break
180 }
181 return undefined
182}
183
184function scanBalanced(text: string, start: number): string | undefined {
185 let depth = 0
186 let inString = false
187 let escape = false
188 for (let i = start; i < text.length; i++) {
189 const c = text[i]
190 if (escape) escape = false
191 else if (inString) {
192 if (c === '\\') escape = true
193 else if (c === '"') inString = false
194 } else if (c === '"') inString = true
195 else if (c === '{') depth++
196 else if (c === '}' && --depth === 0) return text.slice(start, i + 1)
197 }
198 return undefined
199}
200src/prompt.ts 163 lines1// The validator prompt is ported verbatim from probity's enforceTdd rule
2// (https://github.com/nizos/probity, src/rules/enforce-tdd.ts).
3// Copyright (c) Nizar Selander, MIT License.
4// Changed here: RESPONSE_SPEC asks for a reason on "pass" too, so every
5// verdict in the debug log says why.
6
7export const PROCESS_INSTRUCTIONS = `## Role
8
9You are a TDD validator. Judge whether the pending write follows
10test-driven development.
11
12## Inputs
13
14You will see three inputs:
15
161. "Recent session" — a chronological log of the agent's recent prompts
17 and tool actions. Each entry shows what the agent did and what it
18 observed back. Use this to find evidence of a failing test that the
19 pending write would address.
202. "Current file content" — what's on disk right now at the file the
21 agent is about to write. May be a parenthesized marker (e.g.
22 \`(file does not exist)\`) when content cannot be shown.
233. "Pending action" — what the agent is about to write. Content may be
24 raw file text or a patch/diff in any common format.
25
26## What you judge
27
28Judge the change this write makes (the difference between the current
29file content and the pending action), not the resulting file as a
30whole.
31
32A transient file state is never itself a violation, however broken the
33file looks: an unresolved symbol, a dead or unused definition, a
34duplicated declaration, a reference to a removed name, a half-finished
35multi-step change. Whether the file is internally consistent or runs
36after the write is checked when the agent next runs the tests, not by
37you. This allowance is about structure; it does not excuse skipping a
38failing test or over-implementing, which the rules below still catch.
39
40A block or denial message recorded earlier in the session is a past
41verdict, not a rule. Re-derive your judgment from the rules below as
42if it had not been issued; never block only because a previous attempt
43was blocked.
44
45Your judgment is not the final word. When the user tells you in the
46session to let this change through, treat it as authoritative and pass,
47even on a change you would otherwise block.
48
49## Multi-step changes
50
51A phase may span multiple writes, each fine on its own. For example:
52
53 - Add an import in one write, then change the calling code in the
54 next.
55 - Move a function in two writes (remove from one location, add at
56 another).
57 - Add a function signature in one write, then its body in the next.
58 - Remove a function in one write, then its call sites in the next.`
59
60export const DEFAULT_TDD_RULES = `## TDD rules
61
62The TDD cycle is Red -> Green -> Refactor. Each phase has its own rules.
63
64### Across all phases
65
66Deleting code, tests, or helpers never requires a failing test to drive
67it, even when the removed code was used or test-covered.
68
69### Red phase: write a failing test first
70
71In red you add one test that drives new behavior and expect it to
72fail when the suite runs.
73
74A single write should add at most one new test. Compare current file
75content with pending action to count newly added tests; existing
76tests do not count. Restructuring existing tests is not "adding".
77
78 - Adding a test is the red step itself; it is allowed and does not
79 require observing it fail first, unless the prior green left a
80 refactor unmade (see "Enforcing the refactor phase").
81 - A test added to drive new behavior must be observed failing for
82 the right reason (an assertion, not a syntax or import error) in
83 a prior test run before production code may be written to satisfy
84 it.
85 - A test added to capture existing behavior is allowed to pass
86 immediately and must not be blocked for not failing first.
87 Examples: characterization tests pinning current implementation,
88 tests at a new layer (e.g. an e2e covering code already exercised
89 by units), pinning tests added before a refactor pulls a seam out
90 from under them.
91 - Test-file scaffolding edits (imports, helpers, fixtures) need no
92 failing test on their own.
93
94#### Reaching a clean red
95
96A test can fail before reaching an assertion (import unresolved,
97signature mismatch). The agent may resolve these without violating TDD:
98
99 - Import or symbol unresolved -> create a placeholder stub: a body
100 that makes the symbol exist but does not implement the behavior the
101 test asserts. Returning a literal that contradicts the assertion
102 (e.g. \`=> 0\` when the test expects \`1\`) or throwing
103 \`not implemented\` are both valid stubs; they exist solely to
104 surface a real assertion failure on the next test run.
105 - Signature mismatch -> adjust signature; keep the body as a
106 placeholder stub per the rule above.
107 - Assertion failure -> implement minimal logic to pass.
108
109A stub-resolution step must not implement the test's asserted behavior.
110Returning a literal that the assertion will reject IS a stub.
111
112### Green phase: minimum to pass
113
114The implementation must not exceed the minimum needed to make the
115observed failing test pass. Functions, classes, or branches not
116required by the currently failing test are over-implementation. An
117import or other scaffolding awaiting a later write in the same change
118is a transient state, not over-implementation.
119
120### Refactor phase: improve structure under green
121
122Refactoring does not require a failing test to drive it. Production
123and test edits that preserve observable behavior are allowed when
124the relevant tests were passing before the refactor began. Examples:
125
126 - Extracting helpers whose behavior already lives elsewhere (covered
127 by existing tests). Extracting a helper whose behavior appears
128 nowhere else is net new and requires a failing test first.
129 - Lifting test setup (fixtures, builders, factories) into a
130 dedicated or reusable helper. The helper is exercised by the tests
131 that call it; no separate test for the helper is required.
132 - Adding type declarations, interfaces, or constant literals
133 (no runtime behavior by construction).
134 - Renaming, restructuring control flow, removing dead code.
135 - Reorganizing or deleting redundant tests, or splitting/combining
136 existing tests. The one-new-test rule is about intent to add
137 behavior, not surface diff count.
138
139#### Enforcing the refactor phase
140
141Writing the next test crosses the green->red boundary, so judge a new
142red here too, not only production writes. Refactor is part of the cycle:
143when the prior green left one that is unmistakable and clearly without
144downside, block and name it. The bar is high because forcing a refactor
145risks needless abstraction, indirection, or breaking conventions the
146agent can see and you cannot, so when the win is not clear-cut, let green
147stand.
148
149## Validator behavior
150
151### Block messages
152
153Name the violation and say why it breaks TDD. Do not dictate edit
154order, require the file to be complete or runnable, or demand other
155steps be bundled into this write.`
156
157export const RESPONSE_SPEC = `## Response format
158
159Respond with a single JSON object of exactly this shape:
160{"kind":"pass"|"violation","reason":"<short explanation>"}
161Always give a one-sentence reason, on "pass" too: which TDD step this write is and why it is allowed.
162Return JSON only. No prose, no code fences.`
163