Stop-hook correctness gate. Scores every answer's factual claims against what the session actually ran, blocks the turn on anything unbacked, and records each…

A Stop hook that blocks an answer when it claims something the session never checked. The dev-core rules already say "no claim without a file:line, a command output, or a fetched URL". This plugin is the part that enforces it. Before it existed, every correctness mechanism in the repo was advice given before the model acted, and nothing looked at the answer.
It exists for two failures the user reported: answers asserting things nobody checked, and steps skipped mid-task (a "tests pass" with no test run).
The hook reads the session transcript, builds a list of what actually ran, and matches the answer against it. Each claim class is on or off in CONFIG in hooks/claim-patterns.cjs.
| Class | Fires on | Backed by | State |
|---|---|---|---|
path-missing | a file path | a tool having printed the file's name this session, or the file existing on this machine: from cwd, the git root, the tracked-file list, or the repo of any folder the answer names by ~/ or / path | note: shown to you, never sent back (pathMissingBlocks: false) |
command-outcome | "tests pass", "build succeeds", "it works" | a clean test/build/lint run after the last file write | on |
url | a cited URL | a tool having printed that URL, WebFetch or WebSearch of that host, or an agent-browser, curl or wget command naming it | on |
path | a file path | a Read/Edit/codegraph of it | off |
version | v1.2, 1.2.3, "version 4" | find_libs, a manifest read, opensrc | off |
absence | "there is no X" | a Grep/Glob/rg/codegraph search | off |
Guesses stay legal. The gate checks labelling, not certainty: "I'd use Postgres" passes, and "Postgres handles this natively" blocks unless something was fetched. A span in "double quotes" or after "e.g." counts as a mention, not a claim. A path-missing block asks the model to correct the path, not delete it. After 3 blocks on the same claims, the hook stops blocking and tells the model to mark them unverified.
SubagentStop runs the same checks on subagent answers, against the subagent's own transcript.
~/.claude/verified/, deliberately outside the plugin data dir so a reinstall never wipes it. It holds ledger.jsonl (one node per evaluated turn, with answer text, so it never goes in git), manifests/, worlds.lock.json (the pinned replay history), replay-log.jsonl, offsets.json and sessions/ (each session's last verdict and recent blocks, so the Stop hook never reads the ledger). A block's resolution is appended to the ledger as an { "op": "resolve" } line and merged on read.
Any change to what gets flagged is scored before it ships. /verified-replay replays a candidate config over the pinned history of real sessions, with no model calls and nothing re-run:
node hooks/replay.cjs --corpus # current config on the pinned history
node hooks/replay.cjs --corpus --sweep # every single-knob candidate
node hooks/replay.cjs --candidate <file> # a forked claim-patterns.cjs vs current
The score is V = catches − false positives − 0.25 × blocked turns. Sessions are split by a hash of their id, about 70% train and 30% held-out test (hooks/split.cjs), and every run prints V on each side with a 95% bootstrap interval over test sessions. A candidate gets one verdict:
| Verdict | When |
|---|---|
REJECT | V drops on train or on test |
SHIP | no session scores differently, or both sides gain and the test gain's interval stays above zero |
OVERFIT | train gains and test stays flat, including when no test session is affected |
NOISE | test gains, but its interval reaches zero |
Only read flags from train sessions when you design a rule. Reading test flags spends the held-out set. --sweep ranks candidates by train V for the same reason. This follows Dream-RSI (dream-rsi.com): the recorded history works as an exact simulator, the scorer (hooks/labels.cjs) stays fixed, and the current config is always one of the candidates. Don't tune labels.cjs to make a favoured config win. Change it only when you've looked at a labelled flag and shown it wrong.
hooks/gold.cjs checks the labeller against hand labels on real live blocks, drawn from train sessions only:
node hooks/gold.cjs --sample 50 # draw blocks into ~/.claude/verified/gold.json
node hooks/gold.cjs --page # write ~/.claude/verified/gold-review.html to label them
node hooks/gold.cjs --agreement # labeller vs the hand labels, per class
The page preselects each draft label. Change any you disagree with, download the file, and save it over gold.json. Both files carry answer text, so they stay in ~/.claude/verified/.
Scores per commit, the scorer's blind spots and the gap to the best possible score are in RESULTS.md.
A fix that only removes wrong blocks can lower V, because the labeller credited those blocks as catches:
On 2026-09-30 the user overrode the rule for three fixes whose removed blocks were all wrong:
| Fix | V with fix vs before |
|---|---|
Bash agent-browser/curl/wget URLs count as fetched | 371.6 vs 398.1 |
$param and (group) route folders stay part of a path | 369.8 vs 371.6 |
| A relative path is also looked up in the repos the answer names | 369.3 vs 369.8 |
On 2026-10-05 the user overrode it again to make path-missing a note. Replay rejects it (ledger V 69.3 to 42.8, corpus 297.5 to 92.0), but 111.8 of the 116.8 path catches on the ledger are the fixed credit on flags the live gate never issued. On the 49 path blocks with a recorded outcome, the score was 5 catches to 26 wrong blocks, and the hand labels in gold.json have 1 right block in 25. The tradeoff was taken knowingly: about 1 real catch per 25 path blocks is lost.
Overriding again needs the same evidence: every block the fix removes has to be a wrong block. Fixing the labeller so it can see these cases would end the need for overrides.
judge/, off unless VERIFIED_JUDGE_ENABLED=1. Replayed over 476 real turns it fired on 92.5% of them, and each call took 5 to 56 seconds (median about 38) with one call in five failing. The time goes to claude -p starting up, not to the model, so a smaller payload doesn't help. Cost was never the issue at about $0.0017 per turn. Haiku 4.5 got 5 out of 5 on the classification probe, so quality wasn't the problem either.claude -p --bare for the judge. It skips hooks and plugins, but it also breaks login, and the child exits with "Not logged in". The judge uses --settings judge/judge-settings.json plus --strict-mcp-config instead, and VERIFIED_JUDGE=1 stops the child's own Stop hook from recursing.path (was the file read this session). It caught 48 claims over 3260 turns, while blocking 202 the session had already printed. The existence check does the job without those false positives.version. It produced 2 catches out of 67 unresolved flags over 651 turns, the worst ratio of any class. Matching version digits can't tell whether a claim about a release is true.absence. It produced 66 flags over 3260 turns, and no signal ever confirmed one either way.path alongside path-missing. It adds 73 catches but also 179 blocked turns, which scores worse at the 0.25 block weight.command-outcome to a note. It would lose the only class with zero false positives.node hooks/replay.cjs), not only the old-transcript corpus. Since 2026-09-30 a path block there counts as a catch only if the rewrite kept the file.url become a non-blocking note? That would cut the block rate from 18.3% to 16.2%.node --test claude/plugins/verified/hooks/verify-stop.test.cjs
hooks/ui.tsx shows a toast and a band above the prompt listing the flagged claims whenever the Stop hook sends an answer back, since the block reason otherwise reaches only the model. The band clears on the next prompt or on Dismiss.
hooks/ui.tsx 54 lines1import { atom, read, update } from 'claude-code'
2import type { Register } from 'claude-code'
3
4import type { Verdict } from '../types'
5
6const last = atom({ plugin: 'verified', key: 'last' } as const, null as Verdict | null)
7
8// The Stop hook's block reason goes only to the model; this shows the person what was flagged.
9export const register: Register = on => {
10 on('classic.Stop', async ($, e, next) => {
11 const ran = await next(e)
12 if (ran.block?.startsWith('verified:')) {
13 const lines = ran.block.split('\n')
14 const claims = lines.filter(line => line.startsWith(' - [')).map(line => line.slice(4))
15 await update($, last, () => ({ headline: lines[0] ?? '', claims }))
16 $.ui.toast(`verified sent the answer back: ${claims.length} unbacked ${claims.length === 1 ? 'claim' : 'claims'}`)
17 }
18
19 return ran
20 })
21
22 on('prompt.submit', async ($, e, next) => {
23 await update($, last, () => null)
24
25 return next(e)
26 })
27
28 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
29 const verdict = await read($, last)
30 if (verdict === null || e.props.hasSurvey) return next(e)
31 const { Box, Button, Text } = $.ui.resolve(e)
32 const rest = await next(e)
33
34 return (
35 <Box flexDirection="column">
36 <Box flexDirection="column" borderStyle="round" borderColor="yellow" paddingX={1}>
37 <Box justifyContent="space-between">
38 <Text color="yellow" bold>
39 verified redid the last answer
40 </Text>
41 <Button key="dismiss" plain role="dismiss" label="Dismiss" onPress={() => update($, last, () => null)} />
42 </Box>
43 {verdict.claims.map((claim, i) => (
44 <Text key={`claim-${i}`} dimColor wrap="truncate-end">
45 {claim}
46 </Text>
47 ))}
48 </Box>
49 {rest}
50 </Box>
51 )
52 })
53}
54types/index.d.ts 8 lines1export type Verdict = { headline: string; claims: string[] }
2
3declare module 'claude-code' {
4 interface PluginState {
5 verified: { last: Verdict | null }
6 }
7}
8