After a turn that changed things, a forked check asks whether the goal was met, where it cut corners and what to do next, and draws the verdict above the prompt

A second opinion on every turn that did real work. When Claude finishes a turn that changed something, a forked copy of the conversation asks four questions about it:
yes, partly, no)The answer is drawn above the prompt. A clean result takes one line:
✓ Supervisor: Solved · Add retry with backoff to the API client
[ Run the full test suite ] [ Dismiss ]
Anything else expands:
◐ Supervisor: Partly solved · fork: 48k cached, 0.2k new, 0.1k out
Goal: Add retry with backoff to the API client
Gaps
· No retry on 429 responses
Shortcuts taken
· Tests were written but never run
[ Run the client tests and fix failures ] [ Handle 429 with Retry-After ] [ Dismiss ]
Each next-step button sends that step as your next prompt (hotkeys n and m once the band has focus). Dismiss clears the verdict. A new turn clears it too.
None of this enters the main conversation: by $.model.fork's documented contract the fork's question and answer stay outside the transcript, so the main context does not grow. (In the first live run nothing from the fork reached the main conversation; see DECISIONS.md.)
| Command | Does |
|---|---|
/supervisor | Toggle on/off (saved across sessions) |
/supervisor on / off | Set it explicitly |
/supervisor now | Check the last turn right away, even when off or un-gated |
One $.model.fork per qualifying turn, and only then.
fork: 48k cached, 0.2k new, 0.1k out)./model, the fork pays full price for the whole transcript. The cached figure falls to near zero when that happens.{await next(e)} underneath, so the quiz (and anything else below it in the chain) still shows. Neither hides the other. Hotkeys are chosen not to clash: quiz-after uses 1-3, s, x; this mod uses n, m.When a survey holds the band, this mod yields completely and draws only what is below it.
claude --plugin-dir /path/to/next-steps-supervisor
Check it before loading:
claude plugin validate /path/to/next-steps-supervisor
claude plugin test /path/to/next-steps-supervisor
.claude-plugin/plugin.json manifest, types pointer
hooks/hooks.json { "modules": ["./register.tsx"] }
hooks/register.tsx gating, the fork, /supervisor, the band
hooks/verdict.ts fork prompt and a tolerant JSON parser
types/index.d.ts Verdict type and PluginState entry
tests/supervisor.test.tsx gating, final-reply quote, queued check, drawing with what is below, button submithooks/register.tsx 288 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { Phase, Verdict } from '../types'
5import { forkPrompt, isClean, parseVerdict } from './verdict'
6
7const verdictAtom = atom({ plugin: 'next-steps-supervisor', key: 'verdict' } as const, null as Verdict | null)
8const phaseAtom = atom({ plugin: 'next-steps-supervisor', key: 'phase' } as const, 'idle' as Phase)
9
10// A turn earns a check when it changed something, or ran long enough that
11// "did it actually finish?" is a real question.
12const MANY_TOOL_CALLS = 8
13
14// quiz-after takes 1-3, s and x in the same band; these stay clear of them.
15const NEXT_HOTKEYS = ['n', 'm']
16
17const short = (n: number) => (n >= 1000 ? `${+(n / 1000).toFixed(1)}k` : `${n}`)
18
19const isOn = async ($: EngineInterface) => (await $.store.get('isOn').catch(() => true)) !== false
20
21const say = ($: EngineInterface, line: string) => $.ui.log(`next-steps-supervisor: ${line}`, { to: 'debug' })
22
23// Per-turn bookkeeping. Module variables are fine: a reload mid-turn costs
24// one skipped check, and nothing draws from them.
25let turnSeq = 0
26let checkedSeq = -1
27let tools = 0
28let changes = 0
29let isChecking = false
30let lastAnswer = ''
31// A qualifying turn that ended while an older fork was still out; checked when it returns.
32let queuedSeq = -1
33
34// One fork. Resolves a line for /supervisor now; never rejects.
35const check = async ($: EngineInterface, isForced: boolean): Promise<string> => {
36 if (isChecking) return 'A supervisor check is already running.'
37 isChecking = true
38 const seq = turnSeq
39 checkedSeq = seq
40
41 try {
42 await update($, phaseAtom, () => 'checking')
43 const reply = await $.model.fork({ prompt: forkPrompt(lastAnswer) })
44
45 if (!reply.isAnswered) {
46 say($, `no verdict (${reply.reason})`)
47 return `No verdict: the fork did not answer (${reply.reason}).`
48 }
49
50 // A new turn started while the fork ran: this verdict is about stale work.
51 if (seq !== turnSeq) return 'Skipped: a new turn started.'
52 // Turned off while the fork ran: honour that for the automatic check.
53 if (!isForced && !(await isOn($))) return 'Skipped: the supervisor was turned off.'
54
55 const parsed = parseVerdict(reply.text)
56 if (parsed === null) {
57 say($, `unreadable reply: ${reply.text.slice(0, 120)}`)
58 return 'The fork answered, but not with a verdict this mod could read.'
59 }
60
61 const { usage } = reply
62 const cost = {
63 cached: usage.cache_read_input_tokens ?? 0,
64 input: usage.input_tokens + (usage.cache_creation_input_tokens ?? 0),
65 output: usage.output_tokens,
66 }
67 await update($, verdictAtom, () => ({ ...parsed, cost }))
68
69 return `Supervisor: ${parsed.solved === 'yes' ? 'solved' : parsed.solved === 'partly' ? 'partly solved' : 'not solved'}. Verdict above the prompt.`
70 } catch (err) {
71 say($, `check failed: ${err}`)
72 return 'No verdict: the check failed (see the debug log).'
73 } finally {
74 isChecking = false
75 await update($, phaseAtom, () => 'idle').catch(() => undefined)
76 if (queuedSeq === turnSeq && checkedSeq !== turnSeq) {
77 $.clock.after(0, () => {
78 void check($, false)
79 })
80 }
81 }
82}
83
84const dismiss = ($: EngineInterface) => () => {
85 void update($, verdictAtom, () => null).catch(() => undefined)
86}
87
88// Clear first so the band does not show a verdict about the turn being replaced.
89const send = ($: EngineInterface, text: string) => () => {
90 void (async () => {
91 try {
92 await update($, verdictAtom, () => null)
93 await $.prompt.submit({ text, asUser: true })
94 } catch (err) {
95 say($, `submit failed: ${err}`)
96 $.ui.toast('Could not send that next step (see the debug log)')
97 }
98 })()
99}
100
101const LABEL = { yes: 'Solved', partly: 'Partly solved', no: 'Not solved' } as const
102const COLOR = { yes: 'green', partly: 'yellow', no: 'red' } as const
103const ICON = { yes: '✓', partly: '◐', no: '✗' } as const
104
105export const register: Register = on => {
106 on('session.start', async ($, e, next) => {
107 const result = await next(e)
108
109 await $.command
110 .register({
111 name: 'supervisor',
112 description: 'Turn-end supervisor: on, off, or now (next-steps-supervisor)',
113 argumentHint: '[on|off|now]',
114 })
115 .catch(err => say($, `/supervisor not registered: ${err}`))
116
117 return result
118 })
119
120 // A /clear raises no session.start, only this; the old verdict must go with it.
121 on('session.end', async ($, e, next) => {
122 if (e.reason === 'clear') {
123 await update($, verdictAtom, () => null).catch(() => undefined)
124 }
125
126 return next(e)
127 })
128
129 on('turn.start', async ($, e, next) => {
130 turnSeq += 1
131 tools = 0
132 changes = 0
133 lastAnswer = ''
134 await update($, verdictAtom, () => null).catch(() => undefined)
135 // An older fork still out is about the replaced turn; stop saying "checking".
136 await update($, phaseAtom, () => 'idle').catch(() => undefined)
137
138 return next(e)
139 })
140
141 // Only the main loop's calls count: a subagent's work is judged through the
142 // main thread's account of it.
143 on('tool.call', async ($, e, next) => {
144 const result = await next(e)
145
146 if (e.agentId === undefined) {
147 tools += 1
148 if (result.deny === undefined && result.isReadOnly !== true) changes += 1
149 }
150
151 return result
152 })
153
154 on('turn.complete', async ($, e, next) => {
155 const result = await next(e)
156 if (e.agentId === undefined) lastAnswer = e.answer
157
158 const isWorthChecking =
159 e.agentId === undefined &&
160 e.reason === 'answer' &&
161 checkedSeq !== turnSeq &&
162 (changes > 0 || tools >= MANY_TOOL_CALLS)
163
164 if (!isWorthChecking || !(await isOn($))) return result
165
166 // The fork runs on a timer so the turn ends at once instead of waiting on it.
167 if (isChecking) {
168 queuedSeq = turnSeq
169 return result
170 }
171
172 $.clock.after(0, () => {
173 void check($, false)
174 })
175
176 return result
177 })
178
179 on('command.run', { command: 'supervisor' }, async ($, e) => {
180 const arg = e.args.trim().toLowerCase()
181 if (arg === 'now') return { text: await check($, true) }
182
183 const turnOn = arg === 'on' ? true : arg === 'off' ? false : arg === '' ? !(await isOn($)) : undefined
184 if (turnOn === undefined) return { text: 'Usage: /supervisor [on|off|now]' }
185
186 try {
187 await $.store.set('isOn', turnOn)
188 if (!turnOn) await update($, verdictAtom, () => null)
189 } catch (err) {
190 return { text: `next-steps-supervisor: could not save the setting (${err})` }
191 }
192
193 return {
194 text: turnOn
195 ? 'Supervisor is on: after a turn that changed something or ran 8+ tools, one forked check draws a verdict.'
196 : 'Supervisor is off: no forks until /supervisor on. /supervisor now still works.',
197 }
198 })
199
200 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
201 const verdict = await read($, verdictAtom)
202 const phase = await read($, phaseAtom)
203 // Whatever other mods (quiz-after, token-weather...) draw goes under ours.
204 const below = await next(e)
205
206 if (e.props.hasSurvey || (verdict === null && phase === 'idle')) return below
207
208 const { Box, Button, Text } = $.ui.resolve(e)
209
210 if (verdict === null) {
211 return (
212 <Box flexDirection="column">
213 <Text dimColor>Supervisor: checking whether that turn met the goal…</Text>
214 {below}
215 </Box>
216 )
217 }
218
219 const { cached, input, output } = verdict.cost
220 const color = COLOR[verdict.solved]
221 // One button per row: two long step labels and Dismiss on one row ran off the
222 // right edge in a normal-width window, pushing Dismiss out of view.
223 const buttons = (
224 <Box flexDirection="column">
225 {verdict.next.map((step, i) => (
226 <Button
227 key={`supervisor-next-${i}`}
228 hotkey={NEXT_HOTKEYS[i]}
229 label={step.length > 60 ? `${step.slice(0, 59)}…` : step}
230 variant={i === 0 ? 'primary' : undefined}
231 onPress={send($, step)}
232 />
233 ))}
234 <Button key="supervisor-dismiss" role="dismiss" label="Dismiss" dimColor onPress={dismiss($)} />
235 </Box>
236 )
237
238 if (isClean(verdict)) {
239 return (
240 <Box flexDirection="column">
241 <Text wrap="truncate-end">
242 <Text color={color} bold>
243 {ICON.yes} Supervisor: {LABEL.yes}
244 </Text>
245 <Text dimColor> · {verdict.goal}</Text>
246 </Text>
247 {buttons}
248 {below}
249 </Box>
250 )
251 }
252
253 const section = (title: string, items: string[], tint?: string) =>
254 items.length === 0 ? null : (
255 <Box flexDirection="column">
256 <Text bold color={tint}>
257 {title}
258 </Text>
259 {items.map((item, i) => (
260 <Text key={`supervisor-${title}-${i}`} wrap="wrap">
261 {' '}· {item}
262 </Text>
263 ))}
264 </Box>
265 )
266
267 return (
268 <Box flexDirection="column">
269 <Text wrap="truncate-end">
270 <Text color={color} bold>
271 {ICON[verdict.solved]} Supervisor: {LABEL[verdict.solved]}
272 </Text>
273 <Text dimColor>
274 {' '}
275 · fork: {short(cached)} cached, {short(input)} new, {short(output)} out
276 </Text>
277 </Text>
278 {verdict.goal === '' ? null : <Text wrap="wrap">Goal: {verdict.goal}</Text>}
279 {section('Gaps', verdict.gaps, 'yellow')}
280 {section('Shortcuts taken', verdict.shortcuts, 'magenta')}
281 {section('Needs your OK', verdict.needsApproval, 'red')}
282 {buttons}
283 {below}
284 </Box>
285 )
286 })
287}
288hooks/verdict.ts 82 lines1import type { Solved, Verdict } from '../types'
2
3// Past this the final reply costs more full-price input than it adds; the
4// claims a supervisor checks ("tests pass", "done") sit near its end anyway.
5const ANSWER_CHARS = 4000
6
7// The fork replays the main thread's last request, and the final reply is that
8// request's response, so it is most likely not in what the fork sees. Quote it,
9// or the supervisor never reads the claims it is meant to check. No tools:
10// every fork tool is denied.
11export const forkPrompt = (answer: string) =>
12 [
13 'Pause the work. You are now a supervisor reviewing the conversation above, not continuing it.',
14 'Do not call tools. Judge the most recent task against what the user originally asked for.',
15 ...(answer.trim() === ''
16 ? []
17 : [
18 '',
19 'The assistant\'s final reply to the user for this task (it may not appear above) was:',
20 '<final-reply>',
21 answer.length > ANSWER_CHARS ? `…${answer.slice(-ANSWER_CHARS)}` : answer,
22 '</final-reply>',
23 ]),
24 '',
25 'Answer with one JSON object and nothing else:',
26 '{"goal": string, "solved": "yes" | "partly" | "no", "gaps": string[], "shortcuts": string[], "needsApproval": string[], "next": string[]}',
27 '',
28 '- goal: the user\'s original goal for this task, in one sentence.',
29 '- solved: did the work as it stands actually achieve that goal?',
30 '- gaps: parts of the goal still not met.',
31 '- shortcuts: where the work was lazy: verification skipped, tests not run, code stubbed or left TODO, a claim made without checking, a simpler fix chosen over the correct one.',
32 '- needsApproval: things done or proposed that the user should have approved first (destructive commands, pushes, new dependencies, work outside the request).',
33 '- next: at most 2 next steps, each written as a prompt the user could send as is.',
34 '',
35 'Use empty arrays when there is nothing to say. Be specific and brief: each item under 20 words. Do not flatter.',
36 ].join('\n')
37
38const SOLVED: readonly Solved[] = ['yes', 'partly', 'no']
39
40const clip = (s: string, n: number) => (s.length > n ? `${s.slice(0, n - 1)}…` : s)
41
42const strings = (value: unknown, max: number) =>
43 Array.isArray(value)
44 ? value
45 .filter((v): v is string => typeof v === 'string' && v.trim() !== '')
46 .slice(0, max)
47 .map(v => clip(v.trim(), 240))
48 : []
49
50// Models wrap JSON in fences or a sentence often enough that a strict parse
51// would drop good verdicts; take the outermost braces and validate the shape.
52export const parseVerdict = (text: string): Omit<Verdict, 'cost'> | null => {
53 const start = text.indexOf('{')
54 const end = text.lastIndexOf('}')
55 if (start === -1 || end <= start) return null
56
57 let raw: unknown
58 try {
59 raw = JSON.parse(text.slice(start, end + 1))
60 } catch {
61 return null
62 }
63
64 if (typeof raw !== 'object' || raw === null) return null
65 const o = raw as Record<string, unknown>
66 const solved = SOLVED.find(s => s === o.solved)
67 if (solved === undefined) return null
68
69 return {
70 goal: typeof o.goal === 'string' ? clip(o.goal.trim(), 240) : '',
71 solved,
72 gaps: strings(o.gaps, 3),
73 shortcuts: strings(o.shortcuts, 3),
74 needsApproval: strings(o.needsApproval, 3),
75 next: strings(o.next, 2),
76 }
77}
78
79// One line is enough when there is nothing for the person to act on.
80export const isClean = (v: Pick<Verdict, 'solved' | 'gaps' | 'shortcuts' | 'needsApproval'>) =>
81 v.solved === 'yes' && v.gaps.length === 0 && v.shortcuts.length === 0 && v.needsApproval.length === 0
82types/index.d.ts 22 lines1export type Solved = 'yes' | 'partly' | 'no'
2
3export type Cost = { cached: number; input: number; output: number }
4
5export type Verdict = {
6 goal: string
7 solved: Solved
8 gaps: string[]
9 shortcuts: string[]
10 needsApproval: string[]
11 next: string[]
12 cost: Cost
13}
14
15export type Phase = 'idle' | 'checking'
16
17declare module 'claude-code' {
18 interface PluginState {
19 'next-steps-supervisor': { verdict: Verdict | null; phase: Phase }
20 }
21}
22