UNVERIFIED when Claude claims done without a check: a receipt line under every answer that claims done, naming the test or build that ran after the last edit…

A Claude Code mod that says UNVERIFIED when Claude claims done without a check. Under every answer that claims done, a receipt line names the test or build that ran after the last edit, or says that nothing did. It also replaces the wall of tool calls with a progress band that tells the truth.
The name reads like a cost meter, and it isn't one. It's a verification receipt.
It does three jobs:
Tested with Claude Code 2.1.291. Mods need 2.1.287 or later.
claude --plugin-dir ~/projects/claude-receipts
That loads the mod for one session. Nothing opens by itself: the band appears above the prompt when you send a prompt. /receipts opens a pane with the long view if you want it.
Or install it from its marketplace:
claude plugin marketplace add shawnpetros/claude-receipts
claude plugin install receipts@claude-receipts
Use one or the other, not both. With --plugin-dir and the installed copy in one session, two copies of the mod draw the same band and keep separate state.
To check that it's learning, finish a turn and run /receipts stats. The count of finished tasks should go up by one.
While a turn runs, the band has an orange border and these rows:
Step 2 of 4, a full-width bar, the percent done, then the estimate in dim text with its basis, such as ~4 to 9 min · from 3 similar tasks. The percent is weighted by how long each step is expected to take.Working, Next, Later or Done. At most six steps show, around the current one, then +n more.When the turn ends, the band always leaves the working state. It becomes a completion card and stays until your next prompt:
✓ All done badge, only when every step finished. A task-tool plan counts as finished when every task is completed.■ Turn ended · 2 of 5 steps reached badge when steps were left. Those steps read Not reached, never Done.✓ All done · bun test 152 pass badge, every step ticked with a full bar, and took 1m 47s on the right.⚠ Done, unverified badge. The title row shows the receipt, such as claimed done, no test/build/run after the last edit (src/x.ts at 14:02). When that doesn't fit, it shortens to the file and time first.✗ Done, checks failed badge and the run that failed.■ Stopped badge when you interrupt the turn.If agents the turn started are still running when it ends, the title row says waiting on 2 agents. The count drops as each one finishes, and the next prompt clears it.
A plan the mod derived marks a step done when a small model says the assistant's latest message finished it. When the turn ends, one more pass reads the final answer against the steps still open and marks the ones it shows were completed. That pass has 1.5 seconds. If it misses, the card says how far the plan got rather than guess.
When the elapsed time passes the top of the range, the border and bar turn grey and the summary reads over by 2m 10s · ~1 to 4 min more. Running long isn't an error, so it's never red.
Below 110 columns the per-step bars drop and each step keeps its word. Nothing in the band is ever wider than the terminal.
[▾] on the title row folds the band to that one row, and [▸] opens it again. Claude Code draws its own [-] just outside the border, and that one hides the whole band.
Click the band, or press ctrl+x then tab, to give it the keyboard. Then:
| Key | What it does | | :- | :- | | b | Shows a one-line tooltip with the estimate's basis and the calibration score. Press again to hide it. | | t | Opens or closes the settings popover |
t or /receipts tools opens a small bordered panel at the right of the band. It has three sections:
/config model row when this build has one that takes it, and otherwise runs /model <name>./effort. The current level is highlighted once a model request has carried it.Your plugin settings can't change while a session runs, so these choices override them for the rest of the session. /clear puts your settings back.
The [-] at the top right closes it, as does t again. Escape only hands the keyboard back to the prompt, because Claude Code tells a mod nothing when you press it.
These are the rules the estimate follows. Each one exists because some tool, somewhere, broke it.
~4 to 9 min. A countdown never freezes. When the range is too narrow to round to two values, it's widened until it shows two.prior only.from 7 similar tasks, spike guess, prior only, and derived plan when the plan is the mod's own guess.b in the band, the pane and /receipts stats show a line such as calibration: 61% of 18 tasks ended inside the range. After 10 or more tasks, if the score falls below 50%, the band says plainly that the ranges have been missing. That warning stays on the band and can't be hidden.3:52 to 8:40 left appear only when confidence is 0.5 or higher. Below that you get the rounded range and a bar.over by 2m 10s, the bar dims, and the range widens with elapsed time as its floor. It never resets to indeterminate.derived. The mod never passes its own guess off as the assistant's plan.Each task is filed under a shape: task type (build, debug, research, writing, config or refactor), step count (1, 2-3, 4-6 or 7+), a hash of the repository, whether the repository has tests, and the tool mix. Remaining time is the sum of the expected time of each step not yet done. The expectation comes from the most specific shape with at least 3 past tasks. Failing that, it falls back to a coarser shape, and then to a global prior.
The spread counts only the steps that are left, so it shrinks as steps finish. When a shape is new and the plan has 3 or more steps, the mod asks one cheap, read-only subagent for minutes per step. It does this once per task shape per session, with a 60-second limit, and counts the answer as half of one past task. It never does it for research, writing or chat tasks, whose steps are cheap.
Time spent waiting on a permission prompt isn't learned as work. Claude Code reports each tool's own run time without the prompt. The rest of the call is waiting, and it comes out of the step and the total before they're saved. Calibration is still scored on the wall clock, because that's what the range promised.
History lives in the mod's own store. Only finished task records are saved, and the averages are rebuilt from them. The store keeps the newest 500 tasks and stays under 1 MiB.
Each turn keeps a ledger of edits and of shell commands that look like verification: test runners, tsc, linters, builds and make check. At the end of the turn:
| What happened | Line under the answer | | :- | :- | | Edited, claimed done, nothing verified after the last edit | UNVERIFIED · claimed done, no test/build/run after the last edit (src/x.ts at 14:02) | | Edited, claimed done, a verify run after the last edit | receipt · bun test ✓ 152 pass · 1m ago | | A verify run after the last edit that failed | receipt · bun test ✗ exit 1 · 150 pass · 2 fail · 5s ago | | No edits, or the answer doesn't claim done | nothing |
When Claude Code reports usage, the line ends with what the turn cost:
| Who you are | What the line adds | | :- | :- | | A Claude plan user, with rate-limit windows | 5h window 6% used, resets 18:30 | | An API key user, with no windows | $0.42 this turn | | Neither is known | nothing, never a guess |
If the current turn's receipt line is wrong, for example it called an answer done when it wasn't, run /receipts wrong. Each receipt can be marked once. /receipts stats shows claims-done false positives: 1 of 12 receipts (8%). That number is the precondition for ever letting the receipt block a turn. See the roadmap.
A small model decides whether the answer claims done. If its label isn't back within 1.5 seconds, plain completion words in the answer decide instead.
| Command | What it does | | :- | :- | | /receipts | Opens or closes the pane, the long view with the calibration line | | /receipts tools | Opens the settings popover in the band | | /receipts rows [off\|clean\|quiet] | Sets the rows level, or steps to the next one with no argument | | /receipts clean | Switches the rows level to Off, and back to what it was | | /receipts basis | Shows or hides the basis tooltip, as b does | | /receipts stats | Shows the calibration history and task counts by type, with median durations | | /receipts wrong | Marks the current turn's receipt line as a false positive, for the count in stats | | /receipts reset-history | Forgets every learned task |
In the pane, c switches the rows level to Off and back, and b shows where the estimate comes from.
The rows level sets how much of the transcript the mod draws away. Quiet is the default.
| Level | What draws | | :- | :- | | Off | Every row exactly as Claude Code draws it | | Clean | Tool rows as one dim line each, such as ● Edit src/x.ts. Milestone rows in full. | | Quiet | See below |
Quiet goes further:
● Bash bun test ✓.↳ message from @Explore: Found 3 mods.Your own prompts always draw in full. Quiet never touches a question the assistant asks you or a permission prompt.
ctrl+o is the escape hatch. It shows the full transcript, and a hand-back row it expands draws in full. ctrl+o also expands a folded group of reads and searches. A single tool row or a block of assistant text can't be expanded past the level, because Claude Code doesn't tell a mod when one is expanded. Set the level to Off to see those in full.
Set these under pluginConfigs in your Claude Code settings. Use the key receipts@inline for a session started with --plugin-dir.
| Key | Default | What it does | | :- | :- | :- | | spike | true | Allows the sizing subagent for unfamiliar tasks. Set it to false and the mod never spawns one. | | cleanView | true | Starts each session at the quiet rows level. Set it to false to start at Off. The popover and /receipts rows change it for the session. |
The mod calls the model API for its labels and the derived plan, and spawns a subagent for the spike. It makes no other network calls and sends no telemetry of its own. The only thing it writes is its own store.
/receipts and its subcommands aren't there. The band's keys still work.--plugin-dir) and the marketplace install in the same session. Both draw the same band and keep separate state, so the band you see may belong to the copy that missed the turn's end.hooks/hooks.json carries an empty hooks key beside modules, so an older build loads nothing instead of failing.Each release answers one question it can measure.
/receipts help.$.session.usage() already shipped in 0.2.2. The rest is an optional gate that turns UNVERIFIED from a mirror into a block, off by default. It ships only after the claims-done classifier's false-positive rate has been measured over enough live turns with /receipts wrong and /receipts stats. A gate that blocks a finished turn gets the mod uninstalled.Not planned: token meters, burn bars, or a rename.
claude plugin validate . --strict
claude plugin test
bun scripts/simulate.ts
bun scripts/mock-band.ts 155
bun scripts/check-manifest.ts
scripts/mock-band.ts prints the band as plain text in each state: working, over the range, verified, unverified, narrow, collapsed, with the tooltip, and the popover. It uses the same pure view the mod draws from. The argument is the band's width, which is the terminal's width less 5.
src/ holds the logic as plain modules with no mods API dependency. These are the estimator, history, milestones, ledger, shape, spike, view and band modules, plus a session that talks to Claude Code through a small host interface. hooks/register.ts builds that host from $ and wires the hooks.
hooks/register.ts 547 lines1import type { ConfigRow, EngineInterface, On, PluginOptions, RenderElement } from 'claude-code'
2import { atom, read, update } from 'claude-code'
3
4import {
5 ACCENT,
6 bandView,
7 EFFORT_CHOICES,
8 MODEL_CHOICES,
9 ROWS_LEVELS,
10 TOOLS_COLUMNS,
11 toolsView,
12 type Row,
13 type RowsLevel,
14 type Seg,
15} from '../src/band'
16import type { Host } from '../src/host'
17import { firstLineOf, handbackLineOf, QuietLog } from '../src/quiet'
18import { PANE_ID, PANE_TITLE, ReceiptsSession } from '../src/session'
19import { paneLines, toolGroupText, toolResultText, toolRowText, type Line } from '../src/view'
20
21/**
22 * The rows level: `off` draws every row as Claude Code does; `clean` draws
23 * tool rows as one dim line and milestones in full; `quiet` (the default)
24 * draws tool rows as nothing and milestones as one dim line, during the turn
25 * and after it, and folds the chrome a turn scatters: the spinner, progress
26 * pills, notices and command output while working, subagent hand-backs, and
27 * interim assistant text down to its first line.
28 *
29 * Invariant 4, drawing only: these hooks return a drawing and never touch
30 * the transcript, so `off` shows every row as it was. Scar: the /buddy
31 * main-model leak; a mod rewriting content is a different and riskier thing
32 * than a mod redrawing it. Milestones always show (invariant 5).
33 */
34const rows = atom({ plugin: 'receipts', key: 'rows' }, 'quiet')
35
36/** The level `/receipts clean` and the pane's `c` go back to from `off`. */
37const rowsLast = atom({ plugin: 'receipts', key: 'rowsLast' }, 'quiet')
38
39/**
40 * Whether a main-loop turn is running. The quiet sites read it, so they draw
41 * again once at each turn edge, not on every tick.
42 */
43const working = atom({ plugin: 'receipts', key: 'working' }, false)
44
45/** The band folded to its title row by its own `[▾]` (`[▸]` folded). */
46const collapsed = atom({ plugin: 'receipts', key: 'collapsed' }, false)
47
48/** The settings popover, drawn in the band. */
49const toolsOpen = atom({ plugin: 'receipts', key: 'toolsOpen' }, false)
50
51/**
52 * The spike switch as the person last set it. userConfig is read-only at
53 * runtime, so this state overrides it for the session; it starts from it.
54 */
55const spike = atom({ plugin: 'receipts', key: 'spike' }, true)
56
57/**
58 * Read by the pane and band only, so bumping it once a second redraws those
59 * two and not every tool row in the transcript.
60 */
61const tick = atom({ plugin: 'receipts', key: 'tick' }, 0)
62
63const USAGE =
64 'usage: /receipts (toggle the pane), /receipts tools, /receipts rows [off|clean|quiet], /receipts clean, /receipts basis, /receipts stats, /receipts wrong, /receipts reset-history'
65const PANE_MIN_ROWS = 6
66const PANE_MAX_ROWS = 24
67
68/**
69 * The mods API as the session logic sees it. Top level and handed `$`, so
70 * `claude plugin validate` lists every call through it.
71 */
72function hostOf($: EngineInterface): Host {
73 return {
74 now: () => $.clock.now(),
75 every: (ms, fn) => $.clock.every(ms, fn),
76 after: (ms, fn) => $.clock.after(ms, fn),
77 storeGet: key => $.store.get(key),
78 storeSet: (key, value) => $.store.set(key, value),
79 storeDelete: key => $.store.delete(key),
80 redraw: () => {
81 update($, tick, n => n + 1).catch(() => $.ui.invalidate('ui.render'))
82 },
83 classify: (text, labels) => $.model.classify(text, labels),
84 complete: request => $.model.complete(request),
85 spawn: args => $.agent.spawn(args),
86 cwd: () => $.session.cwd(),
87 repo: () => $.session.repo(),
88 list: path => $.fs.list(path),
89 agents: () => $.agent.list(),
90 usage: () => $.session.usage(),
91 }
92}
93
94/**
95 * `off` and back: to the level the person had before, quiet if none.
96 */
97async function toggleClean($: EngineInterface): Promise<RowsLevel> {
98 const level = (await read($, rows)) as RowsLevel
99 if (level === 'off') {
100 const back = (await read($, rowsLast)) as RowsLevel
101 await update($, rows, () => back)
102 return back
103 }
104 await update($, rowsLast, () => level)
105 await update($, rows, () => 'off')
106 return 'off'
107}
108
109async function setRows($: EngineInterface, level: RowsLevel): Promise<void> {
110 if (level !== 'off') await update($, rowsLast, () => level)
111 await update($, rows, () => level)
112}
113
114async function levelOf($: EngineInterface): Promise<RowsLevel> {
115 return (await read($, rows)) as RowsLevel
116}
117
118/**
119 * Sets the model or effort the way the person would: the `/config` row when
120 * the build has one that takes it, else the slash command. Top level and
121 * handed `$`, so validate lists the calls.
122 */
123async function choose($: EngineInterface, key: 'model' | 'effort', value: string): Promise<void> {
124 const rows: ConfigRow[] = await $.config.list().catch(() => [])
125 const row = rows.find(candidate => candidate.key === key)
126 if (row && !row.isLocked && row.kind !== 'boolean' && row.kind !== 'number') {
127 const result = await $.config.set({ key, value }).catch(() => ({ deny: 'failed' }))
128 if (result.deny === undefined) return
129 }
130 await $.command.run({ command: key, args: value })
131}
132
133/**
134 * The userConfig defaults into state, where the band's switches change them.
135 */
136async function applyDefaults($: EngineInterface, session: ReceiptsSession, isClean: boolean, isSpike: boolean): Promise<void> {
137 if (!isClean) await update($, rows, () => 'off')
138 if (!isSpike) await update($, spike, () => false)
139 session.setSpike(isSpike)
140}
141
142async function sessionModelOf($: EngineInterface): Promise<string> {
143 return $.session.model().catch(() => '')
144}
145
146/**
147 * A line of the pane as Text props, leaving out the styles it does not set.
148 */
149function textPropsOf(line: Line) {
150 return {
151 key: line.key,
152 wrap: 'truncate-end' as const,
153 children: [line.text],
154 ...(line.dim ? { dimColor: true } : {}),
155 ...(line.bold ? { bold: true } : {}),
156 ...(line.color ? { color: line.color } : {}),
157 }
158}
159
160type Elements = ReturnType<EngineInterface['ui']['resolve']>
161type Presses = Record<string, () => unknown>
162
163/**
164 * A row of the band as elements: a Text per segment, a plain Button where a
165 * segment carries one. A Button whose key has no handler draws as text.
166 */
167function rowOf(el: Elements, row: Row, presses: Presses): RenderElement {
168 const children = row.segs.map((seg: Seg): RenderElement => {
169 const press = seg.button ? presses[seg.button.key] : undefined
170 if (seg.button && press) {
171 return el.Button({
172 key: seg.button.key,
173 label: seg.button.label,
174 plain: true,
175 ...(seg.button.hotkey ? { hotkey: seg.button.hotkey } : {}),
176 ...(seg.dim ? { dimColor: true } : {}),
177 onPress: () => press(),
178 })
179 }
180 return el.Text({
181 children: [seg.text],
182 ...(seg.color ? { color: seg.color } : {}),
183 ...(seg.dim ? { dimColor: true } : {}),
184 ...(seg.bold ? { bold: true } : {}),
185 ...(seg.inverse ? { inverse: true } : {}),
186 })
187 })
188 return el.Box({ key: row.key, flexDirection: 'row', children })
189}
190
191/**
192 * The one dim line a milestone row folds to while suppressed: the call and
193 * how it ended. A verify run's exit shows as ✓ or ✗ (invariant 5).
194 */
195function outcomeMarkOf(props: { isRunning?: boolean; isErrored?: boolean; isInterrupted?: boolean }): string {
196 if (props.isRunning) return ' …'
197 if (props.isInterrupted) return ' ■'
198 return props.isErrored ? ' ✗' : ' ✓'
199}
200
201export function register(on: On, options: PluginOptions): void {
202 const isSpikeByDefault = options.spike !== false
203 const session = new ReceiptsSession({ spike: isSpikeByDefault })
204 const quiet = new QuietLog()
205 const isCleanByDefault = options.cleanView !== false
206
207 on('session.start', async ($, e, next) => {
208 await session.start(hostOf($))
209 const stored = await $.state.get({ plugin: 'receipts', key: 'rows' })
210 if (stored.version === 0) await applyDefaults($, session, isCleanByDefault, isSpikeByDefault)
211 try {
212 await $.command.register({
213 name: 'receipts',
214 description: 'Toggle the receipts pane, open its settings, or show calibration stats',
215 argumentHint: '[tools|rows|clean|basis|stats|reset-history]',
216 immediate: true,
217 })
218 } catch {
219 // A name clash leaves the mod running without its command
220 }
221 return next(e)
222 })
223
224 // /clear, /resume and /branch reset $.state: put the configured defaults back
225 on('classic.SessionStart', { source: ['clear', 'resume', 'fork'] }, async ($, e, next) => {
226 await applyDefaults($, session, isCleanByDefault, isSpikeByDefault)
227 return next(e)
228 }).catch(($, e, next) => next(e))
229
230 on('turn.start', async ($, e, next) => {
231 await session.turnStart(hostOf($), e.turnId, e.text)
232 await update($, working, () => true)
233 return next(e)
234 })
235
236 on('turn.step', async function* ($, e, next) {
237 if (e.agentId === undefined) session.noteEffort(e.effort)
238 const result = yield* next(e)
239 session.stepResult(hostOf($), e, result)
240 return result
241 })
242
243 on('tool.call', async ($, e, next) => {
244 session.beforeTool(e, e)
245 session.toolStarted(e.tool_use_id, await $.clock.now())
246 const outcome = await next(e)
247 await session.afterTool(hostOf($), e, e, outcome).catch(() => undefined)
248 return outcome
249 }).catch(($, e, next) => next(e))
250
251 // The tool's own run time, which excludes the permission prompt: the rest
252 // of the call's span was waiting on the person, and is not learned as work
253 on('classic.PostToolUse', async ($, e, next) => {
254 session.toolRan(e.tool_use_id, e.duration_ms)
255 return next(e)
256 }).catch(($, e, next) => next(e))
257
258 // Under the answer: the receipt, or UNVERIFIED. A mirror, never a gate
259 on('turn.complete', async ($, e, next) => {
260 const result = await next(e)
261 if (e.agentId !== undefined) {
262 session.subagentComplete(hostOf($), e.agentId, e.answer)
263 return result
264 }
265 const line = await session.turnComplete(hostOf($), e)
266 await update($, working, () => false)
267 // The final answer, drawn as one dim line while it streamed, in full now
268 if (quiet.finish(e.answer)) $.ui.invalidate('ui.render')
269 if (!line) return result
270 const isOwnText = result.text !== '' && result.text !== e.answer
271 return { ...result, text: isOwnText ? `${result.text}\n${line}` : line }
272 })
273
274 on('command.run', { command: 'receipts' }, async ($, e) => {
275 const arg = e.args.trim()
276 if (arg === 'stats') return { text: session.statsText() }
277 if (arg === 'tools') {
278 await update($, toolsOpen, () => true)
279 return {}
280 }
281 if (arg === 'basis') {
282 session.toggleBasis()
283 $.ui.invalidate('ui.render')
284 return { text: session.isBasisShown ? 'basis shown on the band' : 'basis hidden' }
285 }
286 if (arg === 'wrong') return { text: await session.markWrong(hostOf($)) }
287 if (arg === 'clean') return { text: `rows: ${await toggleClean($)}` }
288 if (arg === 'rows' || arg.startsWith('rows ')) {
289 const wanted = arg.slice(4).trim()
290 if (wanted === '') {
291 const now = await levelOf($)
292 const next = ROWS_LEVELS[(ROWS_LEVELS.indexOf(now) + 1) % ROWS_LEVELS.length]!
293 await setRows($, next)
294 return { text: `rows: ${next}` }
295 }
296 if (!(ROWS_LEVELS as readonly string[]).includes(wanted)) return { text: USAGE }
297 await setRows($, wanted as RowsLevel)
298 return { text: `rows: ${wanted}` }
299 }
300 if (arg === 'reset-history') {
301 const count = await session.resetHistory(hostOf($))
302 return { text: `history cleared: ${count} ${count === 1 ? 'task' : 'tasks'} forgotten` }
303 }
304 if (arg !== '') return { text: USAGE }
305 if (session.paneOpen) {
306 await $.ui.close({ id: PANE_ID })
307 session.paneClosed()
308 return {}
309 }
310 // rows: inline (the main screen) it opens as tall as its content, not cut
311 // to a third; a dock ignores it and the render fits its rows instead
312 const wanted = paneLines(session.viewModel(await $.clock.now())).length + 1
313 const placed = await $.ui.open({ id: PANE_ID, title: PANE_TITLE, rows: Math.max(PANE_MIN_ROWS, Math.min(PANE_MAX_ROWS, wanted)) })
314 session.paneOpened(placed.isPlaced)
315 return {}
316 })
317
318 on('ui.close', { id: PANE_ID }, async ($, e, next) => {
319 session.paneClosed()
320 return next(e)
321 }).catch(($, e, next) => next(e))
322
323 // The pane: the long view, opened only by /receipts
324 on('ui.render', { component: 'Pane' }, async ($, e, next) => {
325 if (e.requestId !== PANE_ID) return next(e)
326 const isClean = (await levelOf($)) !== 'off'
327 await read($, tick)
328 const now = await $.clock.now()
329 session.paneDrawn()
330 session.ensureTicking(hostOf($), now)
331 const { Box, Text, Button } = $.ui.resolve(e)
332 // Never more rows than the room has: the header, then the lines that fit
333 const all = paneLines(session.viewModel(now))
334 const room = Math.max(1, e.props.scroll.bodyRows - 1)
335 const lines =
336 all.length <= room ? all : [...all.slice(0, room - 1), { key: 'pane-more', text: `+${all.length - room + 1} more · /receipts stats`, dim: true }]
337 return Box({
338 flexDirection: 'column',
339 children: [
340 Box({
341 flexDirection: 'row',
342 columnGap: 2,
343 children: [
344 Text({ bold: true, children: ['receipts'] }),
345 Button({
346 key: 'clean-toggle',
347 label: isClean ? 'clean view on' : 'clean view off',
348 hotkey: 'c',
349 plain: true,
350 onPress: () => toggleClean($),
351 }),
352 Button({
353 key: 'basis',
354 label: session.isBasisShown ? 'hide basis' : 'estimate basis',
355 hotkey: 'b',
356 plain: true,
357 onPress: () => {
358 session.toggleBasis()
359 $.ui.invalidate('ui.render')
360 },
361 }),
362 ],
363 }),
364 ...lines.map(line => Text(textPropsOf(line))),
365 ],
366 })
367 })
368
369 // The band: the primary surface. Progress while working, the completion
370 // card after, the settings popover under it when asked
371 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
372 if (e.props.hasSurvey) return next(e)
373 const isToolsOpen = await read($, toolsOpen)
374 const isBandShown = session.isBandWanted()
375 if (!isBandShown && !isToolsOpen) return next(e)
376 await read($, tick)
377 const isCollapsed = await read($, collapsed)
378 const now = await $.clock.now()
379 session.ensureTicking(hostOf($), now)
380 const el = $.ui.resolve(e)
381 const presses: Presses = {
382 collapse: () => update($, collapsed, value => !value),
383 basis: () => {
384 session.toggleBasis()
385 $.ui.invalidate('ui.render')
386 },
387 'tools-toggle': () => update($, toolsOpen, value => !value),
388 'tools-close': () => update($, toolsOpen, () => false),
389 'set-spike': () =>
390 update($, spike, value => {
391 session.setSpike(!value)
392 return !value
393 }),
394 }
395 for (const level of ROWS_LEVELS) presses[`rows-${level}`] = () => setRows($, level)
396 for (const choice of MODEL_CHOICES) {
397 presses[`model-${choice.value}`] = async () => {
398 await choose($, 'model', choice.value)
399 $.ui.invalidate('ui.render')
400 }
401 }
402 for (const choice of EFFORT_CHOICES) {
403 presses[`effort-${choice.value}`] = async () => {
404 await choose($, 'effort', choice.value)
405 session.noteEffort(choice.value)
406 $.ui.invalidate('ui.render')
407 }
408 }
409
410 const columns = e.props.bodyColumns
411 let band: RenderElement | null = null
412 if (isBandShown) {
413 const view = bandView(session.viewModel(now), columns, { collapsed: isCollapsed, showBasis: session.isBasisShown })
414 band = el.Box({
415 key: 'band',
416 flexDirection: 'column',
417 borderStyle: 'round',
418 borderColor: view.border,
419 paddingX: 1,
420 width: columns,
421 children: view.rows.map(row => rowOf(el, row, presses)),
422 })
423 }
424 if (!isToolsOpen && band) return band
425
426 const toolsWidth = Math.min(columns, TOOLS_COLUMNS)
427 const rows = toolsView(
428 {
429 model: await sessionModelOf($),
430 ...(session.effortLevel ? { effort: session.effortLevel } : {}),
431 rows: await levelOf($),
432 spike: await read($, spike),
433 },
434 toolsWidth - 4,
435 )
436 const tools = el.Box({
437 key: 'tools',
438 flexDirection: 'column',
439 borderStyle: 'round',
440 borderColor: ACCENT,
441 paddingX: 1,
442 width: toolsWidth,
443 children: rows.map(row => rowOf(el, row, presses)),
444 })
445 return el.Box({
446 flexDirection: 'column',
447 width: columns,
448 children: [...(band ? [band] : []), el.Box({ flexDirection: 'row', justifyContent: 'flex-end', children: [tools] })],
449 })
450 })
451
452 // ctrl+o: a ToolGroup's props say when it is expanded, so an expanded group
453 // draws in full. ToolUse and ToolResult props carry no such flag, so a
454 // single row cannot be expanded past the level; `/receipts rows off` can
455 on('ui.render', { component: 'ToolUse' }, async ($, e, next) => {
456 const level = await levelOf($)
457 if (level === 'off') return next(e)
458 const isMilestone = session.isMilestoneRow(e.props.tool_use_id, e.props.tool, e.props.input)
459 const { Box, Text } = $.ui.resolve(e)
460 if (level === 'quiet') {
461 if (!isMilestone) return Box({})
462 const text = toolRowText(e.props.tool, e.props.input, session.workingDirectory) + outcomeMarkOf(e.props)
463 return Text({ dimColor: true, wrap: 'truncate-end', children: [text] })
464 }
465 if (isMilestone) return next(e)
466 return Text({ dimColor: true, wrap: 'truncate-end', children: [toolRowText(e.props.tool, e.props.input, session.workingDirectory)] })
467 })
468
469 on('ui.render', { component: 'ToolResult' }, async ($, e, next) => {
470 const level = await levelOf($)
471 if (level === 'off') return next(e)
472 const { Box, Text } = $.ui.resolve(e)
473 // Quiet: a milestone's outcome is folded into its ToolUse line
474 if (level === 'quiet') return Box({})
475 if (session.isMilestoneRow(e.props.tool_use_id, e.props.tool, undefined)) return next(e)
476 return Text({ dimColor: true, wrap: 'truncate-end', children: [toolResultText(e.props.output, e.props.isErrored)] })
477 })
478
479 on('ui.render', { component: 'ToolGroup' }, async ($, e, next) => {
480 const level = await levelOf($)
481 if (e.props.isExpanded || level === 'off') return next(e)
482 const { Box, Text } = $.ui.resolve(e)
483 if (level === 'quiet') return Box({})
484 return Text({ dimColor: true, wrap: 'truncate-end', children: [toolGroupText(e.props.calls)] })
485 })
486
487 // ---- quiet: the chrome a turn scatters ------------------------------------
488 // Never hooked: AskUserQuestion and the permission dialogs. A question to
489 // the person is the one thing quiet must never fold.
490
491 // The band shows the elapsed time and the step; the spinner repeats it
492 on('ui.render', { component: 'Spinner' }, async ($, e, next) => {
493 if ((await levelOf($)) !== 'quiet') return next(e)
494 const { Box } = $.ui.resolve(e)
495 return Box({})
496 })
497
498 on('ui.render', { component: 'ToolProgress' }, async ($, e, next) => {
499 if ((await levelOf($)) !== 'quiet' || !(await read($, working))) return next(e)
500 const { Box } = $.ui.resolve(e)
501 return Box({})
502 })
503
504 on('ui.render', { component: 'TurnDuration' }, async ($, e, next) => {
505 if ((await levelOf($)) !== 'quiet' || !(await read($, working))) return next(e)
506 const { Box } = $.ui.resolve(e)
507 return Box({})
508 })
509
510 on('ui.render', { component: 'InfoNotice' }, async ($, e, next) => {
511 if ((await levelOf($)) !== 'quiet' || !(await read($, working))) return next(e)
512 const { Box } = $.ui.resolve(e)
513 return Box({})
514 })
515
516 // Another command's output mid-turn; this mod's own and any error line show
517 on('ui.render', { component: 'CommandOutput' }, async ($, e, next) => {
518 if (e.props.command === 'receipts' || e.props.isErrored) return next(e)
519 if ((await levelOf($)) !== 'quiet' || !(await read($, working))) return next(e)
520 const { Box } = $.ui.resolve(e)
521 return Box({})
522 })
523
524 // A subagent's hand-back, a peer's message, a task notification: nothing
525 // while the turn runs, one dim line after. The person's own prompt and a
526 // ctrl+o expanded row are never touched
527 on('ui.render', { component: 'UserMessage' }, async ($, e, next) => {
528 if (e.props.isExpanded || e.props.origin.kind === 'composer') return next(e)
529 if ((await levelOf($)) !== 'quiet') return next(e)
530 const { Box, Text } = $.ui.resolve(e)
531 if (await read($, working)) return Box({})
532 return Text({ dimColor: true, wrap: 'truncate-end', children: [handbackLineOf(e.props.text, e.props.from?.name)] })
533 })
534
535 // Interim text: its first line, dim. The final answer redraws in full when
536 // turn.complete names it (QuietLog). A block never seen mid-turn is left alone
537 on('ui.render', { component: 'AssistantMessage' }, async ($, e, next) => {
538 if ((await levelOf($)) !== 'quiet') return next(e)
539 const isWorking = await read($, working)
540 if (isWorking) quiet.seen(e.requestId, e.props.text)
541 const kind = quiet.kindOf(e.requestId)
542 if (kind !== 'interim') return next(e)
543 const { Text } = $.ui.resolve(e)
544 return Text({ dimColor: true, wrap: 'truncate-end', children: [firstLineOf(e.props.text)] })
545 })
546}
547src/band.ts 506 lines1/**
2 * The band above the prompt and the settings popover, as rows of styled
3 * segments, each row exactly the width it is given. Pure: no `$`, no
4 * elements. `hooks/register.ts` turns a segment into a Text (or a Button
5 * when it carries one) and a row into a Box; tests and `scripts/mock-band.ts`
6 * read the same rows as plain text.
7 *
8 * Laying out to exact widths here, instead of trusting flex to right-align,
9 * keeps the right-hand column (elapsed, percent, the estimate) where the
10 * tests say it is at every width, and nothing ever wraps the band taller.
11 */
12
13import type { Color } from 'claude-code'
14
15import { basisLines, durationOf, estimateRow, type Estimate } from './estimator'
16import type { Milestone, Plan } from './milestones'
17import { truncate, type ViewModel } from './view'
18
19export const ACCENT: Color = 'claude'
20export const DONE: Color = 'success'
21export const WARN: Color = 'warning'
22/** Over the range: the bar and border go grey, never red. Running long is not an error. */
23export const OVER: Color = 'inactive'
24
25/**
26 * Below this many band body columns the per-step bars drop. The band's body
27 * is the terminal less the engine's five cells for its own `[-]`, so 105 is
28 * a 110-column terminal.
29 */
30export const NARROW_BODY = 105
31/** The round border and one cell of padding each side. */
32export const FRAME = 4
33export const MINI_BAR = 12
34export const MAX_STEP_ROWS = 6
35const MIN_BAR = 8
36const MIN_TAIL = 12
37
38export type SegButton = {
39 key: string
40 label: string
41 hotkey?: string
42}
43
44export type Seg = {
45 text: string
46 color?: Color
47 dim?: boolean
48 bold?: boolean
49 inverse?: boolean
50 /** Shrinks first when the row is too wide. */
51 grow?: boolean
52 /** Drawn as a plain Button; `text` is what the terminal draws for it. */
53 button?: SegButton
54}
55
56export type Row = {
57 key: string
58 segs: Seg[]
59}
60
61export type BandTone = 'working' | 'over' | 'done' | 'ended' | 'unverified' | 'stopped'
62
63export type BandView = {
64 tone: BandTone
65 border: Color
66 rows: Row[]
67}
68
69export type BandOptions = {
70 collapsed: boolean
71 showBasis: boolean
72}
73
74export function innerWidthOf(columns: number): number {
75 return Math.max(1, columns - FRAME)
76}
77
78export function rowText(row: Row): string {
79 return row.segs.map(seg => seg.text).join('')
80}
81
82function widthOf(segs: readonly Seg[]): number {
83 return segs.reduce((sum, seg) => sum + [...seg.text].length, 0)
84}
85
86function space(n: number): Seg {
87 return { text: ' '.repeat(Math.max(0, n)) }
88}
89
90/**
91 * A plain Button as a segment: `b: basis` with a hotkey, the label alone without.
92 */
93function buttonSeg(key: string, label: string, hotkey?: string, dim = true): Seg {
94 return {
95 text: hotkey ? `${hotkey}: ${label}` : label,
96 ...(dim ? { dim: true } : {}),
97 button: { key, label, ...(hotkey ? { hotkey } : {}) },
98 }
99}
100
101/**
102 * Cuts segments to `width` cells from the right, keeping each one's style;
103 * a Button that no longer fits whole is dropped, never drawn half.
104 */
105function clip(segs: readonly Seg[], width: number): Seg[] {
106 const out: Seg[] = []
107 let left = width
108 for (const seg of segs) {
109 const length = [...seg.text].length
110 if (length <= left) {
111 out.push(seg)
112 left -= length
113 continue
114 }
115 if (left > 0 && !seg.button) out.push({ ...seg, text: truncate(seg.text, left) })
116 else if (left > 0) out.push(space(left))
117 left = 0
118 break
119 }
120 return out
121}
122
123/**
124 * One row of exactly `width` cells: `left` at the start, `right` at the end,
125 * at least `gap` spaces between. The `grow` segment of `left` gives way first.
126 */
127function line(width: number, left: readonly Seg[], right: readonly Seg[] = [], gap = 2): Seg[] {
128 const rightWidth = widthOf(right)
129 const room = width - rightWidth - (right.length > 0 ? gap : 0)
130 let segs = [...left]
131 const over = widthOf(segs) - room
132 if (over > 0) {
133 const i = segs.findIndex(seg => seg.grow)
134 if (i !== -1) {
135 const seg = segs[i]!
136 const keep = Math.max(0, [...seg.text].length - over)
137 segs[i] = { ...seg, text: keep === 0 ? '' : truncate(seg.text, keep) }
138 }
139 if (widthOf(segs) > Math.max(0, room)) segs = clip(segs, Math.max(0, room))
140 }
141 if (rightWidth > width) return clip([...right], width)
142 const pad = width - widthOf(segs) - rightWidth
143 return [...segs, space(pad), ...right]
144}
145
146function barSegs(progress: number, width: number, color: Color, isDim = false, full = '━', empty = '─'): Seg[] {
147 const filled = Math.max(0, Math.min(width, Math.round(progress * width)))
148 return [
149 { text: full.repeat(filled), color, ...(isDim ? { dim: true } : {}) },
150 { text: empty.repeat(width - filled), dim: true },
151 ]
152}
153
154// ---- receipts on the card ---------------------------------------------------
155
156type Verdict =
157 | { kind: 'none' }
158 | { kind: 'verified'; label: string }
159 | { kind: 'failed'; text: string }
160 | { kind: 'unverified'; text: string; short: string }
161
162/**
163 * Reads the receipt line `ledger.receiptLine` wrote back into what the card
164 * needs: `receipt · bun test ✓ 152 pass · 1m ago` is verified, labelled
165 * `bun test 152 pass`; a ✗ is a failed check; UNVERIFIED is unverified.
166 */
167export function verdictOf(receipt: string | null): Verdict {
168 if (!receipt) return { kind: 'none' }
169 if (receipt.startsWith('UNVERIFIED')) {
170 const text = receipt.replace(/^UNVERIFIED · /, '')
171 const where = /\((.+) at (\d{1,2}:\d{2})\)$/.exec(text)
172 return { kind: 'unverified', text, short: where ? `${where[1]} at ${where[2]}, no check after` : text }
173 }
174 const body = receipt.replace(/^receipt · /, '').replace(/ · \d+[smh] ago$/, '')
175 if (body.includes('✗')) return { kind: 'failed', text: body }
176 return { kind: 'verified', label: body.replace(' ✓', '').replace(/\s+/g, ' ').trim() }
177}
178
179// ---- the band ---------------------------------------------------------------
180
181function currentIndexOf(plan: Plan): number {
182 return plan.items.findIndex(item => item.state === 'current')
183}
184
185function reachedOf(plan: Plan | null): { done: number; total: number } {
186 const items = plan?.items ?? []
187 return { done: items.filter(item => item.state === 'done').length, total: items.length }
188}
189
190/** Every step done, or no steps at all to fall short of. */
191function isAllDone(plan: Plan | null): boolean {
192 const { done, total } = reachedOf(plan)
193 return done === total
194}
195
196function stepCounterOf(plan: Plan | null): string {
197 if (!plan || plan.items.length === 0) return 'Planning'
198 const n = plan.items.length
199 const current = currentIndexOf(plan)
200 const done = plan.items.filter(item => item.state === 'done').length
201 const i = current === -1 ? Math.min(done + 1, n) : current + 1
202 return `Step ${i} of ${n}` + (plan.source === 'derived' ? ' · derived' : '')
203}
204
205function titleRow(model: ViewModel, width: number, options: BandOptions, tone: BandTone, verdict: Verdict): Row {
206 const title = model.title.replace(/\s+/g, ' ').trim() || 'continuing'
207 const buttons = [buttonSeg('basis', 'basis', 'b'), space(2), buttonSeg('tools-toggle', 'tools', 't'), space(2)]
208 // Not [-]: Claude Code draws its own [-] beside the band, which hides it
209 // whole; this one folds to the title row
210 const collapse = buttonSeg('collapse', options.collapsed ? '[▸]' : '[▾]', undefined, false)
211 if (tone === 'working' || tone === 'over') {
212 const left: Seg[] = [{ text: '✶ ', color: ACCENT, bold: true }, { text: title, bold: true, grow: true }]
213 return { key: 'title', segs: line(width, left, [...buttons, { text: durationOf(model.elapsedMs) }, space(1), collapse]) }
214 }
215 const waiting = model.finished?.waitingAgents ?? 0
216 const took = `took ${durationOf(model.finished?.totalMs ?? model.elapsedMs)}`
217 const tail: Seg[] = waiting > 0 ? [{ text: `waiting on ${waiting} ${waiting === 1 ? 'agent' : 'agents'}`, color: WARN }, space(2)] : []
218 let badge: Seg
219 let text: Seg
220 switch (verdict.kind) {
221 case 'unverified': {
222 // The file is the point: when the whole receipt does not fit, the
223 // path-first form does, so the cut never lands on the path
224 badge = { text: ' ⚠ Done, unverified ', color: WARN, inverse: true, bold: true }
225 const room = width - [...badge.text].length - 1 - 2 - widthOf([...tail, ...buttons, { text: took }, space(1), collapse])
226 text = { text: [...verdict.text].length <= room ? verdict.text : verdict.short, color: WARN, grow: true }
227 break
228 }
229 case 'failed':
230 badge = { text: ' ✗ Done, checks failed ', color: WARN, inverse: true, bold: true }
231 text = { text: verdict.text, color: WARN, grow: true }
232 break
233 case 'verified':
234 badge = { text: ` ✓ All done · ${verdict.label} `, color: DONE, inverse: true, bold: true }
235 text = { text: title, color: DONE, grow: true }
236 break
237 default:
238 if (tone === 'stopped') {
239 badge = { text: ' ■ Stopped ', color: OVER, inverse: true, bold: true }
240 text = { text: title, dim: true, grow: true }
241 } else if (tone === 'ended') {
242 // Steps left when the turn ended: say how far it got, never "done".
243 // Scar: a green card over three steps that never ran, read as finished
244 const { done, total } = reachedOf(model.plan)
245 badge = { text: ` ■ Turn ended · ${done} of ${total} ${total === 1 ? 'step' : 'steps'} reached `, color: OVER, inverse: true, bold: true }
246 text = { text: title, grow: true }
247 } else {
248 badge = { text: ' ✓ All done ', color: DONE, inverse: true, bold: true }
249 text = { text: title, color: DONE, grow: true }
250 }
251 }
252 return { key: 'title', segs: line(width, [badge, space(1), text], [...tail, ...buttons, { text: took }, space(1), collapse]) }
253}
254
255function summaryRow(model: ViewModel, width: number, tone: BandTone): Row {
256 const plan = model.plan
257 if (!model.isWorking) {
258 const n = plan?.items.length ?? 0
259 const done = plan?.items.filter(item => item.state === 'done').length ?? 0
260 const label = `${done} of ${n} ${n === 1 ? 'step' : 'steps'}` + (plan?.source === 'derived' ? ' · derived' : '')
261 const progress = n === 0 ? 1 : done / n
262 const percent = `${Math.round(progress * 100)}%`.padStart(4)
263 const color = tone === 'done' ? DONE : tone === 'stopped' || tone === 'ended' ? OVER : WARN
264 const barWidth = Math.max(1, width - [...label].length - 2 - percent.length)
265 return {
266 key: 'summary',
267 segs: line(width, [{ text: label }, space(1), ...barSegs(progress, barWidth, color), space(1), { text: percent, color, bold: true }], [], 0),
268 }
269 }
270 const est: Estimate = model.estimate ?? { kind: 'indeterminate' }
271 const isOver = est.kind === 'range' && est.isOver
272 const progress = est.kind === 'range' ? est.progress : 0
273 const label = stepCounterOf(plan)
274 const percent = `${Math.round(progress * 100)}%`.padStart(4)
275 let tail = estimateRow(est)
276 const fixed = [...label].length + 1 + 1 + percent.length
277 let barWidth = width - fixed - 2 - [...tail].length
278 if (barWidth < MIN_BAR) {
279 const tailRoom = width - fixed - 2 - MIN_BAR
280 tail = tailRoom >= MIN_TAIL ? truncate(tail, tailRoom) : ''
281 barWidth = width - fixed - (tail ? 2 + [...tail].length : 0)
282 }
283 const barColor = isOver ? OVER : ACCENT
284 const segs: Seg[] = [
285 { text: label },
286 space(1),
287 ...barSegs(progress, Math.max(1, barWidth), barColor),
288 space(1),
289 isOver ? { text: percent, dim: true } : { text: percent, color: ACCENT, bold: true },
290 ...(tail ? [space(2), { text: tail, dim: true }] : []),
291 ]
292 return { key: 'summary', segs: line(width, segs, [], 0) }
293}
294
295function stateWordOf(item: Milestone, index: number, firstPending: number, isWorking: boolean): Seg {
296 if (item.state === 'done') return { text: 'Done', ...(isWorking ? { dim: true } : {}) }
297 if (!isWorking) return { text: 'Not reached', dim: true }
298 if (item.state === 'current') return { text: 'Working', color: ACCENT, bold: true }
299 return { text: index === firstPending ? 'Next' : 'Later', dim: true }
300}
301
302function stepRows(model: ViewModel, width: number, tone: BandTone): Row[] {
303 const plan = model.plan
304 if (!plan || plan.items.length === 0) return []
305 const items = plan.items
306 const current = currentIndexOf(plan)
307 const firstPending = items.findIndex((item, i) => item.state === 'pending' && i > current)
308 const start = items.length <= MAX_STEP_ROWS ? 0 : Math.max(0, Math.min(current - 1, items.length - MAX_STEP_ROWS))
309 const shown = items.slice(start, start + MAX_STEP_ROWS)
310 const isWide = width + FRAME >= NARROW_BODY
311 const longest = Math.max(...shown.map(item => [...item.label].length))
312 const wordWidth = 'Not reached'.length
313 const labelWidth = isWide
314 ? Math.max(8, Math.min(longest, Math.floor(width * 0.4), width - 2 - 2 - MINI_BAR - 2 - wordWidth))
315 : Math.max(4, Math.min(longest, width - 2 - 2 - wordWidth))
316 const stepProgress = model.estimate?.kind === 'range' ? model.estimate.stepProgress : []
317
318 const rows = shown.map((item, offset): Row => {
319 const i = start + offset
320 const glyph: Seg =
321 item.state === 'done'
322 ? { text: '✓', color: DONE }
323 : item.state === 'current' && model.isWorking
324 ? { text: '●', color: ACCENT }
325 : { text: '○', dim: true }
326 const label = truncate(item.label, labelWidth).padEnd(labelWidth)
327 const labelSeg: Seg =
328 item.state === 'current' && model.isWorking ? { text: label, bold: true } : item.state === 'pending' ? { text: label, dim: true } : { text: label }
329 const word = stateWordOf(item, i, firstPending, model.isWorking)
330 let bar: Seg[] = []
331 if (isWide) {
332 const doneColor = tone === 'working' || tone === 'over' || tone === 'done' ? DONE : OVER
333 bar =
334 item.state === 'done'
335 ? barSegs(1, MINI_BAR, doneColor, false, '█', '░')
336 : item.state === 'current' && model.isWorking
337 ? barSegs(stepProgress[i] ?? 0, MINI_BAR, tone === 'over' ? OVER : ACCENT, false, '█', '░')
338 : barSegs(0, MINI_BAR, OVER, false, '█', '░')
339 bar = [...bar, space(2)]
340 }
341 return { key: `step-${i}`, segs: line(width, [glyph, space(1), labelSeg, space(2), ...bar, word], [], 0) }
342 })
343 if (items.length > MAX_STEP_ROWS) {
344 rows.push({ key: 'more', segs: line(width, [{ text: `+${items.length - MAX_STEP_ROWS} more`, dim: true }], [], 0) })
345 }
346 return rows
347}
348
349function toneOf(model: ViewModel, verdict: Verdict): BandTone {
350 if (model.isWorking) return model.estimate?.kind === 'range' && model.estimate.isOver ? 'over' : 'working'
351 if (model.finished?.isAborted) return 'stopped'
352 if (verdict.kind === 'unverified' || verdict.kind === 'failed') return 'unverified'
353 return isAllDone(model.plan) ? 'done' : 'ended'
354}
355
356const BORDER: Record<BandTone, Color> = {
357 working: ACCENT,
358 over: OVER,
359 done: DONE,
360 ended: OVER,
361 unverified: WARN,
362 stopped: OVER,
363}
364
365/**
366 * The band: title, summary, one row per step (six at most, then `+n more`),
367 * the basis tooltip when asked, and a failing calibration score, which is
368 * never folded away (invariant 3: a bad score is displayed, not hidden).
369 * Collapsed, the title row alone. `columns` is the band's body width.
370 */
371export function bandView(model: ViewModel, columns: number, options: BandOptions): BandView {
372 const width = innerWidthOf(columns)
373 const verdict = model.isWorking ? ({ kind: 'none' } as const) : verdictOf(model.finished?.receipt ?? null)
374 const tone = toneOf(model, verdict)
375 const rows: Row[] = [titleRow(model, width, options, tone, verdict)]
376 if (!options.collapsed) {
377 rows.push(summaryRow(model, width, tone), ...stepRows(model, width, tone))
378 if (options.showBasis) {
379 const basis = model.isWorking && model.estimate ? (basisLines(model.estimate)[0] ?? '') : ''
380 const text = [basis, model.calibration[0] ?? ''].filter(Boolean).join(' · ')
381 rows.push({ key: 'basis', segs: line(width, [{ text, dim: true, grow: true }], [], 0) })
382 }
383 const warning = model.calibration[1]
384 if (warning) rows.push({ key: 'calibration', segs: line(width, [{ text: warning, color: WARN, grow: true }], [], 0) })
385 }
386 return { tone, border: BORDER[tone], rows }
387}
388
389// ---- the settings popover ---------------------------------------------------
390
391export const MODEL_CHOICES = [
392 { value: 'haiku', label: 'Haiku' },
393 { value: 'sonnet', label: 'Sonnet' },
394 { value: 'opus', label: 'Opus' },
395 { value: 'fable', label: 'Fable' },
396] as const
397
398export const EFFORT_CHOICES = [
399 { value: 'low', label: 'Low' },
400 { value: 'medium', label: 'Medium' },
401 { value: 'high', label: 'High' },
402 { value: 'xhigh', label: 'XHigh' },
403 { value: 'max', label: 'Max' },
404] as const
405
406/** The popover's widest, border and padding included. */
407export const TOOLS_COLUMNS = 64
408
409/**
410 * How much of the transcript the mod draws away. `off`: every row as Claude
411 * Code draws it. `clean`: tool rows one dim line, milestones in full.
412 * `quiet`: tool rows nothing, milestones one dim line, and the chrome a turn
413 * scatters (spinner, progress, notices, hand-backs, interim text) folded.
414 */
415export type RowsLevel = 'off' | 'clean' | 'quiet'
416export const ROWS_LEVELS: readonly RowsLevel[] = ['off', 'clean', 'quiet']
417
418export const ROWS_CHOICES = [
419 { value: 'off', label: 'Off' },
420 { value: 'clean', label: 'Clean' },
421 { value: 'quiet', label: 'Quiet' },
422] as const
423
424export type ToolsModel = {
425 /** What `$.session.model()` answered. */
426 model: string
427 /** The effort level last seen on a model request, if any. */
428 effort?: string
429 rows: RowsLevel
430 spike: boolean
431}
432
433/**
434 * The alias a model name answers to: `claude-opus-5-5[1m]` is `opus`.
435 */
436export function aliasOf(model: string): string | undefined {
437 const name = model.toLowerCase()
438 return MODEL_CHOICES.find(choice => name.includes(choice.value))?.value
439}
440
441function spaced(text: string): string {
442 return [...text.toUpperCase()].join(' ')
443}
444
445function chipRow(key: string, header: string, choices: readonly { value: string; label: string }[], current: string | undefined, width: number): Row {
446 const segs: Seg[] = [{ text: spaced(header).padEnd(13), dim: true }]
447 choices.forEach((choice, i) => {
448 if (i > 0) segs.push(space(1))
449 if (choice.value === current) segs.push({ text: ` ${choice.label} `, color: ACCENT, inverse: true, bold: true })
450 else segs.push(space(1), buttonSeg(`${key}-${choice.value}`, choice.label, undefined, false), space(1))
451 })
452 return { key, segs: line(width, segs, [], 0) }
453}
454
455function settingRow(key: string, name: string, help: string, isOn: boolean, width: number): Row {
456 const left: Seg[] = [
457 isOn ? { text: '●', color: DONE } : { text: '○', dim: true },
458 space(1),
459 { text: name.padEnd(19), bold: true },
460 { text: help, dim: true, grow: true },
461 ]
462 return { key: `${key}-row`, segs: line(width, left, [buttonSeg(key, isOn ? 'On' : 'Off', undefined, !isOn)]) }
463}
464
465function ruleOf(prefix: string, width: number): Seg[] {
466 const head = truncate(prefix, width)
467 return [{ text: head, dim: true }, { text: '─'.repeat(Math.max(0, width - [...head].length)), dim: true }]
468}
469
470/**
471 * The settings popover: model and effort chips (the current one inverted and
472 * not pressable), letter-spaced headers, and the mod's three switches.
473 * `width` is the popover's inner width.
474 */
475export function toolsView(model: ToolsModel, width: number): Row[] {
476 const alias = aliasOf(model.model)
477 const effort = model.effort?.toLowerCase()
478 const status = [model.model || 'model unknown', effort ? (EFFORT_CHOICES.find(c => c.value === effort)?.label ?? effort) : null].filter(Boolean).join(' · ')
479 return [
480 {
481 key: 'tools-head',
482 segs: line(
483 width,
484 [{ text: '◆ ', color: ACCENT }, { text: spaced('receipts'), color: ACCENT, bold: true }],
485 [{ text: truncate(status, Math.max(4, width - 26)), dim: true }, space(1), buttonSeg('tools-close', '[-]', undefined, false)],
486 ),
487 },
488 chipRow('model', 'model', MODEL_CHOICES, alias, width),
489 chipRow('effort', 'effort', EFFORT_CHOICES, effort, width),
490 chipRow('rows', 'rows', ROWS_CHOICES, model.rows, width),
491 { key: 'tools-rule', segs: ruleOf(`── ${spaced('settings')} `, width) },
492 settingRow('set-spike', 'Spike', 'sizing subagent, new tasks', model.spike, width),
493 ]
494}
495
496/**
497 * Rows inside a round border, as the terminal draws them, for mocks and
498 * docs. Plain text has no fill, so an inverted ` cell ` shows as `[cell]`,
499 * the same width.
500 */
501export function framedText(rows: readonly Row[], width: number): string[] {
502 const plain = (row: Row) =>
503 row.segs.map(seg => (seg.inverse && /^ .* $/.test(seg.text) ? `[${seg.text.slice(1, -1)}]` : seg.text)).join('')
504 return [`╭${'─'.repeat(width + 2)}╮`, ...rows.map(row => `│ ${plain(row)} │`), `╰${'─'.repeat(width + 2)}╯`]
505}
506src/host.ts 39 lines1import type {
2 AgentSpawnArgs,
3 AgentSpawnResult,
4 FsEntry,
5 ModelCompleteRequest,
6 ModelCompleteResult,
7 SessionRepo,
8 Timer,
9} from 'claude-code'
10
11/**
12 * Everything the session logic needs from Claude Code, as plain functions.
13 *
14 * `hooks/register.ts` builds one from `$` in a top-level `hostOf($)`, so
15 * every mods API call stays spelled out where `claude plugin validate` reads
16 * it, and nothing under `src/` touches `$`. A test can hand the session a
17 * fake host with a scripted clock and store.
18 */
19export type Host = {
20 now: () => Promise<number>
21 every: (ms: number, fn: () => void) => Timer
22 after: (ms: number, fn: () => void) => Timer
23 storeGet: (key: string) => Promise<unknown>
24 storeSet: (key: string, value: unknown) => Promise<void>
25 storeDelete: (key: string) => Promise<void>
26 /** Redraws the pane and the band (they read the tick); tool rows are left alone. */
27 redraw: () => void
28 classify: (text: string, labels: readonly string[]) => Promise<string | undefined>
29 complete: (request: ModelCompleteRequest) => Promise<ModelCompleteResult>
30 spawn: (args: AgentSpawnArgs) => Promise<AgentSpawnResult>
31 cwd: () => Promise<string>
32 repo: () => Promise<SessionRepo | null>
33 list: (path: string) => Promise<FsEntry[]>
34 /** The session's agents, as `$.agent.list()` answers. */
35 agents: () => Promise<readonly { id: string; status: string; parentId?: string; spawnedBy?: string }[]>
36 /** The session's usage, as `$.session.usage()` answers: rate-limit windows and cost. */
37 usage: () => Promise<{ rateLimits: readonly { kind: string; percentUsed: number; resetsAt?: string }[]; cost?: { usd: number } }>
38}
39src/quiet.ts 76 lines1/**
2 * The quiet level's memory of assistant text: which blocks were interim
3 * (drawn as their first line, dim) and which was the turn's final answer
4 * (redrawn in full when the turn completes). Keyed by the render's
5 * `requestId`, the block's own id, so a redraw finds the same verdict.
6 *
7 * Drawing only (invariant 4): nothing here touches the stored messages, and
8 * a block the mod never saw during a turn (history, a resumed session) is
9 * left to Claude Code.
10 */
11
12/** Ids kept per kind before the oldest are forgotten; a long session stays small. */
13const MAX_IDS = 2_000
14
15export class QuietLog {
16 private readonly turnBlocks = new Map<string, string>()
17 private readonly interim = new Set<string>()
18 private readonly final = new Set<string>()
19
20 /** A block drawn while a turn runs: interim until the turn says otherwise. */
21 seen(id: string, text: string): void {
22 if (this.final.has(id)) return
23 this.turnBlocks.set(id, text)
24 this.interim.add(id)
25 trim(this.interim)
26 }
27
28 /**
29 * The turn ended with `answer`: the blocks whose text is part of it are the
30 * final answer; failing any, the last block drawn. Returns whether any
31 * verdict changed, so the caller knows to redraw.
32 */
33 finish(answer: string): boolean {
34 const blocks = [...this.turnBlocks.entries()]
35 this.turnBlocks.clear()
36 const flat = answer.trim()
37 let finals = blocks.filter(([, text]) => text.trim() !== '' && flat.includes(text.trim())).map(([id]) => id)
38 if (finals.length === 0 && blocks.length > 0) finals = [blocks[blocks.length - 1]![0]]
39 for (const id of finals) {
40 this.interim.delete(id)
41 this.final.add(id)
42 }
43 trim(this.final)
44 return finals.length > 0
45 }
46
47 kindOf(id: string): 'interim' | 'final' | 'unknown' {
48 if (this.final.has(id)) return 'final'
49 if (this.interim.has(id)) return 'interim'
50 return 'unknown'
51 }
52}
53
54function trim(ids: Set<string>): void {
55 while (ids.size > MAX_IDS) {
56 const oldest = ids.values().next().value
57 if (oldest === undefined) return
58 ids.delete(oldest)
59 }
60}
61
62/** The first non-empty line of a block, markdown left as is. */
63export function firstLineOf(text: string): string {
64 return text.split('\n').find(line => line.trim() !== '')?.trim() ?? ''
65}
66
67/**
68 * The one dim line a hand-back or notification row becomes once its turn
69 * ended: `↳ message from @Explore: Found 3 mods`.
70 */
71export function handbackLineOf(text: string, from: string | undefined): string {
72 const head = firstLineOf(text)
73 const who = from ? `message from @${from}` : 'notification'
74 return head ? `↳ ${who}: ${head}` : `↳ ${who}`
75}
76src/session.ts 830 lines1/**
2 * One session's receipts: the task under way, the history it learns into,
3 * and the view model the pane and band draw. Hooks call in through a Host,
4 * never `$` (see host.ts).
5 *
6 * Invariant 6, hooks return fast: every model call here starts unawaited
7 * (plan derivation, task type, step-done, the spike) except the claims-done
8 * label on turn.complete, which waits at most CLAIMS_DEADLINE_MS and then
9 * falls back to plain words in the answer. Results are cached per step per
10 * turn, so a long turn costs one classify per model step at most.
11 */
12
13import { estimate, type Estimate } from './estimator'
14import {
15 addTask,
16 bucketSamples,
17 calibrationLines,
18 calibrationOf,
19 emptyHistory,
20 HISTORY_KEY,
21 parseHistory,
22 statsOf,
23 type History,
24 type Stats,
25} from './history'
26import type { Host } from './host'
27import {
28 emptyLedger,
29 exitCodeOf,
30 isVerifyCommand,
31 looksDone,
32 receiptLine,
33 recordEdit,
34 recordVerify,
35 type Ledger,
36 type ToolOutcome,
37} from './ledger'
38import {
39 completeAtEnd,
40 completeCurrent,
41 derivedPlan,
42 editedPathOf,
43 emptyPlan,
44 fallbackPlan,
45 finishPlan,
46 isComplete,
47 isEditTool,
48 isMilestoneCommand,
49 isSubagentTool,
50 onTaskCreate,
51 onTaskUpdate,
52 onTodoWrite,
53 parseDerivedSteps,
54 stepDurations,
55 type Pause,
56 type Plan,
57 type TaskUpdateArgs,
58 type Todo,
59} from './milestones'
60import {
61 countTool,
62 NO_TOOLS,
63 repoHashOf,
64 stepBucketOf,
65 TASK_TYPES,
66 toolMixOf,
67 type Shape,
68 type TaskType,
69 type ToolCounts,
70} from './shape'
71import { parseSpike, SPIKE_AGENT, SPIKE_MIN_STEPS, SPIKE_MODEL, SPIKE_TIMEOUT_MS, spikePromptOf } from './spike'
72import type { ViewModel } from './view'
73
74export const PANE_ID = 'receipts'
75export const PANE_TITLE = 'receipts'
76export const CLAIMS_DEADLINE_MS = 1_500
77export const TICK_MS = 1_000
78export const CLAIMS_LABELS = ['claims-done', 'not-claiming-done'] as const
79export const STEP_LABELS = ['step-done', 'step-not-done'] as const
80/** Task types whose steps are cheap: a spike would cost more than it tells. */
81export const NO_SPIKE_TYPES: readonly string[] = ['research', 'writing', 'chat']
82export const AUDIT_KEY = 'audit'
83const STEPS_DONE_SYSTEM =
84 'You read the final message of a coding assistant and a list of planned steps not yet marked done. ' +
85 'Answer with the numbers of the steps the message shows were completed, comma separated, or "none". No other text.'
86const PLAN_MODEL = 'haiku'
87const PLAN_TIMEOUT_MS = 20_000
88const PLAN_SYSTEM =
89 'You read a request to a coding assistant and its first message, and list the plan it is following ' +
90 'as 2 to 8 steps, one per line, each a short noun phrase of at most 8 words. No other text.'
91
92export type SessionOptions = {
93 spike: boolean
94}
95
96type Snapshot = {
97 low: number
98 high: number
99}
100
101type TaskState = {
102 turnId: string
103 startedAt: number
104 prompt: string
105 cwd: string
106 plan: Plan
107 counts: ToolCounts
108 ledger: Ledger
109 taskType: TaskType
110 /** Settles once the task-type label is in (or failed). */
111 typed: Promise<void>
112 isDeriving: boolean
113 isSpiked: boolean
114 spike?: readonly number[]
115 checkedSteps: Set<number>
116 seenEdits: Set<string>
117 claims?: Promise<boolean | undefined>
118 /** The last range shown while 2 or more steps were left. */
119 beforeFinal?: Snapshot
120 firstRange?: Snapshot
121 /** Permission waits inside the task, subtracted from what is learned. */
122 pauses: Pause[]
123 /** The session's cost when the task started, for the turn's share. */
124 usdAtStart?: number
125}
126
127type Audit = {
128 /** Receipt lines shown under an answer that claimed done. */
129 shown: number
130 /** Of those, the ones the person marked wrong with /receipts wrong. */
131 wrong: number
132}
133
134type StepEvent = {
135 turnId: string
136 index: number
137 agentId?: string
138}
139
140type StepResult = {
141 answer: string
142 toolUses: readonly { name: string }[]
143 stopReason: string | null
144}
145
146type ToolEvent = {
147 tool: string
148 tool_use_id: string
149}
150
151type CompleteEvent = {
152 turnId: string
153 answer: string
154 isAborted: boolean
155 reason: string
156}
157
158export class ReceiptsSession {
159 private history: History = emptyHistory()
160 private stats: Stats = statsOf(this.history)
161 private isLoaded = false
162 private probing: Promise<void> | null = null
163 private repoKey = 'unknown'
164 private hasTests = false
165 private repoEntries: string[] = []
166 private task: TaskState | null = null
167 private last: { plan: Plan; prompt: string; totalMs: number; receipt: string | null; isAborted: boolean } | null = null
168 /** Agents the main loop started that were still running when its turn ended. */
169 private waiting = new Set<string>()
170 /** Shapes already spiked this session: one guess per shape is enough. */
171 private readonly spikedShapes = new Set<string>()
172 private readonly toolStarts = new Map<string, number>()
173 private readonly toolEnds = new Map<string, number>()
174 private readonly toolRuns = new Map<string, number>()
175 private lastTickAt = 0
176 private audit: Audit = { shown: 0, wrong: 0 }
177 private lastReceipt: { turnId: string; isMarked: boolean } | null = null
178 private readonly milestoneIds = new Set<string>()
179 private readonly spikeWaiters = new Map<string, TaskState>()
180 private ticker: { cancel: () => void } | null = null
181 private isPaneOpen = false
182 private isPaneShown = false
183 private showBasis = false
184 private cwd = ''
185 private effort: string | undefined
186
187 constructor(private readonly options: SessionOptions) {}
188
189 async start(host: Host): Promise<void> {
190 await this.load(host)
191 this.probing ??= this.probe(host)
192 }
193
194 async turnStart(host: Host, turnId: string, text: string): Promise<void> {
195 await this.load(host)
196 this.probing ??= this.probe(host)
197 const now = await host.now()
198 this.cwd = await host.cwd().catch(() => this.cwd)
199 const task: TaskState = {
200 turnId,
201 startedAt: now,
202 prompt: text,
203 cwd: this.cwd,
204 plan: emptyPlan(),
205 counts: NO_TOOLS,
206 ledger: emptyLedger(),
207 taskType: 'unknown',
208 typed: Promise.resolve(),
209 isDeriving: false,
210 isSpiked: false,
211 checkedSteps: new Set(),
212 seenEdits: new Set(),
213 pauses: [],
214 }
215 this.task = task
216 this.last = null
217 // /receipts wrong is about the current turn's receipt, never an older one
218 this.lastReceipt = null
219 // The fallback for an agent whose end never reached us: the next prompt
220 this.waiting.clear()
221 if (text.trim()) {
222 task.typed = host
223 .classify(text.slice(0, 4_000), TASK_TYPES)
224 .then(label => {
225 task.taskType = (TASK_TYPES as readonly string[]).includes(label ?? '') ? (label as TaskType) : 'unknown'
226 })
227 .catch(() => undefined)
228 }
229 this.startTicker(host, now)
230 const usage = await host.usage().catch(() => null)
231 if (usage?.cost) task.usdAtStart = usage.cost.usd
232 // No pane opens unasked: the band above the prompt is the surface, and
233 // `/receipts` opens the pane for the long view. Scar: the auto-opened
234 // dock took a third of a fullscreen terminal to show four lines.
235 host.redraw()
236 }
237
238 /**
239 * After each model request: derive a plan if the first step brought no
240 * task tools, ask whether the current derived step is done, and start the
241 * claims-done label as soon as the final answer exists.
242 */
243 stepResult(host: Host, e: StepEvent, result: StepResult): void {
244 const task = this.task
245 if (!task || e.agentId !== undefined || e.turnId !== task.turnId) return
246 const usesTaskTools = result.toolUses.some(use => use.name === 'TaskCreate' || use.name === 'TodoWrite')
247 if (task.plan.source === 'none' && !task.isDeriving && !usesTaskTools) {
248 task.isDeriving = true
249 void this.derive(host, task, result.answer)
250 }
251 const answer = result.answer.trim()
252 if (task.plan.source === 'derived' && answer && !task.checkedSteps.has(e.index)) {
253 task.checkedSteps.add(e.index)
254 void this.checkStep(host, task, answer)
255 }
256 if (result.stopReason === 'end_turn' && answer && task.ledger.edits.length > 0) {
257 task.claims = this.claimsDone(host, answer)
258 }
259 }
260
261 /**
262 * Before a tool runs, decide whether its rows are milestone rows, so they
263 * draw in full while running and their result rows (which carry no input)
264 * can be told by id: the first edit of each file this turn, a verify run,
265 * a commit, a pull request, a subagent.
266 */
267 beforeTool(e: ToolEvent, input: unknown): void {
268 if (this.isMilestoneRow(e.tool_use_id, e.tool, input) && !isEditTool(e.tool)) {
269 this.milestoneIds.add(e.tool_use_id)
270 return
271 }
272 const task = this.task
273 if (!task || !isEditTool(e.tool)) return
274 const path = editedPathOf(e.tool, input)
275 if (path === undefined || task.seenEdits.has(path)) return
276 task.seenEdits.add(path)
277 this.milestoneIds.add(e.tool_use_id)
278 }
279
280 async afterTool(host: Host, e: ToolEvent, input: unknown, outcome: ToolOutcome): Promise<void> {
281 const task = this.task
282 if (!task) return
283 const now = await host.now()
284 if (this.toolStarts.has(e.tool_use_id)) {
285 this.toolEnds.set(e.tool_use_id, now)
286 this.settlePause(e.tool_use_id)
287 }
288 task.counts = countTool(task.counts, e.tool)
289 const fields = (typeof input === 'object' && input !== null ? input : {}) as Record<string, unknown>
290 const result = outcome.result as Record<string, unknown> | undefined
291 const isOk = outcome.deny === undefined && !outcome.isError
292 let isPlanChanged = false
293
294 if (e.tool === 'TaskCreate' && isOk) {
295 const created = result?.task as { id?: unknown } | undefined
296 if (typeof created?.id === 'string' && typeof fields.subject === 'string') {
297 task.plan = onTaskCreate(task.plan, created.id, fields.subject, now)
298 isPlanChanged = true
299 }
300 } else if (e.tool === 'TaskUpdate' && isOk && typeof fields.taskId === 'string') {
301 task.plan = onTaskUpdate(task.plan, fields as unknown as TaskUpdateArgs, now)
302 isPlanChanged = true
303 } else if (e.tool === 'TodoWrite' && isOk && Array.isArray(fields.todos)) {
304 task.plan = onTodoWrite(task.plan, fields.todos as Todo[], now)
305 isPlanChanged = true
306 } else if (isEditTool(e.tool) && isOk) {
307 const path = editedPathOf(e.tool, input)
308 if (path !== undefined) task.ledger = recordEdit(task.ledger, path, now)
309 } else if (e.tool === 'Bash' && outcome.deny === undefined && typeof fields.command === 'string') {
310 const interrupted = (result as { interrupted?: unknown } | undefined)?.interrupted === true
311 if (isVerifyCommand(fields.command) && !interrupted) {
312 const stdout = typeof result?.stdout === 'string' ? result.stdout : ''
313 const stderr = typeof result?.stderr === 'string' ? result.stderr : ''
314 const output = outcome.text ?? [stdout, stderr].join('\n')
315 task.ledger = recordVerify(task.ledger, fields.command, exitCodeOf(outcome), output, now)
316 }
317 }
318
319 if (isPlanChanged) {
320 this.observe(now)
321 void this.maybeSpike(host, task)
322 host.redraw()
323 }
324 }
325
326 /**
327 * The turn ended: the receipt line, if any, then learning. Never blocks or
328 * aborts the turn; at worst it waits CLAIMS_DEADLINE_MS for a label.
329 */
330 async turnComplete(host: Host, e: CompleteEvent): Promise<string | null> {
331 const task = this.task
332 if (!task) return null
333 // Any main-loop turn.complete ends the task the band shows, even one whose
334 // id the band never saw start (a reload, a second copy of the mod). Scar:
335 // a band left in Working after the turn had ended. Only a matching id is
336 // learned from.
337 const isOwn = task.turnId === e.turnId
338 this.task = null
339 this.ticker?.cancel()
340 this.ticker = null
341 const now = await host.now()
342
343 // Both labels run at once, under one deadline: the claims-done label, and
344 // for a derived plan which of its open steps the final answer completed
345 const hasEdits = task.ledger.edits.length > 0 && !e.isAborted
346 const [claims, finishedSteps] = await Promise.all([
347 hasEdits ? this.withDeadline(host, task.claims ?? this.claimsDone(host, e.answer), CLAIMS_DEADLINE_MS) : Promise.resolve(undefined),
348 e.isAborted ? Promise.resolve(undefined) : this.withDeadline(host, this.stepsDoneBy(host, task, e.answer), CLAIMS_DEADLINE_MS),
349 ])
350 let line: string | null = null
351 if (hasEdits) line = receiptLine(task.ledger, claims ?? looksDone(e.answer), now, task.cwd)
352
353 const answered = finishedSteps && finishedSteps.length > 0 ? completeAtEnd(task.plan, finishedSteps, now, task.startedAt) : task.plan
354 const plan = finishPlan(answered.source === 'none' ? fallbackPlan(task.startedAt) : answered, now)
355 const totalMs = Math.max(0, now - task.startedAt)
356 this.last = { plan, prompt: task.prompt, totalMs, receipt: line, isAborted: e.isAborted }
357 let shown = line
358 if (line) {
359 this.audit = { ...this.audit, shown: this.audit.shown + 1 }
360 this.lastReceipt = { turnId: task.turnId, isMarked: false }
361 await host.storeSet(AUDIT_KEY, this.audit).catch(() => undefined)
362 const cost = await this.costOf(host, task)
363 if (cost) shown = `${line} · ${cost}`
364 }
365
366 // A task-tool plan finished only when every item did; a derived or
367 // fallback plan finished when the turn answered.
368 const isFinished = plan.source === 'tasks' ? isComplete(plan) : true
369 if (isOwn && !e.isAborted && e.reason === 'answer' && isFinished) {
370 const range = task.beforeFinal ?? task.firstRange
371 // Learned as work: permission waits come out of the steps and the total.
372 // Calibration is scored on the wall clock, as the range was shown
373 const steps = stepDurations(plan, task.startedAt, task.pauses)
374 const paused = task.pauses.reduce((sum, pause) => sum + pause.ms, 0)
375 this.history = addTask(this.history, {
376 at: now,
377 shape: { ...this.shapeOf(task), steps: stepBucketOf(Math.max(1, plan.items.length)) },
378 steps,
379 totalMs: Math.max(0, totalMs - paused),
380 ...(range ? { inside: totalMs >= range.low && totalMs <= range.high } : {}),
381 })
382 this.stats = statsOf(this.history)
383 await host.storeSet(HISTORY_KEY, this.history).catch(() => undefined)
384 }
385 await this.countWaiting(host)
386 host.redraw()
387 return shown
388 }
389
390 /**
391 * One pass over the final answer for a derived plan: which of the steps not
392 * yet done does it show were completed? Indexes into the plan; none for a
393 * plan of another source, or when nothing is open.
394 */
395 private async stepsDoneBy(host: Host, task: TaskState, answer: string): Promise<number[]> {
396 if (task.plan.source !== 'derived' || !answer.trim()) return []
397 const open = task.plan.items.map((item, i) => ({ item, i })).filter(({ item }) => item.state === 'pending')
398 if (open.length === 0) return []
399 try {
400 const reply = await host.complete({
401 model: PLAN_MODEL,
402 system: STEPS_DONE_SYSTEM,
403 prompt: `Steps not yet marked done:\n${open.map(({ item, i }) => `${i + 1}. ${item.label}`).join('\n')}\n\nThe final message:\n${answer.slice(0, 4_000)}`,
404 maxTokens: 40,
405 timeoutMs: CLAIMS_DEADLINE_MS,
406 })
407 if (!reply.isAnswered) return []
408 const wanted = new Set(open.map(({ i }) => i))
409 return [...reply.text.matchAll(/\d+/g)].map(match => Number(match[0]) - 1).filter(i => wanted.has(i))
410 } catch {
411 return []
412 }
413 }
414
415 /**
416 * What the turn cost, for the receipt line: a plan user's five-hour window
417 * and its reset, an API user's dollars this turn, or nothing when neither
418 * is known. Never a guess.
419 */
420 private async costOf(host: Host, task: TaskState): Promise<string | null> {
421 const usage = await host.usage().catch(() => null)
422 if (!usage) return null
423 const window = usage.rateLimits.find(limit => limit.kind === 'five_hour')
424 if (window) {
425 const resets = window.resetsAt ? new Date(window.resetsAt) : null
426 const at = resets && !Number.isNaN(resets.getTime()) ? `, resets ${String(resets.getHours()).padStart(2, '0')}:${String(resets.getMinutes()).padStart(2, '0')}` : ''
427 return `5h window ${window.percentUsed}% used${at}`
428 }
429 if (usage.rateLimits.length === 0 && usage.cost && task.usdAtStart !== undefined) {
430 const spent = usage.cost.usd - task.usdAtStart
431 if (spent > 0) return `$${spent.toFixed(2)} this turn`
432 }
433 return null
434 }
435
436 /** The person says this turn's receipt was wrong: one mark per receipt. */
437 async markWrong(host: Host): Promise<string> {
438 if (!this.lastReceipt) return 'no receipt this session to mark'
439 if (this.lastReceipt.isMarked) return 'already marked wrong'
440 this.lastReceipt.isMarked = true
441 this.audit = { ...this.audit, wrong: this.audit.wrong + 1 }
442 await host.storeSet(AUDIT_KEY, this.audit).catch(() => undefined)
443 return `marked wrong · ${this.auditLine()}`
444 }
445
446 private auditLine(): string {
447 const { shown, wrong } = this.audit
448 const percent = shown === 0 ? 0 : Math.round((wrong / shown) * 100)
449 return `claims-done false positives: ${wrong} of ${shown} ${shown === 1 ? 'receipt' : 'receipts'} (${percent}%)`
450 }
451
452 /** A tool call began: its start, for the permission wait. */
453 toolStarted(toolUseId: string, at: number): void {
454 this.toolStarts.set(toolUseId, at)
455 }
456
457 /** The tool's own run time, from classic PostToolUse (no prompt or hook time). */
458 toolRan(toolUseId: string, ms: number | undefined): void {
459 if (typeof ms !== 'number' || !Number.isFinite(ms)) return
460 this.toolRuns.set(toolUseId, ms)
461 this.settlePause(toolUseId)
462 }
463
464 /**
465 * Once a call's span and run time are both in: the rest is waiting. Only a
466 * wait of a second or more counts, so hook overhead is never a pause.
467 */
468 private settlePause(toolUseId: string): void {
469 const start = this.toolStarts.get(toolUseId)
470 const end = this.toolEnds.get(toolUseId)
471 const run = this.toolRuns.get(toolUseId)
472 if (start === undefined || end === undefined || run === undefined) return
473 this.toolStarts.delete(toolUseId)
474 this.toolEnds.delete(toolUseId)
475 this.toolRuns.delete(toolUseId)
476 const wait = end - start - run
477 if (wait >= 1_000 && this.task) this.task.pauses.push({ at: end, ms: wait })
478 }
479
480 /**
481 * After a hot reload the module's timers die while a drawing stays up. A
482 * render calls this: when a task is under way and nothing has ticked for
483 * two intervals, the ticker starts again.
484 */
485 ensureTicking(host: Host, now: number): void {
486 if (!this.task || now - this.lastTickAt <= 2 * TICK_MS) return
487 this.startTicker(host, now)
488 }
489
490 private startTicker(host: Host, now: number): void {
491 this.ticker?.cancel()
492 this.lastTickAt = now
493 this.ticker = host.every(TICK_MS, () => {
494 void this.tick(host)
495 })
496 }
497
498 /**
499 * The main loop's agents still going: started by the model or the person
500 * (not by this mod, so not the spike), not finished.
501 */
502 private async countWaiting(host: Host): Promise<void> {
503 const agents = await host.agents().catch(() => [])
504 this.waiting = new Set(
505 agents
506 .filter(agent => agent.parentId === undefined && agent.spawnedBy !== 'receipts')
507 .filter(agent => agent.status === 'pending' || agent.status === 'running' || agent.status === 'waiting')
508 .map(agent => agent.id),
509 )
510 }
511
512 /**
513 * A subagent's turn ended; if it was this mod's spike, read its guess.
514 */
515 subagentComplete(host: Host, agentId: string, answer: string): void {
516 if (this.waiting.delete(agentId)) host.redraw()
517 const task = this.spikeWaiters.get(agentId)
518 if (!task) return
519 this.spikeWaiters.delete(agentId)
520 const guess = parseSpike(answer, task.plan.items.length)
521 if (guess && this.task === task) {
522 task.spike = guess.stepsMs
523 host.redraw()
524 }
525 }
526
527 // ---- views ---------------------------------------------------------------
528
529 viewModel(now: number): ViewModel {
530 const task = this.task
531 const est = task ? this.observe(now) : null
532 return {
533 plan: task ? task.plan : (this.last?.plan ?? null),
534 estimate: est,
535 isWorking: task !== null,
536 showBasis: this.showBasis,
537 calibration: calibrationLines(calibrationOf(this.history)),
538 title: task ? task.prompt : (this.last?.prompt ?? ''),
539 elapsedMs: task ? Math.max(0, now - task.startedAt) : (this.last?.totalMs ?? 0),
540 finished: this.last
541 ? { totalMs: this.last.totalMs, receipt: this.last.receipt, isAborted: this.last.isAborted, waitingAgents: this.waiting.size }
542 : null,
543 }
544 }
545
546 isMilestoneRow(toolUseId: string, tool: string, input: unknown): boolean {
547 if (this.milestoneIds.has(toolUseId)) return true
548 if (isEditTool(tool)) return false
549 if (isSubagentTool(tool)) return true
550 if (tool === 'Bash') {
551 const command = (input as { command?: unknown } | null | undefined)?.command
552 return typeof command === 'string' && isMilestoneCommand(command)
553 }
554 return false
555 }
556
557 get workingDirectory(): string {
558 return this.cwd
559 }
560
561 /**
562 * The band shows while a task runs, and after it as the completion card
563 * until the next prompt, whenever no pane is placed.
564 */
565 isBandWanted(): boolean {
566 return (this.task !== null || this.last !== null) && !this.isPaneShown
567 }
568
569 get isWorking(): boolean {
570 return this.task !== null
571 }
572
573 /** The user's switch for the spike, over the userConfig default. */
574 setSpike(isOn: boolean): void {
575 this.options.spike = isOn
576 }
577
578 get isSpikeOn(): boolean {
579 return this.options.spike
580 }
581
582 /** The effort level the last main-loop model request carried. */
583 noteEffort(effort: string | number | undefined): void {
584 if (typeof effort === 'string') this.effort = effort
585 }
586
587 get effortLevel(): string | undefined {
588 return this.effort
589 }
590
591 paneDrawn(): void {
592 this.isPaneShown = true
593 this.isPaneOpen = true
594 }
595
596 paneClosed(): void {
597 this.isPaneShown = false
598 this.isPaneOpen = false
599 }
600
601 get paneOpen(): boolean {
602 return this.isPaneOpen
603 }
604
605 paneOpened(isPlaced: boolean): void {
606 this.isPaneOpen = true
607 this.isPaneShown = isPlaced
608 }
609
610 toggleBasis(): void {
611 this.showBasis = !this.showBasis
612 }
613
614 get isBasisShown(): boolean {
615 return this.showBasis
616 }
617
618 // ---- commands ------------------------------------------------------------
619
620 statsText(): string {
621 const tasks = this.history.tasks
622 const audit = `${this.auditLine()}; mark a wrong one with /receipts wrong`
623 if (tasks.length === 0) return ['no finished tasks yet; the estimate runs on its prior until a few land', audit].join('\n')
624 const byType = new Map<string, number[]>()
625 for (const task of tasks) {
626 const list = byType.get(task.shape.taskType) ?? []
627 list.push(task.totalMs)
628 byType.set(task.shape.taskType, list)
629 }
630 const rows = [...byType.entries()]
631 .sort((a, b) => b[1].length - a[1].length)
632 .map(([type, list]) => `${type} ${list.length} (median ${minutesOf(medianOf(list))})`)
633 return [
634 `${tasks.length} finished ${tasks.length === 1 ? 'task' : 'tasks'} in history`,
635 ...calibrationLines(calibrationOf(this.history)),
636 `by type: ${rows.join(', ')}`,
637 audit,
638 ].join('\n')
639 }
640
641 async resetHistory(host: Host): Promise<number> {
642 const count = this.history.tasks.length
643 await host.storeDelete(HISTORY_KEY)
644 this.history = emptyHistory()
645 this.stats = statsOf(this.history)
646 host.redraw()
647 return count
648 }
649
650 // ---- internals -----------------------------------------------------------
651
652 private async load(host: Host): Promise<void> {
653 if (this.isLoaded) return
654 this.isLoaded = true
655 const raw = await host.storeGet(HISTORY_KEY).catch(() => undefined)
656 this.history = parseHistory(raw)
657 const audit = (await host.storeGet(AUDIT_KEY).catch(() => undefined)) as Partial<Audit> | undefined
658 if (typeof audit?.shown === 'number' && typeof audit.wrong === 'number') this.audit = { shown: audit.shown, wrong: audit.wrong }
659 this.stats = statsOf(this.history)
660 }
661
662 private async probe(host: Host): Promise<void> {
663 try {
664 const repo = await host.repo().catch(() => null)
665 const root = repo?.root ?? (await host.cwd())
666 this.repoKey = repoHashOf(root)
667 const entries = await host.list(root).catch(() => [])
668 const names = entries.map(entry => entry.name)
669 this.repoEntries = names.slice(0, 40)
670 this.hasTests = names.some(name => /^(tests?|__tests__|specs?)$/.test(name) || /\.(test|spec)\.\w+$/.test(name))
671 } catch {
672 // Unknown repo: the shape still works, it just learns under 'unknown'
673 }
674 }
675
676 private shapeOf(task: TaskState): Shape {
677 return {
678 taskType: task.taskType,
679 steps: stepBucketOf(Math.max(1, task.plan.items.length)),
680 repo: this.repoKey,
681 hasTests: this.hasTests,
682 mix: toolMixOf(task.counts),
683 }
684 }
685
686 private estimateOf(task: TaskState, now: number): Estimate {
687 return estimate({
688 now,
689 startedAt: task.startedAt,
690 plan: task.plan,
691 stats: this.stats,
692 shape: this.shapeOf(task),
693 ...(task.spike ? { spike: task.spike } : {}),
694 })
695 }
696
697 /**
698 * Computes the estimate for now and keeps the ranges calibration is scored
699 * against: the first one shown, and the last one shown before the final step.
700 */
701 private observe(now: number): Estimate | null {
702 const task = this.task
703 if (!task) return null
704 const est = this.estimateOf(task, now)
705 if (est.kind === 'range') {
706 const snapshot = { low: est.totalLowMs, high: est.totalHighMs }
707 task.firstRange ??= snapshot
708 const left = task.plan.items.filter(item => item.state !== 'done').length
709 if (left >= 2) task.beforeFinal = snapshot
710 }
711 return est
712 }
713
714 private async tick(host: Host): Promise<void> {
715 if (!this.task) return
716 const now = await host.now()
717 this.lastTickAt = now
718 this.observe(now)
719 host.redraw()
720 }
721
722 private async derive(host: Host, task: TaskState, firstText: string): Promise<void> {
723 let labels: string[] = []
724 if (task.prompt.trim()) {
725 try {
726 const reply = await host.complete({
727 model: PLAN_MODEL,
728 system: PLAN_SYSTEM,
729 prompt: `The request:\n${task.prompt.slice(0, 4_000)}\n\nThe assistant's first message:\n${firstText.slice(0, 2_000) || '(none yet)'}`,
730 maxTokens: 300,
731 timeoutMs: PLAN_TIMEOUT_MS,
732 })
733 if (reply.isAnswered) labels = parseDerivedSteps(reply.text)
734 } catch {
735 // No plan from the model: fall back to the single milestone
736 }
737 }
738 // Task tools may have arrived meanwhile; they win
739 if (this.task !== task || task.plan.source !== 'none') return
740 task.plan = labels.length > 0 ? derivedPlan(labels, task.startedAt) : fallbackPlan(task.startedAt)
741 this.observe(await host.now())
742 void this.maybeSpike(host, task)
743 host.redraw()
744 }
745
746 private async checkStep(host: Host, task: TaskState, answer: string): Promise<void> {
747 const current = task.plan.items.find(item => item.state === 'current')
748 if (!current) return
749 const label = await host
750 .classify(`Planned step: "${current.label}"\n\nThe assistant's latest message:\n${answer.slice(0, 3_000)}`, STEP_LABELS)
751 .catch(() => undefined)
752 if (label !== 'step-done' || this.task !== task) return
753 if (task.plan.items.find(item => item.state === 'current')?.id !== current.id) return
754 const now = await host.now()
755 task.plan = completeCurrent(task.plan, now)
756 this.observe(now)
757 host.redraw()
758 }
759
760 private claimsDone(host: Host, answer: string): Promise<boolean | undefined> {
761 return host
762 .classify(`The final message of a coding assistant to its user:\n${answer.slice(0, 4_000)}`, CLAIMS_LABELS)
763 .then(label => (label === 'claims-done' ? true : label === 'not-claiming-done' ? false : undefined))
764 .catch(() => undefined)
765 }
766
767 /**
768 * Once per task, for an unfamiliar shape with a plan of SPIKE_MIN_STEPS or
769 * more, and only when the user has not turned the spike off.
770 */
771 private async maybeSpike(host: Host, task: TaskState): Promise<void> {
772 if (!this.options.spike || task.isSpiked) return
773 if (task.plan.items.length < SPIKE_MIN_STEPS) return
774 if (task.plan.source === 'none' || task.plan.source === 'fallback') return
775 task.isSpiked = true
776 await Promise.all([this.probing, task.typed])
777 if (bucketSamples(this.stats, this.shapeOf(task)) >= 3) return
778 // Once per shape per session, and never for cheap task types. Scar: a
779 // cheap subagent per prompt in an unfamiliar repo (the 0.2.0 live run)
780 if (NO_SPIKE_TYPES.includes(task.taskType)) return
781 const shape = this.shapeOf(task)
782 const shapeKey = `${shape.taskType}|${shape.steps}|${shape.repo}`
783 if (this.spikedShapes.has(shapeKey)) return
784 this.spikedShapes.add(shapeKey)
785 try {
786 const spawned = await host.spawn({
787 prompt: spikePromptOf(task.prompt, task.plan.items.map(item => item.label), this.repoEntries),
788 description: 'receipts: size this task',
789 model: SPIKE_MODEL,
790 subagentType: SPIKE_AGENT,
791 })
792 if (spawned.deny !== undefined || spawned.agentId === undefined) return
793 const agentId = spawned.agentId
794 this.spikeWaiters.set(agentId, task)
795 host.after(SPIKE_TIMEOUT_MS, () => {
796 this.spikeWaiters.delete(agentId)
797 })
798 } catch {
799 // A refused or failed spike leaves the prior in place
800 }
801 }
802
803 private withDeadline<T>(host: Host, promise: Promise<T>, ms: number): Promise<T | undefined> {
804 return new Promise(resolve => {
805 const timer = host.after(ms, () => resolve(undefined))
806 promise.then(
807 value => {
808 timer.cancel()
809 resolve(value)
810 },
811 () => {
812 timer.cancel()
813 resolve(undefined)
814 },
815 )
816 })
817 }
818}
819
820function medianOf(values: readonly number[]): number {
821 const sorted = [...values].sort((a, b) => a - b)
822 const mid = Math.floor(sorted.length / 2)
823 return sorted.length % 2 === 1 ? sorted[mid]! : (sorted[mid - 1]! + sorted[mid]!) / 2
824}
825
826function minutesOf(ms: number): string {
827 const seconds = Math.round(ms / 1000)
828 return seconds < 60 ? `${seconds}s` : `${Math.floor(seconds / 60)}m ${seconds % 60}s`
829}
830src/view.ts 168 lines1/**
2 * What the pane and the clean-view rows say, as plain text lines, and the
3 * view model the band draws from (the band itself is band.ts).
4 * `hooks/register.ts` turns each line into a Text element; nothing here knows
5 * about elements or `$`.
6 */
7
8import { basisLines, estimateRow, type Estimate } from './estimator'
9import { relativeOf } from './ledger'
10import { glyphOf, type Plan } from './milestones'
11
12export type Line = {
13 key: string
14 text: string
15 dim?: boolean
16 bold?: boolean
17 color?: string
18}
19
20export type ViewModel = {
21 /** The plan being worked, or the last finished one. */
22 plan: Plan | null
23 estimate: Estimate | null
24 isWorking: boolean
25 showBasis: boolean
26 calibration: readonly string[]
27 /** The user's prompt for the task shown, as typed. */
28 title: string
29 /** Time since the task shown started. */
30 elapsedMs: number
31 /** Set once a task finished: how long it took, its receipt line, and whether it was cut short. */
32 finished: { totalMs: number; receipt: string | null; isAborted: boolean; waitingAgents: number } | null
33}
34
35export const BAR_WIDTH = 24
36
37/**
38 * Cut to `columns` code points, with an ellipsis when cut.
39 */
40export function truncate(text: string, columns: number): string {
41 const points = [...text]
42 if (points.length <= columns) return text
43 if (columns <= 1) return points.slice(0, Math.max(0, columns)).join('')
44 return points.slice(0, columns - 1).join('') + '…'
45}
46
47export function barOf(progress: number, width: number): string {
48 const filled = Math.max(0, Math.min(width, Math.round(progress * width)))
49 return '█'.repeat(filled) + '░'.repeat(width - filled)
50}
51
52function headerOf(plan: Plan): Line {
53 switch (plan.source) {
54 case 'derived':
55 return { key: 'plan-header', text: 'plan · derived (the mod\'s guess, not the assistant\'s)', dim: true }
56 case 'fallback':
57 return { key: 'plan-header', text: 'plan · none given', dim: true }
58 case 'none':
59 return { key: 'plan-header', text: 'plan · waiting for the first step', dim: true }
60 default:
61 return { key: 'plan-header', text: 'plan', dim: true }
62 }
63}
64
65function durationText(ms: number): string {
66 const seconds = Math.round(ms / 1000)
67 if (seconds < 60) return `${seconds}s`
68 return `${Math.floor(seconds / 60)}m ${seconds % 60}s`
69}
70
71/**
72 * The pane, top to bottom under its title row: milestones, the progress bar,
73 * the estimate, the basis detail when asked, the calibration line, and the
74 * finished task's receipt.
75 */
76export function paneLines(model: ViewModel): Line[] {
77 const lines: Line[] = []
78 const { plan, estimate } = model
79 if (!plan) {
80 lines.push({ key: 'idle', text: 'no task yet', dim: true })
81 } else {
82 lines.push(headerOf(plan))
83 plan.items.forEach((item, i) => {
84 lines.push({
85 key: `m-${i}`,
86 text: `${glyphOf(item.state)} ${item.label}`,
87 ...(item.state === 'done' ? { dim: true } : {}),
88 ...(item.state === 'current' ? { bold: true } : {}),
89 })
90 })
91 }
92 if (model.isWorking && estimate) {
93 const progress = estimate.kind === 'range' ? estimate.progress : 0
94 const isOver = estimate.kind === 'range' && estimate.isOver
95 lines.push({ key: 'bar', text: `${barOf(progress, BAR_WIDTH)} ${Math.round(progress * 100)}%`, ...(isOver ? { dim: true } : {}) })
96 lines.push({ key: 'estimate', text: estimateRow(estimate), ...(isOver ? { color: 'yellow' } : {}) })
97 if (model.showBasis) {
98 basisLines(estimate).forEach((text, i) => lines.push({ key: `basis-${i}`, text, dim: true }))
99 }
100 }
101 if (model.finished) {
102 lines.push({ key: 'finished', text: `finished in ${durationText(model.finished.totalMs)}`, dim: true })
103 if (model.finished.receipt) {
104 const isBad = model.finished.receipt.startsWith('UNVERIFIED')
105 lines.push({ key: 'receipt', text: model.finished.receipt, ...(isBad ? { color: 'yellow' } : { color: 'green' }) })
106 }
107 }
108 model.calibration.forEach((text, i) =>
109 lines.push({ key: `cal-${i}`, text, ...(i === 0 ? { dim: true } : { color: 'yellow' }) }),
110 )
111 return lines
112}
113
114function fieldOf(input: unknown, name: string): string | undefined {
115 if (typeof input !== 'object' || input === null) return undefined
116 const value = (input as Record<string, unknown>)[name]
117 return typeof value === 'string' ? value : undefined
118}
119
120/**
121 * The one dim line a tool call becomes in clean view: `● Edit src/x.ts`.
122 */
123export function toolRowText(tool: string, input: unknown, cwd: string): string {
124 const path = fieldOf(input, 'file_path') ?? fieldOf(input, 'notebook_path') ?? fieldOf(input, 'path')
125 const arg =
126 (path !== undefined ? relativeOf(path, cwd) : undefined) ??
127 fieldOf(input, 'command') ??
128 fieldOf(input, 'pattern') ??
129 fieldOf(input, 'url') ??
130 fieldOf(input, 'query') ??
131 fieldOf(input, 'description') ??
132 fieldOf(input, 'subject') ??
133 ''
134 const flat = arg.replace(/\s+/g, ' ').trim()
135 return truncate(flat ? `● ${tool} ${flat}` : `● ${tool}`, 160)
136}
137
138function textOf(output: unknown): string | undefined {
139 if (typeof output === 'string') return output
140 if (typeof output !== 'object' || output === null) return undefined
141 const fields = output as { stdout?: unknown; stderr?: unknown; file?: { content?: unknown }; content?: unknown }
142 if (typeof fields.stdout === 'string') return [fields.stdout, typeof fields.stderr === 'string' ? fields.stderr : ''].filter(Boolean).join('\n')
143 if (typeof fields.file?.content === 'string') return fields.file.content
144 if (typeof fields.content === 'string') return fields.content
145 return undefined
146}
147
148/**
149 * The one dim line a tool result becomes in clean view: `⎿ 12 lines`.
150 */
151export function toolResultText(output: unknown, isErrored: boolean): string {
152 if (isErrored) return '⎿ error'
153 const text = textOf(output)
154 if (text === undefined) return '⎿ done'
155 const trimmed = text.replace(/\n+$/, '')
156 if (trimmed.length === 0) return '⎿ no output'
157 const count = trimmed.split('\n').length
158 return `⎿ ${count} ${count === 1 ? 'line' : 'lines'}`
159}
160
161/**
162 * A folded run of reads and searches, as one dim line.
163 */
164export function toolGroupText(calls: readonly { tool: string }[]): string {
165 const tools = [...new Set(calls.map(call => call.tool))]
166 return truncate(`● ${calls.length} ${calls.length === 1 ? 'call' : 'calls'}: ${tools.join(', ')}`, 160)
167}
168src/estimator.ts 310 lines1/**
2 * The estimate (SPEC 3): remaining time as a range with its basis named.
3 *
4 * The honesty contract, and the scar behind each rule:
5 * - Never a point, never a frozen countdown (invariant 1; every "ETA" in
6 * every tool ever). `estimateRow` always prints two different numbers, and
7 * the over-by branch keeps moving with the clock.
8 * - Indeterminate is time-boxed (invariant 2; "don't cheat and just stay
9 * indeterminate"): only while there is no plan AND no history, and never
10 * past the first completed milestone or 90 seconds.
11 * - The basis is always named (invariant 3).
12 * - Countdown digits only at confidence 0.5 or more, and even then as a
13 * range: `3:52 to 8:40 left`.
14 *
15 * Remaining time = sum over the steps not done of their expected duration,
16 * from the most specific shape bucket with 3 or more samples, else a coarser
17 * one, else the global prior. sigma sums only the steps not yet done, so it
18 * shrinks as steps complete; each step's sigma also shrinks with samples.
19 */
20
21import { emptyBucket, type Bucket, type Stats, bucketAdd } from './history'
22import { hasCompleted, type Plan } from './milestones'
23import { GLOBAL_KEY, levelKeysOf, planKeyOf, planlessKeysOf, type Shape } from './shape'
24
25export const INDETERMINATE_CAP_MS = 90_000
26/** Interval half-width in sigmas: about an 87% band for a normal. */
27export const K = 1.5
28export const MIN_SAMPLES = 3
29export const COUNTDOWN_CONFIDENCE = 0.5
30
31/** The prior when nothing has been learned: a step, and a whole task. */
32const DEFAULT_STEP = { mean: 120_000, sd: 90_000 }
33const DEFAULT_TOTAL = { mean: 300_000, sd: 180_000 }
34/** How much wider the prior is than its own sd. */
35const PRIOR_SPREAD = 1.25
36/** No step is ever known to better than a quarter of its length, or 15s. */
37const SD_FLOOR_RATIO = 0.25
38const SD_FLOOR_MS = 15_000
39/** A step under way always has a tenth of its expected time left. */
40const MIN_LEFT_RATIO = 0.1
41/** Past the range, re-widen by half of what the estimate missed by. */
42const OVER_WIDEN = 0.5
43
44export type EstimateInput = {
45 now: number
46 startedAt: number
47 plan: Plan
48 stats: Stats
49 shape: Shape
50 /** The spike subagent's minutes per step, in ms; one sample at weight 0.5. */
51 spike?: readonly number[]
52}
53
54export type RangeEstimate = {
55 kind: 'range'
56 basis: string
57 remainingLowMs: number
58 remainingHighMs: number
59 totalLowMs: number
60 totalHighMs: number
61 sigmaMs: number
62 confidence: number
63 isOver: boolean
64 overByMs: number
65 /** Fraction done, weighted by expected step durations. */
66 progress: number
67 samples: number
68 /** Each plan step's own fraction done, in plan order: 1 done, 0 not started. */
69 stepProgress: readonly number[]
70}
71
72export type Estimate = { kind: 'indeterminate' } | RangeEstimate
73
74type Basis = {
75 bucket: Bucket | null
76 label: string
77 /** Effective samples for confidence; 0 for any prior. */
78 samples: number
79}
80
81type Expectation = {
82 mean: number
83 sd: number
84}
85
86export function estimate(input: EstimateInput): Estimate {
87 const { now, startedAt, plan, stats } = input
88 const elapsed = Math.max(0, now - startedAt)
89 const isPlanless = plan.source === 'none' || plan.source === 'fallback'
90 const hasHistory = (stats.get(GLOBAL_KEY)?.total.n ?? 0) > 0
91
92 if (isPlanless && !hasHistory && elapsed < INDETERMINATE_CAP_MS && !hasCompleted(plan)) {
93 return { kind: 'indeterminate' }
94 }
95
96 return isPlanless ? totalMode(input, elapsed) : stepMode(input, elapsed)
97}
98
99/**
100 * No plan to walk: the whole task's learned duration against elapsed time.
101 */
102function totalMode(input: EstimateInput, elapsed: number): RangeEstimate {
103 const basis = basisOf(input.stats, planlessKeysOf(input.shape))
104 const exp = basis.bucket ? spreadOf(basis.bucket.total.mean, basis.bucket.total.var, basis.samples) : priorOf(DEFAULT_TOTAL)
105 const remMean = Math.max(exp.mean - elapsed, MIN_LEFT_RATIO * exp.mean)
106 const plannedHigh = exp.mean + K * exp.sd
107 const progress = Math.min(elapsed / exp.mean, 0.95)
108 return rangeOf({ elapsed, remMean, sigma: exp.sd, plannedHigh, progress, basis, suffix: '', stepProgress: input.plan.items.map(() => progress) })
109}
110
111/**
112 * A plan to walk: each step not done contributes its expected duration, the
113 * one under way less what it has used.
114 */
115function stepMode(input: EstimateInput, elapsed: number): RangeEstimate {
116 const { now, startedAt, plan, stats, shape, spike } = input
117 const basis = basisOf(stats, levelKeysOf(shape), spike, planKeyOf(shape))
118 const items = plan.items
119 const stepCount = items.length
120 const hasCurrent = items.some(item => item.state === 'current')
121 const implicit = hasCurrent ? -1 : items.findIndex(item => item.state === 'pending')
122
123 let remMean = 0
124 let remVar = 0
125 let doneWeight = 0
126 let totalWeight = 0
127 let plannedEnd: number | null = null
128 let pendingMean = 0
129 let previousEnd = startedAt
130 const stepProgress: number[] = []
131
132 items.forEach((item, i) => {
133 const exp = basis.bucket ? stepExpectation(basis.bucket, i, stepCount, basis.samples) : priorOf(DEFAULT_STEP)
134 totalWeight += exp.mean
135 if (item.state === 'done') {
136 doneWeight += exp.mean
137 previousEnd = item.doneAt ?? previousEnd
138 stepProgress.push(1)
139 return
140 }
141 remVar += exp.sd * exp.sd
142 if (item.state === 'current' || i === implicit) {
143 const start = item.startedAt ?? previousEnd
144 const inStep = Math.max(0, now - start)
145 remMean += Math.max(exp.mean - inStep, MIN_LEFT_RATIO * exp.mean)
146 doneWeight += Math.min(inStep / exp.mean, 0.95) * exp.mean
147 stepProgress.push(Math.min(inStep / exp.mean, 0.95))
148 plannedEnd = Math.max(plannedEnd ?? -Infinity, start + exp.mean)
149 return
150 }
151 remMean += exp.mean
152 pendingMean += exp.mean
153 stepProgress.push(0)
154 })
155
156 const sigma = Math.sqrt(remVar)
157 const plannedHigh = (plannedEnd ?? now) - startedAt + pendingMean + K * sigma
158 const progress = totalWeight > 0 ? doneWeight / totalWeight : 0
159 const suffix = plan.source === 'derived' ? ' · derived plan' : ''
160 return rangeOf({ elapsed, remMean, sigma, plannedHigh, progress, basis, suffix, stepProgress })
161}
162
163function rangeOf(args: {
164 elapsed: number
165 remMean: number
166 sigma: number
167 plannedHigh: number
168 progress: number
169 basis: Basis
170 suffix: string
171 stepProgress: readonly number[]
172}): RangeEstimate {
173 const { elapsed, remMean, sigma, plannedHigh, progress, basis } = args
174 const isOver = elapsed > plannedHigh
175 const overByMs = isOver ? elapsed - plannedHigh : 0
176 const remainingLowMs = Math.max(0, remMean - K * sigma)
177 const remainingHighMs = remMean + K * sigma + OVER_WIDEN * overByMs
178 const n = basis.samples
179 const cv = sigma / Math.max(remMean, 1)
180 const confidence = n <= 0 ? 0 : (n / (n + MIN_SAMPLES)) * (0.6 + 0.4 * progress) / (1 + cv)
181 return {
182 kind: 'range',
183 basis: basis.label + args.suffix,
184 remainingLowMs,
185 remainingHighMs,
186 totalLowMs: elapsed + remainingLowMs,
187 totalHighMs: elapsed + remainingHighMs,
188 sigmaMs: sigma,
189 confidence: Math.max(0, Math.min(1, confidence)),
190 isOver,
191 overByMs,
192 progress: Math.max(0, Math.min(1, progress)),
193 samples: n,
194 stepProgress: args.stepProgress,
195 }
196}
197
198/**
199 * The most specific bucket with MIN_SAMPLES or more; else, with a spike, the
200 * plan-time bucket plus the spike at weight 0.5; else the global prior.
201 */
202function basisOf(stats: Stats, keys: readonly string[], spike?: readonly number[], spikeKey?: string): Basis {
203 for (const key of keys) {
204 const bucket = stats.get(key)
205 if (bucket && bucket.total.n >= MIN_SAMPLES) {
206 return { bucket, label: `from ${countOf(bucket.total.n)} similar tasks`, samples: bucket.total.n }
207 }
208 }
209 if (spike && spike.length > 0 && spikeKey !== undefined) {
210 const real = stats.get(spikeKey) ?? emptyBucket()
211 const bucket = bucketAdd(real, spike, spike.reduce((a, b) => a + b, 0), 0.5)
212 const n = real.total.n
213 return { bucket, label: n > 0 ? `spike guess + ${countOf(n)} similar` : 'spike guess', samples: bucket.total.n }
214 }
215 const global = stats.get(GLOBAL_KEY)
216 if (global && global.total.n > 0) return { bucket: global, label: 'prior only', samples: 0 }
217 return { bucket: null, label: 'prior only', samples: 0 }
218}
219
220function stepExpectation(bucket: Bucket, index: number, stepCount: number, samples: number): Expectation {
221 const own = bucket.steps[index]
222 if (own && own.n > 0) return spreadOf(own.mean, own.var, samples)
223 if (bucket.pooled.n > 0) return spreadOf(bucket.pooled.mean, bucket.pooled.var, samples)
224 return spreadOf(bucket.total.mean / Math.max(1, stepCount), bucket.total.var / Math.max(1, stepCount), samples)
225}
226
227/**
228 * A learned mean and variance as an expectation: the sd floored, then
229 * widened for how few samples stand behind it (predictive sd, sqrt(1 + 1/n));
230 * a bucket used only as a prior gets the prior's spread.
231 */
232function spreadOf(mean: number, variance: number, samples: number): Expectation {
233 const safeMean = Math.max(mean, 1_000)
234 const sd = Math.max(Math.sqrt(Math.max(variance, 0)), SD_FLOOR_RATIO * safeMean, SD_FLOOR_MS)
235 const widen = samples > 0 ? Math.sqrt(1 + 1 / samples) : PRIOR_SPREAD
236 return { mean: safeMean, sd: sd * widen }
237}
238
239function priorOf(prior: Expectation): Expectation {
240 return { mean: prior.mean, sd: prior.sd * PRIOR_SPREAD }
241}
242
243function countOf(n: number): string {
244 return Number.isInteger(n) ? String(n) : n.toFixed(1)
245}
246
247/**
248 * The estimate row the pane and band print.
249 */
250export function estimateRow(est: Estimate): string {
251 if (est.kind === 'indeterminate') return 'indeterminate'
252 if (est.isOver) {
253 return `over by ${durationOf(est.overByMs)} · ${approxRangeOf(est.remainingLowMs, est.remainingHighMs)} more · ${est.basis}`
254 }
255 const range =
256 est.confidence >= COUNTDOWN_CONFIDENCE
257 ? `${digitsOf(est.remainingLowMs)} to ${digitsOf(Math.max(est.remainingHighMs, est.remainingLowMs + 1_000))} left`
258 : approxRangeOf(est.remainingLowMs, est.remainingHighMs)
259 return `${range} · ${est.basis}`
260}
261
262/**
263 * `~4 to 9 min`, or `~20 to 45 sec` when the top is under a minute. The two
264 * numbers always differ: a range that rounds to one value is widened. The
265 * low end never shows 0: `~0 to 10 min` is a range in name only, so it
266 * floors at one unit (1 min, 5 sec) and the top widens to stay above it.
267 */
268export function approxRangeOf(lowMs: number, highMs: number): string {
269 if (highMs < 60_000) {
270 const low = Math.max(5, Math.floor(lowMs / 5_000) * 5)
271 let high = Math.ceil(highMs / 5_000) * 5
272 if (high <= low) high = low + 5
273 return `~${low} to ${high} sec`
274 }
275 const low = Math.max(1, Math.floor(lowMs / 60_000))
276 let high = Math.ceil(highMs / 60_000)
277 if (high <= low) high = low + 1
278 return `~${low} to ${high} min`
279}
280
281function digitsOf(ms: number): string {
282 const seconds = Math.max(0, Math.round(ms / 1000))
283 return `${Math.floor(seconds / 60)}:${String(seconds % 60).padStart(2, '0')}`
284}
285
286/**
287 * `2m 10s`, `45s`, `1h 5m`.
288 */
289export function durationOf(ms: number): string {
290 const seconds = Math.max(0, Math.round(ms / 1000))
291 if (seconds < 60) return `${seconds}s`
292 const minutes = Math.floor(seconds / 60)
293 if (minutes < 60) return `${minutes}m ${seconds % 60}s`
294 return `${Math.floor(minutes / 60)}h ${minutes % 60}m`
295}
296
297/**
298 * The basis button's detail: where the numbers come from, in plain words.
299 */
300export function basisLines(est: Estimate): string[] {
301 if (est.kind === 'indeterminate') {
302 return ['basis: no plan and no history yet; a range from the prior shows by 90s or the first finished milestone']
303 }
304 return [
305 `basis: ${est.basis}`,
306 `sigma ±${durationOf(est.sigmaMs)} over the steps left · confidence ${est.confidence.toFixed(2)}`,
307 `whole task: ${durationOf(est.totalLowMs)} to ${durationOf(est.totalHighMs)}`,
308 ]
309}
310src/milestones.ts 271 lines1/**
2 * Milestones: what the turn is working on, what is finished, what is left
3 * (SPEC 2), and which tool rows count as milestone events.
4 *
5 * Sources, in order of preference: the task tools the assistant called
6 * (TaskCreate / TaskUpdate, or TodoWrite), a plan derived by a model call
7 * (labelled `derived`, invariant 9: the mod never passes its own guess off as
8 * the assistant's plan), or a single fallback milestone, "the task".
9 */
10
11import { isVerifyCommand } from './ledger'
12
13export type MilestoneState = 'done' | 'current' | 'pending'
14
15export type Milestone = {
16 id: string
17 label: string
18 state: MilestoneState
19 startedAt?: number
20 doneAt?: number
21}
22
23export type PlanSource = 'none' | 'tasks' | 'derived' | 'fallback'
24
25export type Plan = {
26 source: PlanSource
27 items: readonly Milestone[]
28}
29
30export type TaskUpdateArgs = {
31 taskId: string
32 status?: 'pending' | 'in_progress' | 'completed' | 'deleted'
33 subject?: string
34}
35
36export type Todo = {
37 content: string
38 status: 'pending' | 'in_progress' | 'completed'
39 activeForm?: string
40}
41
42export const FALLBACK_LABEL = 'the task'
43const MIN_DERIVED = 2
44const MAX_DERIVED = 8
45const MAX_LABEL = 80
46
47export function emptyPlan(): Plan {
48 return { source: 'none', items: [] }
49}
50
51/**
52 * A task tool beats a derived or fallback plan: the first TaskCreate drops
53 * the mod's guess and starts the assistant's own list.
54 */
55export function onTaskCreate(plan: Plan, id: string, subject: string, now: number): Plan {
56 const kept = plan.source === 'tasks' ? plan.items : []
57 if (kept.some(item => item.id === id)) return plan
58 return { source: 'tasks', items: [...kept, { id, label: labelOf(subject), state: 'pending' }] }
59}
60
61export function onTaskUpdate(plan: Plan, args: TaskUpdateArgs, now: number): Plan {
62 const index = plan.items.findIndex(item => item.id === args.taskId)
63 if (plan.source !== 'tasks' || index === -1) return plan
64 if (args.status === 'deleted') {
65 return { ...plan, items: plan.items.filter((_, i) => i !== index) }
66 }
67 const items = plan.items.map((item, i): Milestone => {
68 if (i !== index) return item
69 const label = args.subject === undefined ? item.label : labelOf(args.subject)
70 switch (args.status) {
71 case 'in_progress':
72 return { ...item, label, state: 'current', startedAt: item.startedAt ?? now, doneAt: undefined }
73 case 'completed':
74 return { ...item, label, state: 'done', startedAt: item.startedAt ?? startOf(plan.items, i, now), doneAt: now }
75 case 'pending':
76 return { ...item, label, state: 'pending', doneAt: undefined }
77 default:
78 return { ...item, label }
79 }
80 })
81 return { ...plan, items }
82}
83
84/**
85 * TodoWrite sends the whole list each time: replace it, keeping the times of
86 * items whose text did not change.
87 */
88export function onTodoWrite(plan: Plan, todos: readonly Todo[], now: number): Plan {
89 const before = new Map(plan.items.map(item => [item.label, item]))
90 const items = todos.map((todo, i): Milestone => {
91 const label = labelOf(todo.content)
92 const was = before.get(label)
93 const state: MilestoneState = todo.status === 'completed' ? 'done' : todo.status === 'in_progress' ? 'current' : 'pending'
94 const startedAt = state === 'pending' ? was?.startedAt : (was?.startedAt ?? now)
95 const doneAt = state === 'done' ? (was?.doneAt ?? now) : undefined
96 return { id: `todo-${i}`, label, state, startedAt, doneAt }
97 })
98 return { source: 'tasks', items }
99}
100
101export function derivedPlan(labels: readonly string[], now: number): Plan {
102 return {
103 source: 'derived',
104 items: labels.map((label, i) => ({
105 id: `derived-${i}`,
106 label: labelOf(label),
107 state: i === 0 ? 'current' : 'pending',
108 ...(i === 0 ? { startedAt: now } : {}),
109 })),
110 }
111}
112
113export function fallbackPlan(now: number): Plan {
114 return { source: 'fallback', items: [{ id: 'task', label: FALLBACK_LABEL, state: 'current', startedAt: now }] }
115}
116
117/**
118 * Marks the first current milestone done and starts the next pending one; a
119 * derived step the classifier called done, or the fallback at turn end.
120 */
121export function completeCurrent(plan: Plan, now: number): Plan {
122 let index = plan.items.findIndex(item => item.state === 'current')
123 if (index === -1) index = plan.items.findIndex(item => item.state === 'pending')
124 if (index === -1) return plan
125 const items = plan.items.map((item, i): Milestone => {
126 if (i === index) return { ...item, state: 'done', startedAt: item.startedAt ?? startOf(plan.items, i, now), doneAt: now }
127 if (i === index + 1 && item.state === 'pending') return { ...item, state: 'current', startedAt: now }
128 return item
129 })
130 return { ...plan, items }
131}
132
133/**
134 * The turn ended: whatever is under way is done now. Steps never started stay
135 * pending and are not learned from.
136 */
137export function finishPlan(plan: Plan, now: number): Plan {
138 const items = plan.items.map((item, i): Milestone =>
139 item.state === 'current' ? { ...item, state: 'done', startedAt: item.startedAt ?? startOf(plan.items, i, now), doneAt: now } : item,
140 )
141 return { ...plan, items }
142}
143
144/**
145 * Whether the plan finished: every milestone done.
146 */
147export function isComplete(plan: Plan): boolean {
148 return plan.items.length > 0 && plan.items.every(item => item.state === 'done')
149}
150
151export function hasCompleted(plan: Plan): boolean {
152 return plan.items.some(item => item.state === 'done')
153}
154
155/**
156 * How long each finished step took, in plan order, measured from its own
157 * start or, failing that, from the previous step's end or the task's start.
158 */
159export function stepDurations(plan: Plan, startedAt: number, pauses: readonly Pause[] = []): number[] {
160 const durations: number[] = []
161 let previousEnd = startedAt
162 for (const item of plan.items) {
163 if (item.state !== 'done' || item.doneAt === undefined) continue
164 const start = item.startedAt ?? previousEnd
165 const doneAt = item.doneAt
166 const paused = pauses.filter(pause => pause.at > start && pause.at <= doneAt).reduce((sum, pause) => sum + pause.ms, 0)
167 durations.push(Math.max(0, doneAt - start - paused))
168 previousEnd = doneAt
169 }
170 return durations
171}
172
173/**
174 * Time inside a step that was not work: a permission prompt waiting on the
175 * person, measured as a tool call's span less its run time. `at` is when the
176 * call ended, which places it in its step.
177 */
178export type Pause = {
179 at: number
180 ms: number
181}
182
183/**
184 * The turn ended and a pass over the final answer said it completed the
185 * steps at `indexes` too. They and the step under way share the time since
186 * the last finished step in equal slices: the mod knows they finished, not
187 * when, and an even split teaches less wrong than zero-length steps.
188 */
189export function completeAtEnd(plan: Plan, indexes: readonly number[], now: number, startedAt: number): Plan {
190 const current = plan.items.findIndex(item => item.state === 'current')
191 const chosen = indexes.filter(i => plan.items[i] !== undefined && plan.items[i]!.state === 'pending')
192 if (chosen.length === 0) return plan
193 const span = [...new Set([...(current === -1 ? [] : [current]), ...chosen])].sort((a, b) => a - b)
194 const lastDone = plan.items.reduce((at, item) => (item.doneAt !== undefined && item.doneAt > at ? item.doneAt : at), startedAt)
195 const from = Math.min(current === -1 ? lastDone : (plan.items[current]!.startedAt ?? lastDone), now)
196 const slice = (now - from) / span.length
197 const items = plan.items.map((item, i): Milestone => {
198 const k = span.indexOf(i)
199 if (k === -1) return item
200 return { ...item, state: 'done', startedAt: from + k * slice, doneAt: from + (k + 1) * slice }
201 })
202 return { ...plan, items }
203}
204
205/**
206 * Lines of a model's plan reply, numbering and bullets stripped; 2..8 steps
207 * or none at all.
208 */
209export function parseDerivedSteps(text: string): string[] {
210 const steps = text
211 .split('\n')
212 .map(line => line.replace(/^\s*(?:\d+\s*[.):]|[-*•])\s*/, '').trim())
213 .filter(line => line.length > 0)
214 .slice(0, MAX_DERIVED)
215 .map(labelOf)
216 return steps.length >= MIN_DERIVED ? steps : []
217}
218
219export function glyphOf(state: MilestoneState): string {
220 return state === 'done' ? '✓' : state === 'current' ? '▸' : '·'
221}
222
223function labelOf(text: string): string {
224 const flat = text.replace(/\s+/g, ' ').trim()
225 return flat.length > MAX_LABEL ? flat.slice(0, MAX_LABEL - 1) + '…' : flat
226}
227
228function startOf(items: readonly Milestone[], index: number, now: number): number {
229 for (let i = index - 1; i >= 0; i -= 1) {
230 const doneAt = items[i]?.doneAt
231 if (doneAt !== undefined) return doneAt
232 }
233 return now
234}
235
236export const EDIT_TOOLS = ['Edit', 'Write', 'MultiEdit', 'NotebookEdit'] as const
237
238export function isEditTool(tool: string): boolean {
239 return (EDIT_TOOLS as readonly string[]).includes(tool)
240}
241
242/**
243 * The file an edit tool call changes, or undefined for any other call.
244 */
245export function editedPathOf(tool: string, input: unknown): string | undefined {
246 if (!isEditTool(tool) || typeof input !== 'object' || input === null) return undefined
247 const fields = input as { file_path?: unknown; notebook_path?: unknown }
248 const path = fields.file_path ?? fields.notebook_path
249 return typeof path === 'string' ? path : undefined
250}
251
252export function isCommitCommand(command: string): boolean {
253 return /\bgit\s+(?:-\S+\s+)*commit\b/.test(command)
254}
255
256export function isPrCreateCommand(command: string): boolean {
257 return /\bgh\s+pr\s+create\b/.test(command)
258}
259
260/**
261 * A shell command whose row always draws in full: a verify run (its exit code
262 * and test counts), a commit, or a pull request (invariant 5).
263 */
264export function isMilestoneCommand(command: string): boolean {
265 return isVerifyCommand(command) || isCommitCommand(command) || isPrCreateCommand(command)
266}
267
268export function isSubagentTool(tool: string): boolean {
269 return tool === 'Agent' || tool === 'Task'
270}
271src/history.ts 206 lines1/**
2 * The learning store: finished tasks, the statistics replayed from them, and
3 * the mod's own calibration score (SPEC 3, "History" and "Learning").
4 *
5 * Only finished task records are persisted. EMA mean and variance per shape
6 * key and step index are rebuilt by replaying those records oldest first, so
7 * pruning the oldest tasks also ages them out of the statistics; there is no
8 * second copy of the numbers to drift from the records.
9 *
10 * Invariant 8, the store stays under 1 MiB: `addTask` prunes to the newest
11 * MAX_TASKS and then drops the oldest until the JSON fits MAX_BYTES.
12 */
13
14import { GLOBAL_KEY, levelKeysOf, planKeyOf, type Shape } from './shape'
15
16export const HISTORY_KEY = 'history'
17export const MAX_TASKS = 500
18/** Under 1 MiB (1,048,576) with room for the store's own framing. */
19export const MAX_BYTES = 1_000_000
20
21export type TaskRecord = {
22 /** When the task ended, ms. */
23 at: number
24 shape: Shape
25 /** How long each step that ran took, in order, ms. */
26 steps: number[]
27 totalMs: number
28 /**
29 * Whether the total landed inside the last range shown before the final
30 * step; absent when no range was ever shown.
31 */
32 inside?: boolean
33}
34
35export type History = {
36 v: 1
37 tasks: TaskRecord[]
38}
39
40export type Ema = {
41 /** Sum of sample weights. A spike counts 0.5. */
42 n: number
43 mean: number
44 var: number
45}
46
47export type Bucket = {
48 total: Ema
49 /** Per step index. */
50 steps: Ema[]
51 /** Every step duration, any index. */
52 pooled: Ema
53}
54
55export type Stats = ReadonlyMap<string, Bucket>
56
57export type Calibration = {
58 inside: number
59 total: number
60}
61
62/** The EMA's floor rate: early samples average, later ones decay at 10%. */
63const ALPHA = 0.1
64
65export const EMPTY_EMA: Ema = { n: 0, mean: 0, var: 0 }
66
67export function emptyHistory(): History {
68 return { v: 1, tasks: [] }
69}
70
71/**
72 * Adds one sample at weight `w`. While few samples are in, this is the plain
73 * running mean and population variance; past 1/ALPHA samples it decays.
74 */
75export function emaAdd(ema: Ema, x: number, w = 1): Ema {
76 const n = ema.n + w
77 const a = Math.min(1, Math.max(w / n, ALPHA * w))
78 const d = x - ema.mean
79 return { n, mean: ema.mean + a * d, var: (1 - a) * (ema.var + a * d * d) }
80}
81
82export function emptyBucket(): Bucket {
83 return { total: EMPTY_EMA, steps: [], pooled: EMPTY_EMA }
84}
85
86/**
87 * Folds one task into a bucket at a weight.
88 */
89export function bucketAdd(bucket: Bucket, steps: readonly number[], totalMs: number, w = 1): Bucket {
90 const next: Ema[] = [...bucket.steps]
91 let pooled = bucket.pooled
92 steps.forEach((ms, i) => {
93 next[i] = emaAdd(next[i] ?? EMPTY_EMA, ms, w)
94 pooled = emaAdd(pooled, ms, w)
95 })
96 return { total: emaAdd(bucket.total, totalMs, w), steps: next, pooled }
97}
98
99/**
100 * Replays every retained task, oldest first, into a bucket per shape level
101 * and one global bucket.
102 */
103export function statsOf(history: History): Stats {
104 const stats = new Map<string, Bucket>()
105 for (const task of history.tasks) {
106 for (const key of [...levelKeysOf(task.shape), GLOBAL_KEY]) {
107 stats.set(key, bucketAdd(stats.get(key) ?? emptyBucket(), task.steps, task.totalMs))
108 }
109 }
110 return stats
111}
112
113/**
114 * Sample weight behind the plan-time shape (every field but tool mix), with
115 * a spike counted at 0.5. The spike gate fires below 3.
116 */
117export function bucketSamples(stats: Stats, shape: Shape, spike?: readonly number[]): number {
118 const real = stats.get(planKeyOf(shape))?.total.n ?? 0
119 return real + (spike && spike.length > 0 ? 0.5 : 0)
120}
121
122export function sizeOf(history: History): number {
123 return new TextEncoder().encode(JSON.stringify(history)).length
124}
125
126/**
127 * The newest MAX_TASKS tasks, then the oldest dropped until the JSON fits.
128 */
129export function prune(history: History, maxBytes = MAX_BYTES): History {
130 let tasks = history.tasks.slice(-MAX_TASKS)
131 let size = sizeOf({ v: 1, tasks })
132 while (size > maxBytes && tasks.length > 0) {
133 const perTask = size / tasks.length
134 const drop = Math.max(1, Math.ceil((size - maxBytes) / perTask))
135 tasks = tasks.slice(drop)
136 size = sizeOf({ v: 1, tasks })
137 }
138 return { v: 1, tasks }
139}
140
141export function addTask(history: History, task: TaskRecord): History {
142 return prune({ v: 1, tasks: [...history.tasks, task] })
143}
144
145export function calibrationOf(history: History): Calibration {
146 let inside = 0
147 let total = 0
148 for (const task of history.tasks) {
149 if (task.inside === undefined) continue
150 total += 1
151 if (task.inside) inside += 1
152 }
153 return { inside, total }
154}
155
156/**
157 * The calibration line, always shown, and a plain-words warning once the mod
158 * has missed more often than not over 10 or more tasks (invariant 3).
159 */
160export function calibrationLines(calibration: Calibration): string[] {
161 if (calibration.total === 0) return ['calibration: no finished tasks yet']
162 const percent = Math.round((calibration.inside / calibration.total) * 100)
163 const noun = calibration.total === 1 ? 'task' : 'tasks'
164 const lines = [`calibration: ${percent}% of ${calibration.total} ${noun} ended inside the range`]
165 if (calibration.total >= 10 && percent < 50) {
166 lines.push('these ranges have missed more often than not; treat them as rough')
167 }
168 return lines
169}
170
171function isRecord(value: unknown): value is Record<string, unknown> {
172 return typeof value === 'object' && value !== null && !Array.isArray(value)
173}
174
175function isShape(value: unknown): value is Shape {
176 return (
177 isRecord(value) &&
178 typeof value.taskType === 'string' &&
179 typeof value.steps === 'string' &&
180 typeof value.repo === 'string' &&
181 typeof value.hasTests === 'boolean' &&
182 typeof value.mix === 'string'
183 )
184}
185
186function isTask(value: unknown): value is TaskRecord {
187 return (
188 isRecord(value) &&
189 typeof value.at === 'number' &&
190 isShape(value.shape) &&
191 Array.isArray(value.steps) &&
192 value.steps.every(ms => typeof ms === 'number' && Number.isFinite(ms)) &&
193 typeof value.totalMs === 'number' &&
194 (value.inside === undefined || typeof value.inside === 'boolean')
195 )
196}
197
198/**
199 * Reads what the store holds, keeping only well-formed records: a corrupt or
200 * foreign value is an empty history, never a crash in a hook.
201 */
202export function parseHistory(raw: unknown): History {
203 if (!isRecord(raw) || !Array.isArray(raw.tasks)) return emptyHistory()
204 return { v: 1, tasks: raw.tasks.filter(isTask) }
205}
206src/ledger.ts 153 lines1/**
2 * The per-turn ledger behind the receipt (SPEC 4): every edit, every verify
3 * run with its exit code, and the one line printed under the answer.
4 *
5 * Order is by sequence number, not clock: "a verify after the last edit"
6 * must hold even when two calls land in the same millisecond.
7 *
8 * This is a mirror, not a gate, in v1: nothing here blocks or aborts a turn.
9 */
10
11/**
12 * SPEC 4's verify list, verbatim, with word boundaries added so `latest` or
13 * `contest` do not read as `test`. `claude plugin test` matches on `test`.
14 */
15export const VERIFY_PATTERN =
16 /(?<![\w-])(test|spec|jest|vitest|pytest|cargo (test|check|clippy)|go test|bun test|npm (test|run (test|build|lint))|pnpm|tsc|eslint|ruff|mypy|make (test|check)|build)(?![\w-])/
17
18export type TestCounts = {
19 pass?: number
20 fail?: number
21}
22
23export type EditEntry = {
24 seq: number
25 at: number
26 path: string
27}
28
29export type VerifyEntry = {
30 seq: number
31 at: number
32 command: string
33 exitCode: number
34 counts?: TestCounts
35}
36
37export type Ledger = {
38 seq: number
39 edits: readonly EditEntry[]
40 verifies: readonly VerifyEntry[]
41}
42
43/**
44 * What a tool call resolved to, as far as the ledger reads it: Bash carries
45 * no exit code field, so a failed run is `isError` with `Exit code N` text.
46 */
47export type ToolOutcome = {
48 result?: unknown
49 isError?: boolean
50 text?: string
51 deny?: string
52}
53
54export function emptyLedger(): Ledger {
55 return { seq: 0, edits: [], verifies: [] }
56}
57
58export function isVerifyCommand(command: string): boolean {
59 return VERIFY_PATTERN.test(command)
60}
61
62export function recordEdit(ledger: Ledger, path: string, at: number): Ledger {
63 const seq = ledger.seq + 1
64 return { ...ledger, seq, edits: [...ledger.edits, { seq, at, path }] }
65}
66
67export function recordVerify(ledger: Ledger, command: string, exitCode: number, output: string, at: number): Ledger {
68 const seq = ledger.seq + 1
69 const counts = testCountsOf(output)
70 const entry: VerifyEntry = { seq, at, command, exitCode, ...(counts ? { counts } : {}) }
71 return { ...ledger, seq, verifies: [...ledger.verifies, entry] }
72}
73
74export function exitCodeOf(outcome: ToolOutcome): number {
75 if (!outcome.isError) return 0
76 const match = /Exit code (\d+)/.exec(outcome.text ?? '')
77 return match ? Number(match[1]) : 1
78}
79
80/**
81 * Pass and fail counts from a test runner's output: bun, jest, vitest,
82 * pytest, cargo and go all print `N pass(ed)` and `N fail(ed)` somewhere.
83 */
84export function testCountsOf(text: string): TestCounts | undefined {
85 const pass = /(\d+) (?:pass|passed|passing)\b/.exec(text)
86 const fail = /(\d+) (?:fail|failed|failing)\b/.exec(text)
87 if (!pass && !fail) return undefined
88 return {
89 ...(pass ? { pass: Number(pass[1]) } : {}),
90 ...(fail ? { fail: Number(fail[1]) } : {}),
91 }
92}
93
94/**
95 * The receipt line, or null when there is nothing to say: no edits this
96 * turn, or the answer does not claim done.
97 */
98export function receiptLine(ledger: Ledger, claimsDone: boolean, now: number, cwd: string): string | null {
99 if (ledger.edits.length === 0 || !claimsDone) return null
100 const lastEdit = ledger.edits.reduce((a, b) => (b.seq > a.seq ? b : a))
101 const after = ledger.verifies.filter(run => run.seq > lastEdit.seq)
102 const verify = after[after.length - 1]
103 if (!verify) {
104 return `UNVERIFIED · claimed done, no test/build/run after the last edit (${relativeOf(lastEdit.path, cwd)} at ${clockOf(lastEdit.at)})`
105 }
106 return `receipt · ${commandLabelOf(verify.command)} ${outcomeOf(verify)} · ${agoOf(now - verify.at)}`
107}
108
109function outcomeOf(run: VerifyEntry): string {
110 const pass = run.counts?.pass
111 const fail = run.counts?.fail
112 if (run.exitCode === 0) {
113 return '✓' + (pass !== undefined ? ` ${pass} pass` : '') + (fail ? ` · ${fail} fail` : '')
114 }
115 return `✗ exit ${run.exitCode}` + (pass !== undefined ? ` · ${pass} pass` : '') + (fail ? ` · ${fail} fail` : '')
116}
117
118/**
119 * The part of a compound command that matched, as the user would name it.
120 */
121export function commandLabelOf(command: string): string {
122 const parts = command.split(/&&|\|\||;|\|/).map(part => part.trim())
123 const label = parts.find(part => isVerifyCommand(part)) ?? command.trim()
124 return label.length > 40 ? label.slice(0, 39) + '…' : label
125}
126
127export function relativeOf(path: string, cwd: string): string {
128 const base = cwd.endsWith('/') ? cwd : cwd + '/'
129 return path.startsWith(base) ? path.slice(base.length) : path
130}
131
132function clockOf(ms: number): string {
133 const date = new Date(ms)
134 return `${String(date.getHours()).padStart(2, '0')}:${String(date.getMinutes()).padStart(2, '0')}`
135}
136
137export function agoOf(ms: number): string {
138 const seconds = Math.max(0, Math.floor(ms / 1000))
139 if (seconds < 60) return `${seconds}s ago`
140 const minutes = Math.floor(seconds / 60)
141 if (minutes < 60) return `${minutes}m ago`
142 return `${Math.floor(minutes / 60)}h ago`
143}
144
145/**
146 * The fallback when the claims-done label is not back in time: plain words
147 * of completion in the answer. Used only when the model call misses its
148 * deadline or fails.
149 */
150export function looksDone(answer: string): boolean {
151 return /\b(done|complete[ds]?|finished|fixed|implemented|all (?:\w+ )?(?:tests? )?pass(?:es|ing)?|ready|shipped)\b/i.test(answer)
152}
153src/shape.ts 116 lines1/**
2 * Task shape: the key the estimator learns under (SPEC 3).
3 *
4 * `{ task_type, step_count_bucket, repo, has_tests, tool_mix_bucket }`, plus
5 * the ladder of coarser keys the estimator falls back through when the exact
6 * shape has too few samples. No `$` here: plain data in, plain data out.
7 */
8
9export const TASK_TYPES = ['build', 'debug', 'research', 'writing', 'config', 'refactor', 'chat'] as const
10
11export type TaskType = (typeof TASK_TYPES)[number] | 'unknown'
12
13export type StepBucket = '1' | '2-3' | '4-6' | '7+'
14
15/**
16 * What kind of tools dominated the task. Known only as the task runs, so the
17 * plan-time lookups skip it (see `planKeyOf`).
18 */
19export type ToolMix = 'none' | 'edit' | 'read' | 'shell' | 'mixed'
20
21export type Shape = {
22 taskType: TaskType
23 steps: StepBucket
24 /** A short hash of the repository root or working directory, never the path. */
25 repo: string
26 hasTests: boolean
27 mix: ToolMix
28}
29
30export type ToolCounts = {
31 edit: number
32 read: number
33 shell: number
34 other: number
35}
36
37export const NO_TOOLS: ToolCounts = { edit: 0, read: 0, shell: 0, other: 0 }
38
39/**
40 * The key every task shares: the global prior's bucket.
41 */
42export const GLOBAL_KEY = '*'
43
44export function stepBucketOf(count: number): StepBucket {
45 if (count <= 1) return '1'
46 if (count <= 3) return '2-3'
47 if (count <= 6) return '4-6'
48 return '7+'
49}
50
51/**
52 * The dominant kind of tool, when one kind is at least 60% of the calls.
53 */
54export function toolMixOf(counts: ToolCounts): ToolMix {
55 const total = counts.edit + counts.read + counts.shell + counts.other
56 if (total === 0) return 'none'
57 const kinds: [ToolMix, number][] = [
58 ['edit', counts.edit],
59 ['read', counts.read],
60 ['shell', counts.shell],
61 ]
62 for (const [kind, count] of kinds) {
63 if (count / total >= 0.6) return kind
64 }
65 return 'mixed'
66}
67
68export function countTool(counts: ToolCounts, tool: string): ToolCounts {
69 if (/^(Edit|Write|MultiEdit|NotebookEdit)$/.test(tool)) return { ...counts, edit: counts.edit + 1 }
70 if (/^(Read|Grep|Glob|LS|WebFetch|WebSearch)$/.test(tool)) return { ...counts, read: counts.read + 1 }
71 if (tool === 'Bash') return { ...counts, shell: counts.shell + 1 }
72 return { ...counts, other: counts.other + 1 }
73}
74
75/**
76 * FNV-1a over the text, as 8 hex digits: enough to tell repos apart in one
77 * person's history without writing their paths into the store.
78 */
79export function repoHashOf(text: string): string {
80 let hash = 0x811c9dc5
81 for (let i = 0; i < text.length; i += 1) {
82 hash ^= text.charCodeAt(i)
83 hash = Math.imul(hash, 0x01000193) >>> 0
84 }
85 return hash.toString(16).padStart(8, '0')
86}
87
88/**
89 * The shape's keys from most to least specific, the global key not included:
90 * full shape, shape without tool mix, type and steps, type alone.
91 */
92export function levelKeysOf(shape: Shape): string[] {
93 const tests = shape.hasTests ? 't' : 'n'
94 return [
95 `${shape.taskType}|${shape.steps}|${shape.repo}|${tests}|${shape.mix}`,
96 planKeyOf(shape),
97 `${shape.taskType}|${shape.steps}`,
98 `${shape.taskType}`,
99 ]
100}
101
102/**
103 * The most specific key known when the plan is: every field but the tool
104 * mix, which only the finished task can say. The spike gate reads this one.
105 */
106export function planKeyOf(shape: Shape): string {
107 return `${shape.taskType}|${shape.steps}|${shape.repo}|${shape.hasTests ? 't' : 'n'}`
108}
109
110/**
111 * The keys that do not depend on a step count, for a task with no plan yet.
112 */
113export function planlessKeysOf(shape: Shape): string[] {
114 return [`${shape.taskType}`]
115}
116src/spike.ts 63 lines1/**
2 * The spike (SPEC 3): when a task's shape is unfamiliar (under 3 samples) and
3 * its plan has 3 or more steps, ask one cheap, read-only subagent for minutes
4 * per step, once per task, capped at 60 seconds. Its answer counts as one
5 * sample at weight 0.5, and the pane says "spike guess" while it is the only
6 * basis.
7 */
8
9export const SPIKE_TIMEOUT_MS = 60_000
10export const SPIKE_MIN_STEPS = 3
11export const SPIKE_MODEL = 'haiku'
12/** Read-only, so a sizing question can never edit the user's files. */
13export const SPIKE_AGENT = 'Explore'
14
15const MAX_MINUTES = 240
16const MIN_MINUTES = 0.1
17
18export function spikePromptOf(request: string, steps: readonly string[], repoEntries: readonly string[]): string {
19 return [
20 'Estimate how long a coding assistant will take for each step of this plan.',
21 'Do not change any file. Do not run long commands. Answer within a minute.',
22 '',
23 'The request:',
24 request.slice(0, 2_000),
25 '',
26 'The plan:',
27 ...steps.map((step, i) => `${i + 1}. ${step}`),
28 '',
29 'Top level of the repository:',
30 repoEntries.length > 0 ? repoEntries.join(', ') : '(unknown)',
31 '',
32 `Reply with JSON only, one number of minutes per step (${steps.length} numbers) and your confidence from 1 to 5:`,
33 '{"minutes": [2, 5, 3], "confidence": 3}',
34 ].join('\n')
35}
36
37export type SpikeGuess = {
38 stepsMs: number[]
39 confidence: number
40}
41
42/**
43 * The guess in a subagent's reply, or null when it gave none that fits the
44 * plan. Each step is clamped to 6 seconds .. 4 hours.
45 */
46export function parseSpike(answer: string, stepCount: number): SpikeGuess | null {
47 const match = /\{[\s\S]*\}/.exec(answer)
48 if (!match) return null
49 let parsed: unknown
50 try {
51 parsed = JSON.parse(match[0])
52 } catch {
53 return null
54 }
55 if (typeof parsed !== 'object' || parsed === null) return null
56 const fields = parsed as { minutes?: unknown; confidence?: unknown }
57 if (!Array.isArray(fields.minutes) || fields.minutes.length !== stepCount) return null
58 if (!fields.minutes.every(m => typeof m === 'number' && Number.isFinite(m))) return null
59 const stepsMs = (fields.minutes as number[]).map(m => Math.min(MAX_MINUTES, Math.max(MIN_MINUTES, m)) * 60_000)
60 const confidence = typeof fields.confidence === 'number' ? Math.min(5, Math.max(1, fields.confidence)) : 1
61 return { stepsMs, confidence }
62}
63