SLOPSHOPPER

receipts

UNVERIFIED when Claude claims done without a check: a receipt line under every answer that claims done, naming the test or build that ran after the last edit…

newpanebandspinnerrowsguard
v0.2.3MITupdated 2026-10-07shawnpetros/claude-receipts
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · receipts
│ ┃ receipts ✕ › fix the failing auth test and add an audit log call │ ┃ receipts c: clean view on b: estimate basi │ ┃ plan · none given ⏺ Read(src/auth.ts) │ ┃ ✓ the task ⎿ Read 6 lines │ ┃ finished in 0s ⏺ Bash(bun test) │ ┃ receipt · rm -rf build ✓ · 0s ago ⎿ 3 pass, 1 fail │ ┃ calibration: no finished tasks yet │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ receipt · rm -rf build ✓ · 0s ago · 5h window 31% used, resets 09:53 │ │ › /receipts │ │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · receipts
receipts c: clean view on b: estimate basis plan · none given ✓ the task finished in 0s receipt · rm -rf build ✓ · 0s ago calibration: no finished tasks yet
README

receipts

A Claude Code mod that says UNVERIFIED when Claude claims done without a check. Under every answer that claims done, a receipt line names the test or build that ran after the last edit, or says that nothing did. It also replaces the wall of tool calls with a progress band that tells the truth.

The name reads like a cost meter, and it isn't one. It's a verification receipt.

It does three jobs:

  1. Receipts. When the turn ends, if the answer claims done and nothing verified the work after the last edit, a line says so under the answer. If something did, the line shows what ran and how it went, and what the turn cost when Claude Code knows.
  2. Milestones instead of tool walls. A bordered band directly above the prompt shows what the turn is working on, what's finished and what's left. At the default quiet level, ordinary tool rows draw nothing, during the turn and after it, and so does most of the chrome a turn scatters. The rows that matter still show, one dim line each with how they ended: the first edit of each file, test and build runs, commits, pull requests, and finished subagents.
  3. An estimate that learns and doesn't lie. It stays indeterminate until there's a basis. Then it shows a range, names where the range comes from, and shows how often its past ranges were right. The range narrows as steps finish.

Tested with Claude Code 2.1.291. Mods need 2.1.287 or later.

Install

claude --plugin-dir ~/projects/claude-receipts

That loads the mod for one session. Nothing opens by itself: the band appears above the prompt when you send a prompt. /receipts opens a pane with the long view if you want it.

Or install it from its marketplace:

claude plugin marketplace add shawnpetros/claude-receipts
claude plugin install receipts@claude-receipts

Use one or the other, not both. With --plugin-dir and the installed copy in one session, two copies of the mod draw the same band and keep separate state.

To check that it's learning, finish a turn and run /receipts stats. The count of finished tasks should go up by one.

The band

While a turn runs, the band has an orange border and these rows:

  • Title. A ✶, your prompt cut to fit, and the elapsed time on the right.
  • Summary. Step 2 of 4, a full-width bar, the percent done, then the estimate in dim text with its basis, such as ~4 to 9 min · from 3 similar tasks. The percent is weighted by how long each step is expected to take.
  • One row per step. A 12-cell bar for the step and a word: Working, Next, Later or Done. At most six steps show, around the current one, then +n more.

When the turn ends, the band always leaves the working state. It becomes a completion card and stays until your next prompt:

  • All done. A green border and a ✓ All done badge, only when every step finished. A task-tool plan counts as finished when every task is completed.
  • Turn ended short. A grey ■ Turn ended · 2 of 5 steps reached badge when steps were left. Those steps read Not reached, never Done.
  • Verified. A green border and a ✓ All done · bun test 152 pass badge, every step ticked with a full bar, and took 1m 47s on the right.
  • Unverified. A yellow border and a ⚠ Done, unverified badge. The title row shows the receipt, such as claimed done, no test/build/run after the last edit (src/x.ts at 14:02). When that doesn't fit, it shortens to the file and time first.
  • A failed check. A yellow ✗ Done, checks failed badge and the run that failed.
  • Stopped. A grey ■ Stopped badge when you interrupt the turn.

If agents the turn started are still running when it ends, the title row says waiting on 2 agents. The count drops as each one finishes, and the next prompt clears it.

A plan the mod derived marks a step done when a small model says the assistant's latest message finished it. When the turn ends, one more pass reads the final answer against the steps still open and marks the ones it shows were completed. That pass has 1.5 seconds. If it misses, the card says how far the plan got rather than guess.

When the elapsed time passes the top of the range, the border and bar turn grey and the summary reads over by 2m 10s · ~1 to 4 min more. Running long isn't an error, so it's never red.

Below 110 columns the per-step bars drop and each step keeps its word. Nothing in the band is ever wider than the terminal.

[▾] on the title row folds the band to that one row, and [▸] opens it again. Claude Code draws its own [-] just outside the border, and that one hides the whole band.

Keys

Click the band, or press ctrl+x then tab, to give it the keyboard. Then:

| Key | What it does | | :- | :- | | b | Shows a one-line tooltip with the estimate's basis and the calibration score. Press again to hide it. | | t | Opens or closes the settings popover |

Settings popover

t or /receipts tools opens a small bordered panel at the right of the band. It has three sections:

  • M O D E L. Haiku, Sonnet, Opus and Fable. The current model is highlighted. Picking one sets the /config model row when this build has one that takes it, and otherwise runs /model <name>.
  • E F F O R T. Low, Medium, High, XHigh and Max, the same way, through /effort. The current level is highlighted once a model request has carried it.
  • R O W S. Off, Clean or Quiet: how much of the transcript the mod draws away. See below.
  • S E T T I N G S. An On and Off switch for the spike.

Your plugin settings can't change while a session runs, so these choices override them for the rest of the session. /clear puts your settings back.

The [-] at the top right closes it, as does t again. Escape only hands the keyboard back to the prompt, because Claude Code tells a mod nothing when you press it.

The honesty contract

These are the rules the estimate follows. Each one exists because some tool, somewhere, broke it.

  • Never a single number. You always get a range such as ~4 to 9 min. A countdown never freezes. When the range is too narrow to round to two values, it's widened until it shows two.
  • "Indeterminate" has a time limit. You only see it while there's no plan and no history. It ends at the first finished milestone or after 90 seconds, whichever comes first. After that you get a range from the prior, labelled prior only.
  • The source is always named. Every range says where it comes from: from 7 similar tasks, spike guess, prior only, and derived plan when the plan is the mod's own guess.
  • The mod reports its own score. b in the band, the pane and /receipts stats show a line such as calibration: 61% of 18 tasks ended inside the range. After 10 or more tasks, if the score falls below 50%, the band says plainly that the ranges have been missing. That warning stays on the band and can't be hidden.
  • Countdown digits require confidence. Digits such as 3:52 to 8:40 left appear only when confidence is 0.5 or higher. Below that you get the rounded range and a bar.
  • Running over is reported. When elapsed time passes the top of the range, the row reads over by 2m 10s, the bar dims, and the range widens with elapsed time as its floor. It never resets to indeterminate.
  • The progress bar is weighted by time. It weights each step by how long that step is expected to take, not by the count of steps.
  • A derived plan is labelled. When the assistant didn't make a task list, the mod asks a small model for one and marks it derived. The mod never passes its own guess off as the assistant's plan.
  • The rows level only changes drawing. The transcript is never modified. Set it to Off and every row is back exactly as it was.
  • The receipt reports and never blocks. It never stops a turn or holds one back. The worst it does is wait up to 1.5 seconds for a label.

How the estimate works

Each task is filed under a shape: task type (build, debug, research, writing, config or refactor), step count (1, 2-3, 4-6 or 7+), a hash of the repository, whether the repository has tests, and the tool mix. Remaining time is the sum of the expected time of each step not yet done. The expectation comes from the most specific shape with at least 3 past tasks. Failing that, it falls back to a coarser shape, and then to a global prior.

The spread counts only the steps that are left, so it shrinks as steps finish. When a shape is new and the plan has 3 or more steps, the mod asks one cheap, read-only subagent for minutes per step. It does this once per task shape per session, with a 60-second limit, and counts the answer as half of one past task. It never does it for research, writing or chat tasks, whose steps are cheap.

Time spent waiting on a permission prompt isn't learned as work. Claude Code reports each tool's own run time without the prompt. The rest of the call is waiting, and it comes out of the step and the total before they're saved. Calibration is still scored on the wall clock, because that's what the range promised.

History lives in the mod's own store. Only finished task records are saved, and the averages are rebuilt from them. The store keeps the newest 500 tasks and stays under 1 MiB.

Receipts

Each turn keeps a ledger of edits and of shell commands that look like verification: test runners, tsc, linters, builds and make check. At the end of the turn:

| What happened | Line under the answer | | :- | :- | | Edited, claimed done, nothing verified after the last edit | UNVERIFIED · claimed done, no test/build/run after the last edit (src/x.ts at 14:02) | | Edited, claimed done, a verify run after the last edit | receipt · bun test ✓ 152 pass · 1m ago | | A verify run after the last edit that failed | receipt · bun test ✗ exit 1 · 150 pass · 2 fail · 5s ago | | No edits, or the answer doesn't claim done | nothing |

When Claude Code reports usage, the line ends with what the turn cost:

| Who you are | What the line adds | | :- | :- | | A Claude plan user, with rate-limit windows | 5h window 6% used, resets 18:30 | | An API key user, with no windows | $0.42 this turn | | Neither is known | nothing, never a guess |

If the current turn's receipt line is wrong, for example it called an answer done when it wasn't, run /receipts wrong. Each receipt can be marked once. /receipts stats shows claims-done false positives: 1 of 12 receipts (8%). That number is the precondition for ever letting the receipt block a turn. See the roadmap.

A small model decides whether the answer claims done. If its label isn't back within 1.5 seconds, plain completion words in the answer decide instead.

Commands

| Command | What it does | | :- | :- | | /receipts | Opens or closes the pane, the long view with the calibration line | | /receipts tools | Opens the settings popover in the band | | /receipts rows [off\|clean\|quiet] | Sets the rows level, or steps to the next one with no argument | | /receipts clean | Switches the rows level to Off, and back to what it was | | /receipts basis | Shows or hides the basis tooltip, as b does | | /receipts stats | Shows the calibration history and task counts by type, with median durations | | /receipts wrong | Marks the current turn's receipt line as a false positive, for the count in stats | | /receipts reset-history | Forgets every learned task |

In the pane, c switches the rows level to Off and back, and b shows where the estimate comes from.

Rows level

The rows level sets how much of the transcript the mod draws away. Quiet is the default.

| Level | What draws | | :- | :- | | Off | Every row exactly as Claude Code draws it | | Clean | Tool rows as one dim line each, such as ● Edit src/x.ts. Milestone rows in full. | | Quiet | See below |

Quiet goes further:

  • Tool rows. Ordinary tool rows, their results and folded groups draw nothing, during the turn and after it. Each milestone becomes one dim line with how it ended, such as ● Bash bun test ✓.
  • The spinner. It draws nothing, because the band already shows the elapsed time and the step.
  • While a turn runs. Progress pills, the turn-duration line, status notices and other commands' output draw nothing. This mod's own output and any error line always show.
  • Hand-backs. A subagent's hand-back, a message from another session or a task notification draws nothing while the turn runs. After the turn it draws one dim line, such as ↳ message from @Explore: Found 3 mods.
  • Assistant text. Text between tool calls draws as its first line, dim. The final answer draws in full once the turn completes.

Your own prompts always draw in full. Quiet never touches a question the assistant asks you or a permission prompt.

ctrl+o is the escape hatch. It shows the full transcript, and a hand-back row it expands draws in full. ctrl+o also expands a folded group of reads and searches. A single tool row or a block of assistant text can't be expanded past the level, because Claude Code doesn't tell a mod when one is expanded. Set the level to Off to see those in full.

Configuration

Set these under pluginConfigs in your Claude Code settings. Use the key receipts@inline for a session started with --plugin-dir.

| Key | Default | What it does | | :- | :- | :- | | spike | true | Allows the sizing subagent for unfamiliar tasks. Set it to false and the mod never spawns one. | | cleanView | true | Starts each session at the quiet rows level. Set it to false to start at Off. The popover and /receipts rows change it for the session. |

What it reaches

The mod calls the model API for its labels and the derived plan, and spawns a subagent for the spike. It makes no other network calls and sends no telemetry of its own. The only thing it writes is its own store.

Known issues

  • VS Code draws no pane and no band, so only the receipt line and the transcript rows show. This is tracked upstream as anthropics/claude-code #99423 and #99691.
  • Claude Code Desktop drops commands a mod registers, so /receipts and its subcommands aren't there. The band's keys still work.
  • Two copies at once. Don't load the inline copy (--plugin-dir) and the marketplace install in the same session. Both draw the same band and keep separate state, so the band you see may belong to the copy that missed the turn's end.
  • Builds before 2.1.287 don't run mods. hooks/hooks.json carries an empty hooks key beside modules, so an older build loads nothing instead of failing.

Roadmap

Each release answers one question it can measure.

  • 0.3: can someone use it without a manual? Everything from the 0.2 brief that hasn't landed yet. The test is a new user reading only the band and /receipts help.
  • 0.4: does the receipt change behaviour? The cost line from $.session.usage() already shipped in 0.2.2. The rest is an optional gate that turns UNVERIFIED from a mirror into a block, off by default. It ships only after the claims-done classifier's false-positive rate has been measured over enough live turns with /receipts wrong and /receipts stats. A gate that blocks a finished turn gets the mod uninstalled.

Not planned: token meters, burn bars, or a rename.

Development

claude plugin validate . --strict
claude plugin test
bun scripts/simulate.ts
bun scripts/mock-band.ts 155
bun scripts/check-manifest.ts

scripts/mock-band.ts prints the band as plain text in each state: working, over the range, verified, unverified, narrow, collapsed, with the tooltip, and the popover. It uses the same pure view the mod draws from. The argument is the band's width, which is the terminal's width less 5.

src/ holds the logic as plain modules with no mods API dependency. These are the estimator, history, milestones, ledger, shape, spike, view and band modules, plus a session that talks to Claude Code through a small host interface. hooks/register.ts builds that host from $ and wires the hooks.

Source 13 files
hooks/register.ts 547 lines
1import type { ConfigRow, EngineInterface, On, PluginOptions, RenderElement } from 'claude-code'
2import { atom, read, update } from 'claude-code'
3
4import {
5  ACCENT,
6  bandView,
7  EFFORT_CHOICES,
8  MODEL_CHOICES,
9  ROWS_LEVELS,
10  TOOLS_COLUMNS,
11  toolsView,
12  type Row,
13  type RowsLevel,
14  type Seg,
15} from '../src/band'
16import type { Host } from '../src/host'
17import { firstLineOf, handbackLineOf, QuietLog } from '../src/quiet'
18import { PANE_ID, PANE_TITLE, ReceiptsSession } from '../src/session'
19import { paneLines, toolGroupText, toolResultText, toolRowText, type Line } from '../src/view'
20
21/**
22 * The rows level: `off` draws every row as Claude Code does; `clean` draws
23 * tool rows as one dim line and milestones in full; `quiet` (the default)
24 * draws tool rows as nothing and milestones as one dim line, during the turn
25 * and after it, and folds the chrome a turn scatters: the spinner, progress
26 * pills, notices and command output while working, subagent hand-backs, and
27 * interim assistant text down to its first line.
28 *
29 * Invariant 4, drawing only: these hooks return a drawing and never touch
30 * the transcript, so `off` shows every row as it was. Scar: the /buddy
31 * main-model leak; a mod rewriting content is a different and riskier thing
32 * than a mod redrawing it. Milestones always show (invariant 5).
33 */
34const rows = atom({ plugin: 'receipts', key: 'rows' }, 'quiet')
35
36/** The level `/receipts clean` and the pane's `c` go back to from `off`. */
37const rowsLast = atom({ plugin: 'receipts', key: 'rowsLast' }, 'quiet')
38
39/**
40 * Whether a main-loop turn is running. The quiet sites read it, so they draw
41 * again once at each turn edge, not on every tick.
42 */
43const working = atom({ plugin: 'receipts', key: 'working' }, false)
44
45/** The band folded to its title row by its own `[▾]` (`[▸]` folded). */
46const collapsed = atom({ plugin: 'receipts', key: 'collapsed' }, false)
47
48/** The settings popover, drawn in the band. */
49const toolsOpen = atom({ plugin: 'receipts', key: 'toolsOpen' }, false)
50
51/**
52 * The spike switch as the person last set it. userConfig is read-only at
53 * runtime, so this state overrides it for the session; it starts from it.
54 */
55const spike = atom({ plugin: 'receipts', key: 'spike' }, true)
56
57/**
58 * Read by the pane and band only, so bumping it once a second redraws those
59 * two and not every tool row in the transcript.
60 */
61const tick = atom({ plugin: 'receipts', key: 'tick' }, 0)
62
63const USAGE =
64  'usage: /receipts (toggle the pane), /receipts tools, /receipts rows [off|clean|quiet], /receipts clean, /receipts basis, /receipts stats, /receipts wrong, /receipts reset-history'
65const PANE_MIN_ROWS = 6
66const PANE_MAX_ROWS = 24
67
68/**
69 * The mods API as the session logic sees it. Top level and handed `$`, so
70 * `claude plugin validate` lists every call through it.
71 */
72function hostOf($: EngineInterface): Host {
73  return {
74    now: () => $.clock.now(),
75    every: (ms, fn) => $.clock.every(ms, fn),
76    after: (ms, fn) => $.clock.after(ms, fn),
77    storeGet: key => $.store.get(key),
78    storeSet: (key, value) => $.store.set(key, value),
79    storeDelete: key => $.store.delete(key),
80    redraw: () => {
81      update($, tick, n => n + 1).catch(() => $.ui.invalidate('ui.render'))
82    },
83    classify: (text, labels) => $.model.classify(text, labels),
84    complete: request => $.model.complete(request),
85    spawn: args => $.agent.spawn(args),
86    cwd: () => $.session.cwd(),
87    repo: () => $.session.repo(),
88    list: path => $.fs.list(path),
89    agents: () => $.agent.list(),
90    usage: () => $.session.usage(),
91  }
92}
93
94/**
95 * `off` and back: to the level the person had before, quiet if none.
96 */
97async function toggleClean($: EngineInterface): Promise<RowsLevel> {
98  const level = (await read($, rows)) as RowsLevel
99  if (level === 'off') {
100    const back = (await read($, rowsLast)) as RowsLevel
101    await update($, rows, () => back)
102    return back
103  }
104  await update($, rowsLast, () => level)
105  await update($, rows, () => 'off')
106  return 'off'
107}
108
109async function setRows($: EngineInterface, level: RowsLevel): Promise<void> {
110  if (level !== 'off') await update($, rowsLast, () => level)
111  await update($, rows, () => level)
112}
113
114async function levelOf($: EngineInterface): Promise<RowsLevel> {
115  return (await read($, rows)) as RowsLevel
116}
117
118/**
119 * Sets the model or effort the way the person would: the `/config` row when
120 * the build has one that takes it, else the slash command. Top level and
121 * handed `$`, so validate lists the calls.
122 */
123async function choose($: EngineInterface, key: 'model' | 'effort', value: string): Promise<void> {
124  const rows: ConfigRow[] = await $.config.list().catch(() => [])
125  const row = rows.find(candidate => candidate.key === key)
126  if (row && !row.isLocked && row.kind !== 'boolean' && row.kind !== 'number') {
127    const result = await $.config.set({ key, value }).catch(() => ({ deny: 'failed' }))
128    if (result.deny === undefined) return
129  }
130  await $.command.run({ command: key, args: value })
131}
132
133/**
134 * The userConfig defaults into state, where the band's switches change them.
135 */
136async function applyDefaults($: EngineInterface, session: ReceiptsSession, isClean: boolean, isSpike: boolean): Promise<void> {
137  if (!isClean) await update($, rows, () => 'off')
138  if (!isSpike) await update($, spike, () => false)
139  session.setSpike(isSpike)
140}
141
142async function sessionModelOf($: EngineInterface): Promise<string> {
143  return $.session.model().catch(() => '')
144}
145
146/**
147 * A line of the pane as Text props, leaving out the styles it does not set.
148 */
149function textPropsOf(line: Line) {
150  return {
151    key: line.key,
152    wrap: 'truncate-end' as const,
153    children: [line.text],
154    ...(line.dim ? { dimColor: true } : {}),
155    ...(line.bold ? { bold: true } : {}),
156    ...(line.color ? { color: line.color } : {}),
157  }
158}
159
160type Elements = ReturnType<EngineInterface['ui']['resolve']>
161type Presses = Record<string, () => unknown>
162
163/**
164 * A row of the band as elements: a Text per segment, a plain Button where a
165 * segment carries one. A Button whose key has no handler draws as text.
166 */
167function rowOf(el: Elements, row: Row, presses: Presses): RenderElement {
168  const children = row.segs.map((seg: Seg): RenderElement => {
169    const press = seg.button ? presses[seg.button.key] : undefined
170    if (seg.button && press) {
171      return el.Button({
172        key: seg.button.key,
173        label: seg.button.label,
174        plain: true,
175        ...(seg.button.hotkey ? { hotkey: seg.button.hotkey } : {}),
176        ...(seg.dim ? { dimColor: true } : {}),
177        onPress: () => press(),
178      })
179    }
180    return el.Text({
181      children: [seg.text],
182      ...(seg.color ? { color: seg.color } : {}),
183      ...(seg.dim ? { dimColor: true } : {}),
184      ...(seg.bold ? { bold: true } : {}),
185      ...(seg.inverse ? { inverse: true } : {}),
186    })
187  })
188  return el.Box({ key: row.key, flexDirection: 'row', children })
189}
190
191/**
192 * The one dim line a milestone row folds to while suppressed: the call and
193 * how it ended. A verify run's exit shows as ✓ or ✗ (invariant 5).
194 */
195function outcomeMarkOf(props: { isRunning?: boolean; isErrored?: boolean; isInterrupted?: boolean }): string {
196  if (props.isRunning) return ' …'
197  if (props.isInterrupted) return ' ■'
198  return props.isErrored ? ' ✗' : ' ✓'
199}
200
201export function register(on: On, options: PluginOptions): void {
202  const isSpikeByDefault = options.spike !== false
203  const session = new ReceiptsSession({ spike: isSpikeByDefault })
204  const quiet = new QuietLog()
205  const isCleanByDefault = options.cleanView !== false
206
207  on('session.start', async ($, e, next) => {
208    await session.start(hostOf($))
209    const stored = await $.state.get({ plugin: 'receipts', key: 'rows' })
210    if (stored.version === 0) await applyDefaults($, session, isCleanByDefault, isSpikeByDefault)
211    try {
212      await $.command.register({
213        name: 'receipts',
214        description: 'Toggle the receipts pane, open its settings, or show calibration stats',
215        argumentHint: '[tools|rows|clean|basis|stats|reset-history]',
216        immediate: true,
217      })
218    } catch {
219      // A name clash leaves the mod running without its command
220    }
221    return next(e)
222  })
223
224  // /clear, /resume and /branch reset $.state: put the configured defaults back
225  on('classic.SessionStart', { source: ['clear', 'resume', 'fork'] }, async ($, e, next) => {
226    await applyDefaults($, session, isCleanByDefault, isSpikeByDefault)
227    return next(e)
228  }).catch(($, e, next) => next(e))
229
230  on('turn.start', async ($, e, next) => {
231    await session.turnStart(hostOf($), e.turnId, e.text)
232    await update($, working, () => true)
233    return next(e)
234  })
235
236  on('turn.step', async function* ($, e, next) {
237    if (e.agentId === undefined) session.noteEffort(e.effort)
238    const result = yield* next(e)
239    session.stepResult(hostOf($), e, result)
240    return result
241  })
242
243  on('tool.call', async ($, e, next) => {
244    session.beforeTool(e, e)
245    session.toolStarted(e.tool_use_id, await $.clock.now())
246    const outcome = await next(e)
247    await session.afterTool(hostOf($), e, e, outcome).catch(() => undefined)
248    return outcome
249  }).catch(($, e, next) => next(e))
250
251  // The tool's own run time, which excludes the permission prompt: the rest
252  // of the call's span was waiting on the person, and is not learned as work
253  on('classic.PostToolUse', async ($, e, next) => {
254    session.toolRan(e.tool_use_id, e.duration_ms)
255    return next(e)
256  }).catch(($, e, next) => next(e))
257
258  // Under the answer: the receipt, or UNVERIFIED. A mirror, never a gate
259  on('turn.complete', async ($, e, next) => {
260    const result = await next(e)
261    if (e.agentId !== undefined) {
262      session.subagentComplete(hostOf($), e.agentId, e.answer)
263      return result
264    }
265    const line = await session.turnComplete(hostOf($), e)
266    await update($, working, () => false)
267    // The final answer, drawn as one dim line while it streamed, in full now
268    if (quiet.finish(e.answer)) $.ui.invalidate('ui.render')
269    if (!line) return result
270    const isOwnText = result.text !== '' && result.text !== e.answer
271    return { ...result, text: isOwnText ? `${result.text}\n${line}` : line }
272  })
273
274  on('command.run', { command: 'receipts' }, async ($, e) => {
275    const arg = e.args.trim()
276    if (arg === 'stats') return { text: session.statsText() }
277    if (arg === 'tools') {
278      await update($, toolsOpen, () => true)
279      return {}
280    }
281    if (arg === 'basis') {
282      session.toggleBasis()
283      $.ui.invalidate('ui.render')
284      return { text: session.isBasisShown ? 'basis shown on the band' : 'basis hidden' }
285    }
286    if (arg === 'wrong') return { text: await session.markWrong(hostOf($)) }
287    if (arg === 'clean') return { text: `rows: ${await toggleClean($)}` }
288    if (arg === 'rows' || arg.startsWith('rows ')) {
289      const wanted = arg.slice(4).trim()
290      if (wanted === '') {
291        const now = await levelOf($)
292        const next = ROWS_LEVELS[(ROWS_LEVELS.indexOf(now) + 1) % ROWS_LEVELS.length]!
293        await setRows($, next)
294        return { text: `rows: ${next}` }
295      }
296      if (!(ROWS_LEVELS as readonly string[]).includes(wanted)) return { text: USAGE }
297      await setRows($, wanted as RowsLevel)
298      return { text: `rows: ${wanted}` }
299    }
300    if (arg === 'reset-history') {
301      const count = await session.resetHistory(hostOf($))
302      return { text: `history cleared: ${count} ${count === 1 ? 'task' : 'tasks'} forgotten` }
303    }
304    if (arg !== '') return { text: USAGE }
305    if (session.paneOpen) {
306      await $.ui.close({ id: PANE_ID })
307      session.paneClosed()
308      return {}
309    }
310    // rows: inline (the main screen) it opens as tall as its content, not cut
311    // to a third; a dock ignores it and the render fits its rows instead
312    const wanted = paneLines(session.viewModel(await $.clock.now())).length + 1
313    const placed = await $.ui.open({ id: PANE_ID, title: PANE_TITLE, rows: Math.max(PANE_MIN_ROWS, Math.min(PANE_MAX_ROWS, wanted)) })
314    session.paneOpened(placed.isPlaced)
315    return {}
316  })
317
318  on('ui.close', { id: PANE_ID }, async ($, e, next) => {
319    session.paneClosed()
320    return next(e)
321  }).catch(($, e, next) => next(e))
322
323  // The pane: the long view, opened only by /receipts
324  on('ui.render', { component: 'Pane' }, async ($, e, next) => {
325    if (e.requestId !== PANE_ID) return next(e)
326    const isClean = (await levelOf($)) !== 'off'
327    await read($, tick)
328    const now = await $.clock.now()
329    session.paneDrawn()
330    session.ensureTicking(hostOf($), now)
331    const { Box, Text, Button } = $.ui.resolve(e)
332    // Never more rows than the room has: the header, then the lines that fit
333    const all = paneLines(session.viewModel(now))
334    const room = Math.max(1, e.props.scroll.bodyRows - 1)
335    const lines =
336      all.length <= room ? all : [...all.slice(0, room - 1), { key: 'pane-more', text: `+${all.length - room + 1} more · /receipts stats`, dim: true }]
337    return Box({
338      flexDirection: 'column',
339      children: [
340        Box({
341          flexDirection: 'row',
342          columnGap: 2,
343          children: [
344            Text({ bold: true, children: ['receipts'] }),
345            Button({
346              key: 'clean-toggle',
347              label: isClean ? 'clean view on' : 'clean view off',
348              hotkey: 'c',
349              plain: true,
350              onPress: () => toggleClean($),
351            }),
352            Button({
353              key: 'basis',
354              label: session.isBasisShown ? 'hide basis' : 'estimate basis',
355              hotkey: 'b',
356              plain: true,
357              onPress: () => {
358                session.toggleBasis()
359                $.ui.invalidate('ui.render')
360              },
361            }),
362          ],
363        }),
364        ...lines.map(line => Text(textPropsOf(line))),
365      ],
366    })
367  })
368
369  // The band: the primary surface. Progress while working, the completion
370  // card after, the settings popover under it when asked
371  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
372    if (e.props.hasSurvey) return next(e)
373    const isToolsOpen = await read($, toolsOpen)
374    const isBandShown = session.isBandWanted()
375    if (!isBandShown && !isToolsOpen) return next(e)
376    await read($, tick)
377    const isCollapsed = await read($, collapsed)
378    const now = await $.clock.now()
379    session.ensureTicking(hostOf($), now)
380    const el = $.ui.resolve(e)
381    const presses: Presses = {
382      collapse: () => update($, collapsed, value => !value),
383      basis: () => {
384        session.toggleBasis()
385        $.ui.invalidate('ui.render')
386      },
387      'tools-toggle': () => update($, toolsOpen, value => !value),
388      'tools-close': () => update($, toolsOpen, () => false),
389      'set-spike': () =>
390        update($, spike, value => {
391          session.setSpike(!value)
392          return !value
393        }),
394    }
395    for (const level of ROWS_LEVELS) presses[`rows-${level}`] = () => setRows($, level)
396    for (const choice of MODEL_CHOICES) {
397      presses[`model-${choice.value}`] = async () => {
398        await choose($, 'model', choice.value)
399        $.ui.invalidate('ui.render')
400      }
401    }
402    for (const choice of EFFORT_CHOICES) {
403      presses[`effort-${choice.value}`] = async () => {
404        await choose($, 'effort', choice.value)
405        session.noteEffort(choice.value)
406        $.ui.invalidate('ui.render')
407      }
408    }
409
410    const columns = e.props.bodyColumns
411    let band: RenderElement | null = null
412    if (isBandShown) {
413      const view = bandView(session.viewModel(now), columns, { collapsed: isCollapsed, showBasis: session.isBasisShown })
414      band = el.Box({
415        key: 'band',
416        flexDirection: 'column',
417        borderStyle: 'round',
418        borderColor: view.border,
419        paddingX: 1,
420        width: columns,
421        children: view.rows.map(row => rowOf(el, row, presses)),
422      })
423    }
424    if (!isToolsOpen && band) return band
425
426    const toolsWidth = Math.min(columns, TOOLS_COLUMNS)
427    const rows = toolsView(
428      {
429        model: await sessionModelOf($),
430        ...(session.effortLevel ? { effort: session.effortLevel } : {}),
431        rows: await levelOf($),
432        spike: await read($, spike),
433      },
434      toolsWidth - 4,
435    )
436    const tools = el.Box({
437      key: 'tools',
438      flexDirection: 'column',
439      borderStyle: 'round',
440      borderColor: ACCENT,
441      paddingX: 1,
442      width: toolsWidth,
443      children: rows.map(row => rowOf(el, row, presses)),
444    })
445    return el.Box({
446      flexDirection: 'column',
447      width: columns,
448      children: [...(band ? [band] : []), el.Box({ flexDirection: 'row', justifyContent: 'flex-end', children: [tools] })],
449    })
450  })
451
452  // ctrl+o: a ToolGroup's props say when it is expanded, so an expanded group
453  // draws in full. ToolUse and ToolResult props carry no such flag, so a
454  // single row cannot be expanded past the level; `/receipts rows off` can
455  on('ui.render', { component: 'ToolUse' }, async ($, e, next) => {
456    const level = await levelOf($)
457    if (level === 'off') return next(e)
458    const isMilestone = session.isMilestoneRow(e.props.tool_use_id, e.props.tool, e.props.input)
459    const { Box, Text } = $.ui.resolve(e)
460    if (level === 'quiet') {
461      if (!isMilestone) return Box({})
462      const text = toolRowText(e.props.tool, e.props.input, session.workingDirectory) + outcomeMarkOf(e.props)
463      return Text({ dimColor: true, wrap: 'truncate-end', children: [text] })
464    }
465    if (isMilestone) return next(e)
466    return Text({ dimColor: true, wrap: 'truncate-end', children: [toolRowText(e.props.tool, e.props.input, session.workingDirectory)] })
467  })
468
469  on('ui.render', { component: 'ToolResult' }, async ($, e, next) => {
470    const level = await levelOf($)
471    if (level === 'off') return next(e)
472    const { Box, Text } = $.ui.resolve(e)
473    // Quiet: a milestone's outcome is folded into its ToolUse line
474    if (level === 'quiet') return Box({})
475    if (session.isMilestoneRow(e.props.tool_use_id, e.props.tool, undefined)) return next(e)
476    return Text({ dimColor: true, wrap: 'truncate-end', children: [toolResultText(e.props.output, e.props.isErrored)] })
477  })
478
479  on('ui.render', { component: 'ToolGroup' }, async ($, e, next) => {
480    const level = await levelOf($)
481    if (e.props.isExpanded || level === 'off') return next(e)
482    const { Box, Text } = $.ui.resolve(e)
483    if (level === 'quiet') return Box({})
484    return Text({ dimColor: true, wrap: 'truncate-end', children: [toolGroupText(e.props.calls)] })
485  })
486
487  // ---- quiet: the chrome a turn scatters ------------------------------------
488  // Never hooked: AskUserQuestion and the permission dialogs. A question to
489  // the person is the one thing quiet must never fold.
490
491  // The band shows the elapsed time and the step; the spinner repeats it
492  on('ui.render', { component: 'Spinner' }, async ($, e, next) => {
493    if ((await levelOf($)) !== 'quiet') return next(e)
494    const { Box } = $.ui.resolve(e)
495    return Box({})
496  })
497
498  on('ui.render', { component: 'ToolProgress' }, async ($, e, next) => {
499    if ((await levelOf($)) !== 'quiet' || !(await read($, working))) return next(e)
500    const { Box } = $.ui.resolve(e)
501    return Box({})
502  })
503
504  on('ui.render', { component: 'TurnDuration' }, async ($, e, next) => {
505    if ((await levelOf($)) !== 'quiet' || !(await read($, working))) return next(e)
506    const { Box } = $.ui.resolve(e)
507    return Box({})
508  })
509
510  on('ui.render', { component: 'InfoNotice' }, async ($, e, next) => {
511    if ((await levelOf($)) !== 'quiet' || !(await read($, working))) return next(e)
512    const { Box } = $.ui.resolve(e)
513    return Box({})
514  })
515
516  // Another command's output mid-turn; this mod's own and any error line show
517  on('ui.render', { component: 'CommandOutput' }, async ($, e, next) => {
518    if (e.props.command === 'receipts' || e.props.isErrored) return next(e)
519    if ((await levelOf($)) !== 'quiet' || !(await read($, working))) return next(e)
520    const { Box } = $.ui.resolve(e)
521    return Box({})
522  })
523
524  // A subagent's hand-back, a peer's message, a task notification: nothing
525  // while the turn runs, one dim line after. The person's own prompt and a
526  // ctrl+o expanded row are never touched
527  on('ui.render', { component: 'UserMessage' }, async ($, e, next) => {
528    if (e.props.isExpanded || e.props.origin.kind === 'composer') return next(e)
529    if ((await levelOf($)) !== 'quiet') return next(e)
530    const { Box, Text } = $.ui.resolve(e)
531    if (await read($, working)) return Box({})
532    return Text({ dimColor: true, wrap: 'truncate-end', children: [handbackLineOf(e.props.text, e.props.from?.name)] })
533  })
534
535  // Interim text: its first line, dim. The final answer redraws in full when
536  // turn.complete names it (QuietLog). A block never seen mid-turn is left alone
537  on('ui.render', { component: 'AssistantMessage' }, async ($, e, next) => {
538    if ((await levelOf($)) !== 'quiet') return next(e)
539    const isWorking = await read($, working)
540    if (isWorking) quiet.seen(e.requestId, e.props.text)
541    const kind = quiet.kindOf(e.requestId)
542    if (kind !== 'interim') return next(e)
543    const { Text } = $.ui.resolve(e)
544    return Text({ dimColor: true, wrap: 'truncate-end', children: [firstLineOf(e.props.text)] })
545  })
546}
547
src/band.ts 506 lines
1/**
2 * The band above the prompt and the settings popover, as rows of styled
3 * segments, each row exactly the width it is given. Pure: no `$`, no
4 * elements. `hooks/register.ts` turns a segment into a Text (or a Button
5 * when it carries one) and a row into a Box; tests and `scripts/mock-band.ts`
6 * read the same rows as plain text.
7 *
8 * Laying out to exact widths here, instead of trusting flex to right-align,
9 * keeps the right-hand column (elapsed, percent, the estimate) where the
10 * tests say it is at every width, and nothing ever wraps the band taller.
11 */
12
13import type { Color } from 'claude-code'
14
15import { basisLines, durationOf, estimateRow, type Estimate } from './estimator'
16import type { Milestone, Plan } from './milestones'
17import { truncate, type ViewModel } from './view'
18
19export const ACCENT: Color = 'claude'
20export const DONE: Color = 'success'
21export const WARN: Color = 'warning'
22/** Over the range: the bar and border go grey, never red. Running long is not an error. */
23export const OVER: Color = 'inactive'
24
25/**
26 * Below this many band body columns the per-step bars drop. The band's body
27 * is the terminal less the engine's five cells for its own `[-]`, so 105 is
28 * a 110-column terminal.
29 */
30export const NARROW_BODY = 105
31/** The round border and one cell of padding each side. */
32export const FRAME = 4
33export const MINI_BAR = 12
34export const MAX_STEP_ROWS = 6
35const MIN_BAR = 8
36const MIN_TAIL = 12
37
38export type SegButton = {
39  key: string
40  label: string
41  hotkey?: string
42}
43
44export type Seg = {
45  text: string
46  color?: Color
47  dim?: boolean
48  bold?: boolean
49  inverse?: boolean
50  /** Shrinks first when the row is too wide. */
51  grow?: boolean
52  /** Drawn as a plain Button; `text` is what the terminal draws for it. */
53  button?: SegButton
54}
55
56export type Row = {
57  key: string
58  segs: Seg[]
59}
60
61export type BandTone = 'working' | 'over' | 'done' | 'ended' | 'unverified' | 'stopped'
62
63export type BandView = {
64  tone: BandTone
65  border: Color
66  rows: Row[]
67}
68
69export type BandOptions = {
70  collapsed: boolean
71  showBasis: boolean
72}
73
74export function innerWidthOf(columns: number): number {
75  return Math.max(1, columns - FRAME)
76}
77
78export function rowText(row: Row): string {
79  return row.segs.map(seg => seg.text).join('')
80}
81
82function widthOf(segs: readonly Seg[]): number {
83  return segs.reduce((sum, seg) => sum + [...seg.text].length, 0)
84}
85
86function space(n: number): Seg {
87  return { text: ' '.repeat(Math.max(0, n)) }
88}
89
90/**
91 * A plain Button as a segment: `b: basis` with a hotkey, the label alone without.
92 */
93function buttonSeg(key: string, label: string, hotkey?: string, dim = true): Seg {
94  return {
95    text: hotkey ? `${hotkey}: ${label}` : label,
96    ...(dim ? { dim: true } : {}),
97    button: { key, label, ...(hotkey ? { hotkey } : {}) },
98  }
99}
100
101/**
102 * Cuts segments to `width` cells from the right, keeping each one's style;
103 * a Button that no longer fits whole is dropped, never drawn half.
104 */
105function clip(segs: readonly Seg[], width: number): Seg[] {
106  const out: Seg[] = []
107  let left = width
108  for (const seg of segs) {
109    const length = [...seg.text].length
110    if (length <= left) {
111      out.push(seg)
112      left -= length
113      continue
114    }
115    if (left > 0 && !seg.button) out.push({ ...seg, text: truncate(seg.text, left) })
116    else if (left > 0) out.push(space(left))
117    left = 0
118    break
119  }
120  return out
121}
122
123/**
124 * One row of exactly `width` cells: `left` at the start, `right` at the end,
125 * at least `gap` spaces between. The `grow` segment of `left` gives way first.
126 */
127function line(width: number, left: readonly Seg[], right: readonly Seg[] = [], gap = 2): Seg[] {
128  const rightWidth = widthOf(right)
129  const room = width - rightWidth - (right.length > 0 ? gap : 0)
130  let segs = [...left]
131  const over = widthOf(segs) - room
132  if (over > 0) {
133    const i = segs.findIndex(seg => seg.grow)
134    if (i !== -1) {
135      const seg = segs[i]!
136      const keep = Math.max(0, [...seg.text].length - over)
137      segs[i] = { ...seg, text: keep === 0 ? '' : truncate(seg.text, keep) }
138    }
139    if (widthOf(segs) > Math.max(0, room)) segs = clip(segs, Math.max(0, room))
140  }
141  if (rightWidth > width) return clip([...right], width)
142  const pad = width - widthOf(segs) - rightWidth
143  return [...segs, space(pad), ...right]
144}
145
146function barSegs(progress: number, width: number, color: Color, isDim = false, full = '━', empty = '─'): Seg[] {
147  const filled = Math.max(0, Math.min(width, Math.round(progress * width)))
148  return [
149    { text: full.repeat(filled), color, ...(isDim ? { dim: true } : {}) },
150    { text: empty.repeat(width - filled), dim: true },
151  ]
152}
153
154// ---- receipts on the card ---------------------------------------------------
155
156type Verdict =
157  | { kind: 'none' }
158  | { kind: 'verified'; label: string }
159  | { kind: 'failed'; text: string }
160  | { kind: 'unverified'; text: string; short: string }
161
162/**
163 * Reads the receipt line `ledger.receiptLine` wrote back into what the card
164 * needs: `receipt · bun test ✓ 152 pass · 1m ago` is verified, labelled
165 * `bun test 152 pass`; a ✗ is a failed check; UNVERIFIED is unverified.
166 */
167export function verdictOf(receipt: string | null): Verdict {
168  if (!receipt) return { kind: 'none' }
169  if (receipt.startsWith('UNVERIFIED')) {
170    const text = receipt.replace(/^UNVERIFIED · /, '')
171    const where = /\((.+) at (\d{1,2}:\d{2})\)$/.exec(text)
172    return { kind: 'unverified', text, short: where ? `${where[1]} at ${where[2]}, no check after` : text }
173  }
174  const body = receipt.replace(/^receipt · /, '').replace(/ · \d+[smh] ago$/, '')
175  if (body.includes('✗')) return { kind: 'failed', text: body }
176  return { kind: 'verified', label: body.replace(' ✓', '').replace(/\s+/g, ' ').trim() }
177}
178
179// ---- the band ---------------------------------------------------------------
180
181function currentIndexOf(plan: Plan): number {
182  return plan.items.findIndex(item => item.state === 'current')
183}
184
185function reachedOf(plan: Plan | null): { done: number; total: number } {
186  const items = plan?.items ?? []
187  return { done: items.filter(item => item.state === 'done').length, total: items.length }
188}
189
190/** Every step done, or no steps at all to fall short of. */
191function isAllDone(plan: Plan | null): boolean {
192  const { done, total } = reachedOf(plan)
193  return done === total
194}
195
196function stepCounterOf(plan: Plan | null): string {
197  if (!plan || plan.items.length === 0) return 'Planning'
198  const n = plan.items.length
199  const current = currentIndexOf(plan)
200  const done = plan.items.filter(item => item.state === 'done').length
201  const i = current === -1 ? Math.min(done + 1, n) : current + 1
202  return `Step ${i} of ${n}` + (plan.source === 'derived' ? ' · derived' : '')
203}
204
205function titleRow(model: ViewModel, width: number, options: BandOptions, tone: BandTone, verdict: Verdict): Row {
206  const title = model.title.replace(/\s+/g, ' ').trim() || 'continuing'
207  const buttons = [buttonSeg('basis', 'basis', 'b'), space(2), buttonSeg('tools-toggle', 'tools', 't'), space(2)]
208  // Not [-]: Claude Code draws its own [-] beside the band, which hides it
209  // whole; this one folds to the title row
210  const collapse = buttonSeg('collapse', options.collapsed ? '[▸]' : '[▾]', undefined, false)
211  if (tone === 'working' || tone === 'over') {
212    const left: Seg[] = [{ text: '✶ ', color: ACCENT, bold: true }, { text: title, bold: true, grow: true }]
213    return { key: 'title', segs: line(width, left, [...buttons, { text: durationOf(model.elapsedMs) }, space(1), collapse]) }
214  }
215  const waiting = model.finished?.waitingAgents ?? 0
216  const took = `took ${durationOf(model.finished?.totalMs ?? model.elapsedMs)}`
217  const tail: Seg[] = waiting > 0 ? [{ text: `waiting on ${waiting} ${waiting === 1 ? 'agent' : 'agents'}`, color: WARN }, space(2)] : []
218  let badge: Seg
219  let text: Seg
220  switch (verdict.kind) {
221    case 'unverified': {
222      // The file is the point: when the whole receipt does not fit, the
223      // path-first form does, so the cut never lands on the path
224      badge = { text: ' ⚠ Done, unverified ', color: WARN, inverse: true, bold: true }
225      const room = width - [...badge.text].length - 1 - 2 - widthOf([...tail, ...buttons, { text: took }, space(1), collapse])
226      text = { text: [...verdict.text].length <= room ? verdict.text : verdict.short, color: WARN, grow: true }
227      break
228    }
229    case 'failed':
230      badge = { text: ' ✗ Done, checks failed ', color: WARN, inverse: true, bold: true }
231      text = { text: verdict.text, color: WARN, grow: true }
232      break
233    case 'verified':
234      badge = { text: ` ✓ All done · ${verdict.label} `, color: DONE, inverse: true, bold: true }
235      text = { text: title, color: DONE, grow: true }
236      break
237    default:
238      if (tone === 'stopped') {
239        badge = { text: ' ■ Stopped ', color: OVER, inverse: true, bold: true }
240        text = { text: title, dim: true, grow: true }
241      } else if (tone === 'ended') {
242        // Steps left when the turn ended: say how far it got, never "done".
243        // Scar: a green card over three steps that never ran, read as finished
244        const { done, total } = reachedOf(model.plan)
245        badge = { text: ` ■ Turn ended · ${done} of ${total} ${total === 1 ? 'step' : 'steps'} reached `, color: OVER, inverse: true, bold: true }
246        text = { text: title, grow: true }
247      } else {
248        badge = { text: ' ✓ All done ', color: DONE, inverse: true, bold: true }
249        text = { text: title, color: DONE, grow: true }
250      }
251  }
252  return { key: 'title', segs: line(width, [badge, space(1), text], [...tail, ...buttons, { text: took }, space(1), collapse]) }
253}
254
255function summaryRow(model: ViewModel, width: number, tone: BandTone): Row {
256  const plan = model.plan
257  if (!model.isWorking) {
258    const n = plan?.items.length ?? 0
259    const done = plan?.items.filter(item => item.state === 'done').length ?? 0
260    const label = `${done} of ${n} ${n === 1 ? 'step' : 'steps'}` + (plan?.source === 'derived' ? ' · derived' : '')
261    const progress = n === 0 ? 1 : done / n
262    const percent = `${Math.round(progress * 100)}%`.padStart(4)
263    const color = tone === 'done' ? DONE : tone === 'stopped' || tone === 'ended' ? OVER : WARN
264    const barWidth = Math.max(1, width - [...label].length - 2 - percent.length)
265    return {
266      key: 'summary',
267      segs: line(width, [{ text: label }, space(1), ...barSegs(progress, barWidth, color), space(1), { text: percent, color, bold: true }], [], 0),
268    }
269  }
270  const est: Estimate = model.estimate ?? { kind: 'indeterminate' }
271  const isOver = est.kind === 'range' && est.isOver
272  const progress = est.kind === 'range' ? est.progress : 0
273  const label = stepCounterOf(plan)
274  const percent = `${Math.round(progress * 100)}%`.padStart(4)
275  let tail = estimateRow(est)
276  const fixed = [...label].length + 1 + 1 + percent.length
277  let barWidth = width - fixed - 2 - [...tail].length
278  if (barWidth < MIN_BAR) {
279    const tailRoom = width - fixed - 2 - MIN_BAR
280    tail = tailRoom >= MIN_TAIL ? truncate(tail, tailRoom) : ''
281    barWidth = width - fixed - (tail ? 2 + [...tail].length : 0)
282  }
283  const barColor = isOver ? OVER : ACCENT
284  const segs: Seg[] = [
285    { text: label },
286    space(1),
287    ...barSegs(progress, Math.max(1, barWidth), barColor),
288    space(1),
289    isOver ? { text: percent, dim: true } : { text: percent, color: ACCENT, bold: true },
290    ...(tail ? [space(2), { text: tail, dim: true }] : []),
291  ]
292  return { key: 'summary', segs: line(width, segs, [], 0) }
293}
294
295function stateWordOf(item: Milestone, index: number, firstPending: number, isWorking: boolean): Seg {
296  if (item.state === 'done') return { text: 'Done', ...(isWorking ? { dim: true } : {}) }
297  if (!isWorking) return { text: 'Not reached', dim: true }
298  if (item.state === 'current') return { text: 'Working', color: ACCENT, bold: true }
299  return { text: index === firstPending ? 'Next' : 'Later', dim: true }
300}
301
302function stepRows(model: ViewModel, width: number, tone: BandTone): Row[] {
303  const plan = model.plan
304  if (!plan || plan.items.length === 0) return []
305  const items = plan.items
306  const current = currentIndexOf(plan)
307  const firstPending = items.findIndex((item, i) => item.state === 'pending' && i > current)
308  const start = items.length <= MAX_STEP_ROWS ? 0 : Math.max(0, Math.min(current - 1, items.length - MAX_STEP_ROWS))
309  const shown = items.slice(start, start + MAX_STEP_ROWS)
310  const isWide = width + FRAME >= NARROW_BODY
311  const longest = Math.max(...shown.map(item => [...item.label].length))
312  const wordWidth = 'Not reached'.length
313  const labelWidth = isWide
314    ? Math.max(8, Math.min(longest, Math.floor(width * 0.4), width - 2 - 2 - MINI_BAR - 2 - wordWidth))
315    : Math.max(4, Math.min(longest, width - 2 - 2 - wordWidth))
316  const stepProgress = model.estimate?.kind === 'range' ? model.estimate.stepProgress : []
317
318  const rows = shown.map((item, offset): Row => {
319    const i = start + offset
320    const glyph: Seg =
321      item.state === 'done'
322        ? { text: '✓', color: DONE }
323        : item.state === 'current' && model.isWorking
324          ? { text: '●', color: ACCENT }
325          : { text: '○', dim: true }
326    const label = truncate(item.label, labelWidth).padEnd(labelWidth)
327    const labelSeg: Seg =
328      item.state === 'current' && model.isWorking ? { text: label, bold: true } : item.state === 'pending' ? { text: label, dim: true } : { text: label }
329    const word = stateWordOf(item, i, firstPending, model.isWorking)
330    let bar: Seg[] = []
331    if (isWide) {
332      const doneColor = tone === 'working' || tone === 'over' || tone === 'done' ? DONE : OVER
333      bar =
334        item.state === 'done'
335          ? barSegs(1, MINI_BAR, doneColor, false, '█', '░')
336          : item.state === 'current' && model.isWorking
337            ? barSegs(stepProgress[i] ?? 0, MINI_BAR, tone === 'over' ? OVER : ACCENT, false, '█', '░')
338            : barSegs(0, MINI_BAR, OVER, false, '█', '░')
339      bar = [...bar, space(2)]
340    }
341    return { key: `step-${i}`, segs: line(width, [glyph, space(1), labelSeg, space(2), ...bar, word], [], 0) }
342  })
343  if (items.length > MAX_STEP_ROWS) {
344    rows.push({ key: 'more', segs: line(width, [{ text: `+${items.length - MAX_STEP_ROWS} more`, dim: true }], [], 0) })
345  }
346  return rows
347}
348
349function toneOf(model: ViewModel, verdict: Verdict): BandTone {
350  if (model.isWorking) return model.estimate?.kind === 'range' && model.estimate.isOver ? 'over' : 'working'
351  if (model.finished?.isAborted) return 'stopped'
352  if (verdict.kind === 'unverified' || verdict.kind === 'failed') return 'unverified'
353  return isAllDone(model.plan) ? 'done' : 'ended'
354}
355
356const BORDER: Record<BandTone, Color> = {
357  working: ACCENT,
358  over: OVER,
359  done: DONE,
360  ended: OVER,
361  unverified: WARN,
362  stopped: OVER,
363}
364
365/**
366 * The band: title, summary, one row per step (six at most, then `+n more`),
367 * the basis tooltip when asked, and a failing calibration score, which is
368 * never folded away (invariant 3: a bad score is displayed, not hidden).
369 * Collapsed, the title row alone. `columns` is the band's body width.
370 */
371export function bandView(model: ViewModel, columns: number, options: BandOptions): BandView {
372  const width = innerWidthOf(columns)
373  const verdict = model.isWorking ? ({ kind: 'none' } as const) : verdictOf(model.finished?.receipt ?? null)
374  const tone = toneOf(model, verdict)
375  const rows: Row[] = [titleRow(model, width, options, tone, verdict)]
376  if (!options.collapsed) {
377    rows.push(summaryRow(model, width, tone), ...stepRows(model, width, tone))
378    if (options.showBasis) {
379      const basis = model.isWorking && model.estimate ? (basisLines(model.estimate)[0] ?? '') : ''
380      const text = [basis, model.calibration[0] ?? ''].filter(Boolean).join(' · ')
381      rows.push({ key: 'basis', segs: line(width, [{ text, dim: true, grow: true }], [], 0) })
382    }
383    const warning = model.calibration[1]
384    if (warning) rows.push({ key: 'calibration', segs: line(width, [{ text: warning, color: WARN, grow: true }], [], 0) })
385  }
386  return { tone, border: BORDER[tone], rows }
387}
388
389// ---- the settings popover ---------------------------------------------------
390
391export const MODEL_CHOICES = [
392  { value: 'haiku', label: 'Haiku' },
393  { value: 'sonnet', label: 'Sonnet' },
394  { value: 'opus', label: 'Opus' },
395  { value: 'fable', label: 'Fable' },
396] as const
397
398export const EFFORT_CHOICES = [
399  { value: 'low', label: 'Low' },
400  { value: 'medium', label: 'Medium' },
401  { value: 'high', label: 'High' },
402  { value: 'xhigh', label: 'XHigh' },
403  { value: 'max', label: 'Max' },
404] as const
405
406/** The popover's widest, border and padding included. */
407export const TOOLS_COLUMNS = 64
408
409/**
410 * How much of the transcript the mod draws away. `off`: every row as Claude
411 * Code draws it. `clean`: tool rows one dim line, milestones in full.
412 * `quiet`: tool rows nothing, milestones one dim line, and the chrome a turn
413 * scatters (spinner, progress, notices, hand-backs, interim text) folded.
414 */
415export type RowsLevel = 'off' | 'clean' | 'quiet'
416export const ROWS_LEVELS: readonly RowsLevel[] = ['off', 'clean', 'quiet']
417
418export const ROWS_CHOICES = [
419  { value: 'off', label: 'Off' },
420  { value: 'clean', label: 'Clean' },
421  { value: 'quiet', label: 'Quiet' },
422] as const
423
424export type ToolsModel = {
425  /** What `$.session.model()` answered. */
426  model: string
427  /** The effort level last seen on a model request, if any. */
428  effort?: string
429  rows: RowsLevel
430  spike: boolean
431}
432
433/**
434 * The alias a model name answers to: `claude-opus-5-5[1m]` is `opus`.
435 */
436export function aliasOf(model: string): string | undefined {
437  const name = model.toLowerCase()
438  return MODEL_CHOICES.find(choice => name.includes(choice.value))?.value
439}
440
441function spaced(text: string): string {
442  return [...text.toUpperCase()].join(' ')
443}
444
445function chipRow(key: string, header: string, choices: readonly { value: string; label: string }[], current: string | undefined, width: number): Row {
446  const segs: Seg[] = [{ text: spaced(header).padEnd(13), dim: true }]
447  choices.forEach((choice, i) => {
448    if (i > 0) segs.push(space(1))
449    if (choice.value === current) segs.push({ text: ` ${choice.label} `, color: ACCENT, inverse: true, bold: true })
450    else segs.push(space(1), buttonSeg(`${key}-${choice.value}`, choice.label, undefined, false), space(1))
451  })
452  return { key, segs: line(width, segs, [], 0) }
453}
454
455function settingRow(key: string, name: string, help: string, isOn: boolean, width: number): Row {
456  const left: Seg[] = [
457    isOn ? { text: '●', color: DONE } : { text: '○', dim: true },
458    space(1),
459    { text: name.padEnd(19), bold: true },
460    { text: help, dim: true, grow: true },
461  ]
462  return { key: `${key}-row`, segs: line(width, left, [buttonSeg(key, isOn ? 'On' : 'Off', undefined, !isOn)]) }
463}
464
465function ruleOf(prefix: string, width: number): Seg[] {
466  const head = truncate(prefix, width)
467  return [{ text: head, dim: true }, { text: '─'.repeat(Math.max(0, width - [...head].length)), dim: true }]
468}
469
470/**
471 * The settings popover: model and effort chips (the current one inverted and
472 * not pressable), letter-spaced headers, and the mod's three switches.
473 * `width` is the popover's inner width.
474 */
475export function toolsView(model: ToolsModel, width: number): Row[] {
476  const alias = aliasOf(model.model)
477  const effort = model.effort?.toLowerCase()
478  const status = [model.model || 'model unknown', effort ? (EFFORT_CHOICES.find(c => c.value === effort)?.label ?? effort) : null].filter(Boolean).join(' · ')
479  return [
480    {
481      key: 'tools-head',
482      segs: line(
483        width,
484        [{ text: '◆ ', color: ACCENT }, { text: spaced('receipts'), color: ACCENT, bold: true }],
485        [{ text: truncate(status, Math.max(4, width - 26)), dim: true }, space(1), buttonSeg('tools-close', '[-]', undefined, false)],
486      ),
487    },
488    chipRow('model', 'model', MODEL_CHOICES, alias, width),
489    chipRow('effort', 'effort', EFFORT_CHOICES, effort, width),
490    chipRow('rows', 'rows', ROWS_CHOICES, model.rows, width),
491    { key: 'tools-rule', segs: ruleOf(`── ${spaced('settings')} `, width) },
492    settingRow('set-spike', 'Spike', 'sizing subagent, new tasks', model.spike, width),
493  ]
494}
495
496/**
497 * Rows inside a round border, as the terminal draws them, for mocks and
498 * docs. Plain text has no fill, so an inverted ` cell ` shows as `[cell]`,
499 * the same width.
500 */
501export function framedText(rows: readonly Row[], width: number): string[] {
502  const plain = (row: Row) =>
503    row.segs.map(seg => (seg.inverse && /^ .* $/.test(seg.text) ? `[${seg.text.slice(1, -1)}]` : seg.text)).join('')
504  return [`╭${'─'.repeat(width + 2)}╮`, ...rows.map(row => `│ ${plain(row)} │`), `╰${'─'.repeat(width + 2)}╯`]
505}
506
src/host.ts 39 lines
1import type {
2  AgentSpawnArgs,
3  AgentSpawnResult,
4  FsEntry,
5  ModelCompleteRequest,
6  ModelCompleteResult,
7  SessionRepo,
8  Timer,
9} from 'claude-code'
10
11/**
12 * Everything the session logic needs from Claude Code, as plain functions.
13 *
14 * `hooks/register.ts` builds one from `$` in a top-level `hostOf($)`, so
15 * every mods API call stays spelled out where `claude plugin validate` reads
16 * it, and nothing under `src/` touches `$`. A test can hand the session a
17 * fake host with a scripted clock and store.
18 */
19export type Host = {
20  now: () => Promise<number>
21  every: (ms: number, fn: () => void) => Timer
22  after: (ms: number, fn: () => void) => Timer
23  storeGet: (key: string) => Promise<unknown>
24  storeSet: (key: string, value: unknown) => Promise<void>
25  storeDelete: (key: string) => Promise<void>
26  /** Redraws the pane and the band (they read the tick); tool rows are left alone. */
27  redraw: () => void
28  classify: (text: string, labels: readonly string[]) => Promise<string | undefined>
29  complete: (request: ModelCompleteRequest) => Promise<ModelCompleteResult>
30  spawn: (args: AgentSpawnArgs) => Promise<AgentSpawnResult>
31  cwd: () => Promise<string>
32  repo: () => Promise<SessionRepo | null>
33  list: (path: string) => Promise<FsEntry[]>
34  /** The session's agents, as `$.agent.list()` answers. */
35  agents: () => Promise<readonly { id: string; status: string; parentId?: string; spawnedBy?: string }[]>
36  /** The session's usage, as `$.session.usage()` answers: rate-limit windows and cost. */
37  usage: () => Promise<{ rateLimits: readonly { kind: string; percentUsed: number; resetsAt?: string }[]; cost?: { usd: number } }>
38}
39
src/quiet.ts 76 lines
1/**
2 * The quiet level's memory of assistant text: which blocks were interim
3 * (drawn as their first line, dim) and which was the turn's final answer
4 * (redrawn in full when the turn completes). Keyed by the render's
5 * `requestId`, the block's own id, so a redraw finds the same verdict.
6 *
7 * Drawing only (invariant 4): nothing here touches the stored messages, and
8 * a block the mod never saw during a turn (history, a resumed session) is
9 * left to Claude Code.
10 */
11
12/** Ids kept per kind before the oldest are forgotten; a long session stays small. */
13const MAX_IDS = 2_000
14
15export class QuietLog {
16  private readonly turnBlocks = new Map<string, string>()
17  private readonly interim = new Set<string>()
18  private readonly final = new Set<string>()
19
20  /** A block drawn while a turn runs: interim until the turn says otherwise. */
21  seen(id: string, text: string): void {
22    if (this.final.has(id)) return
23    this.turnBlocks.set(id, text)
24    this.interim.add(id)
25    trim(this.interim)
26  }
27
28  /**
29   * The turn ended with `answer`: the blocks whose text is part of it are the
30   * final answer; failing any, the last block drawn. Returns whether any
31   * verdict changed, so the caller knows to redraw.
32   */
33  finish(answer: string): boolean {
34    const blocks = [...this.turnBlocks.entries()]
35    this.turnBlocks.clear()
36    const flat = answer.trim()
37    let finals = blocks.filter(([, text]) => text.trim() !== '' && flat.includes(text.trim())).map(([id]) => id)
38    if (finals.length === 0 && blocks.length > 0) finals = [blocks[blocks.length - 1]![0]]
39    for (const id of finals) {
40      this.interim.delete(id)
41      this.final.add(id)
42    }
43    trim(this.final)
44    return finals.length > 0
45  }
46
47  kindOf(id: string): 'interim' | 'final' | 'unknown' {
48    if (this.final.has(id)) return 'final'
49    if (this.interim.has(id)) return 'interim'
50    return 'unknown'
51  }
52}
53
54function trim(ids: Set<string>): void {
55  while (ids.size > MAX_IDS) {
56    const oldest = ids.values().next().value
57    if (oldest === undefined) return
58    ids.delete(oldest)
59  }
60}
61
62/** The first non-empty line of a block, markdown left as is. */
63export function firstLineOf(text: string): string {
64  return text.split('\n').find(line => line.trim() !== '')?.trim() ?? ''
65}
66
67/**
68 * The one dim line a hand-back or notification row becomes once its turn
69 * ended: `↳ message from @Explore: Found 3 mods`.
70 */
71export function handbackLineOf(text: string, from: string | undefined): string {
72  const head = firstLineOf(text)
73  const who = from ? `message from @${from}` : 'notification'
74  return head ? `↳ ${who}: ${head}` : `↳ ${who}`
75}
76
src/session.ts 830 lines
1/**
2 * One session's receipts: the task under way, the history it learns into,
3 * and the view model the pane and band draw. Hooks call in through a Host,
4 * never `$` (see host.ts).
5 *
6 * Invariant 6, hooks return fast: every model call here starts unawaited
7 * (plan derivation, task type, step-done, the spike) except the claims-done
8 * label on turn.complete, which waits at most CLAIMS_DEADLINE_MS and then
9 * falls back to plain words in the answer. Results are cached per step per
10 * turn, so a long turn costs one classify per model step at most.
11 */
12
13import { estimate, type Estimate } from './estimator'
14import {
15  addTask,
16  bucketSamples,
17  calibrationLines,
18  calibrationOf,
19  emptyHistory,
20  HISTORY_KEY,
21  parseHistory,
22  statsOf,
23  type History,
24  type Stats,
25} from './history'
26import type { Host } from './host'
27import {
28  emptyLedger,
29  exitCodeOf,
30  isVerifyCommand,
31  looksDone,
32  receiptLine,
33  recordEdit,
34  recordVerify,
35  type Ledger,
36  type ToolOutcome,
37} from './ledger'
38import {
39  completeAtEnd,
40  completeCurrent,
41  derivedPlan,
42  editedPathOf,
43  emptyPlan,
44  fallbackPlan,
45  finishPlan,
46  isComplete,
47  isEditTool,
48  isMilestoneCommand,
49  isSubagentTool,
50  onTaskCreate,
51  onTaskUpdate,
52  onTodoWrite,
53  parseDerivedSteps,
54  stepDurations,
55  type Pause,
56  type Plan,
57  type TaskUpdateArgs,
58  type Todo,
59} from './milestones'
60import {
61  countTool,
62  NO_TOOLS,
63  repoHashOf,
64  stepBucketOf,
65  TASK_TYPES,
66  toolMixOf,
67  type Shape,
68  type TaskType,
69  type ToolCounts,
70} from './shape'
71import { parseSpike, SPIKE_AGENT, SPIKE_MIN_STEPS, SPIKE_MODEL, SPIKE_TIMEOUT_MS, spikePromptOf } from './spike'
72import type { ViewModel } from './view'
73
74export const PANE_ID = 'receipts'
75export const PANE_TITLE = 'receipts'
76export const CLAIMS_DEADLINE_MS = 1_500
77export const TICK_MS = 1_000
78export const CLAIMS_LABELS = ['claims-done', 'not-claiming-done'] as const
79export const STEP_LABELS = ['step-done', 'step-not-done'] as const
80/** Task types whose steps are cheap: a spike would cost more than it tells. */
81export const NO_SPIKE_TYPES: readonly string[] = ['research', 'writing', 'chat']
82export const AUDIT_KEY = 'audit'
83const STEPS_DONE_SYSTEM =
84  'You read the final message of a coding assistant and a list of planned steps not yet marked done. ' +
85  'Answer with the numbers of the steps the message shows were completed, comma separated, or "none". No other text.'
86const PLAN_MODEL = 'haiku'
87const PLAN_TIMEOUT_MS = 20_000
88const PLAN_SYSTEM =
89  'You read a request to a coding assistant and its first message, and list the plan it is following ' +
90  'as 2 to 8 steps, one per line, each a short noun phrase of at most 8 words. No other text.'
91
92export type SessionOptions = {
93  spike: boolean
94}
95
96type Snapshot = {
97  low: number
98  high: number
99}
100
101type TaskState = {
102  turnId: string
103  startedAt: number
104  prompt: string
105  cwd: string
106  plan: Plan
107  counts: ToolCounts
108  ledger: Ledger
109  taskType: TaskType
110  /** Settles once the task-type label is in (or failed). */
111  typed: Promise<void>
112  isDeriving: boolean
113  isSpiked: boolean
114  spike?: readonly number[]
115  checkedSteps: Set<number>
116  seenEdits: Set<string>
117  claims?: Promise<boolean | undefined>
118  /** The last range shown while 2 or more steps were left. */
119  beforeFinal?: Snapshot
120  firstRange?: Snapshot
121  /** Permission waits inside the task, subtracted from what is learned. */
122  pauses: Pause[]
123  /** The session's cost when the task started, for the turn's share. */
124  usdAtStart?: number
125}
126
127type Audit = {
128  /** Receipt lines shown under an answer that claimed done. */
129  shown: number
130  /** Of those, the ones the person marked wrong with /receipts wrong. */
131  wrong: number
132}
133
134type StepEvent = {
135  turnId: string
136  index: number
137  agentId?: string
138}
139
140type StepResult = {
141  answer: string
142  toolUses: readonly { name: string }[]
143  stopReason: string | null
144}
145
146type ToolEvent = {
147  tool: string
148  tool_use_id: string
149}
150
151type CompleteEvent = {
152  turnId: string
153  answer: string
154  isAborted: boolean
155  reason: string
156}
157
158export class ReceiptsSession {
159  private history: History = emptyHistory()
160  private stats: Stats = statsOf(this.history)
161  private isLoaded = false
162  private probing: Promise<void> | null = null
163  private repoKey = 'unknown'
164  private hasTests = false
165  private repoEntries: string[] = []
166  private task: TaskState | null = null
167  private last: { plan: Plan; prompt: string; totalMs: number; receipt: string | null; isAborted: boolean } | null = null
168  /** Agents the main loop started that were still running when its turn ended. */
169  private waiting = new Set<string>()
170  /** Shapes already spiked this session: one guess per shape is enough. */
171  private readonly spikedShapes = new Set<string>()
172  private readonly toolStarts = new Map<string, number>()
173  private readonly toolEnds = new Map<string, number>()
174  private readonly toolRuns = new Map<string, number>()
175  private lastTickAt = 0
176  private audit: Audit = { shown: 0, wrong: 0 }
177  private lastReceipt: { turnId: string; isMarked: boolean } | null = null
178  private readonly milestoneIds = new Set<string>()
179  private readonly spikeWaiters = new Map<string, TaskState>()
180  private ticker: { cancel: () => void } | null = null
181  private isPaneOpen = false
182  private isPaneShown = false
183  private showBasis = false
184  private cwd = ''
185  private effort: string | undefined
186
187  constructor(private readonly options: SessionOptions) {}
188
189  async start(host: Host): Promise<void> {
190    await this.load(host)
191    this.probing ??= this.probe(host)
192  }
193
194  async turnStart(host: Host, turnId: string, text: string): Promise<void> {
195    await this.load(host)
196    this.probing ??= this.probe(host)
197    const now = await host.now()
198    this.cwd = await host.cwd().catch(() => this.cwd)
199    const task: TaskState = {
200      turnId,
201      startedAt: now,
202      prompt: text,
203      cwd: this.cwd,
204      plan: emptyPlan(),
205      counts: NO_TOOLS,
206      ledger: emptyLedger(),
207      taskType: 'unknown',
208      typed: Promise.resolve(),
209      isDeriving: false,
210      isSpiked: false,
211      checkedSteps: new Set(),
212      seenEdits: new Set(),
213      pauses: [],
214    }
215    this.task = task
216    this.last = null
217    // /receipts wrong is about the current turn's receipt, never an older one
218    this.lastReceipt = null
219    // The fallback for an agent whose end never reached us: the next prompt
220    this.waiting.clear()
221    if (text.trim()) {
222      task.typed = host
223        .classify(text.slice(0, 4_000), TASK_TYPES)
224        .then(label => {
225          task.taskType = (TASK_TYPES as readonly string[]).includes(label ?? '') ? (label as TaskType) : 'unknown'
226        })
227        .catch(() => undefined)
228    }
229    this.startTicker(host, now)
230    const usage = await host.usage().catch(() => null)
231    if (usage?.cost) task.usdAtStart = usage.cost.usd
232    // No pane opens unasked: the band above the prompt is the surface, and
233    // `/receipts` opens the pane for the long view. Scar: the auto-opened
234    // dock took a third of a fullscreen terminal to show four lines.
235    host.redraw()
236  }
237
238  /**
239   * After each model request: derive a plan if the first step brought no
240   * task tools, ask whether the current derived step is done, and start the
241   * claims-done label as soon as the final answer exists.
242   */
243  stepResult(host: Host, e: StepEvent, result: StepResult): void {
244    const task = this.task
245    if (!task || e.agentId !== undefined || e.turnId !== task.turnId) return
246    const usesTaskTools = result.toolUses.some(use => use.name === 'TaskCreate' || use.name === 'TodoWrite')
247    if (task.plan.source === 'none' && !task.isDeriving && !usesTaskTools) {
248      task.isDeriving = true
249      void this.derive(host, task, result.answer)
250    }
251    const answer = result.answer.trim()
252    if (task.plan.source === 'derived' && answer && !task.checkedSteps.has(e.index)) {
253      task.checkedSteps.add(e.index)
254      void this.checkStep(host, task, answer)
255    }
256    if (result.stopReason === 'end_turn' && answer && task.ledger.edits.length > 0) {
257      task.claims = this.claimsDone(host, answer)
258    }
259  }
260
261  /**
262   * Before a tool runs, decide whether its rows are milestone rows, so they
263   * draw in full while running and their result rows (which carry no input)
264   * can be told by id: the first edit of each file this turn, a verify run,
265   * a commit, a pull request, a subagent.
266   */
267  beforeTool(e: ToolEvent, input: unknown): void {
268    if (this.isMilestoneRow(e.tool_use_id, e.tool, input) && !isEditTool(e.tool)) {
269      this.milestoneIds.add(e.tool_use_id)
270      return
271    }
272    const task = this.task
273    if (!task || !isEditTool(e.tool)) return
274    const path = editedPathOf(e.tool, input)
275    if (path === undefined || task.seenEdits.has(path)) return
276    task.seenEdits.add(path)
277    this.milestoneIds.add(e.tool_use_id)
278  }
279
280  async afterTool(host: Host, e: ToolEvent, input: unknown, outcome: ToolOutcome): Promise<void> {
281    const task = this.task
282    if (!task) return
283    const now = await host.now()
284    if (this.toolStarts.has(e.tool_use_id)) {
285      this.toolEnds.set(e.tool_use_id, now)
286      this.settlePause(e.tool_use_id)
287    }
288    task.counts = countTool(task.counts, e.tool)
289    const fields = (typeof input === 'object' && input !== null ? input : {}) as Record<string, unknown>
290    const result = outcome.result as Record<string, unknown> | undefined
291    const isOk = outcome.deny === undefined && !outcome.isError
292    let isPlanChanged = false
293
294    if (e.tool === 'TaskCreate' && isOk) {
295      const created = result?.task as { id?: unknown } | undefined
296      if (typeof created?.id === 'string' && typeof fields.subject === 'string') {
297        task.plan = onTaskCreate(task.plan, created.id, fields.subject, now)
298        isPlanChanged = true
299      }
300    } else if (e.tool === 'TaskUpdate' && isOk && typeof fields.taskId === 'string') {
301      task.plan = onTaskUpdate(task.plan, fields as unknown as TaskUpdateArgs, now)
302      isPlanChanged = true
303    } else if (e.tool === 'TodoWrite' && isOk && Array.isArray(fields.todos)) {
304      task.plan = onTodoWrite(task.plan, fields.todos as Todo[], now)
305      isPlanChanged = true
306    } else if (isEditTool(e.tool) && isOk) {
307      const path = editedPathOf(e.tool, input)
308      if (path !== undefined) task.ledger = recordEdit(task.ledger, path, now)
309    } else if (e.tool === 'Bash' && outcome.deny === undefined && typeof fields.command === 'string') {
310      const interrupted = (result as { interrupted?: unknown } | undefined)?.interrupted === true
311      if (isVerifyCommand(fields.command) && !interrupted) {
312        const stdout = typeof result?.stdout === 'string' ? result.stdout : ''
313        const stderr = typeof result?.stderr === 'string' ? result.stderr : ''
314        const output = outcome.text ?? [stdout, stderr].join('\n')
315        task.ledger = recordVerify(task.ledger, fields.command, exitCodeOf(outcome), output, now)
316      }
317    }
318
319    if (isPlanChanged) {
320      this.observe(now)
321      void this.maybeSpike(host, task)
322      host.redraw()
323    }
324  }
325
326  /**
327   * The turn ended: the receipt line, if any, then learning. Never blocks or
328   * aborts the turn; at worst it waits CLAIMS_DEADLINE_MS for a label.
329   */
330  async turnComplete(host: Host, e: CompleteEvent): Promise<string | null> {
331    const task = this.task
332    if (!task) return null
333    // Any main-loop turn.complete ends the task the band shows, even one whose
334    // id the band never saw start (a reload, a second copy of the mod). Scar:
335    // a band left in Working after the turn had ended. Only a matching id is
336    // learned from.
337    const isOwn = task.turnId === e.turnId
338    this.task = null
339    this.ticker?.cancel()
340    this.ticker = null
341    const now = await host.now()
342
343    // Both labels run at once, under one deadline: the claims-done label, and
344    // for a derived plan which of its open steps the final answer completed
345    const hasEdits = task.ledger.edits.length > 0 && !e.isAborted
346    const [claims, finishedSteps] = await Promise.all([
347      hasEdits ? this.withDeadline(host, task.claims ?? this.claimsDone(host, e.answer), CLAIMS_DEADLINE_MS) : Promise.resolve(undefined),
348      e.isAborted ? Promise.resolve(undefined) : this.withDeadline(host, this.stepsDoneBy(host, task, e.answer), CLAIMS_DEADLINE_MS),
349    ])
350    let line: string | null = null
351    if (hasEdits) line = receiptLine(task.ledger, claims ?? looksDone(e.answer), now, task.cwd)
352
353    const answered = finishedSteps && finishedSteps.length > 0 ? completeAtEnd(task.plan, finishedSteps, now, task.startedAt) : task.plan
354    const plan = finishPlan(answered.source === 'none' ? fallbackPlan(task.startedAt) : answered, now)
355    const totalMs = Math.max(0, now - task.startedAt)
356    this.last = { plan, prompt: task.prompt, totalMs, receipt: line, isAborted: e.isAborted }
357    let shown = line
358    if (line) {
359      this.audit = { ...this.audit, shown: this.audit.shown + 1 }
360      this.lastReceipt = { turnId: task.turnId, isMarked: false }
361      await host.storeSet(AUDIT_KEY, this.audit).catch(() => undefined)
362      const cost = await this.costOf(host, task)
363      if (cost) shown = `${line} · ${cost}`
364    }
365
366    // A task-tool plan finished only when every item did; a derived or
367    // fallback plan finished when the turn answered.
368    const isFinished = plan.source === 'tasks' ? isComplete(plan) : true
369    if (isOwn && !e.isAborted && e.reason === 'answer' && isFinished) {
370      const range = task.beforeFinal ?? task.firstRange
371      // Learned as work: permission waits come out of the steps and the total.
372      // Calibration is scored on the wall clock, as the range was shown
373      const steps = stepDurations(plan, task.startedAt, task.pauses)
374      const paused = task.pauses.reduce((sum, pause) => sum + pause.ms, 0)
375      this.history = addTask(this.history, {
376        at: now,
377        shape: { ...this.shapeOf(task), steps: stepBucketOf(Math.max(1, plan.items.length)) },
378        steps,
379        totalMs: Math.max(0, totalMs - paused),
380        ...(range ? { inside: totalMs >= range.low && totalMs <= range.high } : {}),
381      })
382      this.stats = statsOf(this.history)
383      await host.storeSet(HISTORY_KEY, this.history).catch(() => undefined)
384    }
385    await this.countWaiting(host)
386    host.redraw()
387    return shown
388  }
389
390  /**
391   * One pass over the final answer for a derived plan: which of the steps not
392   * yet done does it show were completed? Indexes into the plan; none for a
393   * plan of another source, or when nothing is open.
394   */
395  private async stepsDoneBy(host: Host, task: TaskState, answer: string): Promise<number[]> {
396    if (task.plan.source !== 'derived' || !answer.trim()) return []
397    const open = task.plan.items.map((item, i) => ({ item, i })).filter(({ item }) => item.state === 'pending')
398    if (open.length === 0) return []
399    try {
400      const reply = await host.complete({
401        model: PLAN_MODEL,
402        system: STEPS_DONE_SYSTEM,
403        prompt: `Steps not yet marked done:\n${open.map(({ item, i }) => `${i + 1}. ${item.label}`).join('\n')}\n\nThe final message:\n${answer.slice(0, 4_000)}`,
404        maxTokens: 40,
405        timeoutMs: CLAIMS_DEADLINE_MS,
406      })
407      if (!reply.isAnswered) return []
408      const wanted = new Set(open.map(({ i }) => i))
409      return [...reply.text.matchAll(/\d+/g)].map(match => Number(match[0]) - 1).filter(i => wanted.has(i))
410    } catch {
411      return []
412    }
413  }
414
415  /**
416   * What the turn cost, for the receipt line: a plan user's five-hour window
417   * and its reset, an API user's dollars this turn, or nothing when neither
418   * is known. Never a guess.
419   */
420  private async costOf(host: Host, task: TaskState): Promise<string | null> {
421    const usage = await host.usage().catch(() => null)
422    if (!usage) return null
423    const window = usage.rateLimits.find(limit => limit.kind === 'five_hour')
424    if (window) {
425      const resets = window.resetsAt ? new Date(window.resetsAt) : null
426      const at = resets && !Number.isNaN(resets.getTime()) ? `, resets ${String(resets.getHours()).padStart(2, '0')}:${String(resets.getMinutes()).padStart(2, '0')}` : ''
427      return `5h window ${window.percentUsed}% used${at}`
428    }
429    if (usage.rateLimits.length === 0 && usage.cost && task.usdAtStart !== undefined) {
430      const spent = usage.cost.usd - task.usdAtStart
431      if (spent > 0) return `$${spent.toFixed(2)} this turn`
432    }
433    return null
434  }
435
436  /** The person says this turn's receipt was wrong: one mark per receipt. */
437  async markWrong(host: Host): Promise<string> {
438    if (!this.lastReceipt) return 'no receipt this session to mark'
439    if (this.lastReceipt.isMarked) return 'already marked wrong'
440    this.lastReceipt.isMarked = true
441    this.audit = { ...this.audit, wrong: this.audit.wrong + 1 }
442    await host.storeSet(AUDIT_KEY, this.audit).catch(() => undefined)
443    return `marked wrong · ${this.auditLine()}`
444  }
445
446  private auditLine(): string {
447    const { shown, wrong } = this.audit
448    const percent = shown === 0 ? 0 : Math.round((wrong / shown) * 100)
449    return `claims-done false positives: ${wrong} of ${shown} ${shown === 1 ? 'receipt' : 'receipts'} (${percent}%)`
450  }
451
452  /** A tool call began: its start, for the permission wait. */
453  toolStarted(toolUseId: string, at: number): void {
454    this.toolStarts.set(toolUseId, at)
455  }
456
457  /** The tool's own run time, from classic PostToolUse (no prompt or hook time). */
458  toolRan(toolUseId: string, ms: number | undefined): void {
459    if (typeof ms !== 'number' || !Number.isFinite(ms)) return
460    this.toolRuns.set(toolUseId, ms)
461    this.settlePause(toolUseId)
462  }
463
464  /**
465   * Once a call's span and run time are both in: the rest is waiting. Only a
466   * wait of a second or more counts, so hook overhead is never a pause.
467   */
468  private settlePause(toolUseId: string): void {
469    const start = this.toolStarts.get(toolUseId)
470    const end = this.toolEnds.get(toolUseId)
471    const run = this.toolRuns.get(toolUseId)
472    if (start === undefined || end === undefined || run === undefined) return
473    this.toolStarts.delete(toolUseId)
474    this.toolEnds.delete(toolUseId)
475    this.toolRuns.delete(toolUseId)
476    const wait = end - start - run
477    if (wait >= 1_000 && this.task) this.task.pauses.push({ at: end, ms: wait })
478  }
479
480  /**
481   * After a hot reload the module's timers die while a drawing stays up. A
482   * render calls this: when a task is under way and nothing has ticked for
483   * two intervals, the ticker starts again.
484   */
485  ensureTicking(host: Host, now: number): void {
486    if (!this.task || now - this.lastTickAt <= 2 * TICK_MS) return
487    this.startTicker(host, now)
488  }
489
490  private startTicker(host: Host, now: number): void {
491    this.ticker?.cancel()
492    this.lastTickAt = now
493    this.ticker = host.every(TICK_MS, () => {
494      void this.tick(host)
495    })
496  }
497
498  /**
499   * The main loop's agents still going: started by the model or the person
500   * (not by this mod, so not the spike), not finished.
501   */
502  private async countWaiting(host: Host): Promise<void> {
503    const agents = await host.agents().catch(() => [])
504    this.waiting = new Set(
505      agents
506        .filter(agent => agent.parentId === undefined && agent.spawnedBy !== 'receipts')
507        .filter(agent => agent.status === 'pending' || agent.status === 'running' || agent.status === 'waiting')
508        .map(agent => agent.id),
509    )
510  }
511
512  /**
513   * A subagent's turn ended; if it was this mod's spike, read its guess.
514   */
515  subagentComplete(host: Host, agentId: string, answer: string): void {
516    if (this.waiting.delete(agentId)) host.redraw()
517    const task = this.spikeWaiters.get(agentId)
518    if (!task) return
519    this.spikeWaiters.delete(agentId)
520    const guess = parseSpike(answer, task.plan.items.length)
521    if (guess && this.task === task) {
522      task.spike = guess.stepsMs
523      host.redraw()
524    }
525  }
526
527  // ---- views ---------------------------------------------------------------
528
529  viewModel(now: number): ViewModel {
530    const task = this.task
531    const est = task ? this.observe(now) : null
532    return {
533      plan: task ? task.plan : (this.last?.plan ?? null),
534      estimate: est,
535      isWorking: task !== null,
536      showBasis: this.showBasis,
537      calibration: calibrationLines(calibrationOf(this.history)),
538      title: task ? task.prompt : (this.last?.prompt ?? ''),
539      elapsedMs: task ? Math.max(0, now - task.startedAt) : (this.last?.totalMs ?? 0),
540      finished: this.last
541        ? { totalMs: this.last.totalMs, receipt: this.last.receipt, isAborted: this.last.isAborted, waitingAgents: this.waiting.size }
542        : null,
543    }
544  }
545
546  isMilestoneRow(toolUseId: string, tool: string, input: unknown): boolean {
547    if (this.milestoneIds.has(toolUseId)) return true
548    if (isEditTool(tool)) return false
549    if (isSubagentTool(tool)) return true
550    if (tool === 'Bash') {
551      const command = (input as { command?: unknown } | null | undefined)?.command
552      return typeof command === 'string' && isMilestoneCommand(command)
553    }
554    return false
555  }
556
557  get workingDirectory(): string {
558    return this.cwd
559  }
560
561  /**
562   * The band shows while a task runs, and after it as the completion card
563   * until the next prompt, whenever no pane is placed.
564   */
565  isBandWanted(): boolean {
566    return (this.task !== null || this.last !== null) && !this.isPaneShown
567  }
568
569  get isWorking(): boolean {
570    return this.task !== null
571  }
572
573  /** The user's switch for the spike, over the userConfig default. */
574  setSpike(isOn: boolean): void {
575    this.options.spike = isOn
576  }
577
578  get isSpikeOn(): boolean {
579    return this.options.spike
580  }
581
582  /** The effort level the last main-loop model request carried. */
583  noteEffort(effort: string | number | undefined): void {
584    if (typeof effort === 'string') this.effort = effort
585  }
586
587  get effortLevel(): string | undefined {
588    return this.effort
589  }
590
591  paneDrawn(): void {
592    this.isPaneShown = true
593    this.isPaneOpen = true
594  }
595
596  paneClosed(): void {
597    this.isPaneShown = false
598    this.isPaneOpen = false
599  }
600
601  get paneOpen(): boolean {
602    return this.isPaneOpen
603  }
604
605  paneOpened(isPlaced: boolean): void {
606    this.isPaneOpen = true
607    this.isPaneShown = isPlaced
608  }
609
610  toggleBasis(): void {
611    this.showBasis = !this.showBasis
612  }
613
614  get isBasisShown(): boolean {
615    return this.showBasis
616  }
617
618  // ---- commands ------------------------------------------------------------
619
620  statsText(): string {
621    const tasks = this.history.tasks
622    const audit = `${this.auditLine()}; mark a wrong one with /receipts wrong`
623    if (tasks.length === 0) return ['no finished tasks yet; the estimate runs on its prior until a few land', audit].join('\n')
624    const byType = new Map<string, number[]>()
625    for (const task of tasks) {
626      const list = byType.get(task.shape.taskType) ?? []
627      list.push(task.totalMs)
628      byType.set(task.shape.taskType, list)
629    }
630    const rows = [...byType.entries()]
631      .sort((a, b) => b[1].length - a[1].length)
632      .map(([type, list]) => `${type} ${list.length} (median ${minutesOf(medianOf(list))})`)
633    return [
634      `${tasks.length} finished ${tasks.length === 1 ? 'task' : 'tasks'} in history`,
635      ...calibrationLines(calibrationOf(this.history)),
636      `by type: ${rows.join(', ')}`,
637      audit,
638    ].join('\n')
639  }
640
641  async resetHistory(host: Host): Promise<number> {
642    const count = this.history.tasks.length
643    await host.storeDelete(HISTORY_KEY)
644    this.history = emptyHistory()
645    this.stats = statsOf(this.history)
646    host.redraw()
647    return count
648  }
649
650  // ---- internals -----------------------------------------------------------
651
652  private async load(host: Host): Promise<void> {
653    if (this.isLoaded) return
654    this.isLoaded = true
655    const raw = await host.storeGet(HISTORY_KEY).catch(() => undefined)
656    this.history = parseHistory(raw)
657    const audit = (await host.storeGet(AUDIT_KEY).catch(() => undefined)) as Partial<Audit> | undefined
658    if (typeof audit?.shown === 'number' && typeof audit.wrong === 'number') this.audit = { shown: audit.shown, wrong: audit.wrong }
659    this.stats = statsOf(this.history)
660  }
661
662  private async probe(host: Host): Promise<void> {
663    try {
664      const repo = await host.repo().catch(() => null)
665      const root = repo?.root ?? (await host.cwd())
666      this.repoKey = repoHashOf(root)
667      const entries = await host.list(root).catch(() => [])
668      const names = entries.map(entry => entry.name)
669      this.repoEntries = names.slice(0, 40)
670      this.hasTests = names.some(name => /^(tests?|__tests__|specs?)$/.test(name) || /\.(test|spec)\.\w+$/.test(name))
671    } catch {
672      // Unknown repo: the shape still works, it just learns under 'unknown'
673    }
674  }
675
676  private shapeOf(task: TaskState): Shape {
677    return {
678      taskType: task.taskType,
679      steps: stepBucketOf(Math.max(1, task.plan.items.length)),
680      repo: this.repoKey,
681      hasTests: this.hasTests,
682      mix: toolMixOf(task.counts),
683    }
684  }
685
686  private estimateOf(task: TaskState, now: number): Estimate {
687    return estimate({
688      now,
689      startedAt: task.startedAt,
690      plan: task.plan,
691      stats: this.stats,
692      shape: this.shapeOf(task),
693      ...(task.spike ? { spike: task.spike } : {}),
694    })
695  }
696
697  /**
698   * Computes the estimate for now and keeps the ranges calibration is scored
699   * against: the first one shown, and the last one shown before the final step.
700   */
701  private observe(now: number): Estimate | null {
702    const task = this.task
703    if (!task) return null
704    const est = this.estimateOf(task, now)
705    if (est.kind === 'range') {
706      const snapshot = { low: est.totalLowMs, high: est.totalHighMs }
707      task.firstRange ??= snapshot
708      const left = task.plan.items.filter(item => item.state !== 'done').length
709      if (left >= 2) task.beforeFinal = snapshot
710    }
711    return est
712  }
713
714  private async tick(host: Host): Promise<void> {
715    if (!this.task) return
716    const now = await host.now()
717    this.lastTickAt = now
718    this.observe(now)
719    host.redraw()
720  }
721
722  private async derive(host: Host, task: TaskState, firstText: string): Promise<void> {
723    let labels: string[] = []
724    if (task.prompt.trim()) {
725      try {
726        const reply = await host.complete({
727          model: PLAN_MODEL,
728          system: PLAN_SYSTEM,
729          prompt: `The request:\n${task.prompt.slice(0, 4_000)}\n\nThe assistant's first message:\n${firstText.slice(0, 2_000) || '(none yet)'}`,
730          maxTokens: 300,
731          timeoutMs: PLAN_TIMEOUT_MS,
732        })
733        if (reply.isAnswered) labels = parseDerivedSteps(reply.text)
734      } catch {
735        // No plan from the model: fall back to the single milestone
736      }
737    }
738    // Task tools may have arrived meanwhile; they win
739    if (this.task !== task || task.plan.source !== 'none') return
740    task.plan = labels.length > 0 ? derivedPlan(labels, task.startedAt) : fallbackPlan(task.startedAt)
741    this.observe(await host.now())
742    void this.maybeSpike(host, task)
743    host.redraw()
744  }
745
746  private async checkStep(host: Host, task: TaskState, answer: string): Promise<void> {
747    const current = task.plan.items.find(item => item.state === 'current')
748    if (!current) return
749    const label = await host
750      .classify(`Planned step: "${current.label}"\n\nThe assistant's latest message:\n${answer.slice(0, 3_000)}`, STEP_LABELS)
751      .catch(() => undefined)
752    if (label !== 'step-done' || this.task !== task) return
753    if (task.plan.items.find(item => item.state === 'current')?.id !== current.id) return
754    const now = await host.now()
755    task.plan = completeCurrent(task.plan, now)
756    this.observe(now)
757    host.redraw()
758  }
759
760  private claimsDone(host: Host, answer: string): Promise<boolean | undefined> {
761    return host
762      .classify(`The final message of a coding assistant to its user:\n${answer.slice(0, 4_000)}`, CLAIMS_LABELS)
763      .then(label => (label === 'claims-done' ? true : label === 'not-claiming-done' ? false : undefined))
764      .catch(() => undefined)
765  }
766
767  /**
768   * Once per task, for an unfamiliar shape with a plan of SPIKE_MIN_STEPS or
769   * more, and only when the user has not turned the spike off.
770   */
771  private async maybeSpike(host: Host, task: TaskState): Promise<void> {
772    if (!this.options.spike || task.isSpiked) return
773    if (task.plan.items.length < SPIKE_MIN_STEPS) return
774    if (task.plan.source === 'none' || task.plan.source === 'fallback') return
775    task.isSpiked = true
776    await Promise.all([this.probing, task.typed])
777    if (bucketSamples(this.stats, this.shapeOf(task)) >= 3) return
778    // Once per shape per session, and never for cheap task types. Scar: a
779    // cheap subagent per prompt in an unfamiliar repo (the 0.2.0 live run)
780    if (NO_SPIKE_TYPES.includes(task.taskType)) return
781    const shape = this.shapeOf(task)
782    const shapeKey = `${shape.taskType}|${shape.steps}|${shape.repo}`
783    if (this.spikedShapes.has(shapeKey)) return
784    this.spikedShapes.add(shapeKey)
785    try {
786      const spawned = await host.spawn({
787        prompt: spikePromptOf(task.prompt, task.plan.items.map(item => item.label), this.repoEntries),
788        description: 'receipts: size this task',
789        model: SPIKE_MODEL,
790        subagentType: SPIKE_AGENT,
791      })
792      if (spawned.deny !== undefined || spawned.agentId === undefined) return
793      const agentId = spawned.agentId
794      this.spikeWaiters.set(agentId, task)
795      host.after(SPIKE_TIMEOUT_MS, () => {
796        this.spikeWaiters.delete(agentId)
797      })
798    } catch {
799      // A refused or failed spike leaves the prior in place
800    }
801  }
802
803  private withDeadline<T>(host: Host, promise: Promise<T>, ms: number): Promise<T | undefined> {
804    return new Promise(resolve => {
805      const timer = host.after(ms, () => resolve(undefined))
806      promise.then(
807        value => {
808          timer.cancel()
809          resolve(value)
810        },
811        () => {
812          timer.cancel()
813          resolve(undefined)
814        },
815      )
816    })
817  }
818}
819
820function medianOf(values: readonly number[]): number {
821  const sorted = [...values].sort((a, b) => a - b)
822  const mid = Math.floor(sorted.length / 2)
823  return sorted.length % 2 === 1 ? sorted[mid]! : (sorted[mid - 1]! + sorted[mid]!) / 2
824}
825
826function minutesOf(ms: number): string {
827  const seconds = Math.round(ms / 1000)
828  return seconds < 60 ? `${seconds}s` : `${Math.floor(seconds / 60)}m ${seconds % 60}s`
829}
830
src/view.ts 168 lines
1/**
2 * What the pane and the clean-view rows say, as plain text lines, and the
3 * view model the band draws from (the band itself is band.ts).
4 * `hooks/register.ts` turns each line into a Text element; nothing here knows
5 * about elements or `$`.
6 */
7
8import { basisLines, estimateRow, type Estimate } from './estimator'
9import { relativeOf } from './ledger'
10import { glyphOf, type Plan } from './milestones'
11
12export type Line = {
13  key: string
14  text: string
15  dim?: boolean
16  bold?: boolean
17  color?: string
18}
19
20export type ViewModel = {
21  /** The plan being worked, or the last finished one. */
22  plan: Plan | null
23  estimate: Estimate | null
24  isWorking: boolean
25  showBasis: boolean
26  calibration: readonly string[]
27  /** The user's prompt for the task shown, as typed. */
28  title: string
29  /** Time since the task shown started. */
30  elapsedMs: number
31  /** Set once a task finished: how long it took, its receipt line, and whether it was cut short. */
32  finished: { totalMs: number; receipt: string | null; isAborted: boolean; waitingAgents: number } | null
33}
34
35export const BAR_WIDTH = 24
36
37/**
38 * Cut to `columns` code points, with an ellipsis when cut.
39 */
40export function truncate(text: string, columns: number): string {
41  const points = [...text]
42  if (points.length <= columns) return text
43  if (columns <= 1) return points.slice(0, Math.max(0, columns)).join('')
44  return points.slice(0, columns - 1).join('') + '…'
45}
46
47export function barOf(progress: number, width: number): string {
48  const filled = Math.max(0, Math.min(width, Math.round(progress * width)))
49  return '█'.repeat(filled) + '░'.repeat(width - filled)
50}
51
52function headerOf(plan: Plan): Line {
53  switch (plan.source) {
54    case 'derived':
55      return { key: 'plan-header', text: 'plan · derived (the mod\'s guess, not the assistant\'s)', dim: true }
56    case 'fallback':
57      return { key: 'plan-header', text: 'plan · none given', dim: true }
58    case 'none':
59      return { key: 'plan-header', text: 'plan · waiting for the first step', dim: true }
60    default:
61      return { key: 'plan-header', text: 'plan', dim: true }
62  }
63}
64
65function durationText(ms: number): string {
66  const seconds = Math.round(ms / 1000)
67  if (seconds < 60) return `${seconds}s`
68  return `${Math.floor(seconds / 60)}m ${seconds % 60}s`
69}
70
71/**
72 * The pane, top to bottom under its title row: milestones, the progress bar,
73 * the estimate, the basis detail when asked, the calibration line, and the
74 * finished task's receipt.
75 */
76export function paneLines(model: ViewModel): Line[] {
77  const lines: Line[] = []
78  const { plan, estimate } = model
79  if (!plan) {
80    lines.push({ key: 'idle', text: 'no task yet', dim: true })
81  } else {
82    lines.push(headerOf(plan))
83    plan.items.forEach((item, i) => {
84      lines.push({
85        key: `m-${i}`,
86        text: `${glyphOf(item.state)} ${item.label}`,
87        ...(item.state === 'done' ? { dim: true } : {}),
88        ...(item.state === 'current' ? { bold: true } : {}),
89      })
90    })
91  }
92  if (model.isWorking && estimate) {
93    const progress = estimate.kind === 'range' ? estimate.progress : 0
94    const isOver = estimate.kind === 'range' && estimate.isOver
95    lines.push({ key: 'bar', text: `${barOf(progress, BAR_WIDTH)} ${Math.round(progress * 100)}%`, ...(isOver ? { dim: true } : {}) })
96    lines.push({ key: 'estimate', text: estimateRow(estimate), ...(isOver ? { color: 'yellow' } : {}) })
97    if (model.showBasis) {
98      basisLines(estimate).forEach((text, i) => lines.push({ key: `basis-${i}`, text, dim: true }))
99    }
100  }
101  if (model.finished) {
102    lines.push({ key: 'finished', text: `finished in ${durationText(model.finished.totalMs)}`, dim: true })
103    if (model.finished.receipt) {
104      const isBad = model.finished.receipt.startsWith('UNVERIFIED')
105      lines.push({ key: 'receipt', text: model.finished.receipt, ...(isBad ? { color: 'yellow' } : { color: 'green' }) })
106    }
107  }
108  model.calibration.forEach((text, i) =>
109    lines.push({ key: `cal-${i}`, text, ...(i === 0 ? { dim: true } : { color: 'yellow' }) }),
110  )
111  return lines
112}
113
114function fieldOf(input: unknown, name: string): string | undefined {
115  if (typeof input !== 'object' || input === null) return undefined
116  const value = (input as Record<string, unknown>)[name]
117  return typeof value === 'string' ? value : undefined
118}
119
120/**
121 * The one dim line a tool call becomes in clean view: `● Edit src/x.ts`.
122 */
123export function toolRowText(tool: string, input: unknown, cwd: string): string {
124  const path = fieldOf(input, 'file_path') ?? fieldOf(input, 'notebook_path') ?? fieldOf(input, 'path')
125  const arg =
126    (path !== undefined ? relativeOf(path, cwd) : undefined) ??
127    fieldOf(input, 'command') ??
128    fieldOf(input, 'pattern') ??
129    fieldOf(input, 'url') ??
130    fieldOf(input, 'query') ??
131    fieldOf(input, 'description') ??
132    fieldOf(input, 'subject') ??
133    ''
134  const flat = arg.replace(/\s+/g, ' ').trim()
135  return truncate(flat ? `● ${tool} ${flat}` : `● ${tool}`, 160)
136}
137
138function textOf(output: unknown): string | undefined {
139  if (typeof output === 'string') return output
140  if (typeof output !== 'object' || output === null) return undefined
141  const fields = output as { stdout?: unknown; stderr?: unknown; file?: { content?: unknown }; content?: unknown }
142  if (typeof fields.stdout === 'string') return [fields.stdout, typeof fields.stderr === 'string' ? fields.stderr : ''].filter(Boolean).join('\n')
143  if (typeof fields.file?.content === 'string') return fields.file.content
144  if (typeof fields.content === 'string') return fields.content
145  return undefined
146}
147
148/**
149 * The one dim line a tool result becomes in clean view: `⎿ 12 lines`.
150 */
151export function toolResultText(output: unknown, isErrored: boolean): string {
152  if (isErrored) return '⎿ error'
153  const text = textOf(output)
154  if (text === undefined) return '⎿ done'
155  const trimmed = text.replace(/\n+$/, '')
156  if (trimmed.length === 0) return '⎿ no output'
157  const count = trimmed.split('\n').length
158  return `⎿ ${count} ${count === 1 ? 'line' : 'lines'}`
159}
160
161/**
162 * A folded run of reads and searches, as one dim line.
163 */
164export function toolGroupText(calls: readonly { tool: string }[]): string {
165  const tools = [...new Set(calls.map(call => call.tool))]
166  return truncate(`● ${calls.length} ${calls.length === 1 ? 'call' : 'calls'}: ${tools.join(', ')}`, 160)
167}
168
src/estimator.ts 310 lines
1/**
2 * The estimate (SPEC 3): remaining time as a range with its basis named.
3 *
4 * The honesty contract, and the scar behind each rule:
5 * - Never a point, never a frozen countdown (invariant 1; every "ETA" in
6 *   every tool ever). `estimateRow` always prints two different numbers, and
7 *   the over-by branch keeps moving with the clock.
8 * - Indeterminate is time-boxed (invariant 2; "don't cheat and just stay
9 *   indeterminate"): only while there is no plan AND no history, and never
10 *   past the first completed milestone or 90 seconds.
11 * - The basis is always named (invariant 3).
12 * - Countdown digits only at confidence 0.5 or more, and even then as a
13 *   range: `3:52 to 8:40 left`.
14 *
15 * Remaining time = sum over the steps not done of their expected duration,
16 * from the most specific shape bucket with 3 or more samples, else a coarser
17 * one, else the global prior. sigma sums only the steps not yet done, so it
18 * shrinks as steps complete; each step's sigma also shrinks with samples.
19 */
20
21import { emptyBucket, type Bucket, type Stats, bucketAdd } from './history'
22import { hasCompleted, type Plan } from './milestones'
23import { GLOBAL_KEY, levelKeysOf, planKeyOf, planlessKeysOf, type Shape } from './shape'
24
25export const INDETERMINATE_CAP_MS = 90_000
26/** Interval half-width in sigmas: about an 87% band for a normal. */
27export const K = 1.5
28export const MIN_SAMPLES = 3
29export const COUNTDOWN_CONFIDENCE = 0.5
30
31/** The prior when nothing has been learned: a step, and a whole task. */
32const DEFAULT_STEP = { mean: 120_000, sd: 90_000 }
33const DEFAULT_TOTAL = { mean: 300_000, sd: 180_000 }
34/** How much wider the prior is than its own sd. */
35const PRIOR_SPREAD = 1.25
36/** No step is ever known to better than a quarter of its length, or 15s. */
37const SD_FLOOR_RATIO = 0.25
38const SD_FLOOR_MS = 15_000
39/** A step under way always has a tenth of its expected time left. */
40const MIN_LEFT_RATIO = 0.1
41/** Past the range, re-widen by half of what the estimate missed by. */
42const OVER_WIDEN = 0.5
43
44export type EstimateInput = {
45  now: number
46  startedAt: number
47  plan: Plan
48  stats: Stats
49  shape: Shape
50  /** The spike subagent's minutes per step, in ms; one sample at weight 0.5. */
51  spike?: readonly number[]
52}
53
54export type RangeEstimate = {
55  kind: 'range'
56  basis: string
57  remainingLowMs: number
58  remainingHighMs: number
59  totalLowMs: number
60  totalHighMs: number
61  sigmaMs: number
62  confidence: number
63  isOver: boolean
64  overByMs: number
65  /** Fraction done, weighted by expected step durations. */
66  progress: number
67  samples: number
68  /** Each plan step's own fraction done, in plan order: 1 done, 0 not started. */
69  stepProgress: readonly number[]
70}
71
72export type Estimate = { kind: 'indeterminate' } | RangeEstimate
73
74type Basis = {
75  bucket: Bucket | null
76  label: string
77  /** Effective samples for confidence; 0 for any prior. */
78  samples: number
79}
80
81type Expectation = {
82  mean: number
83  sd: number
84}
85
86export function estimate(input: EstimateInput): Estimate {
87  const { now, startedAt, plan, stats } = input
88  const elapsed = Math.max(0, now - startedAt)
89  const isPlanless = plan.source === 'none' || plan.source === 'fallback'
90  const hasHistory = (stats.get(GLOBAL_KEY)?.total.n ?? 0) > 0
91
92  if (isPlanless && !hasHistory && elapsed < INDETERMINATE_CAP_MS && !hasCompleted(plan)) {
93    return { kind: 'indeterminate' }
94  }
95
96  return isPlanless ? totalMode(input, elapsed) : stepMode(input, elapsed)
97}
98
99/**
100 * No plan to walk: the whole task's learned duration against elapsed time.
101 */
102function totalMode(input: EstimateInput, elapsed: number): RangeEstimate {
103  const basis = basisOf(input.stats, planlessKeysOf(input.shape))
104  const exp = basis.bucket ? spreadOf(basis.bucket.total.mean, basis.bucket.total.var, basis.samples) : priorOf(DEFAULT_TOTAL)
105  const remMean = Math.max(exp.mean - elapsed, MIN_LEFT_RATIO * exp.mean)
106  const plannedHigh = exp.mean + K * exp.sd
107  const progress = Math.min(elapsed / exp.mean, 0.95)
108  return rangeOf({ elapsed, remMean, sigma: exp.sd, plannedHigh, progress, basis, suffix: '', stepProgress: input.plan.items.map(() => progress) })
109}
110
111/**
112 * A plan to walk: each step not done contributes its expected duration, the
113 * one under way less what it has used.
114 */
115function stepMode(input: EstimateInput, elapsed: number): RangeEstimate {
116  const { now, startedAt, plan, stats, shape, spike } = input
117  const basis = basisOf(stats, levelKeysOf(shape), spike, planKeyOf(shape))
118  const items = plan.items
119  const stepCount = items.length
120  const hasCurrent = items.some(item => item.state === 'current')
121  const implicit = hasCurrent ? -1 : items.findIndex(item => item.state === 'pending')
122
123  let remMean = 0
124  let remVar = 0
125  let doneWeight = 0
126  let totalWeight = 0
127  let plannedEnd: number | null = null
128  let pendingMean = 0
129  let previousEnd = startedAt
130  const stepProgress: number[] = []
131
132  items.forEach((item, i) => {
133    const exp = basis.bucket ? stepExpectation(basis.bucket, i, stepCount, basis.samples) : priorOf(DEFAULT_STEP)
134    totalWeight += exp.mean
135    if (item.state === 'done') {
136      doneWeight += exp.mean
137      previousEnd = item.doneAt ?? previousEnd
138      stepProgress.push(1)
139      return
140    }
141    remVar += exp.sd * exp.sd
142    if (item.state === 'current' || i === implicit) {
143      const start = item.startedAt ?? previousEnd
144      const inStep = Math.max(0, now - start)
145      remMean += Math.max(exp.mean - inStep, MIN_LEFT_RATIO * exp.mean)
146      doneWeight += Math.min(inStep / exp.mean, 0.95) * exp.mean
147      stepProgress.push(Math.min(inStep / exp.mean, 0.95))
148      plannedEnd = Math.max(plannedEnd ?? -Infinity, start + exp.mean)
149      return
150    }
151    remMean += exp.mean
152    pendingMean += exp.mean
153    stepProgress.push(0)
154  })
155
156  const sigma = Math.sqrt(remVar)
157  const plannedHigh = (plannedEnd ?? now) - startedAt + pendingMean + K * sigma
158  const progress = totalWeight > 0 ? doneWeight / totalWeight : 0
159  const suffix = plan.source === 'derived' ? ' · derived plan' : ''
160  return rangeOf({ elapsed, remMean, sigma, plannedHigh, progress, basis, suffix, stepProgress })
161}
162
163function rangeOf(args: {
164  elapsed: number
165  remMean: number
166  sigma: number
167  plannedHigh: number
168  progress: number
169  basis: Basis
170  suffix: string
171  stepProgress: readonly number[]
172}): RangeEstimate {
173  const { elapsed, remMean, sigma, plannedHigh, progress, basis } = args
174  const isOver = elapsed > plannedHigh
175  const overByMs = isOver ? elapsed - plannedHigh : 0
176  const remainingLowMs = Math.max(0, remMean - K * sigma)
177  const remainingHighMs = remMean + K * sigma + OVER_WIDEN * overByMs
178  const n = basis.samples
179  const cv = sigma / Math.max(remMean, 1)
180  const confidence = n <= 0 ? 0 : (n / (n + MIN_SAMPLES)) * (0.6 + 0.4 * progress) / (1 + cv)
181  return {
182    kind: 'range',
183    basis: basis.label + args.suffix,
184    remainingLowMs,
185    remainingHighMs,
186    totalLowMs: elapsed + remainingLowMs,
187    totalHighMs: elapsed + remainingHighMs,
188    sigmaMs: sigma,
189    confidence: Math.max(0, Math.min(1, confidence)),
190    isOver,
191    overByMs,
192    progress: Math.max(0, Math.min(1, progress)),
193    samples: n,
194    stepProgress: args.stepProgress,
195  }
196}
197
198/**
199 * The most specific bucket with MIN_SAMPLES or more; else, with a spike, the
200 * plan-time bucket plus the spike at weight 0.5; else the global prior.
201 */
202function basisOf(stats: Stats, keys: readonly string[], spike?: readonly number[], spikeKey?: string): Basis {
203  for (const key of keys) {
204    const bucket = stats.get(key)
205    if (bucket && bucket.total.n >= MIN_SAMPLES) {
206      return { bucket, label: `from ${countOf(bucket.total.n)} similar tasks`, samples: bucket.total.n }
207    }
208  }
209  if (spike && spike.length > 0 && spikeKey !== undefined) {
210    const real = stats.get(spikeKey) ?? emptyBucket()
211    const bucket = bucketAdd(real, spike, spike.reduce((a, b) => a + b, 0), 0.5)
212    const n = real.total.n
213    return { bucket, label: n > 0 ? `spike guess + ${countOf(n)} similar` : 'spike guess', samples: bucket.total.n }
214  }
215  const global = stats.get(GLOBAL_KEY)
216  if (global && global.total.n > 0) return { bucket: global, label: 'prior only', samples: 0 }
217  return { bucket: null, label: 'prior only', samples: 0 }
218}
219
220function stepExpectation(bucket: Bucket, index: number, stepCount: number, samples: number): Expectation {
221  const own = bucket.steps[index]
222  if (own && own.n > 0) return spreadOf(own.mean, own.var, samples)
223  if (bucket.pooled.n > 0) return spreadOf(bucket.pooled.mean, bucket.pooled.var, samples)
224  return spreadOf(bucket.total.mean / Math.max(1, stepCount), bucket.total.var / Math.max(1, stepCount), samples)
225}
226
227/**
228 * A learned mean and variance as an expectation: the sd floored, then
229 * widened for how few samples stand behind it (predictive sd, sqrt(1 + 1/n));
230 * a bucket used only as a prior gets the prior's spread.
231 */
232function spreadOf(mean: number, variance: number, samples: number): Expectation {
233  const safeMean = Math.max(mean, 1_000)
234  const sd = Math.max(Math.sqrt(Math.max(variance, 0)), SD_FLOOR_RATIO * safeMean, SD_FLOOR_MS)
235  const widen = samples > 0 ? Math.sqrt(1 + 1 / samples) : PRIOR_SPREAD
236  return { mean: safeMean, sd: sd * widen }
237}
238
239function priorOf(prior: Expectation): Expectation {
240  return { mean: prior.mean, sd: prior.sd * PRIOR_SPREAD }
241}
242
243function countOf(n: number): string {
244  return Number.isInteger(n) ? String(n) : n.toFixed(1)
245}
246
247/**
248 * The estimate row the pane and band print.
249 */
250export function estimateRow(est: Estimate): string {
251  if (est.kind === 'indeterminate') return 'indeterminate'
252  if (est.isOver) {
253    return `over by ${durationOf(est.overByMs)} · ${approxRangeOf(est.remainingLowMs, est.remainingHighMs)} more · ${est.basis}`
254  }
255  const range =
256    est.confidence >= COUNTDOWN_CONFIDENCE
257      ? `${digitsOf(est.remainingLowMs)} to ${digitsOf(Math.max(est.remainingHighMs, est.remainingLowMs + 1_000))} left`
258      : approxRangeOf(est.remainingLowMs, est.remainingHighMs)
259  return `${range} · ${est.basis}`
260}
261
262/**
263 * `~4 to 9 min`, or `~20 to 45 sec` when the top is under a minute. The two
264 * numbers always differ: a range that rounds to one value is widened. The
265 * low end never shows 0: `~0 to 10 min` is a range in name only, so it
266 * floors at one unit (1 min, 5 sec) and the top widens to stay above it.
267 */
268export function approxRangeOf(lowMs: number, highMs: number): string {
269  if (highMs < 60_000) {
270    const low = Math.max(5, Math.floor(lowMs / 5_000) * 5)
271    let high = Math.ceil(highMs / 5_000) * 5
272    if (high <= low) high = low + 5
273    return `~${low} to ${high} sec`
274  }
275  const low = Math.max(1, Math.floor(lowMs / 60_000))
276  let high = Math.ceil(highMs / 60_000)
277  if (high <= low) high = low + 1
278  return `~${low} to ${high} min`
279}
280
281function digitsOf(ms: number): string {
282  const seconds = Math.max(0, Math.round(ms / 1000))
283  return `${Math.floor(seconds / 60)}:${String(seconds % 60).padStart(2, '0')}`
284}
285
286/**
287 * `2m 10s`, `45s`, `1h 5m`.
288 */
289export function durationOf(ms: number): string {
290  const seconds = Math.max(0, Math.round(ms / 1000))
291  if (seconds < 60) return `${seconds}s`
292  const minutes = Math.floor(seconds / 60)
293  if (minutes < 60) return `${minutes}m ${seconds % 60}s`
294  return `${Math.floor(minutes / 60)}h ${minutes % 60}m`
295}
296
297/**
298 * The basis button's detail: where the numbers come from, in plain words.
299 */
300export function basisLines(est: Estimate): string[] {
301  if (est.kind === 'indeterminate') {
302    return ['basis: no plan and no history yet; a range from the prior shows by 90s or the first finished milestone']
303  }
304  return [
305    `basis: ${est.basis}`,
306    `sigma ±${durationOf(est.sigmaMs)} over the steps left · confidence ${est.confidence.toFixed(2)}`,
307    `whole task: ${durationOf(est.totalLowMs)} to ${durationOf(est.totalHighMs)}`,
308  ]
309}
310
src/milestones.ts 271 lines
1/**
2 * Milestones: what the turn is working on, what is finished, what is left
3 * (SPEC 2), and which tool rows count as milestone events.
4 *
5 * Sources, in order of preference: the task tools the assistant called
6 * (TaskCreate / TaskUpdate, or TodoWrite), a plan derived by a model call
7 * (labelled `derived`, invariant 9: the mod never passes its own guess off as
8 * the assistant's plan), or a single fallback milestone, "the task".
9 */
10
11import { isVerifyCommand } from './ledger'
12
13export type MilestoneState = 'done' | 'current' | 'pending'
14
15export type Milestone = {
16  id: string
17  label: string
18  state: MilestoneState
19  startedAt?: number
20  doneAt?: number
21}
22
23export type PlanSource = 'none' | 'tasks' | 'derived' | 'fallback'
24
25export type Plan = {
26  source: PlanSource
27  items: readonly Milestone[]
28}
29
30export type TaskUpdateArgs = {
31  taskId: string
32  status?: 'pending' | 'in_progress' | 'completed' | 'deleted'
33  subject?: string
34}
35
36export type Todo = {
37  content: string
38  status: 'pending' | 'in_progress' | 'completed'
39  activeForm?: string
40}
41
42export const FALLBACK_LABEL = 'the task'
43const MIN_DERIVED = 2
44const MAX_DERIVED = 8
45const MAX_LABEL = 80
46
47export function emptyPlan(): Plan {
48  return { source: 'none', items: [] }
49}
50
51/**
52 * A task tool beats a derived or fallback plan: the first TaskCreate drops
53 * the mod's guess and starts the assistant's own list.
54 */
55export function onTaskCreate(plan: Plan, id: string, subject: string, now: number): Plan {
56  const kept = plan.source === 'tasks' ? plan.items : []
57  if (kept.some(item => item.id === id)) return plan
58  return { source: 'tasks', items: [...kept, { id, label: labelOf(subject), state: 'pending' }] }
59}
60
61export function onTaskUpdate(plan: Plan, args: TaskUpdateArgs, now: number): Plan {
62  const index = plan.items.findIndex(item => item.id === args.taskId)
63  if (plan.source !== 'tasks' || index === -1) return plan
64  if (args.status === 'deleted') {
65    return { ...plan, items: plan.items.filter((_, i) => i !== index) }
66  }
67  const items = plan.items.map((item, i): Milestone => {
68    if (i !== index) return item
69    const label = args.subject === undefined ? item.label : labelOf(args.subject)
70    switch (args.status) {
71      case 'in_progress':
72        return { ...item, label, state: 'current', startedAt: item.startedAt ?? now, doneAt: undefined }
73      case 'completed':
74        return { ...item, label, state: 'done', startedAt: item.startedAt ?? startOf(plan.items, i, now), doneAt: now }
75      case 'pending':
76        return { ...item, label, state: 'pending', doneAt: undefined }
77      default:
78        return { ...item, label }
79    }
80  })
81  return { ...plan, items }
82}
83
84/**
85 * TodoWrite sends the whole list each time: replace it, keeping the times of
86 * items whose text did not change.
87 */
88export function onTodoWrite(plan: Plan, todos: readonly Todo[], now: number): Plan {
89  const before = new Map(plan.items.map(item => [item.label, item]))
90  const items = todos.map((todo, i): Milestone => {
91    const label = labelOf(todo.content)
92    const was = before.get(label)
93    const state: MilestoneState = todo.status === 'completed' ? 'done' : todo.status === 'in_progress' ? 'current' : 'pending'
94    const startedAt = state === 'pending' ? was?.startedAt : (was?.startedAt ?? now)
95    const doneAt = state === 'done' ? (was?.doneAt ?? now) : undefined
96    return { id: `todo-${i}`, label, state, startedAt, doneAt }
97  })
98  return { source: 'tasks', items }
99}
100
101export function derivedPlan(labels: readonly string[], now: number): Plan {
102  return {
103    source: 'derived',
104    items: labels.map((label, i) => ({
105      id: `derived-${i}`,
106      label: labelOf(label),
107      state: i === 0 ? 'current' : 'pending',
108      ...(i === 0 ? { startedAt: now } : {}),
109    })),
110  }
111}
112
113export function fallbackPlan(now: number): Plan {
114  return { source: 'fallback', items: [{ id: 'task', label: FALLBACK_LABEL, state: 'current', startedAt: now }] }
115}
116
117/**
118 * Marks the first current milestone done and starts the next pending one; a
119 * derived step the classifier called done, or the fallback at turn end.
120 */
121export function completeCurrent(plan: Plan, now: number): Plan {
122  let index = plan.items.findIndex(item => item.state === 'current')
123  if (index === -1) index = plan.items.findIndex(item => item.state === 'pending')
124  if (index === -1) return plan
125  const items = plan.items.map((item, i): Milestone => {
126    if (i === index) return { ...item, state: 'done', startedAt: item.startedAt ?? startOf(plan.items, i, now), doneAt: now }
127    if (i === index + 1 && item.state === 'pending') return { ...item, state: 'current', startedAt: now }
128    return item
129  })
130  return { ...plan, items }
131}
132
133/**
134 * The turn ended: whatever is under way is done now. Steps never started stay
135 * pending and are not learned from.
136 */
137export function finishPlan(plan: Plan, now: number): Plan {
138  const items = plan.items.map((item, i): Milestone =>
139    item.state === 'current' ? { ...item, state: 'done', startedAt: item.startedAt ?? startOf(plan.items, i, now), doneAt: now } : item,
140  )
141  return { ...plan, items }
142}
143
144/**
145 * Whether the plan finished: every milestone done.
146 */
147export function isComplete(plan: Plan): boolean {
148  return plan.items.length > 0 && plan.items.every(item => item.state === 'done')
149}
150
151export function hasCompleted(plan: Plan): boolean {
152  return plan.items.some(item => item.state === 'done')
153}
154
155/**
156 * How long each finished step took, in plan order, measured from its own
157 * start or, failing that, from the previous step's end or the task's start.
158 */
159export function stepDurations(plan: Plan, startedAt: number, pauses: readonly Pause[] = []): number[] {
160  const durations: number[] = []
161  let previousEnd = startedAt
162  for (const item of plan.items) {
163    if (item.state !== 'done' || item.doneAt === undefined) continue
164    const start = item.startedAt ?? previousEnd
165    const doneAt = item.doneAt
166    const paused = pauses.filter(pause => pause.at > start && pause.at <= doneAt).reduce((sum, pause) => sum + pause.ms, 0)
167    durations.push(Math.max(0, doneAt - start - paused))
168    previousEnd = doneAt
169  }
170  return durations
171}
172
173/**
174 * Time inside a step that was not work: a permission prompt waiting on the
175 * person, measured as a tool call's span less its run time. `at` is when the
176 * call ended, which places it in its step.
177 */
178export type Pause = {
179  at: number
180  ms: number
181}
182
183/**
184 * The turn ended and a pass over the final answer said it completed the
185 * steps at `indexes` too. They and the step under way share the time since
186 * the last finished step in equal slices: the mod knows they finished, not
187 * when, and an even split teaches less wrong than zero-length steps.
188 */
189export function completeAtEnd(plan: Plan, indexes: readonly number[], now: number, startedAt: number): Plan {
190  const current = plan.items.findIndex(item => item.state === 'current')
191  const chosen = indexes.filter(i => plan.items[i] !== undefined && plan.items[i]!.state === 'pending')
192  if (chosen.length === 0) return plan
193  const span = [...new Set([...(current === -1 ? [] : [current]), ...chosen])].sort((a, b) => a - b)
194  const lastDone = plan.items.reduce((at, item) => (item.doneAt !== undefined && item.doneAt > at ? item.doneAt : at), startedAt)
195  const from = Math.min(current === -1 ? lastDone : (plan.items[current]!.startedAt ?? lastDone), now)
196  const slice = (now - from) / span.length
197  const items = plan.items.map((item, i): Milestone => {
198    const k = span.indexOf(i)
199    if (k === -1) return item
200    return { ...item, state: 'done', startedAt: from + k * slice, doneAt: from + (k + 1) * slice }
201  })
202  return { ...plan, items }
203}
204
205/**
206 * Lines of a model's plan reply, numbering and bullets stripped; 2..8 steps
207 * or none at all.
208 */
209export function parseDerivedSteps(text: string): string[] {
210  const steps = text
211    .split('\n')
212    .map(line => line.replace(/^\s*(?:\d+\s*[.):]|[-*•])\s*/, '').trim())
213    .filter(line => line.length > 0)
214    .slice(0, MAX_DERIVED)
215    .map(labelOf)
216  return steps.length >= MIN_DERIVED ? steps : []
217}
218
219export function glyphOf(state: MilestoneState): string {
220  return state === 'done' ? '✓' : state === 'current' ? '▸' : '·'
221}
222
223function labelOf(text: string): string {
224  const flat = text.replace(/\s+/g, ' ').trim()
225  return flat.length > MAX_LABEL ? flat.slice(0, MAX_LABEL - 1) + '…' : flat
226}
227
228function startOf(items: readonly Milestone[], index: number, now: number): number {
229  for (let i = index - 1; i >= 0; i -= 1) {
230    const doneAt = items[i]?.doneAt
231    if (doneAt !== undefined) return doneAt
232  }
233  return now
234}
235
236export const EDIT_TOOLS = ['Edit', 'Write', 'MultiEdit', 'NotebookEdit'] as const
237
238export function isEditTool(tool: string): boolean {
239  return (EDIT_TOOLS as readonly string[]).includes(tool)
240}
241
242/**
243 * The file an edit tool call changes, or undefined for any other call.
244 */
245export function editedPathOf(tool: string, input: unknown): string | undefined {
246  if (!isEditTool(tool) || typeof input !== 'object' || input === null) return undefined
247  const fields = input as { file_path?: unknown; notebook_path?: unknown }
248  const path = fields.file_path ?? fields.notebook_path
249  return typeof path === 'string' ? path : undefined
250}
251
252export function isCommitCommand(command: string): boolean {
253  return /\bgit\s+(?:-\S+\s+)*commit\b/.test(command)
254}
255
256export function isPrCreateCommand(command: string): boolean {
257  return /\bgh\s+pr\s+create\b/.test(command)
258}
259
260/**
261 * A shell command whose row always draws in full: a verify run (its exit code
262 * and test counts), a commit, or a pull request (invariant 5).
263 */
264export function isMilestoneCommand(command: string): boolean {
265  return isVerifyCommand(command) || isCommitCommand(command) || isPrCreateCommand(command)
266}
267
268export function isSubagentTool(tool: string): boolean {
269  return tool === 'Agent' || tool === 'Task'
270}
271
src/history.ts 206 lines
1/**
2 * The learning store: finished tasks, the statistics replayed from them, and
3 * the mod's own calibration score (SPEC 3, "History" and "Learning").
4 *
5 * Only finished task records are persisted. EMA mean and variance per shape
6 * key and step index are rebuilt by replaying those records oldest first, so
7 * pruning the oldest tasks also ages them out of the statistics; there is no
8 * second copy of the numbers to drift from the records.
9 *
10 * Invariant 8, the store stays under 1 MiB: `addTask` prunes to the newest
11 * MAX_TASKS and then drops the oldest until the JSON fits MAX_BYTES.
12 */
13
14import { GLOBAL_KEY, levelKeysOf, planKeyOf, type Shape } from './shape'
15
16export const HISTORY_KEY = 'history'
17export const MAX_TASKS = 500
18/** Under 1 MiB (1,048,576) with room for the store's own framing. */
19export const MAX_BYTES = 1_000_000
20
21export type TaskRecord = {
22  /** When the task ended, ms. */
23  at: number
24  shape: Shape
25  /** How long each step that ran took, in order, ms. */
26  steps: number[]
27  totalMs: number
28  /**
29   * Whether the total landed inside the last range shown before the final
30   * step; absent when no range was ever shown.
31   */
32  inside?: boolean
33}
34
35export type History = {
36  v: 1
37  tasks: TaskRecord[]
38}
39
40export type Ema = {
41  /** Sum of sample weights. A spike counts 0.5. */
42  n: number
43  mean: number
44  var: number
45}
46
47export type Bucket = {
48  total: Ema
49  /** Per step index. */
50  steps: Ema[]
51  /** Every step duration, any index. */
52  pooled: Ema
53}
54
55export type Stats = ReadonlyMap<string, Bucket>
56
57export type Calibration = {
58  inside: number
59  total: number
60}
61
62/** The EMA's floor rate: early samples average, later ones decay at 10%. */
63const ALPHA = 0.1
64
65export const EMPTY_EMA: Ema = { n: 0, mean: 0, var: 0 }
66
67export function emptyHistory(): History {
68  return { v: 1, tasks: [] }
69}
70
71/**
72 * Adds one sample at weight `w`. While few samples are in, this is the plain
73 * running mean and population variance; past 1/ALPHA samples it decays.
74 */
75export function emaAdd(ema: Ema, x: number, w = 1): Ema {
76  const n = ema.n + w
77  const a = Math.min(1, Math.max(w / n, ALPHA * w))
78  const d = x - ema.mean
79  return { n, mean: ema.mean + a * d, var: (1 - a) * (ema.var + a * d * d) }
80}
81
82export function emptyBucket(): Bucket {
83  return { total: EMPTY_EMA, steps: [], pooled: EMPTY_EMA }
84}
85
86/**
87 * Folds one task into a bucket at a weight.
88 */
89export function bucketAdd(bucket: Bucket, steps: readonly number[], totalMs: number, w = 1): Bucket {
90  const next: Ema[] = [...bucket.steps]
91  let pooled = bucket.pooled
92  steps.forEach((ms, i) => {
93    next[i] = emaAdd(next[i] ?? EMPTY_EMA, ms, w)
94    pooled = emaAdd(pooled, ms, w)
95  })
96  return { total: emaAdd(bucket.total, totalMs, w), steps: next, pooled }
97}
98
99/**
100 * Replays every retained task, oldest first, into a bucket per shape level
101 * and one global bucket.
102 */
103export function statsOf(history: History): Stats {
104  const stats = new Map<string, Bucket>()
105  for (const task of history.tasks) {
106    for (const key of [...levelKeysOf(task.shape), GLOBAL_KEY]) {
107      stats.set(key, bucketAdd(stats.get(key) ?? emptyBucket(), task.steps, task.totalMs))
108    }
109  }
110  return stats
111}
112
113/**
114 * Sample weight behind the plan-time shape (every field but tool mix), with
115 * a spike counted at 0.5. The spike gate fires below 3.
116 */
117export function bucketSamples(stats: Stats, shape: Shape, spike?: readonly number[]): number {
118  const real = stats.get(planKeyOf(shape))?.total.n ?? 0
119  return real + (spike && spike.length > 0 ? 0.5 : 0)
120}
121
122export function sizeOf(history: History): number {
123  return new TextEncoder().encode(JSON.stringify(history)).length
124}
125
126/**
127 * The newest MAX_TASKS tasks, then the oldest dropped until the JSON fits.
128 */
129export function prune(history: History, maxBytes = MAX_BYTES): History {
130  let tasks = history.tasks.slice(-MAX_TASKS)
131  let size = sizeOf({ v: 1, tasks })
132  while (size > maxBytes && tasks.length > 0) {
133    const perTask = size / tasks.length
134    const drop = Math.max(1, Math.ceil((size - maxBytes) / perTask))
135    tasks = tasks.slice(drop)
136    size = sizeOf({ v: 1, tasks })
137  }
138  return { v: 1, tasks }
139}
140
141export function addTask(history: History, task: TaskRecord): History {
142  return prune({ v: 1, tasks: [...history.tasks, task] })
143}
144
145export function calibrationOf(history: History): Calibration {
146  let inside = 0
147  let total = 0
148  for (const task of history.tasks) {
149    if (task.inside === undefined) continue
150    total += 1
151    if (task.inside) inside += 1
152  }
153  return { inside, total }
154}
155
156/**
157 * The calibration line, always shown, and a plain-words warning once the mod
158 * has missed more often than not over 10 or more tasks (invariant 3).
159 */
160export function calibrationLines(calibration: Calibration): string[] {
161  if (calibration.total === 0) return ['calibration: no finished tasks yet']
162  const percent = Math.round((calibration.inside / calibration.total) * 100)
163  const noun = calibration.total === 1 ? 'task' : 'tasks'
164  const lines = [`calibration: ${percent}% of ${calibration.total} ${noun} ended inside the range`]
165  if (calibration.total >= 10 && percent < 50) {
166    lines.push('these ranges have missed more often than not; treat them as rough')
167  }
168  return lines
169}
170
171function isRecord(value: unknown): value is Record<string, unknown> {
172  return typeof value === 'object' && value !== null && !Array.isArray(value)
173}
174
175function isShape(value: unknown): value is Shape {
176  return (
177    isRecord(value) &&
178    typeof value.taskType === 'string' &&
179    typeof value.steps === 'string' &&
180    typeof value.repo === 'string' &&
181    typeof value.hasTests === 'boolean' &&
182    typeof value.mix === 'string'
183  )
184}
185
186function isTask(value: unknown): value is TaskRecord {
187  return (
188    isRecord(value) &&
189    typeof value.at === 'number' &&
190    isShape(value.shape) &&
191    Array.isArray(value.steps) &&
192    value.steps.every(ms => typeof ms === 'number' && Number.isFinite(ms)) &&
193    typeof value.totalMs === 'number' &&
194    (value.inside === undefined || typeof value.inside === 'boolean')
195  )
196}
197
198/**
199 * Reads what the store holds, keeping only well-formed records: a corrupt or
200 * foreign value is an empty history, never a crash in a hook.
201 */
202export function parseHistory(raw: unknown): History {
203  if (!isRecord(raw) || !Array.isArray(raw.tasks)) return emptyHistory()
204  return { v: 1, tasks: raw.tasks.filter(isTask) }
205}
206
src/ledger.ts 153 lines
1/**
2 * The per-turn ledger behind the receipt (SPEC 4): every edit, every verify
3 * run with its exit code, and the one line printed under the answer.
4 *
5 * Order is by sequence number, not clock: "a verify after the last edit"
6 * must hold even when two calls land in the same millisecond.
7 *
8 * This is a mirror, not a gate, in v1: nothing here blocks or aborts a turn.
9 */
10
11/**
12 * SPEC 4's verify list, verbatim, with word boundaries added so `latest` or
13 * `contest` do not read as `test`. `claude plugin test` matches on `test`.
14 */
15export const VERIFY_PATTERN =
16  /(?<![\w-])(test|spec|jest|vitest|pytest|cargo (test|check|clippy)|go test|bun test|npm (test|run (test|build|lint))|pnpm|tsc|eslint|ruff|mypy|make (test|check)|build)(?![\w-])/
17
18export type TestCounts = {
19  pass?: number
20  fail?: number
21}
22
23export type EditEntry = {
24  seq: number
25  at: number
26  path: string
27}
28
29export type VerifyEntry = {
30  seq: number
31  at: number
32  command: string
33  exitCode: number
34  counts?: TestCounts
35}
36
37export type Ledger = {
38  seq: number
39  edits: readonly EditEntry[]
40  verifies: readonly VerifyEntry[]
41}
42
43/**
44 * What a tool call resolved to, as far as the ledger reads it: Bash carries
45 * no exit code field, so a failed run is `isError` with `Exit code N` text.
46 */
47export type ToolOutcome = {
48  result?: unknown
49  isError?: boolean
50  text?: string
51  deny?: string
52}
53
54export function emptyLedger(): Ledger {
55  return { seq: 0, edits: [], verifies: [] }
56}
57
58export function isVerifyCommand(command: string): boolean {
59  return VERIFY_PATTERN.test(command)
60}
61
62export function recordEdit(ledger: Ledger, path: string, at: number): Ledger {
63  const seq = ledger.seq + 1
64  return { ...ledger, seq, edits: [...ledger.edits, { seq, at, path }] }
65}
66
67export function recordVerify(ledger: Ledger, command: string, exitCode: number, output: string, at: number): Ledger {
68  const seq = ledger.seq + 1
69  const counts = testCountsOf(output)
70  const entry: VerifyEntry = { seq, at, command, exitCode, ...(counts ? { counts } : {}) }
71  return { ...ledger, seq, verifies: [...ledger.verifies, entry] }
72}
73
74export function exitCodeOf(outcome: ToolOutcome): number {
75  if (!outcome.isError) return 0
76  const match = /Exit code (\d+)/.exec(outcome.text ?? '')
77  return match ? Number(match[1]) : 1
78}
79
80/**
81 * Pass and fail counts from a test runner's output: bun, jest, vitest,
82 * pytest, cargo and go all print `N pass(ed)` and `N fail(ed)` somewhere.
83 */
84export function testCountsOf(text: string): TestCounts | undefined {
85  const pass = /(\d+) (?:pass|passed|passing)\b/.exec(text)
86  const fail = /(\d+) (?:fail|failed|failing)\b/.exec(text)
87  if (!pass && !fail) return undefined
88  return {
89    ...(pass ? { pass: Number(pass[1]) } : {}),
90    ...(fail ? { fail: Number(fail[1]) } : {}),
91  }
92}
93
94/**
95 * The receipt line, or null when there is nothing to say: no edits this
96 * turn, or the answer does not claim done.
97 */
98export function receiptLine(ledger: Ledger, claimsDone: boolean, now: number, cwd: string): string | null {
99  if (ledger.edits.length === 0 || !claimsDone) return null
100  const lastEdit = ledger.edits.reduce((a, b) => (b.seq > a.seq ? b : a))
101  const after = ledger.verifies.filter(run => run.seq > lastEdit.seq)
102  const verify = after[after.length - 1]
103  if (!verify) {
104    return `UNVERIFIED · claimed done, no test/build/run after the last edit (${relativeOf(lastEdit.path, cwd)} at ${clockOf(lastEdit.at)})`
105  }
106  return `receipt · ${commandLabelOf(verify.command)} ${outcomeOf(verify)} · ${agoOf(now - verify.at)}`
107}
108
109function outcomeOf(run: VerifyEntry): string {
110  const pass = run.counts?.pass
111  const fail = run.counts?.fail
112  if (run.exitCode === 0) {
113    return '✓' + (pass !== undefined ? ` ${pass} pass` : '') + (fail ? ` · ${fail} fail` : '')
114  }
115  return `✗ exit ${run.exitCode}` + (pass !== undefined ? ` · ${pass} pass` : '') + (fail ? ` · ${fail} fail` : '')
116}
117
118/**
119 * The part of a compound command that matched, as the user would name it.
120 */
121export function commandLabelOf(command: string): string {
122  const parts = command.split(/&&|\|\||;|\|/).map(part => part.trim())
123  const label = parts.find(part => isVerifyCommand(part)) ?? command.trim()
124  return label.length > 40 ? label.slice(0, 39) + '…' : label
125}
126
127export function relativeOf(path: string, cwd: string): string {
128  const base = cwd.endsWith('/') ? cwd : cwd + '/'
129  return path.startsWith(base) ? path.slice(base.length) : path
130}
131
132function clockOf(ms: number): string {
133  const date = new Date(ms)
134  return `${String(date.getHours()).padStart(2, '0')}:${String(date.getMinutes()).padStart(2, '0')}`
135}
136
137export function agoOf(ms: number): string {
138  const seconds = Math.max(0, Math.floor(ms / 1000))
139  if (seconds < 60) return `${seconds}s ago`
140  const minutes = Math.floor(seconds / 60)
141  if (minutes < 60) return `${minutes}m ago`
142  return `${Math.floor(minutes / 60)}h ago`
143}
144
145/**
146 * The fallback when the claims-done label is not back in time: plain words
147 * of completion in the answer. Used only when the model call misses its
148 * deadline or fails.
149 */
150export function looksDone(answer: string): boolean {
151  return /\b(done|complete[ds]?|finished|fixed|implemented|all (?:\w+ )?(?:tests? )?pass(?:es|ing)?|ready|shipped)\b/i.test(answer)
152}
153
src/shape.ts 116 lines
1/**
2 * Task shape: the key the estimator learns under (SPEC 3).
3 *
4 * `{ task_type, step_count_bucket, repo, has_tests, tool_mix_bucket }`, plus
5 * the ladder of coarser keys the estimator falls back through when the exact
6 * shape has too few samples. No `$` here: plain data in, plain data out.
7 */
8
9export const TASK_TYPES = ['build', 'debug', 'research', 'writing', 'config', 'refactor', 'chat'] as const
10
11export type TaskType = (typeof TASK_TYPES)[number] | 'unknown'
12
13export type StepBucket = '1' | '2-3' | '4-6' | '7+'
14
15/**
16 * What kind of tools dominated the task. Known only as the task runs, so the
17 * plan-time lookups skip it (see `planKeyOf`).
18 */
19export type ToolMix = 'none' | 'edit' | 'read' | 'shell' | 'mixed'
20
21export type Shape = {
22  taskType: TaskType
23  steps: StepBucket
24  /** A short hash of the repository root or working directory, never the path. */
25  repo: string
26  hasTests: boolean
27  mix: ToolMix
28}
29
30export type ToolCounts = {
31  edit: number
32  read: number
33  shell: number
34  other: number
35}
36
37export const NO_TOOLS: ToolCounts = { edit: 0, read: 0, shell: 0, other: 0 }
38
39/**
40 * The key every task shares: the global prior's bucket.
41 */
42export const GLOBAL_KEY = '*'
43
44export function stepBucketOf(count: number): StepBucket {
45  if (count <= 1) return '1'
46  if (count <= 3) return '2-3'
47  if (count <= 6) return '4-6'
48  return '7+'
49}
50
51/**
52 * The dominant kind of tool, when one kind is at least 60% of the calls.
53 */
54export function toolMixOf(counts: ToolCounts): ToolMix {
55  const total = counts.edit + counts.read + counts.shell + counts.other
56  if (total === 0) return 'none'
57  const kinds: [ToolMix, number][] = [
58    ['edit', counts.edit],
59    ['read', counts.read],
60    ['shell', counts.shell],
61  ]
62  for (const [kind, count] of kinds) {
63    if (count / total >= 0.6) return kind
64  }
65  return 'mixed'
66}
67
68export function countTool(counts: ToolCounts, tool: string): ToolCounts {
69  if (/^(Edit|Write|MultiEdit|NotebookEdit)$/.test(tool)) return { ...counts, edit: counts.edit + 1 }
70  if (/^(Read|Grep|Glob|LS|WebFetch|WebSearch)$/.test(tool)) return { ...counts, read: counts.read + 1 }
71  if (tool === 'Bash') return { ...counts, shell: counts.shell + 1 }
72  return { ...counts, other: counts.other + 1 }
73}
74
75/**
76 * FNV-1a over the text, as 8 hex digits: enough to tell repos apart in one
77 * person's history without writing their paths into the store.
78 */
79export function repoHashOf(text: string): string {
80  let hash = 0x811c9dc5
81  for (let i = 0; i < text.length; i += 1) {
82    hash ^= text.charCodeAt(i)
83    hash = Math.imul(hash, 0x01000193) >>> 0
84  }
85  return hash.toString(16).padStart(8, '0')
86}
87
88/**
89 * The shape's keys from most to least specific, the global key not included:
90 * full shape, shape without tool mix, type and steps, type alone.
91 */
92export function levelKeysOf(shape: Shape): string[] {
93  const tests = shape.hasTests ? 't' : 'n'
94  return [
95    `${shape.taskType}|${shape.steps}|${shape.repo}|${tests}|${shape.mix}`,
96    planKeyOf(shape),
97    `${shape.taskType}|${shape.steps}`,
98    `${shape.taskType}`,
99  ]
100}
101
102/**
103 * The most specific key known when the plan is: every field but the tool
104 * mix, which only the finished task can say. The spike gate reads this one.
105 */
106export function planKeyOf(shape: Shape): string {
107  return `${shape.taskType}|${shape.steps}|${shape.repo}|${shape.hasTests ? 't' : 'n'}`
108}
109
110/**
111 * The keys that do not depend on a step count, for a task with no plan yet.
112 */
113export function planlessKeysOf(shape: Shape): string[] {
114  return [`${shape.taskType}`]
115}
116
src/spike.ts 63 lines
1/**
2 * The spike (SPEC 3): when a task's shape is unfamiliar (under 3 samples) and
3 * its plan has 3 or more steps, ask one cheap, read-only subagent for minutes
4 * per step, once per task, capped at 60 seconds. Its answer counts as one
5 * sample at weight 0.5, and the pane says "spike guess" while it is the only
6 * basis.
7 */
8
9export const SPIKE_TIMEOUT_MS = 60_000
10export const SPIKE_MIN_STEPS = 3
11export const SPIKE_MODEL = 'haiku'
12/** Read-only, so a sizing question can never edit the user's files. */
13export const SPIKE_AGENT = 'Explore'
14
15const MAX_MINUTES = 240
16const MIN_MINUTES = 0.1
17
18export function spikePromptOf(request: string, steps: readonly string[], repoEntries: readonly string[]): string {
19  return [
20    'Estimate how long a coding assistant will take for each step of this plan.',
21    'Do not change any file. Do not run long commands. Answer within a minute.',
22    '',
23    'The request:',
24    request.slice(0, 2_000),
25    '',
26    'The plan:',
27    ...steps.map((step, i) => `${i + 1}. ${step}`),
28    '',
29    'Top level of the repository:',
30    repoEntries.length > 0 ? repoEntries.join(', ') : '(unknown)',
31    '',
32    `Reply with JSON only, one number of minutes per step (${steps.length} numbers) and your confidence from 1 to 5:`,
33    '{"minutes": [2, 5, 3], "confidence": 3}',
34  ].join('\n')
35}
36
37export type SpikeGuess = {
38  stepsMs: number[]
39  confidence: number
40}
41
42/**
43 * The guess in a subagent's reply, or null when it gave none that fits the
44 * plan. Each step is clamped to 6 seconds .. 4 hours.
45 */
46export function parseSpike(answer: string, stepCount: number): SpikeGuess | null {
47  const match = /\{[\s\S]*\}/.exec(answer)
48  if (!match) return null
49  let parsed: unknown
50  try {
51    parsed = JSON.parse(match[0])
52  } catch {
53    return null
54  }
55  if (typeof parsed !== 'object' || parsed === null) return null
56  const fields = parsed as { minutes?: unknown; confidence?: unknown }
57  if (!Array.isArray(fields.minutes) || fields.minutes.length !== stepCount) return null
58  if (!fields.minutes.every(m => typeof m === 'number' && Number.isFinite(m))) return null
59  const stepsMs = (fields.minutes as number[]).map(m => Math.min(MAX_MINUTES, Math.max(MIN_MINUTES, m)) * 60_000)
60  const confidence = typeof fields.confidence === 'number' ? Math.min(5, Math.max(1, fields.confidence)) : 1
61  return { stepsMs, confidence }
62}
63