Token-efficient delegation: routes sub-agents to the cheapest capable model tier, escalates on failure, learns from outcomes locally, and shows live savings.

A Claude Code mod with one job: spend fewer tokens without lowering the quality of the result.
It scores each task with cheap, deterministic heuristics, sends sub-agents to the cheapest model tier that can do the work, retries one tier up when a result fails its checks, and learns from the outcomes. It works in the Claude Code CLI and in the Code tab of the desktop app.
<sub>These pictures are generated by npm run preview: the mod's real renderers, run over a sample session. The first is the animated strip exactly as the desktop app receives it; the other two are the dashboard and status line text painted as SVG. They are not screenshots of a live session.</sub>
In a Claude Code session:
/plugin install tokensaver --marketplace AGregDev/claude-tokensaver
Or from your shell:
claude plugin install tokensaver --marketplace AGregDev/claude-tokensaver
Answer y to add the marketplace and pick a scope. The mod is active at once. Mods need Claude Code v2.1.287 or later; the shell form of the command needs v2.1.292.
A mod is code that runs with your permissions. Before you install any mod, you can list what it hooks and calls without running it: clone the repository and run claude plugin validate . in it. What it can reach has tokensaver's list.
When Claude starts a sub-agent, tokensaver scores the sub-agent's task and picks its model:
| Tier | Used for | | :- | :- | | Haiku | Simple, mechanical work: searches, bulk reads, repetitive edits | | Sonnet | Medium work, and anything Haiku-sized that the heuristics are unsure about | | Opus | Hard reasoning, architecture, and every risky task |
The score comes from five signals read off the task text with regular expressions and counts: scope (how broad), files touched, ambiguity, reasoning depth, and risk. No model is called to classify anything, so scoring costs no tokens.
Every adjustment after the first pick moves up, never down:
An agent type can name its own model in its definition, and a spawn does not say whether it does. So tokensaver leaves the first spawn of each agent type in a session untouched and watches which model Claude Code resolves for it:
model: in its definition does, and tokensaver never reroutes it.Forks, teammates and workflow agents are left alone, and so is any session whose model family tokensaver does not recognise.
The main conversation's own model is not switched. Changing models mid-session invalidates the prompt cache, which costs more than routing would save.
You do not have to ask for sub-agents. For each prompt you type, tokensaver works out whether handing the reading to a sub-agent would cost less than Claude doing it inline, and when it would, it tells Claude to delegate.
The comparison uses measured numbers, not guesses:
How much reading a task means is estimated from the files the prompt names and, for a broad task, from the size of the repository, which tokensaver samples once per session (file names only, never contents). So "find every usage across the repo" earns a suggestion in a 600-file repository and none in a 40-file one, where a single search is cheaper than any sub-agent.
When delegating wins by a clear margin, tokensaver adds one line to the prompt: about how many files, and to use one Explore sub-agent or up to three in parallel. Claude makes the final call, because it can see the code and tokensaver cannot. That line is the only thing tokensaver ever adds to the model's context for a prompt, and its cost is subtracted from the savings counter.
tokensaver never starts a sub-agent itself. A sub-agent that a mod starts has not been through the safety review Claude's own actions get, and Claude Code refuses it in auto mode. Delegation stays Claude's own action.
When a sub-agent's result fails its checks on a tier tokensaver picked, the same call is retried once on the next tier up, and Claude only sees the retry's result. The checks are:
If your next prompt opens by rejecting the result ("no, that's wrong", "try again"), every sub-agent that turn is routed one tier up.
A denied call is never retried, and neither is a turn you interrupted.
tokensaver changes how a session looks while work is being delegated:
/tokensaver dashboard opens it. Close it and it stays closed until you ask for it again.Animation runs only while something is happening, and the timer stops when things settle.
NO_COLOR (set to anything non-empty) removes every colour; state is still shown by glyphs and words.reducedMotion option, or /tokensaver motion off, stops all movement and shows final figures at once.None of the drawing reaches the model. A /tokensaver command does print one short line into the transcript, which Claude can read on later turns; that is what makes a command visible on every surface.
Both are estimates, and they are kept apart on purpose because adding them would count the same tokens twice.
The cost weights are ratios for estimating, not prices. The token meter counts new work (input, output and cache writes) and leaves out cache reads.
| Command | What it does | | :- | :- | | /tokensaver | Show status | | /tokensaver on / off | Master switch. Off, every event passes through untouched | | /tokensaver dry-run [on\|off] | Only show what would have been delegated or rerouted | | /tokensaver dashboard | Open the dashboard pane | | /tokensaver motion [on\|off] | Animations | | /tokensaver doctor | What tokensaver knows (model, repository size, agent types) and what the display reported. Run this first if nothing shows | | /tokensaver export [path] | Write the outcome log as JSON to a local file (default tokensaver-export.json) | | /tokensaver reset | Clear the outcome log and the learned thresholds | | /tokensaver help | List the commands |
Switches set with a command are remembered across sessions.
The most used ones are also listed in the / menu as /tokensaver:status, /tokensaver:dashboard, /tokensaver:doctor, /tokensaver:on, /tokensaver:off, /tokensaver:dry-run and /tokensaver:help. They do the same thing. They exist because /tokensaver itself is registered when a session starts running, so a brand-new session does not offer it in the menu until you have sent something; you can still type it in full.
Set these when you install, or later with /plugin configure tokensaver@claude-tokensaver.
| Option | Default | What it does | | :- | :- | :- | | enabled | true | Master switch | | dryRun | false | Change nothing; show what would have happened | | contextHint | true | Add the one-line delegation hint to prompts where delegating pays | | statusLine | true | Show the status line under the prompt | | reducedMotion | false | Turn off the spinner and the counting animation |
Each finished delegation that tokensaver routed is appended to a local outcome log: the kind of task, the tier used, tokens, retries, whether it was escalated, and whether it passed its checks. The log holds counts and labels only. No prompt, file name or answer text is stored.
Routing is driven by two cutoffs per kind of task, and the log tunes them:
The log and the thresholds live in the mod's own store on your machine (Claude Code keeps it under your Claude Code configuration directory). The log is capped at 500 entries. A corrupted log is read entry by entry; whatever does not parse is dropped and the rest is kept.
The bar a delegation suggestion must clear is learned the same lopsided way. When a turn ends, the suggestion made in it is judged by what happened: if Claude declined and then read only a few files, or the sub-agent barely did more than start, the suggestion was not worth making and the bar rises by a quarter at once. If the sub-agent did real work, the bar drops by a few percent.
Nothing is sent anywhere. The mod makes no network call at all. /tokensaver export writes a file on your disk, and /tokensaver reset deletes the log and the thresholds.
Each guarantee below is enforced by a test, and the tests are part of CI.
| Guarantee | How it is held | | :- | :- | | Risky, destructive or production-touching work is never downgraded | A property test over every risky prompt, parent model and named model, under default thresholds and under the loosest thresholds learning could ever produce | | Escalation only moves up | A property test over every tier, check result, retry count and flag: the retry is exactly one tier higher or does not happen | | An agent type's own model is never overridden | A property test over every task and parent model, for a type known to bring its own model and for one not yet seen; a test inside Claude Code confirms it against the real agent.spawn event | | Permission decisions are never overridden | The mod has no tool.check hook, so it cannot approve anything; a test inside Claude Code confirms deny, ask and allow come back unchanged, and that a denied call is returned untouched and not retried | | It stacks with other mods | Every hook passes the event on with next, and every hook that could gate an event has a .catch that passes it on unchanged if the hook fails | | It degrades instead of breaking | A failed store, environment or display call never fails the hook that made it, and a test runs the mod with none of them available. When no per-request usage arrives, the meter falls back to turn totals |
One thing it cannot soften: Claude Code checks a mod's event names when it loads the mod. On a build that lacks one of the events above, Claude Code refuses to load tokensaver and says which event, and your session carries on without it. tokensaver is typed and tested against Claude Code v2.1.292.
tokensaver runs after Claude Code's built-in sec-default guard and after any mod your organization lists ahead of user mods. It does nothing to work around either.
This is the complete list, as claude plugin validate reads it from the source. The build fails if the list changes, so it cannot grow without a reviewed edit to scripts/build.mjs.
| It hooks | To | | :- | :- | | session.start | Load the log and settings, register /tokensaver, sample the repository's size | | session.attach | Redraw when an app connects to the session | | prompt.submit | Score the prompt; add the delegation line when it pays | | agent.spawn | Choose the sub-agent's model | | tool.call on Agent only | Read the result; retry once one tier up when it fails | | turn.step, turn.complete, session.measure | Read token usage for the meter, and count how much the main loop reads itself | | command.run for tokensaver | Answer its own command | | ui.render for its own pane, the strip above the prompt, the spinner, and tool rows | Draw the dashboard and the strip; add to the spinner; mark sub-agent rows. Every other tool row is passed on untouched | | ui.close for its own pane | Remember that you closed it |
| It calls | For | | :- | :- | | $.store.get, set, delete | The local log, thresholds and switches | | $.fs.write | /tokensaver export, to the path you give | | $.fs.list | Counting the repository's files once per session: names only, a few folders deep, capped | | $.env.get | Reading NO_COLOR, and nothing else | | $.ui.status, toast, log, open, resolve | Display | | $.state.get, set | Redrawing the pane | | $.clock.every | The animation timer | | $.command.register, $.session.surfaces, $.session.model | Its command; telling the CLI from the desktop app; knowing which model the session runs on |
It does not call a model, start a sub-agent or a process, make a network request, read the contents of your files, or change a permission.
git clone https://github.com/AGregDev/claude-tokensaver
cd claude-tokensaver
npm install
npm run types:sync
npm run check
npm run check runs the typecheck, lint, the unit, safety, integration and UI tests with an 80% coverage threshold, Claude Code's validator, the tests that run inside Claude Code, and the build. CONTRIBUTING.md has the layout and the details.
To try a working copy in a session:
claude --plugin-dir .
hooks/register.tsx 628 lines1/**
2 * tokensaver's hooks module: the thin layer between Claude Code's events and
3 * the pure session controller in ../src. Every hook here observes, routes or
4 * draws; none approves a tool call, and there is no `tool.check` hook at all,
5 * so the mod cannot override a permission rule. It never starts a sub-agent
6 * itself either: delegation is always Claude's own, reviewed, action.
7 */
8import { atom, read, update } from 'claude-code'
9import type { EngineInterface, Register, Timer, ToolCallResult, TurnUsage } from 'claude-code'
10
11import { isRecord } from '../src/core/log.ts'
12import {
13 hydrate,
14 isAnimating,
15 onAgentResult,
16 onCommand,
17 onDiagnostic,
18 onFacts,
19 onMeasure,
20 onPrompt,
21 onSpawn,
22 onSpawned,
23 onStepUsage,
24 onTick,
25 onTurnComplete,
26} from '../src/runtime/session.ts'
27import type { AgentResultInput, Effect, Session, Step, Usage } from '../src/runtime/session.ts'
28import { BAND_HEIGHT, bandLine, bandSvg, isBandVisible, routedBadge, spinnerSuffix } from '../src/ui/band.ts'
29import { renderDashboard, sparklineLine } from '../src/ui/dashboard.ts'
30import { sparklineSvg } from '../src/ui/format.ts'
31import { renderStatusLine } from '../src/ui/statusline.ts'
32import type { Line, Span, Tone } from '../src/ui/theme.ts'
33import { toView } from '../src/ui/view.ts'
34
35const NAME = 'tokensaver'
36const TICK_MS = 250
37const STORE_LOG = 'log'
38const STORE_THRESHOLDS = 'thresholds'
39const STORE_SETTINGS = 'settings'
40const STORE_DELEGATION = 'delegation'
41
42/** Tools that read the repository; counted to see how much the main loop reads for itself. */
43const READING_TOOLS = ['Read', 'Grep', 'Glob']
44
45/** Folders that hold no source of the project's own. */
46const SKIPPED_FOLDERS = [
47 'node_modules',
48 '.git',
49 'dist',
50 'build',
51 'out',
52 'coverage',
53 'target',
54 'vendor',
55 '.next',
56 '.venv',
57 '__pycache__',
58]
59const MAX_LISTINGS = 40
60const MAX_COUNTED = 2000
61
62/** About how wide one text cell is in the desktop app, to size a picture to its slot. */
63const CELL_PX = 7.4
64
65const revision = atom({ plugin: 'tokensaver', key: 'revision' } as const, 0)
66
67type TextStyle = { color?: Tone; bold?: true; dimColor?: true }
68
69const styleOf = (part: Span): TextStyle => ({
70 ...(part.tone === undefined ? {} : { color: part.tone }),
71 ...(part.isBold === true ? { bold: true } : {}),
72 ...(part.isDim === true ? { dimColor: true } : {}),
73})
74
75const usageOf = (usage: TurnUsage): Usage => ({
76 model: usage.model,
77 input: usage.input_tokens,
78 output: usage.output_tokens,
79 cacheWrite: usage.cache_creation_input_tokens,
80})
81
82/** The Agent tool's own token total. A denial, an error and a background launch carry none. */
83const totalTokensOf = (result: unknown): number | undefined => {
84 const total = isRecord(result) ? result['totalTokens'] : undefined
85
86 return typeof total === 'number' ? total : undefined
87}
88
89const resultOf = (toolUseId: string, result: ToolCallResult<'Agent'>): AgentResultInput => ({
90 toolUseId,
91 isDenied: result.deny !== undefined,
92 isError: result.isError === true,
93 text: result.text ?? '',
94 totalTokens: totalTokensOf(result.result),
95 now: Date.now(),
96})
97
98/** A call that falls back instead of failing its hook when it is not available. */
99const attempt = async <Value,>(call: () => Promise<Value>, fallback: Value): Promise<Value> => {
100 try {
101 return await call()
102 } catch {
103 return fallback
104 }
105}
106
107const reasonOf = (error: unknown): string => (error instanceof Error ? error.message : String(error))
108
109// Module state: a session's worth of bookkeeping. A reload starts it over, and
110// `session.start` then rebuilds it from the store.
111let session: Session = hydrate({
112 options: {},
113 storedLog: undefined,
114 storedThresholds: undefined,
115 storedSettings: undefined,
116 noColor: undefined,
117}).session
118let timer: Timer | undefined
119let shownStatus: string | undefined
120let isPaneOpen = false
121let wasPaneClosedByPerson = false
122
123/** Opens the dashboard pane and keeps what Claude Code answered, for `/tokensaver doctor`. */
124const openPane = async ($: EngineInterface, why: string): Promise<void> => {
125 try {
126 const opened = await $.ui.open({ id: NAME, title: NAME })
127
128 isPaneOpen = opened.isPlaced
129 session = onDiagnostic(
130 session,
131 opened.isPlaced ? `pane opened (${why})` : `pane waiting (${why}): ${opened.reason}`,
132 ).session
133 } catch (error) {
134 isPaneOpen = false
135 session = onDiagnostic(session, `pane did not open (${why}): ${reasonOf(error)}`).session
136 }
137}
138
139const run = async ($: EngineInterface, effect: Effect): Promise<void> => {
140 switch (effect.kind) {
141 case 'toast':
142 $.ui.toast(effect.text)
143 break
144 case 'log':
145 $.ui.log(effect.text)
146 break
147 case 'save-log':
148 await $.store.set(STORE_LOG, session.outcomes)
149 await $.store.set(STORE_THRESHOLDS, session.thresholds)
150 await $.store.set(STORE_DELEGATION, { minSaving: session.minSaving })
151 break
152 case 'save-settings':
153 await $.store.set(STORE_SETTINGS, session.overrides)
154 break
155 case 'clear-log':
156 await $.store.delete(STORE_LOG)
157 await $.store.delete(STORE_THRESHOLDS)
158 await $.store.delete(STORE_DELEGATION)
159 break
160 case 'open-dashboard':
161 wasPaneClosedByPerson = false
162 await openPane($, 'asked')
163 break
164 case 'export':
165 await $.fs.write(effect.path, effect.json)
166 break
167 }
168}
169
170/** Redraws the status line and everything the mod draws. Display only: nothing here reaches the model. */
171const refresh = ($: EngineInterface): void => {
172 const status = renderStatusLine(toView(session))
173
174 if (status !== shownStatus) {
175 shownStatus = status
176 $.ui.status(status)
177 }
178
179 void attempt(() => update($, revision, count => count + 1), undefined)
180
181 if (timer === undefined && isAnimating(session)) {
182 timer = $.clock.every(TICK_MS, () => {
183 session = onTick(session, Date.now()).session
184
185 if (!isAnimating(session)) {
186 timer?.cancel()
187 timer = undefined
188 }
189
190 refresh($)
191 })
192 }
193}
194
195/**
196 * Puts the display back in step with the session. What is sent while no
197 * surface is attached is dropped, so this runs again whenever one may have
198 * appeared: when an app attaches, and at the start of every turn.
199 */
200const sync = async ($: EngineInterface, why: string): Promise<void> => {
201 const surfaces = await attempt(() => $.session.surfaces(), [])
202
203 // Send the status line again even when its text has not changed.
204 shownStatus = undefined
205
206 // A surface with room beside the transcript gets the dashboard; the terminal gets the status line.
207 if (
208 session.config.isEnabled &&
209 !isPaneOpen &&
210 !wasPaneClosedByPerson &&
211 surfaces.some(surface => surface !== 'terminal')
212 ) {
213 await openPane($, why)
214 }
215
216 refresh($)
217}
218
219/** Applies a transition: keeps its session, carries out its effects, redraws. */
220const commit = async <Reply,>($: EngineInterface, step: Step<Reply>): Promise<Reply> => {
221 session = { ...step.session, now: Math.max(step.session.now, Date.now()) }
222
223 for (const effect of step.effects) {
224 // A display or storage effect that fails must never fail the hook that caused it.
225 await attempt(() => run($, effect), undefined)
226 }
227
228 refresh($)
229
230 return step.reply
231}
232
233/**
234 * Counts the project's files, a few folders deep, to judge how much reading a
235 * broad task means. A sample with a ceiling: it reads names only, never contents.
236 */
237const countFiles = async ($: EngineInterface): Promise<number | undefined> => {
238 const queue: { path: string; depth: number }[] = [{ path: '.', depth: 0 }]
239 let listings = 0
240 let files = 0
241
242 while (queue.length > 0 && listings < MAX_LISTINGS && files < MAX_COUNTED) {
243 const folder = queue.shift()
244
245 if (folder === undefined) {
246 break
247 }
248
249 const entries = await attempt(() => $.fs.list(folder.path), undefined)
250
251 if (entries === undefined) {
252 // The first listing failing means the count is unknown, not zero.
253 if (listings === 0) {
254 return undefined
255 }
256
257 continue
258 }
259
260 listings += 1
261
262 for (const entry of entries) {
263 if (entry.kind === 'file') {
264 files += 1
265 } else if (entry.kind === 'dir' && folder.depth < 3 && !SKIPPED_FOLDERS.includes(entry.name)) {
266 queue.push({
267 path: folder.path === '.' ? entry.name : `${folder.path}/${entry.name}`,
268 depth: folder.depth + 1,
269 })
270 }
271 }
272 }
273
274 return files
275}
276
277export const register: Register = (on, options) => {
278 session = hydrate({
279 options,
280 storedLog: undefined,
281 storedThresholds: undefined,
282 storedSettings: undefined,
283 noColor: undefined,
284 }).session
285 isPaneOpen = false
286 wasPaneClosedByPerson = false
287
288 on('session.start', async ($, e, next) => {
289 const [storedLog, storedThresholds, storedSettings, storedDelegation, noColor] =
290 await Promise.all([
291 attempt(() => $.store.get(STORE_LOG), undefined),
292 attempt(() => $.store.get(STORE_THRESHOLDS), undefined),
293 attempt(() => $.store.get(STORE_SETTINGS), undefined),
294 attempt(() => $.store.get(STORE_DELEGATION), undefined),
295 attempt(() => $.env.get('NO_COLOR'), undefined),
296 ])
297
298 await commit(
299 $,
300 hydrate({ options, storedLog, storedThresholds, storedSettings, storedDelegation, noColor }),
301 )
302 await attempt(
303 () =>
304 $.command.register({
305 name: 'tokensaver',
306 description: 'Token-saving delegation: status, on/off, dry-run, dashboard, doctor, export, reset',
307 argumentHint: '[on|off|dry-run|dashboard|doctor|motion|export|reset|help]',
308 immediate: true,
309 }),
310 undefined,
311 )
312
313 const [model, repoFiles] = await Promise.all([
314 attempt(() => $.session.model(), undefined),
315 countFiles($),
316 ])
317
318 await commit($, onFacts(session, { model, repoFiles }))
319 await sync($, 'session start')
320
321 return next(e)
322 })
323
324 on('session.attach', async ($, e, next) => {
325 const attached = await next(e)
326
327 session = onDiagnostic(session, `${e.surface} attached`).session
328 await sync($, `${e.surface} attached`)
329
330 return attached
331 })
332
333 on('ui.close', { id: 'tokensaver' }, async (_$, e, next) => {
334 const closed = await next(e)
335
336 isPaneOpen = false
337
338 // A pane the person closed stays closed until they ask for it again.
339 if (e.origin.kind === 'person') {
340 wasPaneClosedByPerson = true
341 }
342
343 return closed
344 }).catch((_$, e, next) => next(e))
345
346 on('command.run', { command: 'tokensaver' }, async ($, e) => {
347 const reply = await commit($, onCommand(session, e.args, Date.now()))
348
349 return { text: reply.text }
350 })
351
352 // The same commands, declared as files in commands/ so the menu lists them before this
353 // module has run: `/tokensaver:doctor` is `/tokensaver doctor`. Answered here, they
354 // never reach the model; their file text is only what Claude reads if this hook is absent.
355 on(
356 'command.run',
357 {
358 command: [
359 'tokensaver:status',
360 'tokensaver:dashboard',
361 'tokensaver:doctor',
362 'tokensaver:on',
363 'tokensaver:off',
364 'tokensaver:dry-run',
365 'tokensaver:help',
366 ],
367 },
368 async ($, e) => {
369 const word = e.command.slice('tokensaver:'.length)
370 const reply = await commit($, onCommand(session, `${word} ${e.args}`, Date.now()))
371
372 return { text: reply.text }
373 },
374 ).catch((_$, e, next) => next(e))
375
376 on('prompt.submit', async ($, e, next) => {
377 const kind = e.origin.kind
378 const model = await attempt(() => $.session.model(), undefined)
379
380 if (model !== undefined) {
381 session = onFacts(session, { model }).session
382 }
383
384 await sync($, 'turn start')
385
386 const reply = await commit(
387 $,
388 onPrompt(session, {
389 text: e.text,
390 isFromUser: kind === 'composer' || kind === 'bridge' || kind === 'sdk',
391 now: Date.now(),
392 }),
393 )
394
395 return reply.context === undefined
396 ? next(e)
397 : next({ ...e, context: [...(e.context ?? []), reply.context] })
398 }).catch((_$, e, next) => next(e))
399
400 let unnamed = 0
401
402 on('agent.spawn', async ($, e, next) => {
403 unnamed += 1
404
405 const toolUseId = e.tool_use_id === '' ? `spawn-${String(unnamed)}` : e.tool_use_id
406 const reply = await commit(
407 $,
408 onSpawn(session, {
409 toolUseId,
410 prompt: e.prompt,
411 description: e.description,
412 subagentType: e.subagentType,
413 explicitModel: e.model,
414 parentModel: e.parentModel,
415 parentAgentId: e.parentAgentId,
416 isBackground: e.background,
417 isPinned: e.fork || e.isTeammate === true || e.workflow !== undefined,
418 now: Date.now(),
419 }),
420 )
421 const started = await next(reply.model === undefined ? e : { ...e, model: reply.model })
422
423 await commit(
424 $,
425 onSpawned(session, {
426 toolUseId,
427 agentId: started.agentId,
428 isDenied: started.deny !== undefined,
429 resolvedModel: started.model,
430 }),
431 )
432
433 return started
434 }).catch((_$, e, next) => next(e))
435
436 on('tool.call', { tool: 'Agent' }, async ($, e, next) => {
437 const first = await next(e)
438 const reply = await commit($, onAgentResult(session, resultOf(e.tool_use_id, first)))
439
440 // A denied call, a passing result and a dry run all come back with nothing to retry.
441 if (reply.retryWith === undefined) {
442 return first
443 }
444
445 const second = await next({ ...e, model: reply.retryWith })
446
447 await commit($, onAgentResult(session, resultOf(e.tool_use_id, second)))
448
449 return second
450 }).catch((_$, e, next) => next(e))
451
452 on('turn.step', async function* ($, e, next) {
453 const result = yield* next(e)
454
455 if (result.usage !== null) {
456 await commit(
457 $,
458 onStepUsage(session, {
459 turnId: e.turnId,
460 agentId: e.agentId,
461 usage: usageOf(result.usage),
462 reads: result.toolUses.filter(use => READING_TOOLS.includes(use.name)).length,
463 }),
464 )
465 }
466
467 return result
468 })
469
470 on('turn.complete', async ($, e, next) => {
471 await commit(
472 $,
473 onTurnComplete(session, {
474 turnId: e.turnId,
475 agentId: e.agentId,
476 usage: e.usage === undefined ? undefined : usageOf(e.usage),
477 isFailed: e.reason === 'error' || e.reason === 'refusal',
478 isAborted: e.isAborted,
479 answer: e.answer,
480 now: Date.now(),
481 }),
482 )
483
484 return next(e)
485 }).catch((_$, e, next) => next(e))
486
487 on('session.measure', async ($, e, next) => {
488 await commit($, onMeasure(session, { tokens: e.context.tokens, window: e.context.window }))
489
490 return next(e)
491 }).catch((_$, e, next) => next(e))
492
493 // --- Drawing ---------------------------------------------------------------
494
495 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
496 // Reading the revision subscribes the strip to every later change.
497 await read($, revision)
498
499 const view = toView(session)
500
501 if (!isBandVisible(view) || e.props.hasSurvey) {
502 return next(e)
503 }
504
505 const { Box, Text } = $.ui.resolve(e)
506 // Whatever another mod draws in the strip stays, beneath tokensaver's line.
507 const beneath = await next(e)
508
509 if (e.surface === 'desktop') {
510 const width = Math.round(Math.max(40, e.props.bodyColumns) * CELL_PX)
511
512 return (
513 <Box flexDirection="column">
514 {h($.ui.resolve({ ...e, surface: 'desktop' as const }).Svg, {
515 source: bandSvg(view, width),
516 alt: bandLine(view, e.props.bodyColumns)
517 .map(part => part.text)
518 .join(''),
519 width,
520 height: BAND_HEIGHT,
521 })}
522 {beneath}
523 </Box>
524 )
525 }
526
527 return (
528 <Box flexDirection="column">
529 <Box>
530 {bandLine(view, e.props.bodyColumns).map(part => (
531 <Text {...styleOf(part)}>{part.text}</Text>
532 ))}
533 </Box>
534 {beneath}
535 </Box>
536 )
537 })
538
539 on('ui.render', { component: 'Spinner' }, async ($, e, next) => {
540 await read($, revision)
541
542 const suffix = spinnerSuffix(toView(session))
543
544 return suffix === undefined
545 ? next(e)
546 : next({ ...e, props: { ...e.props, suffix: `${e.props.suffix}${suffix}` } })
547 })
548
549 on('ui.render', { component: 'ToolUse' }, async ($, e, next) => {
550 if (e.props.tool !== 'Agent') {
551 return next(e)
552 }
553
554 // Only a running call is kept in step: its model is decided just after its row appears.
555 if (e.props.isRunning) {
556 await read($, revision)
557 }
558
559 const badge = routedBadge(toView(session), e.props.tool_use_id)
560 const row = await next(e)
561
562 if (badge === undefined) {
563 return row
564 }
565
566 const { Box, Text } = $.ui.resolve(e)
567
568 return (
569 <Box flexDirection="column">
570 <Box>
571 {badge.map(part => (
572 <Text {...styleOf(part)}>{part.text}</Text>
573 ))}
574 </Box>
575 {row}
576 </Box>
577 )
578 })
579
580 on('ui.render', { component: 'Pane', requestId: 'tokensaver' }, async ($, e) => {
581 await read($, revision)
582 isPaneOpen = true
583
584 const { Box, Text, Button } = $.ui.resolve(e)
585 const view = toView(session)
586 const columns = e.props.bodyColumns
587 const dashboard = renderDashboard(view, columns)
588 const row = (line: Line) => (
589 <Box>
590 {line.map(part => (
591 <Text {...styleOf(part)}>{part.text}</Text>
592 ))}
593 </Box>
594 )
595 const command = (args: string) => () => commit($, onCommand(session, args, Date.now()))
596
597 return (
598 <Box flexDirection="column">
599 {dashboard.top.map(row)}
600 {e.surface === 'desktop' && view.history.length >= 2 ? (
601 <Box>
602 <Text>{'Trend '}</Text>
603 {h($.ui.resolve({ ...e, surface: 'desktop' as const }).Svg, {
604 source: sparklineSvg(view.history),
605 alt: `Estimated savings over the last ${String(view.history.length)} turns`,
606 })}
607 </Box>
608 ) : (
609 row(sparklineLine(view, columns))
610 )}
611 {dashboard.bottom.map(row)}
612 <Box gap={2}>
613 <Button
614 key="power"
615 label={view.mode === 'off' ? 'Turn on' : 'Turn off'}
616 onPress={command(view.mode === 'off' ? 'on' : 'off')}
617 />
618 <Button
619 key="dry-run"
620 label={view.mode === 'dry-run' ? 'Leave dry-run' : 'Dry-run'}
621 onPress={command(view.mode === 'dry-run' ? 'dry-run off' : 'dry-run on')}
622 />
623 </Box>
624 </Box>
625 )
626 })
627}
628src/core/log.ts 177 lines1import { isCheckSignal } from './escalate.ts'
2import type { CheckSignal } from './escalate.ts'
3import { DEFAULT_THRESHOLDS } from './route.ts'
4import type { Cutoffs, Thresholds } from './route.ts'
5import { TASK_TYPES, isTaskType } from './score.ts'
6import type { TaskType } from './score.ts'
7import { TIERS, isTier } from './tiers.ts'
8import type { Tier } from './tiers.ts'
9
10/** One finished delegation, as the local outcome log keeps it. No prompt text is stored. */
11export type Outcome = {
12 seq: number
13 at: number
14 taskType: TaskType
15 tier: Tier
16 tokens: number
17 retries: number
18 wasEscalated: boolean
19 isSuccess: boolean
20 signal: CheckSignal
21}
22
23/** The log is capped so it stays far below the store's size limit. */
24export const MAX_OUTCOMES = 500
25
26export const isRecord = (value: unknown): value is Readonly<Record<string, unknown>> =>
27 typeof value === 'object' && value !== null && !Array.isArray(value)
28
29const isCount = (value: unknown): value is number =>
30 typeof value === 'number' && Number.isFinite(value) && value >= 0
31
32/** Reads one stored entry, or undefined when it is not a well-formed outcome. */
33export const parseOutcome = (value: unknown): Outcome | undefined => {
34 if (!isRecord(value)) {
35 return undefined
36 }
37
38 const { seq, at, taskType, tier, tokens, retries, wasEscalated, isSuccess, signal } = value
39
40 if (
41 !isCount(seq) ||
42 !isCount(at) ||
43 !isTaskType(taskType) ||
44 !isTier(tier) ||
45 !isCount(tokens) ||
46 !isCount(retries) ||
47 typeof wasEscalated !== 'boolean' ||
48 typeof isSuccess !== 'boolean' ||
49 !isCheckSignal(signal)
50 ) {
51 return undefined
52 }
53
54 return { seq, at, taskType, tier, tokens, retries, wasEscalated, isSuccess, signal }
55}
56
57export type ParsedLog = {
58 outcomes: readonly Outcome[]
59 /** Entries that were present but unreadable, including a log that was not a list at all. */
60 dropped: number
61}
62
63/**
64 * Reads the stored log defensively. A missing log is empty; a corrupted one
65 * keeps every entry that still parses and counts the rest as dropped.
66 */
67export const parseLog = (value: unknown): ParsedLog => {
68 if (value === undefined || value === null) {
69 return { outcomes: [], dropped: 0 }
70 }
71
72 if (!Array.isArray(value)) {
73 return { outcomes: [], dropped: 1 }
74 }
75
76 const entries: readonly unknown[] = value
77 const outcomes = entries
78 .map(parseOutcome)
79 .filter(outcome => outcome !== undefined)
80 .slice(-MAX_OUTCOMES)
81
82 return { outcomes, dropped: entries.length - outcomes.length }
83}
84
85export const nextSeq = (outcomes: readonly Outcome[]): number =>
86 outcomes.reduce((highest, outcome) => Math.max(highest, outcome.seq), 0) + 1
87
88export const appendOutcome = (outcomes: readonly Outcome[], outcome: Outcome): readonly Outcome[] =>
89 [...outcomes, outcome].slice(-MAX_OUTCOMES)
90
91/** A clean outcome is the only evidence that counts toward loosening a rule. */
92export const isCleanSuccess = (outcome: Outcome): boolean =>
93 outcome.isSuccess && !outcome.wasEscalated && outcome.retries === 0
94
95const clampCutoff = (value: unknown, fallback: number): number =>
96 typeof value === 'number' && Number.isFinite(value)
97 ? Math.min(100, Math.max(0, Math.round(value)))
98 : fallback
99
100const parseCutoffs = (value: unknown, fallback: Cutoffs): Cutoffs => {
101 if (!isRecord(value)) {
102 return fallback
103 }
104
105 const haikuMax = clampCutoff(value['haikuMax'], fallback.haikuMax)
106 const sonnetMax = Math.max(haikuMax, clampCutoff(value['sonnetMax'], fallback.sonnetMax))
107 const tunedAtSeq = isCount(value['tunedAtSeq']) ? value['tunedAtSeq'] : 0
108
109 return { haikuMax, sonnetMax, tunedAtSeq }
110}
111
112/** Reads stored thresholds, falling back to the default for anything missing or malformed. */
113export const parseThresholds = (value: unknown): Thresholds => {
114 const stored = isRecord(value) ? value : {}
115
116 return {
117 search: parseCutoffs(stored['search'], DEFAULT_THRESHOLDS.search),
118 'bulk-read': parseCutoffs(stored['bulk-read'], DEFAULT_THRESHOLDS['bulk-read']),
119 'repetitive-edit': parseCutoffs(
120 stored['repetitive-edit'],
121 DEFAULT_THRESHOLDS['repetitive-edit'],
122 ),
123 summarize: parseCutoffs(stored['summarize'], DEFAULT_THRESHOLDS.summarize),
124 edit: parseCutoffs(stored['edit'], DEFAULT_THRESHOLDS.edit),
125 reasoning: parseCutoffs(stored['reasoning'], DEFAULT_THRESHOLDS.reasoning),
126 general: parseCutoffs(stored['general'], DEFAULT_THRESHOLDS.general),
127 }
128}
129
130export type LogSummary = {
131 count: number
132 cleanRate: number
133 escalations: number
134 failures: number
135 byTier: Readonly<Record<Tier, number>>
136}
137
138export const summarizeLog = (outcomes: readonly Outcome[]): LogSummary => {
139 const clean = outcomes.filter(isCleanSuccess).length
140 const countAt = (tier: Tier): number => outcomes.filter(outcome => outcome.tier === tier).length
141
142 return {
143 count: outcomes.length,
144 cleanRate: outcomes.length === 0 ? 1 : clean / outcomes.length,
145 escalations: outcomes.filter(outcome => outcome.wasEscalated).length,
146 failures: outcomes.filter(outcome => !outcome.isSuccess).length,
147 byTier: { haiku: countAt('haiku'), sonnet: countAt('sonnet'), opus: countAt('opus') },
148 }
149}
150
151export type ExportBundle = {
152 format: 'tokensaver-export'
153 version: 1
154 exportedAt: number
155 taskTypes: readonly TaskType[]
156 tiers: readonly Tier[]
157 thresholds: Thresholds
158 summary: LogSummary
159 outcomes: readonly Outcome[]
160}
161
162/** Everything `/tokensaver export` writes: the log, the tuned thresholds and a summary. */
163export const buildExport = (
164 outcomes: readonly Outcome[],
165 thresholds: Thresholds,
166 exportedAt: number,
167): ExportBundle => ({
168 format: 'tokensaver-export',
169 version: 1,
170 exportedAt,
171 taskTypes: TASK_TYPES,
172 tiers: TIERS,
173 thresholds,
174 summary: summarizeLog(outcomes),
175 outcomes,
176})
177src/runtime/session.ts 1002 lines1/**
2 * The session controller: every mod event becomes one pure transition,
3 * `(session, input) -> { session, effects, reply }`. Nothing here touches
4 * Claude Code; hooks/register.tsx feeds events in and carries effects out.
5 */
6import { checkResult, isUserRejection, planEscalation } from '../core/escalate.ts'
7import type { CheckSignal } from '../core/escalate.ts'
8import {
9 appendOutcome,
10 buildExport,
11 nextSeq,
12 parseLog,
13 parseThresholds,
14 summarizeLog,
15} from '../core/log.ts'
16import type { Outcome } from '../core/log.ts'
17import {
18 DEFAULT_MIN_SAVING,
19 decideDelegation,
20 judgeHint,
21 parseMinSaving,
22 tuneMinSaving,
23} from '../core/delegate.ts'
24import { DEFAULT_THRESHOLDS, resolveSpawnModel, routeTier } from '../core/route.ts'
25import type { Thresholds } from '../core/route.ts'
26import { contextSavings, estimateTokens, routingSavings } from '../core/savings.ts'
27import { scoreTask } from '../core/score.ts'
28import type { TaskType } from '../core/score.ts'
29import { rankOf, rankOfModel, tierOfModel } from '../core/tiers.ts'
30import type { Tier } from '../core/tiers.ts'
31import { learnFromOutcome, tighten, tuneFromLog } from '../core/tune.ts'
32import { DEFAULT_EXPORT_PATH, HELP_TEXT, parseCommand } from './commands.ts'
33import { applyOverrides, isNoColor, parseOptions, parseOverrides } from './config.ts'
34import type { Config, Overrides } from './config.ts'
35import { formatTokens } from '../ui/format.ts'
36
37export type AgentStatus = 'running' | 'retrying' | 'done' | 'failed' | 'planned'
38
39/**
40 * How an agent type gets its model when nobody names one: from its parent, or
41 * from its own definition. Only the first kind is tokensaver's to route.
42 */
43export type AgentTypeModel = 'inherits' | 'own-model'
44
45/** One sub-agent in the delegation tree. */
46export type AgentNode = {
47 /** The Agent tool call it belongs to. */
48 key: string
49 subagentType: string
50 parentModel: string
51 /** True for the first, untouched spawn of an agent type: it shows how the type picks its model. */
52 isProbe: boolean
53 agentId: string | undefined
54 parentAgentId: string | undefined
55 label: string
56 taskType: TaskType
57 /** The tier it runs on; undefined when it inherits a model tokensaver cannot name. */
58 tier: Tier | undefined
59 /** True when tokensaver chose its model. */
60 isRouted: boolean
61 escalatedFrom: Tier | undefined
62 /** Set between a failed attempt and its retry: the tier the retry must use. */
63 pendingTier: Tier | undefined
64 status: AgentStatus
65 tokens: number
66 /** Tokens a failed attempt spent before its retry. */
67 wastedTokens: number
68 answerTokens: number
69 attempts: number
70 isBackground: boolean
71 isDryRun: boolean
72 /** True when tokensaver's hint asked for this delegation. */
73 isHinted: boolean
74 isRecorded: boolean
75 /** The rank of the model it would have run on without tokensaver. */
76 baselineRank: number | undefined
77}
78
79export type Spend = Readonly<Record<Tier | 'other', number>>
80
81export type Session = {
82 config: Config
83 overrides: Overrides
84 hasNoColor: boolean
85 thresholds: Thresholds
86 outcomes: readonly Outcome[]
87 spend: Spend
88 context: { tokens: number; window: number }
89 /**
90 * Estimated saving from running work on cheaper tiers, in tokens of the
91 * model it would otherwise have run on, net of failed attempts and of
92 * tokensaver's own hint. Negative when the overhead outweighs the gain.
93 */
94 saved: number
95 /** Tokens of delegated work the main context never had to hold: the work less its summary. */
96 keptOut: number
97 savedHistory: readonly number[]
98 agents: readonly AgentNode[]
99 /** What this session has seen of each agent type; a type not listed has not been observed yet. */
100 agentTypes: Readonly<Record<string, AgentTypeModel>>
101 /** Delegations of the turn in progress, by key; the next prompt reads them to judge that turn. */
102 turnKeys: readonly string[]
103 /** Tiers to add to every routing decision this turn (after a user rejection). */
104 boost: number
105 isHintActive: boolean
106 /** Turns whose usage arrived step by step, so their totals are not counted twice. */
107 steppedTurns: readonly string[]
108 /** The bar a delegation suggestion must clear, tuned from how suggestions turn out. */
109 minSaving: number
110 /** What the session knows about where it runs. */
111 facts: SessionFacts
112 /** Files the main loop read or searched itself in the turn in progress. */
113 mainReads: number
114 /** The newest event worth showing above the prompt, for a few seconds. */
115 flash: Flash | undefined
116 /** The clock as of the last event or animation frame. */
117 now: number
118 /** What the display layer reported, for `/tokensaver doctor`. */
119 diagnostics: readonly string[]
120 notes: readonly string[]
121 frame: number
122 shownSaved: number
123}
124
125export type SessionFacts = {
126 /** Files under the working directory, sampled once; undefined until counted. */
127 repoFiles: number | undefined
128 /** The session's model id; undefined until Claude Code has said. */
129 model: string | undefined
130}
131
132export type FlashTone = 'route' | 'escalate' | 'done' | 'fail' | 'info'
133
134export type Flash = {
135 text: string
136 tone: FlashTone
137 at: number
138}
139
140/** How long an event stays above the prompt after it happens. */
141export const FLASH_MS = 6000
142
143export type Effect =
144 | { kind: 'toast'; text: string }
145 | { kind: 'log'; text: string }
146 | { kind: 'save-log' }
147 | { kind: 'save-settings' }
148 | { kind: 'clear-log' }
149 | { kind: 'open-dashboard' }
150 | { kind: 'export'; path: string; json: string }
151
152export type Step<Reply> = {
153 session: Session
154 effects: readonly Effect[]
155 reply: Reply
156}
157
158const MAX_AGENTS = 24
159const MAX_HISTORY = 32
160const MAX_NOTES = 4
161const MAX_STEPPED_TURNS = 64
162const MAX_DIAGNOSTICS = 12
163const HINT_NOUNS: Readonly<Record<TaskType, string>> = {
164 search: 'search',
165 'bulk-read': 'read-through',
166 'repetitive-edit': 'repetitive edit',
167 summarize: 'summarising job',
168 edit: 'edit',
169 reasoning: 'problem',
170 general: 'task',
171}
172
173const withFlash = (session: Session, text: string, tone: FlashTone, at: number): Session => ({
174 ...session,
175 flash: { text, tone, at },
176 now: Math.max(session.now, at),
177})
178
179const still = <Reply>(session: Session, reply: Reply): Step<Reply> => ({
180 session,
181 effects: [],
182 reply,
183})
184
185const shorten = (text: string, limit = 40): string => {
186 const line = text.replace(/\s+/g, ' ').trim()
187
188 return line.length <= limit ? line : `${line.slice(0, limit - 1)}…`
189}
190
191const withNote = (session: Session, note: string): Session => ({
192 ...session,
193 notes: [...session.notes, note].slice(-MAX_NOTES),
194})
195
196const replaceAgent = (session: Session, node: AgentNode): Session => ({
197 ...session,
198 agents: session.agents.map(agent => (agent.key === node.key ? node : agent)),
199})
200
201const signalText = (signal: CheckSignal): string => signal.replace('-', ' ')
202
203export type HydrateInput = {
204 options: Readonly<Record<string, unknown>>
205 storedLog: unknown
206 storedThresholds: unknown
207 storedSettings: unknown
208 storedDelegation?: unknown
209 noColor: string | undefined
210}
211
212/** Builds the session from the plugin's options and whatever the local store holds. */
213export const hydrate = (input: HydrateInput): Step<undefined> => {
214 const log = parseLog(input.storedLog)
215 const overrides = parseOverrides(input.storedSettings)
216 const session: Session = {
217 config: applyOverrides(parseOptions(input.options), overrides),
218 overrides,
219 hasNoColor: isNoColor(input.noColor),
220 thresholds: tuneFromLog(parseThresholds(input.storedThresholds), log.outcomes),
221 outcomes: log.outcomes,
222 spend: { haiku: 0, sonnet: 0, opus: 0, other: 0 },
223 context: { tokens: 0, window: 0 },
224 saved: 0,
225 keptOut: 0,
226 savedHistory: [],
227 agents: [],
228 agentTypes: {},
229 turnKeys: [],
230 boost: 0,
231 isHintActive: false,
232 steppedTurns: [],
233 minSaving: parseMinSaving(input.storedDelegation),
234 facts: { repoFiles: undefined, model: undefined },
235 mainReads: 0,
236 flash: undefined,
237 now: 0,
238 diagnostics: [],
239 notes: [],
240 frame: 0,
241 shownSaved: 0,
242 }
243
244 if (log.dropped === 0) {
245 return still(session, undefined)
246 }
247
248 return {
249 session,
250 effects: [
251 {
252 kind: 'log',
253 text: `tokensaver: ${String(log.dropped)} unreadable entr${log.dropped === 1 ? 'y' : 'ies'} dropped from the outcome log`,
254 },
255 { kind: 'save-log' },
256 ],
257 reply: undefined,
258 }
259}
260
261/** What the display layer learned about the session: its model and how big the repository is. */
262export const onFacts = (session: Session, facts: Partial<SessionFacts>): Step<undefined> =>
263 still({ ...session, facts: { ...session.facts, ...facts } }, undefined)
264
265/** Keeps a line for `/tokensaver doctor`: what was asked of the display and what came back. */
266export const onDiagnostic = (session: Session, line: string): Step<undefined> =>
267 still(
268 { ...session, diagnostics: [...session.diagnostics, line].slice(-MAX_DIAGNOSTICS) },
269 undefined,
270 )
271
272export type PromptInput = {
273 text: string
274 /** True for a prompt the person typed; notifications and peers are not scored. */
275 isFromUser: boolean
276 now: number
277}
278
279/** A user rejection counts against every model tokensaver picked in the turn it rejects. */
280const recordRejections = (session: Session, now: number): Session =>
281 session.agents
282 .filter(agent => session.turnKeys.includes(agent.key) && agent.isRouted && !agent.isDryRun)
283 .reduce((current, agent) => {
284 if (agent.tier === undefined) {
285 return current
286 }
287
288 const outcome: Outcome = {
289 seq: nextSeq(current.outcomes),
290 at: now,
291 taskType: agent.taskType,
292 tier: agent.tier,
293 tokens: agent.tokens,
294 retries: agent.attempts,
295 wasEscalated: agent.escalatedFrom !== undefined,
296 isSuccess: false,
297 signal: 'user-rejection',
298 }
299
300 return {
301 ...current,
302 outcomes: appendOutcome(current.outcomes, outcome),
303 thresholds: tighten(current.thresholds, agent.taskType, agent.tier, outcome.seq),
304 }
305 }, session)
306
307/**
308 * A submitted prompt: judges the previous turn if the user rejects it, then
309 * scores the new task and decides whether delegating would save tokens.
310 */
311export const onPrompt = (
312 session: Session,
313 input: PromptInput,
314): Step<{ context: string | undefined }> => {
315 if (!session.config.isEnabled || !input.isFromUser) {
316 return still(session, { context: undefined })
317 }
318
319 const effects: Effect[] = []
320 const hadRouted = session.agents.some(
321 agent => session.turnKeys.includes(agent.key) && agent.isRouted && !agent.isDryRun,
322 )
323 const isRejected = hadRouted && isUserRejection(input.text)
324 let next: Session = isRejected ? recordRejections(session, input.now) : session
325
326 if (isRejected) {
327 effects.push(
328 { kind: 'toast', text: 'tokensaver: result rejected, routing one tier up this turn' },
329 { kind: 'save-log' },
330 )
331 next = withNote(next, 'rejected: one tier up this turn')
332 }
333
334 const score = scoreTask({ text: input.text })
335 const delegation = decideDelegation(score, {
336 repoFiles: session.facts.repoFiles,
337 parentRank: rankOfModel(session.facts.model),
338 minSaving: session.minSaving,
339 })
340 const shouldHint =
341 delegation.shouldDelegate && session.config.hasContextHint && !session.config.isDryRun
342 const helpers =
343 delegation.agents > 1
344 ? `up to ${String(delegation.agents)} Explore sub-agents in parallel, one per independent part,`
345 : 'one Explore sub-agent'
346 // Directive, because Claude only delegates on a clear instruction; with a way out, because
347 // this estimate is made from the wording and Claude can see the code.
348 const hint = `[tokensaver] This ${HINT_NOUNS[score.taskType]} likely means reading about ${String(delegation.estimatedFiles)} files. Before reading them yourself, hand that reading to ${helpers} with the Agent tool, and work from what they report. Skip this only if one or two targeted lookups would answer it.`
349
350 if (delegation.shouldDelegate && session.config.isDryRun) {
351 const text = `tokensaver dry-run: would suggest delegating this ${HINT_NOUNS[score.taskType]} (about ${String(delegation.estimatedFiles)} files, ~${formatTokens(delegation.estimatedSaving)} tokens saved)`
352 effects.push({ kind: 'toast', text })
353 next = withNote(next, text.replace('tokensaver ', ''))
354 }
355
356 if (shouldHint) {
357 next = withFlash(
358 next,
359 `suggested delegating about ${String(delegation.estimatedFiles)} files of reading`,
360 'info',
361 input.now,
362 )
363 }
364
365 return {
366 session: {
367 ...next,
368 turnKeys: [],
369 mainReads: 0,
370 boost: isRejected ? 1 : 0,
371 isHintActive: shouldHint,
372 // The hint is the one thing tokensaver adds to the model's context: count it as a cost.
373 saved: shouldHint ? next.saved - estimateTokens(hint) : next.saved,
374 },
375 effects,
376 reply: { context: shouldHint ? hint : undefined },
377 }
378}
379
380export type SpawnInput = {
381 toolUseId: string
382 prompt: string
383 description: string
384 subagentType: string
385 explicitModel: string | undefined
386 parentModel: string
387 parentAgentId: string | undefined
388 isBackground: boolean
389 /** True for a fork, a teammate or a workflow agent: their model is not tokensaver's to set. */
390 isPinned: boolean
391 now?: number | undefined
392}
393
394/**
395 * A sub-agent is about to start: scores its task and answers with the model
396 * it should run on, or with nothing to leave the spawn untouched.
397 */
398export const onSpawn = (
399 session: Session,
400 input: SpawnInput,
401): Step<{ model: string | undefined }> => {
402 if (!session.config.isEnabled) {
403 return still(session, { model: undefined })
404 }
405
406 const existing = session.agents.find(agent => agent.key === input.toolUseId)
407 const label = shorten(input.description === '' ? input.prompt : input.description)
408 const score = scoreTask({ text: input.prompt, description: input.description })
409 const decision = routeTier(score, session.thresholds, session.boost)
410 const known = session.agentTypes[input.subagentType]
411 // An agent type may name its own model, and a spawn does not say whether it does. So
412 // tokensaver sets a model only where that cannot override the type's choice: on a retry of
413 // its own, on a model the caller named (which it only ever raises), or on a type it has
414 // watched inherit its parent's model.
415 const canRoute =
416 !input.isPinned &&
417 (existing?.pendingTier !== undefined || input.explicitModel !== undefined || known === 'inherits')
418 const routing = canRoute
419 ? resolveSpawnModel(decision, score, {
420 parentModel: input.parentModel,
421 explicitModel: input.explicitModel,
422 forcedTier: existing?.pendingTier,
423 })
424 : { model: undefined, tier: undefined, note: 'the agent type picks its own model' }
425 const isRouted = routing.model !== undefined
426 const isDryRun = session.config.isDryRun
427 const node: AgentNode = {
428 key: input.toolUseId,
429 subagentType: input.subagentType,
430 parentModel: input.parentModel,
431 isProbe:
432 existing === undefined &&
433 known === undefined &&
434 !input.isPinned &&
435 input.explicitModel === undefined,
436 agentId: undefined,
437 parentAgentId: input.parentAgentId,
438 label,
439 taskType: score.taskType,
440 tier: routing.tier,
441 isRouted,
442 escalatedFrom: existing?.escalatedFrom,
443 pendingTier: undefined,
444 status: isDryRun ? 'planned' : 'running',
445 tokens: existing?.tokens ?? 0,
446 wastedTokens: existing?.wastedTokens ?? 0,
447 answerTokens: 0,
448 attempts: existing?.attempts ?? 0,
449 isBackground: input.isBackground,
450 isDryRun,
451 isHinted: existing?.isHinted ?? session.isHintActive,
452 isRecorded: false,
453 // A retry keeps the baseline of the first attempt: what would have run without tokensaver.
454 baselineRank: existing?.baselineRank ?? rankOfModel(input.explicitModel ?? input.parentModel),
455 }
456 const agents =
457 existing === undefined
458 ? [...session.agents, node].slice(-MAX_AGENTS)
459 : session.agents.map(agent => (agent.key === node.key ? node : agent))
460 const tracked: Session = {
461 ...session,
462 agents,
463 turnKeys: session.turnKeys.includes(node.key) ? session.turnKeys : [...session.turnKeys, node.key],
464 }
465
466 if (!isRouted) {
467 return still(tracked, { model: undefined })
468 }
469
470 const target = routing.tier ?? 'the parent model'
471
472 if (isDryRun) {
473 const text = `tokensaver dry-run: would run "${label}" on ${target} (${score.taskType})`
474
475 return {
476 session: withNote(tracked, text.replace('tokensaver ', '')),
477 effects: [{ kind: 'toast', text }],
478 reply: { model: undefined },
479 }
480 }
481
482 const text =
483 existing?.escalatedFrom === undefined
484 ? `tokensaver: "${label}" → ${target} (${score.taskType})`
485 : `tokensaver: escalated "${label}" ${existing.escalatedFrom} → ${target}`
486
487 return {
488 session: withFlash(
489 withNote(tracked, text.replace('tokensaver: ', '')),
490 existing?.escalatedFrom === undefined
491 ? `${label} → ${target}`
492 : `${label} escalated ${existing.escalatedFrom} → ${target}`,
493 existing?.escalatedFrom === undefined ? 'route' : 'escalate',
494 input.now ?? session.now,
495 ),
496 effects: [{ kind: 'toast', text }],
497 reply: { model: routing.model },
498 }
499}
500
501export type SpawnedInput = {
502 toolUseId: string
503 agentId: string | undefined
504 isDenied: boolean
505 /** The model Claude Code resolved for the sub-agent. */
506 resolvedModel?: string | undefined
507}
508
509const isSameModel = (a: string, b: string): boolean => {
510 const rank = rankOfModel(a)
511
512 return a === b || (rank !== undefined && rank === rankOfModel(b))
513}
514
515/**
516 * The spawn settled: remembers the agent's id, and from an untouched first
517 * spawn learns whether its agent type inherits the parent's model. A refused
518 * spawn is dropped.
519 */
520export const onSpawned = (session: Session, input: SpawnedInput): Step<undefined> => {
521 const node = session.agents.find(agent => agent.key === input.toolUseId)
522
523 if (node === undefined) {
524 return still(session, undefined)
525 }
526
527 if (input.isDenied) {
528 return still(replaceAgent(session, { ...node, status: 'failed', isRecorded: true }), undefined)
529 }
530
531 const started: AgentNode = {
532 ...node,
533 agentId: input.agentId,
534 tier: node.isRouted ? node.tier : (tierOfModel(input.resolvedModel) ?? node.tier),
535 }
536
537 if (!node.isProbe || input.resolvedModel === undefined) {
538 return still(replaceAgent(session, started), undefined)
539 }
540
541 const learned: AgentTypeModel = isSameModel(input.resolvedModel, node.parentModel)
542 ? 'inherits'
543 : 'own-model'
544 const note =
545 learned === 'inherits'
546 ? `"${node.subagentType}" agents inherit their model: routing them from now on`
547 : `"${node.subagentType}" agents choose their own model: left alone`
548
549 return still(
550 withNote(
551 {
552 ...replaceAgent(session, started),
553 agentTypes: { ...session.agentTypes, [node.subagentType]: learned },
554 },
555 note,
556 ),
557 undefined,
558 )
559}
560
561/** Books a finished delegation: its outcome, what it teaches, and what it saved. */
562const finalize = (session: Session, node: AgentNode, signal: CheckSignal, now: number): Step<undefined> => {
563 const isSuccess = signal === 'ok'
564 const usefulTokens = Math.max(0, node.tokens - node.wastedTokens)
565 // A dry run changed nothing, so it saved nothing.
566 const routed =
567 node.isRouted && !node.isDryRun && node.tier !== undefined && node.baselineRank !== undefined
568 ? routingSavings(usefulTokens, rankOf(node.tier), node.baselineRank)
569 : 0
570 const kept =
571 node.isHinted && !node.isDryRun && isSuccess
572 ? contextSavings(usefulTokens, node.answerTokens)
573 : 0
574 const done: AgentNode = { ...node, status: isSuccess ? 'done' : 'failed', isRecorded: true }
575 const settled: Session = {
576 ...withFlash(
577 replaceAgent(session, done),
578 `${node.label} ${isSuccess ? 'done' : `failed (${signalText(signal)})`} · ${formatTokens(node.tokens)}${node.tier === undefined ? '' : ` on ${node.tier}`}`,
579 isSuccess ? 'done' : 'fail',
580 now,
581 ),
582 // Two different savings, kept apart: adding them would count the same tokens twice.
583 saved: session.saved + routed - node.wastedTokens,
584 keptOut: session.keptOut + kept,
585 }
586
587 // The log records tokensaver's own decisions; a model it did not pick teaches it nothing.
588 if (!node.isRouted || node.tier === undefined || node.isDryRun) {
589 return still(settled, undefined)
590 }
591
592 const outcome: Outcome = {
593 seq: nextSeq(session.outcomes),
594 at: now,
595 taskType: node.taskType,
596 tier: node.tier,
597 tokens: node.tokens,
598 retries: node.attempts,
599 wasEscalated: node.escalatedFrom !== undefined,
600 isSuccess,
601 signal,
602 }
603 const outcomes = appendOutcome(session.outcomes, outcome)
604
605 return {
606 session: {
607 ...settled,
608 outcomes,
609 thresholds: learnFromOutcome(session.thresholds, outcomes, outcome),
610 },
611 effects: [{ kind: 'save-log' }],
612 reply: undefined,
613 }
614}
615
616export type AgentResultInput = {
617 toolUseId: string
618 isDenied: boolean
619 isError: boolean
620 text: string
621 /** The Agent tool's own token total, used when no step usage was seen. */
622 totalTokens: number | undefined
623 now: number
624}
625
626/**
627 * An Agent call returned. A result that fails its checks on a tier
628 * tokensaver picked is retried once, one tier up; anything else is final.
629 */
630export const onAgentResult = (
631 session: Session,
632 input: AgentResultInput,
633): Step<{ retryWith: Tier | undefined }> => {
634 const found = session.agents.find(agent => agent.key === input.toolUseId)
635
636 if (!session.config.isEnabled || found === undefined || found.isRecorded) {
637 return still(session, { retryWith: undefined })
638 }
639
640 // A background agent's call returns as it starts; its turn.complete settles it.
641 if (found.isBackground && !input.isDenied && !input.isError) {
642 return still(session, { retryWith: undefined })
643 }
644
645 const node: AgentNode =
646 found.tokens === 0 && input.totalTokens !== undefined
647 ? { ...found, tokens: input.totalTokens }
648 : found
649
650 // A denial is a permission decision, not a quality signal: no retry, no outcome.
651 if (input.isDenied) {
652 return still(replaceAgent(session, { ...node, status: 'failed', isRecorded: true }), {
653 retryWith: undefined,
654 })
655 }
656
657 const signal = checkResult({ isError: input.isError, text: input.text })
658 const plan = planEscalation({
659 signal,
660 tier: node.isRouted ? node.tier : undefined,
661 attempts: node.attempts,
662 isDenied: input.isDenied,
663 isDryRun: node.isDryRun,
664 })
665
666 if (!plan.shouldEscalate || node.tier === undefined) {
667 return { ...finalize(session, node, signal, input.now), reply: { retryWith: undefined } }
668 }
669
670 const retrying: AgentNode = {
671 ...node,
672 status: 'retrying',
673 escalatedFrom: node.tier,
674 pendingTier: plan.to,
675 attempts: node.attempts + 1,
676 wastedTokens: node.tokens,
677 }
678 const text = `tokensaver: "${node.label}" failed on ${node.tier} (${signalText(signal)}), retrying on ${plan.to}`
679
680 return {
681 session: withNote(
682 {
683 ...withFlash(
684 replaceAgent(session, retrying),
685 `${node.label} failed on ${node.tier}, retrying on ${plan.to}`,
686 'escalate',
687 input.now,
688 ),
689 // Tighten now, not when the retry finishes: the cheaper tier has already failed.
690 thresholds: tighten(session.thresholds, node.taskType, node.tier, nextSeq(session.outcomes)),
691 },
692 text.replace('tokensaver: ', ''),
693 ),
694 effects: [{ kind: 'toast', text }, { kind: 'save-log' }],
695 reply: { retryWith: plan.to },
696 }
697}
698
699export type Usage = {
700 model: string
701 input: number
702 output: number
703 cacheWrite: number
704}
705
706const applyUsage = (session: Session, agentId: string | undefined, usage: Usage): Session => {
707 // Cache reads are left out: they are re-read context, not new work.
708 const tokens = usage.input + usage.output + usage.cacheWrite
709 const bucket = tierOfModel(usage.model) ?? 'other'
710 const spend = { ...session.spend, [bucket]: session.spend[bucket] + tokens }
711
712 return {
713 ...session,
714 spend,
715 agents:
716 agentId === undefined
717 ? session.agents
718 : session.agents.map(agent =>
719 agent.agentId === agentId ? { ...agent, tokens: agent.tokens + tokens } : agent,
720 ),
721 }
722}
723
724export type StepUsageInput = {
725 turnId: string
726 agentId: string | undefined
727 usage: Usage
728 /** Read, Grep and Glob calls the response asked for. */
729 reads?: number | undefined
730}
731
732/** One model request finished: adds its tokens to the meter and to its agent. */
733export const onStepUsage = (session: Session, input: StepUsageInput): Step<undefined> => {
734 if (!session.config.isEnabled) {
735 return still(session, undefined)
736 }
737
738 const stepKey = `${input.turnId}/${input.agentId ?? 'main'}`
739
740 return still(
741 {
742 ...applyUsage(session, input.agentId, input.usage),
743 mainReads: session.mainReads + (input.agentId === undefined ? (input.reads ?? 0) : 0),
744 steppedTurns: session.steppedTurns.includes(stepKey)
745 ? session.steppedTurns
746 : [...session.steppedTurns, stepKey].slice(-MAX_STEPPED_TURNS),
747 },
748 undefined,
749 )
750}
751
752export type TurnCompleteInput = {
753 turnId: string
754 agentId: string | undefined
755 /** The turn's totals; used only when no step of this turn reported usage. */
756 usage: Usage | undefined
757 isFailed: boolean
758 /** True when the user interrupted the turn. */
759 isAborted: boolean
760 answer: string
761 now: number
762}
763
764/** A turn ended: the main loop's closes the books on the turn, a sub-agent's on itself. */
765export const onTurnComplete = (session: Session, input: TurnCompleteInput): Step<undefined> => {
766 if (!session.config.isEnabled) {
767 return still(session, undefined)
768 }
769
770 const stepKey = `${input.turnId}/${input.agentId ?? 'main'}`
771 const counted =
772 input.usage === undefined || session.steppedTurns.includes(stepKey)
773 ? session
774 : applyUsage(session, input.agentId, input.usage)
775
776 if (input.agentId === undefined) {
777 const closed: Session = {
778 ...counted,
779 savedHistory: [...counted.savedHistory, counted.saved].slice(-MAX_HISTORY),
780 }
781
782 if (!closed.isHintActive || input.isAborted) {
783 return still(closed, undefined)
784 }
785
786 // The suggestion is judged by what the turn then did, and the bar moves with the verdict.
787 const verdict = judgeHint({
788 largestAgentTokens: Math.max(
789 0,
790 ...closed.agents.filter(agent => closed.turnKeys.includes(agent.key)).map(agent => agent.tokens),
791 ),
792 mainReads: closed.mainReads,
793 })
794 const minSaving = tuneMinSaving(closed.minSaving, verdict)
795 const judged: Session = { ...closed, isHintActive: false, minSaving }
796
797 if (minSaving === closed.minSaving) {
798 return still(judged, undefined)
799 }
800
801 return {
802 session: withNote(
803 judged,
804 verdict === 'over-fired'
805 ? 'delegation suggestion was not needed: raising the bar'
806 : 'delegation suggestion paid off',
807 ),
808 effects: [{ kind: 'save-log' }],
809 reply: undefined,
810 }
811 }
812
813 const node = counted.agents.find(agent => agent.agentId === input.agentId)
814
815 if (node === undefined || node.isRecorded) {
816 return still(counted, undefined)
817 }
818
819 // An interrupt is the user's choice, not a quality signal: no outcome, and never a retry.
820 if (input.isAborted) {
821 return still(replaceAgent(counted, { ...node, status: 'failed', isRecorded: true }), undefined)
822 }
823
824 const answered: AgentNode = { ...node, answerTokens: estimateTokens(input.answer) }
825
826 // A foreground agent is settled by its Agent call, which can still retry it.
827 if (!node.isBackground) {
828 return still(replaceAgent(counted, answered), undefined)
829 }
830
831 const signal = checkResult({ isError: input.isFailed, text: input.answer })
832
833 return finalize(replaceAgent(counted, answered), answered, signal, input.now)
834}
835
836/** The context window moved: feeds the live meter. */
837export const onMeasure = (
838 session: Session,
839 input: { tokens: number | undefined; window: number },
840): Step<undefined> =>
841 still({ ...session, context: { tokens: input.tokens ?? session.context.tokens, window: input.window } }, undefined)
842
843export const activeCount = (session: Session): number =>
844 session.agents.filter(agent => agent.status === 'running' || agent.status === 'retrying').length
845
846/** True while an event is recent enough to still be shown above the prompt. */
847export const hasFreshFlash = (session: Session): boolean =>
848 session.flash !== undefined && session.now - session.flash.at < FLASH_MS
849
850/**
851 * True while the animation clock is needed. A fresh event keeps it running
852 * even under reduced motion, because the clock is also what takes the event
853 * down again; reduced motion stops only the movement.
854 */
855export const isAnimating = (session: Session): boolean =>
856 session.config.isEnabled &&
857 (hasFreshFlash(session) ||
858 (!session.config.hasReducedMotion &&
859 (activeCount(session) > 0 || session.shownSaved !== session.saved)))
860
861/** One animation frame: turns the spinner and eases the saved counter toward its value. */
862export const onTick = (session: Session, now = session.now): Step<undefined> => {
863 if (session.config.hasReducedMotion) {
864 return still({ ...session, now, shownSaved: session.saved }, undefined)
865 }
866
867 const gap = session.saved - session.shownSaved
868 const move = Math.abs(gap) <= 3 ? gap : Math.trunc(gap / 3)
869
870 return still(
871 { ...session, now, frame: session.frame + 1, shownSaved: session.shownSaved + move },
872 undefined,
873 )
874}
875
876const modeOf = (config: Config): string => {
877 if (!config.isEnabled) {
878 return 'off'
879 }
880
881 return config.isDryRun ? 'dry-run' : 'on'
882}
883
884const statusReport = (session: Session): string => {
885 const summary = summarizeLog(session.outcomes)
886 const total = session.spend.haiku + session.spend.sonnet + session.spend.opus + session.spend.other
887
888 return [
889 `tokensaver is ${modeOf(session.config)}`,
890 `${formatTokens(total)} tokens this session (haiku ${formatTokens(session.spend.haiku)}, sonnet ${formatTokens(session.spend.sonnet)}, opus ${formatTokens(session.spend.opus)})`,
891 `~${formatTokens(session.saved)} saved by routing (estimate), ${formatTokens(session.keptOut)} kept out of the main context`,
892 `${String(summary.count)} logged outcomes, ${String(Math.round(summary.cleanRate * 100))}% clean, ${String(summary.escalations)} escalations`,
893 ].join('\n')
894}
895
896/** What tokensaver knows and what the display told it: the first thing to read when nothing shows. */
897const doctorReport = (session: Session): string => {
898 const types = Object.entries(session.agentTypes)
899 .map(([name, kind]) => `${name} ${kind === 'inherits' ? 'inherits' : 'has its own model'}`)
900 .join(', ')
901
902 return [
903 `tokensaver is ${modeOf(session.config)}`,
904 `session model: ${session.facts.model ?? 'not reported yet'}`,
905 `repository: ${session.facts.repoFiles === undefined ? 'not counted yet' : `about ${String(session.facts.repoFiles)} files`}`,
906 `agent types seen: ${types === '' ? 'none yet' : types}`,
907 `a delegation suggestion must save at least ${formatTokens(session.minSaving)} tokens`,
908 `colour ${session.hasNoColor ? 'off (NO_COLOR)' : 'on'}, motion ${session.config.hasReducedMotion ? 'off' : 'on'}`,
909 'display:',
910 ...(session.diagnostics.length === 0
911 ? [' nothing reported yet']
912 : session.diagnostics.map(line => ` ${line}`)),
913 ].join('\n')
914}
915
916export type CommandReply = {
917 /** What the command prints. A command's output is a transcript row, so it is kept short. */
918 text: string
919}
920
921const answer = (session: Session, text: string, effects: readonly Effect[] = []): Step<CommandReply> => ({
922 session,
923 effects,
924 reply: { text },
925})
926
927const withOverride = (session: Session, overrides: Overrides, text: string): Step<CommandReply> => {
928 const merged = { ...session.overrides, ...overrides }
929
930 return answer(
931 { ...session, overrides: merged, config: applyOverrides(session.config, merged) },
932 text,
933 [{ kind: 'save-settings' }],
934 )
935}
936
937/**
938 * `/tokensaver <args>`. Every command answers with a line of text, so it is
939 * visible on every surface even where a pane or a toast is not.
940 */
941export const onCommand = (session: Session, args: string, now: number): Step<CommandReply> => {
942 const command = parseCommand(args)
943
944 switch (command.kind) {
945 case 'status':
946 return answer(session, statusReport(session))
947 case 'help':
948 return answer(session, HELP_TEXT)
949 case 'doctor':
950 return answer(session, doctorReport(session))
951 case 'dashboard':
952 return answer(session, 'tokensaver dashboard opened', [{ kind: 'open-dashboard' }])
953 case 'enable':
954 return withOverride(
955 session,
956 { isEnabled: command.value },
957 `tokensaver is ${command.value ? 'on' : 'off'}`,
958 )
959 case 'dry-run': {
960 const isDryRun = command.value ?? !session.config.isDryRun
961
962 return withOverride(session, { isDryRun }, `tokensaver dry-run is ${isDryRun ? 'on' : 'off'}`)
963 }
964 case 'motion': {
965 const hasMotion = command.value ?? session.config.hasReducedMotion
966
967 return withOverride(
968 { ...session, shownSaved: session.saved },
969 { hasReducedMotion: !hasMotion },
970 `tokensaver animations are ${hasMotion ? 'on' : 'off'}`,
971 )
972 }
973 case 'reset':
974 return answer(
975 {
976 ...session,
977 outcomes: [],
978 thresholds: DEFAULT_THRESHOLDS,
979 minSaving: DEFAULT_MIN_SAVING,
980 saved: 0,
981 keptOut: 0,
982 shownSaved: 0,
983 savedHistory: [],
984 },
985 'tokensaver: outcome log and everything learned from it cleared',
986 [{ kind: 'clear-log' }],
987 )
988 case 'export': {
989 const path = command.path ?? DEFAULT_EXPORT_PATH
990 const json = JSON.stringify(buildExport(session.outcomes, session.thresholds, now), null, 2)
991
992 return answer(
993 session,
994 `tokensaver: wrote ${String(session.outcomes.length)} outcomes to ${path} (local file, nothing was sent anywhere)`,
995 [{ kind: 'export', path, json }],
996 )
997 }
998 case 'unknown':
999 return answer(session, `tokensaver: unknown command "${command.word}"\n${HELP_TEXT}`)
1000 }
1001}
1002src/ui/band.ts 226 lines1/**
2 * The strip above the prompt: who is working for whom, right now. It shows
3 * while a sub-agent runs and for a few seconds after an event, and is gone
4 * the rest of the time. Text for the terminal, an animated vector for
5 * surfaces that draw one.
6 */
7import type { Tier } from '../core/tiers.ts'
8import type { FlashTone } from '../runtime/session.ts'
9import { fit, formatTokens, spinner } from './format.ts'
10import { span } from './theme.ts'
11import type { Line, Tone } from './theme.ts'
12import { themeOf } from './view.ts'
13import type { View, ViewAgent } from './view.ts'
14
15const TIER_TONES: Readonly<Record<Tier, Tone>> = {
16 haiku: 'success',
17 sonnet: 'suggestion',
18 opus: 'claude',
19}
20
21const FLASH_TONES: Readonly<Record<FlashTone, Tone>> = {
22 route: 'success',
23 escalate: 'warning',
24 done: 'success',
25 fail: 'error',
26 info: 'suggestion',
27}
28
29const FLASH_GLYPHS: Readonly<Record<FlashTone, string>> = {
30 route: '→',
31 escalate: '↑',
32 done: '✓',
33 fail: '✗',
34 info: '·',
35}
36
37/** Colours that read on a light and on a dark background alike. */
38const HEX: Readonly<Record<Tone, string>> = {
39 success: '#3fa55b',
40 suggestion: '#7b83eb',
41 claude: '#d77757',
42 warning: '#d99a00',
43 error: '#e5484d',
44 inactive: '#8b8b8b',
45}
46
47const MAX_SHOWN = 3
48const LABEL_CHARS = 22
49
50/** True when the strip has something to say. */
51export const isBandVisible = (view: View): boolean =>
52 view.mode !== 'off' && (view.band.agents.length > 0 || view.band.flash !== undefined)
53
54const tierName = (agent: ViewAgent): string => agent.tier ?? 'inherit'
55
56/** The strip as one line of text. */
57export const bandLine = (view: View, columns: number): Line => {
58 const theme = themeOf(view)
59 const { agents, flash } = view.band
60 const isActive = agents.length > 0
61 const parts = [
62 span(theme, `${isActive ? spinner(view.frame, theme.hasMotion) : '●'} `, { tone: 'claude' }),
63 span(theme, 'tokensaver ', { isBold: true }),
64 ]
65
66 if (isActive) {
67 parts.push(span(theme, 'main', { isDim: true }))
68
69 for (const agent of agents.slice(0, MAX_SHOWN)) {
70 parts.push(
71 span(theme, ' ─▸ ', { isDim: true }),
72 span(theme, tierName(agent), agent.tier === undefined ? {} : { tone: TIER_TONES[agent.tier] }),
73 span(theme, ` ${fit(agent.label, LABEL_CHARS)}`),
74 )
75 }
76
77 if (agents.length > MAX_SHOWN) {
78 parts.push(span(theme, ` +${String(agents.length - MAX_SHOWN)}`, { isDim: true }))
79 }
80 } else if (flash !== undefined) {
81 parts.push(
82 span(theme, `${FLASH_GLYPHS[flash.tone]} `, { tone: FLASH_TONES[flash.tone] }),
83 span(theme, fit(flash.text, Math.max(20, columns - 16))),
84 )
85 }
86
87 return parts
88}
89
90const escapeXml = (text: string): string =>
91 text.replace(/&/g, '&').replace(/</g, '<').replace(/>/g, '>').replace(/"/g, '"')
92
93const STYLE =
94 '<style>' +
95 'text{font-family:ui-sans-serif,system-ui,-apple-system,"Segoe UI",sans-serif;font-size:11.5px}' +
96 '.ink{fill:#2b2a27}.dim{fill:#77756f}.rail{stroke:#77756f}' +
97 '@media (prefers-color-scheme: dark){.ink{fill:#ecebe6}.dim{fill:#9c9a92}.rail{stroke:#9c9a92}}' +
98 '</style>'
99
100/** Approximate width of proportional text at the strip's font size. */
101const textWidth = (text: string): number => Math.ceil(text.length * 6.1)
102
103const HEIGHT = 30
104const MID = HEIGHT / 2
105
106/**
107 * The strip as an animated SVG picture: the main loop on the left, a dashed
108 * line flowing out to each running sub-agent, each agent a pill in its
109 * tier's colour with a pulsing dot. Under reduced motion nothing moves, and
110 * under NO_COLOR everything is drawn in the text colour.
111 */
112export const bandSvg = (view: View, width: number): string => {
113 const theme = themeOf(view)
114 const { agents, flash } = view.band
115 const colour = (tone: Tone): string => (theme.hasColor ? `fill="${HEX[tone]}"` : 'class="ink"')
116 const stroke = (tone: Tone): string => (theme.hasColor ? `stroke="${HEX[tone]}"` : 'class="rail"')
117 const moving = (markup: string): string => (theme.hasMotion ? markup : '')
118 const parts: string[] = [STYLE]
119 const isActive = agents.length > 0
120
121 // The main loop.
122 parts.push(
123 `<circle cx="12" cy="${String(MID)}" r="9" ${colour('claude')} opacity="0.18">${moving(
124 isActive
125 ? '<animate attributeName="r" values="6;11;6" dur="1.6s" repeatCount="indefinite"/><animate attributeName="opacity" values="0.3;0.05;0.3" dur="1.6s" repeatCount="indefinite"/>'
126 : '',
127 )}</circle>`,
128 `<circle cx="12" cy="${String(MID)}" r="4.5" ${colour('claude')}/>`,
129 `<text x="24" y="${String(MID + 4)}" class="ink" font-weight="600">tokensaver</text>`,
130 )
131
132 let x = 24 + textWidth('tokensaver') + 14
133
134 if (isActive) {
135 agents.slice(0, MAX_SHOWN).forEach((agent, index) => {
136 const tone: Tone = agent.status === 'retrying' ? 'warning' : agent.tier === undefined ? 'inactive' : TIER_TONES[agent.tier]
137 const label = `${tierName(agent)} · ${fit(agent.label, LABEL_CHARS)}`
138 const pill = textWidth(label) + 30
139 const rail = 26
140 const delay = (index * 0.25).toFixed(2)
141
142 // The line the work flows along.
143 parts.push(
144 `<line x1="${String(x)}" y1="${String(MID)}" x2="${String(x + rail)}" y2="${String(MID)}" ${stroke(tone)} stroke-width="1.6" stroke-dasharray="4 4" stroke-linecap="round" fill="none">${moving(
145 `<animate attributeName="stroke-dashoffset" from="8" to="0" dur="0.6s" begin="${delay}s" repeatCount="indefinite"/>`,
146 )}</line>`,
147 )
148 x += rail + 4
149
150 // The sub-agent.
151 parts.push(
152 `<rect x="${String(x)}" y="5" width="${String(pill)}" height="20" rx="10" ${colour(tone)} opacity="0.16"/>`,
153 `<rect x="${String(x)}" y="5" width="${String(pill)}" height="20" rx="10" fill="none" ${stroke(tone)} stroke-width="1"/>`,
154 `<circle cx="${String(x + 12)}" cy="${String(MID)}" r="3.5" ${colour(tone)}>${moving(
155 `<animate attributeName="opacity" values="1;0.25;1" dur="1.1s" begin="${delay}s" repeatCount="indefinite"/>`,
156 )}</circle>`,
157 `<text x="${String(x + 21)}" y="${String(MID + 4)}" class="ink">${escapeXml(label)}</text>`,
158 )
159 x += pill + 8
160 })
161
162 if (agents.length > MAX_SHOWN) {
163 parts.push(
164 `<text x="${String(x + 2)}" y="${String(MID + 4)}" class="dim">+${String(agents.length - MAX_SHOWN)}</text>`,
165 )
166 x += 30
167 }
168 } else if (flash !== undefined) {
169 const tone = FLASH_TONES[flash.tone]
170 const text = `${FLASH_GLYPHS[flash.tone]} ${fit(flash.text, 70)}`
171
172 parts.push(
173 `<text x="${String(x)}" y="${String(MID + 4)}" ${colour(tone)}>${escapeXml(text)}${moving(
174 '<animate attributeName="opacity" values="0;1" dur="0.35s" fill="freeze"/>',
175 )}</text>`,
176 )
177 x += textWidth(text) + 12
178 }
179
180 // The running total, where there is room for it.
181 const total = `saved ~${formatTokens(view.hasReducedMotion ? view.saved : view.shownSaved)}`
182
183 if (x + textWidth(total) + 8 <= width) {
184 parts.push(
185 `<text x="${String(width - 6)}" y="${String(MID + 4)}" text-anchor="end" class="dim">${escapeXml(total)}</text>`,
186 )
187 }
188
189 return `<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 ${String(width)} ${String(HEIGHT)}" width="${String(width)}" height="${String(HEIGHT)}">${parts.join('')}</svg>`
190}
191
192export const BAND_HEIGHT = HEIGHT
193
194/** What to add after the spinner's word while sub-agents run, or undefined to leave it alone. */
195export const spinnerSuffix = (view: View): string | undefined => {
196 const { agents } = view.band
197
198 if (view.mode === 'off' || agents.length === 0) {
199 return undefined
200 }
201
202 const tiers = [...new Set(agents.map(tierName))].join(', ')
203
204 return ` · ${String(agents.length)} sub-agent${agents.length === 1 ? '' : 's'} on ${tiers}`
205}
206
207/** A one-line badge for an Agent call tokensaver routed, or undefined when it did not. */
208export const routedBadge = (view: View, toolUseId: string): Line | undefined => {
209 const agent = view.routed.find(entry => entry.key === toolUseId)
210
211 if (agent?.tier === undefined) {
212 return undefined
213 }
214
215 const theme = themeOf(view)
216
217 return [
218 span(theme, '◆ tokensaver ', { isDim: true }),
219 span(theme, agent.escalatedFrom === undefined ? '→ ' : `${agent.escalatedFrom} ↑ `, {
220 tone: agent.escalatedFrom === undefined ? TIER_TONES[agent.tier] : 'warning',
221 }),
222 span(theme, agent.tier, { tone: TIER_TONES[agent.tier], isBold: true }),
223 span(theme, ` (${agent.taskType})`, { isDim: true }),
224 ]
225}
226src/ui/dashboard.ts 204 lines1import type { Tier } from '../core/tiers.ts'
2import type { AgentStatus } from '../runtime/session.ts'
3import { bar, fit, formatTokens, padEnd, padStart, sparkline, spinner } from './format.ts'
4import { savedOnScreen } from './statusline.ts'
5import { span } from './theme.ts'
6import type { Line, Theme, Tone } from './theme.ts'
7import { themeOf } from './view.ts'
8import type { View, ViewAgent } from './view.ts'
9
10const MIN_COLUMNS = 34
11const MAX_COLUMNS = 72
12const LABEL_WIDTH = 9
13
14const TIER_TONES: Readonly<Record<Tier, Tone>> = {
15 haiku: 'success',
16 sonnet: 'suggestion',
17 opus: 'claude',
18}
19
20const STATUS_TONES: Readonly<Record<AgentStatus, Tone>> = {
21 running: 'claude',
22 retrying: 'warning',
23 done: 'success',
24 failed: 'error',
25 planned: 'inactive',
26}
27
28const STATUS_GLYPHS: Readonly<Record<Exclude<AgentStatus, 'running'>, string>> = {
29 retrying: '↑',
30 done: '✓',
31 failed: '✗',
32 planned: '○',
33}
34
35const MODE_TEXT: Readonly<Record<View['mode'], { glyph: string; word: string; tone: Tone }>> = {
36 on: { glyph: '●', word: 'on', tone: 'success' },
37 off: { glyph: '○', word: 'off', tone: 'inactive' },
38 'dry-run': { glyph: '◌', word: 'dry-run', tone: 'warning' },
39}
40
41const widthOf = (columns: number): number =>
42 Math.min(MAX_COLUMNS, Math.max(MIN_COLUMNS, Math.floor(columns)))
43
44const heading = (theme: Theme, text: string): Line => [span(theme, text, { isBold: true })]
45
46const headerLine = (view: View, theme: Theme): Line => {
47 const mode = MODE_TEXT[view.mode]
48
49 return [
50 span(theme, 'tokensaver ', { isBold: true }),
51 span(theme, `${mode.glyph} ${mode.word}`, { tone: mode.tone }),
52 ]
53}
54
55const contextLine = (view: View, theme: Theme, barWidth: number): Line => {
56 const { tokens, window } = view.context
57 const fraction = window > 0 ? tokens / window : 0
58 const tone: Tone = fraction >= 0.85 ? 'error' : fraction >= 0.6 ? 'warning' : 'success'
59 const figure =
60 window > 0
61 ? `${formatTokens(tokens)} / ${formatTokens(window)} ${String(Math.round(fraction * 100))}%`
62 : 'not measured yet'
63
64 return [
65 span(theme, padEnd('Context', LABEL_WIDTH)),
66 span(theme, bar(fraction, barWidth), { tone }),
67 span(theme, ` ${figure}`, { isDim: true }),
68 ]
69}
70
71const sessionLine = (view: View, theme: Theme): Line => [
72 span(theme, padEnd('Session', LABEL_WIDTH)),
73 span(theme, `${formatTokens(view.total)} tokens`),
74]
75
76const savedLine = (view: View, theme: Theme): Line => [
77 span(theme, padEnd('Saved', LABEL_WIDTH)),
78 span(theme, `~${formatTokens(savedOnScreen(view))}`, { tone: 'success', isBold: true }),
79 span(theme, ' est. by routing', { isDim: true }),
80]
81
82const keptOutLine = (view: View, theme: Theme): Line => [
83 span(theme, padEnd('Kept out', LABEL_WIDTH)),
84 span(theme, formatTokens(view.keptOut), { tone: 'success' }),
85 span(theme, ' of main context', { isDim: true }),
86]
87
88/** The savings history as one text row: the terminal's sparkline. */
89export const sparklineLine = (view: View, columns: number): Line => {
90 const theme = themeOf(view)
91 const width = widthOf(columns) - LABEL_WIDTH
92
93 return [
94 span(theme, padEnd('Trend', LABEL_WIDTH)),
95 view.history.length < 2
96 ? span(theme, 'builds up over turns', { isDim: true })
97 : span(theme, sparkline(view.history, width), { tone: 'success' }),
98 ]
99}
100
101const tierLines = (view: View, theme: Theme, barWidth: number): readonly Line[] => {
102 const rows: readonly (readonly [string, number, Tone | undefined])[] = [
103 ['haiku', view.spend.haiku, TIER_TONES.haiku],
104 ['sonnet', view.spend.sonnet, TIER_TONES.sonnet],
105 ['opus', view.spend.opus, TIER_TONES.opus],
106 ...(view.spend.other > 0 ? [['other', view.spend.other, undefined] as const] : []),
107 ]
108 const peak = Math.max(1, ...rows.map(([, tokens]) => tokens))
109
110 return rows.map(([name, tokens, tone]) => [
111 span(theme, ` ${padEnd(name, LABEL_WIDTH - 2)}`),
112 span(theme, bar(tokens / peak, barWidth), tone === undefined ? {} : { tone }),
113 span(theme, ` ${padStart(formatTokens(tokens), 6)}`, { isDim: true }),
114 ])
115}
116
117const agentLine = (agent: ViewAgent, view: View, theme: Theme, width: number): Line => {
118 const rails = agent.guides.map(hasGuide => (hasGuide ? '│ ' : ' ')).join('')
119 const indent = ` ${rails}${agent.isLast ? '└─ ' : '├─ '}`
120 const glyph =
121 agent.status === 'running'
122 ? spinner(view.frame, theme.hasMotion)
123 : STATUS_GLYPHS[agent.status]
124 const tier = padEnd(agent.tier ?? 'inherit', 7)
125 const tokens = padStart(formatTokens(agent.tokens), 6)
126 const room = Math.max(8, width - indent.length - 2 - tier.length - 1 - tokens.length - 2)
127 const wasOn = agent.escalatedFrom === undefined ? '' : ` (was ${agent.escalatedFrom})`
128 // The escalation mark is added after the label is cut, so a long label never hides it;
129 // where the row is too narrow for the words, an arrow stands for them.
130 const escalation = wasOn === '' || room - wasOn.length >= 10 ? wasOn : ' ↑'
131
132 return [
133 span(theme, indent, { isDim: true }),
134 span(theme, `${glyph} `, { tone: STATUS_TONES[agent.status] }),
135 span(theme, `${tier} `, agent.tier === undefined ? { isDim: true } : { tone: TIER_TONES[agent.tier] }),
136 span(theme, padEnd(`${fit(agent.label, room - escalation.length)}${escalation}`, room)),
137 span(theme, ` ${tokens}`, { isDim: true }),
138 ]
139}
140
141const treeLines = (view: View, theme: Theme, width: number): readonly Line[] =>
142 view.agents.length === 0
143 ? [[span(theme, ' no sub-agents yet', { isDim: true })]]
144 : [
145 [span(theme, ' main', { isDim: true })],
146 ...view.agents.map(agent => agentLine(agent, view, theme, width)),
147 ]
148
149const learningLine = (view: View, theme: Theme, width: number): Line => [
150 span(theme, padEnd('Learning', LABEL_WIDTH)),
151 span(
152 theme,
153 view.learning.count === 0
154 ? 'no outcomes logged yet'
155 : fit(
156 `${String(view.learning.count)} runs · ${String(Math.round(view.learning.cleanRate * 100))}% clean · ${String(view.learning.escalations)} escalated`,
157 width - LABEL_WIDTH,
158 ),
159 { isDim: true },
160 ),
161]
162
163export type Dashboard = {
164 /** Header, meters and the saved counter, drawn above the savings sparkline. */
165 top: readonly Line[]
166 /** Per-tier spend, the delegation tree, learning and recent events, drawn below it. */
167 bottom: readonly Line[]
168}
169
170/**
171 * The dashboard as styled lines. The savings sparkline sits between the two
172 * halves so each surface can draw it its own way: `sparklineLine` as text on
173 * the terminal, `sparklineSvg` as a vector on the desktop.
174 */
175export const renderDashboard = (view: View, columns: number): Dashboard => {
176 const theme = themeOf(view)
177 const width = widthOf(columns)
178 // The widest row is the context meter: label, bar, then up to 21 cells of figures.
179 const barWidth = Math.max(4, Math.min(24, width - LABEL_WIDTH - 21))
180
181 return {
182 top: [
183 headerLine(view, theme),
184 contextLine(view, theme, barWidth),
185 sessionLine(view, theme),
186 savedLine(view, theme),
187 keptOutLine(view, theme),
188 ],
189 bottom: [
190 heading(theme, 'Per tier'),
191 ...tierLines(view, theme, barWidth),
192 heading(theme, 'Delegation'),
193 ...treeLines(view, theme, width),
194 learningLine(view, theme, width),
195 ...(view.notes.length === 0
196 ? []
197 : [
198 heading(theme, 'Recent'),
199 ...view.notes.map((note): Line => [span(theme, ` ${fit(note, width - 2)}`, { isDim: true })]),
200 ]),
201 ],
202 }
203}
204src/ui/format.ts 85 lines1const BLOCKS = ['▁', '▂', '▃', '▄', '▅', '▆', '▇', '█'] as const
2const SPINNER = ['◐', '◓', '◑', '◒'] as const
3
4/** 1234 -> "1.2k", 950 -> "950". A negative count keeps its sign. */
5export const formatTokens = (tokens: number): string => {
6 const size = Math.abs(tokens)
7 const sign = tokens < 0 ? '-' : ''
8
9 if (size >= 1_000_000) {
10 return `${sign}${(size / 1_000_000).toFixed(1)}M`
11 }
12
13 return size >= 1000 ? `${sign}${(size / 1000).toFixed(1)}k` : `${sign}${String(Math.round(size))}`
14}
15
16/** A fixed-width meter: `fraction` of `width` cells filled. */
17export const bar = (fraction: number, width: number): string => {
18 const cells = Math.max(1, Math.round(width))
19 const ratio = Number.isFinite(fraction) ? Math.min(1, Math.max(0, fraction)) : 0
20 const filled = Math.round(ratio * cells)
21
22 return '█'.repeat(filled) + '░'.repeat(cells - filled)
23}
24
25/** The last `width` values as block characters, scaled between their own low and high. */
26export const sparkline = (values: readonly number[], width: number): string => {
27 const shown = values.slice(-Math.max(1, width))
28
29 if (shown.length === 0) {
30 return ''
31 }
32
33 const low = Math.min(...shown)
34 const range = Math.max(...shown) - low
35
36 return shown
37 .map(value => {
38 const level = range === 0 ? 0 : Math.round(((value - low) / range) * (BLOCKS.length - 1))
39
40 return BLOCKS[level] ?? BLOCKS[0]
41 })
42 .join('')
43}
44
45/** The spinner glyph for a frame; one fixed glyph when motion is off. */
46export const spinner = (frame: number, hasMotion: boolean): string =>
47 hasMotion ? (SPINNER[Math.abs(frame) % SPINNER.length] ?? SPINNER[0]) : '◆'
48
49export const padEnd = (text: string, width: number): string =>
50 text.length >= width ? text : text + ' '.repeat(width - text.length)
51
52export const padStart = (text: string, width: number): string =>
53 text.length >= width ? text : ' '.repeat(width - text.length) + text
54
55/** Cuts text to a width, marking the cut. */
56export const fit = (text: string, width: number): string => {
57 if (text.length <= width) {
58 return text
59 }
60
61 return width <= 1 ? text.slice(0, Math.max(0, width)) : `${text.slice(0, width - 1)}…`
62}
63
64/**
65 * The savings history as a standalone SVG polyline, for surfaces that draw
66 * vectors. `currentColor` keeps it readable in any theme and under NO_COLOR.
67 */
68export const sparklineSvg = (values: readonly number[], width = 240, height = 36): string => {
69 // Fewer than two readings draw as a flat line rather than nothing.
70 const points = values.length < 2 ? [values[0] ?? 0, values[0] ?? 0] : values
71 const low = Math.min(...points)
72 const range = Math.max(...points) - low
73 const stepX = width / (points.length - 1)
74 const path = points
75 .map((value, index) => {
76 const x = (index * stepX).toFixed(1)
77 const y = (height - 2 - (range === 0 ? 0 : ((value - low) / range) * (height - 4))).toFixed(1)
78
79 return `${x},${y}`
80 })
81 .join(' ')
82
83 return `<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 ${String(width)} ${String(height)}" width="${String(width)}" height="${String(height)}"><polyline fill="none" stroke="currentColor" stroke-width="2" stroke-linejoin="round" stroke-linecap="round" points="${path}"/></svg>`
84}
85src/ui/statusline.ts 39 lines1import { formatTokens, sparkline, spinner } from './format.ts'
2import { themeOf } from './view.ts'
3import type { View } from './view.ts'
4
5const SPARK_WIDTH = 8
6
7/** The saved figure on screen: mid-ease while animating, final under reduced motion. */
8export const savedOnScreen = (view: View): number =>
9 view.hasReducedMotion ? view.saved : view.shownSaved
10
11/**
12 * The compact status line for the CLI, or undefined to clear it. Plain text:
13 * it carries no colour codes, so NO_COLOR needs nothing more here.
14 */
15export const renderStatusLine = (view: View): string | undefined => {
16 if (view.mode === 'off' || !view.hasStatusLine) {
17 return undefined
18 }
19
20 const theme = themeOf(view)
21 const glyph = view.activeCount > 0 ? spinner(view.frame, theme.hasMotion) : '●'
22 const label = view.mode === 'dry-run' ? 'tokensaver [dry-run]' : 'tokensaver'
23 const parts = [
24 `${glyph} ${label} ${formatTokens(view.total)} tok`,
25 `saved ~${formatTokens(savedOnScreen(view))}`,
26 `H ${formatTokens(view.spend.haiku)} S ${formatTokens(view.spend.sonnet)} O ${formatTokens(view.spend.opus)}`,
27 ]
28
29 if (view.history.length >= 2) {
30 parts.push(sparkline(view.history, SPARK_WIDTH))
31 }
32
33 if (view.activeCount > 0) {
34 parts.push(`${String(view.activeCount)} agent${view.activeCount === 1 ? '' : 's'}`)
35 }
36
37 return parts.join(' · ')
38}
39src/ui/theme.ts 36 lines1/** Colours are named by meaning, as Claude Code's theme keys are, so they follow the user's theme. */
2export type Tone = 'success' | 'warning' | 'error' | 'claude' | 'suggestion' | 'inactive'
3
4/** One run of text with one style. The renderers build these; a surface paints them. */
5export type Span = {
6 text: string
7 tone?: Tone
8 isBold?: boolean
9 isDim?: boolean
10}
11
12export type Line = readonly Span[]
13
14export type Theme = {
15 /** False under NO_COLOR: no span carries a tone, and meaning rests on glyphs and words. */
16 hasColor: boolean
17 /** False under reduced motion: one static frame, and counters show their final value. */
18 hasMotion: boolean
19}
20
21export type SpanStyle = {
22 tone?: Tone
23 isBold?: boolean
24 isDim?: boolean
25}
26
27/** Builds a span for the theme: under NO_COLOR the tone is left out entirely. */
28export const span = (theme: Theme, text: string, style: SpanStyle = {}): Span => ({
29 text,
30 ...(theme.hasColor && style.tone !== undefined ? { tone: style.tone } : {}),
31 ...(style.isBold === true ? { isBold: true } : {}),
32 ...(style.isDim === true ? { isDim: true } : {}),
33})
34
35export const plain = (line: Line): string => line.map(part => part.text).join('')
36src/ui/view.ts 131 lines1import { summarizeLog } from '../core/log.ts'
2import type { TaskType } from '../core/score.ts'
3import type { Tier } from '../core/tiers.ts'
4import { activeCount, hasFreshFlash } from '../runtime/session.ts'
5import type { AgentNode, AgentStatus, FlashTone, Session, Spend } from '../runtime/session.ts'
6import type { Theme } from './theme.ts'
7
8export type ViewAgent = {
9 key: string
10 label: string
11 tier: Tier | undefined
12 taskType: TaskType
13 status: AgentStatus
14 tokens: number
15 escalatedFrom: Tier | undefined
16 /** One entry per ancestor level: true where a guide line still runs down beside this row. */
17 guides: readonly boolean[]
18 isLast: boolean
19}
20
21/** Everything the status line and the dashboard draw, as plain data. */
22export type View = {
23 mode: 'on' | 'off' | 'dry-run'
24 hasNoColor: boolean
25 hasReducedMotion: boolean
26 hasStatusLine: boolean
27 context: { tokens: number; window: number }
28 spend: Spend
29 total: number
30 saved: number
31 shownSaved: number
32 keptOut: number
33 history: readonly number[]
34 agents: readonly ViewAgent[]
35 activeCount: number
36 learning: { count: number; cleanRate: number; escalations: number }
37 notes: readonly string[]
38 frame: number
39 /** The strip above the prompt: the sub-agents running now and the newest event, while fresh. */
40 band: {
41 agents: readonly ViewAgent[]
42 flash: { text: string; tone: FlashTone } | undefined
43 }
44 /** Every sub-agent whose model tokensaver chose, for marking its row in the transcript. */
45 routed: readonly ViewAgent[]
46}
47
48const MAX_DEPTH = 4
49
50/** Orders agents as a tree walk: each agent, then the agents it spawned. */
51const walk = (
52 nodes: readonly AgentNode[],
53 parentAgentId: string | undefined,
54 guides: readonly boolean[],
55): readonly ViewAgent[] => {
56 const known = new Set(nodes.flatMap(node => (node.agentId === undefined ? [] : [node.agentId])))
57 const children = nodes.filter(node =>
58 parentAgentId === undefined
59 ? node.parentAgentId === undefined || !known.has(node.parentAgentId)
60 : node.parentAgentId === parentAgentId,
61 )
62
63 return children.flatMap((node, index) => [
64 {
65 key: node.key,
66 label: node.label,
67 tier: node.tier,
68 taskType: node.taskType,
69 status: node.status,
70 tokens: node.tokens,
71 escalatedFrom: node.escalatedFrom,
72 guides,
73 isLast: index === children.length - 1,
74 },
75 ...(node.agentId === undefined || guides.length >= MAX_DEPTH
76 ? []
77 : walk(nodes, node.agentId, [...guides, index < children.length - 1])),
78 ])
79}
80
81const modeOf = (session: Session): View['mode'] => {
82 if (!session.config.isEnabled) {
83 return 'off'
84 }
85
86 return session.config.isDryRun ? 'dry-run' : 'on'
87}
88
89export const toView = (session: Session): View => {
90 const summary = summarizeLog(session.outcomes)
91 const agents = walk(session.agents, undefined, [])
92
93 return {
94 mode: modeOf(session),
95 hasNoColor: session.hasNoColor,
96 hasReducedMotion: session.config.hasReducedMotion,
97 hasStatusLine: session.config.hasStatusLine,
98 context: session.context,
99 spend: session.spend,
100 total: session.spend.haiku + session.spend.sonnet + session.spend.opus + session.spend.other,
101 saved: session.saved,
102 shownSaved: session.shownSaved,
103 keptOut: session.keptOut,
104 history: session.savedHistory,
105 agents,
106 activeCount: activeCount(session),
107 learning: {
108 count: summary.count,
109 cleanRate: summary.cleanRate,
110 escalations: summary.escalations,
111 },
112 notes: session.notes,
113 frame: session.frame,
114 band: {
115 agents: agents.filter(agent => agent.status === 'running' || agent.status === 'retrying'),
116 flash:
117 session.flash !== undefined && hasFreshFlash(session)
118 ? { text: session.flash.text, tone: session.flash.tone }
119 : undefined,
120 },
121 routed: agents.filter(agent =>
122 session.agents.some(node => node.key === agent.key && node.isRouted && !node.isDryRun),
123 ),
124 }
125}
126
127export const themeOf = (view: View): Theme => ({
128 hasColor: !view.hasNoColor,
129 hasMotion: !view.hasReducedMotion,
130})
131src/core/escalate.ts 106 lines1import { nextTierUp } from './tiers.ts'
2import type { Tier } from './tiers.ts'
3
4export const CHECK_SIGNALS = [
5 'ok',
6 'error',
7 'failing-tests',
8 'low-confidence',
9 'empty',
10 'refusal',
11 'user-rejection',
12] as const
13
14/** What the checks made of a sub-agent's result. Anything but `ok` is a failure. */
15export type CheckSignal = (typeof CHECK_SIGNALS)[number]
16
17export const isCheckSignal = (value: unknown): value is CheckSignal =>
18 CHECK_SIGNALS.some(signal => signal === value)
19
20const FAILING_TESTS_PATTERN =
21 /\b\d+\s+(?:tests?\s+)?fail(?:ed|ing|ures?)\b|\btests?\s+(?:are\s+)?(?:still\s+)?fail(?:ed|ing)\b|\bAssertionError\b|^\s*FAIL\b/im
22
23/**
24 * Explicit statements of doubt or of not finishing. "Found nothing" is not
25 * here on purpose: an empty search result is a valid answer, not a failure.
26 */
27const LOW_CONFIDENCE_PATTERN =
28 /\b(?:i(?:'m| am) not (?:sure|certain|confident)|not confident|i (?:may|might) be wrong|(?:unable|was not able|wasn't able) to (?:complete|finish|determine)|could(?: not|n't) (?:complete|finish|determine)|ran out of (?:time|context|turns)|gave up|incomplete|best guess)\b/i
29
30const REJECTION_PATTERN =
31 /^\s*(?:no\b|nope\b|wrong\b|incorrect\b|that(?:'s| is| was)? (?:wrong|not right|incorrect|not what)|not what i|this is wrong|undo\b|revert\b|try again\b|redo\b|(?:it|that|this) (?:didn'?t|doesn'?t|did not|does not) work|still (?:broken|failing|wrong|not working))/i
32
33export type ResultCheck = {
34 isError: boolean
35 text: string
36}
37
38/** Runs the deterministic checks over a sub-agent's result. */
39export const checkResult = ({ isError, text }: ResultCheck): CheckSignal => {
40 if (isError) {
41 return 'error'
42 }
43
44 if (text.trim() === '') {
45 return 'empty'
46 }
47
48 if (FAILING_TESTS_PATTERN.test(text)) {
49 return 'failing-tests'
50 }
51
52 return LOW_CONFIDENCE_PATTERN.test(text) ? 'low-confidence' : 'ok'
53}
54
55/** True when the user's next prompt opens by rejecting the previous result. */
56export const isUserRejection = (prompt: string): boolean => REJECTION_PATTERN.test(prompt)
57
58export type EscalationInput = {
59 signal: CheckSignal
60 /** The tier the failed attempt ran on; undefined when tokensaver did not pick it. */
61 tier: Tier | undefined
62 /** Retries already spent on this task. */
63 attempts: number
64 isDenied: boolean
65 isDryRun: boolean
66}
67
68export type EscalationPlan =
69 | { shouldEscalate: true; to: Tier }
70 | { shouldEscalate: false; reason: string }
71
72/** The most retries one task gets: escalation is one step up, once. */
73export const MAX_RETRIES = 1
74
75/**
76 * Decides whether a failed result is retried one tier up. The answer is
77 * either a strictly higher tier or no retry at all: it never moves down.
78 */
79export const planEscalation = (input: EscalationInput): EscalationPlan => {
80 if (input.isDenied) {
81 return { shouldEscalate: false, reason: 'the call was denied, and a denial is final' }
82 }
83
84 if (input.signal === 'ok') {
85 return { shouldEscalate: false, reason: 'the result passed its checks' }
86 }
87
88 if (input.isDryRun) {
89 return { shouldEscalate: false, reason: 'dry run' }
90 }
91
92 if (input.tier === undefined) {
93 return { shouldEscalate: false, reason: 'tokensaver did not pick this model' }
94 }
95
96 if (input.attempts >= MAX_RETRIES) {
97 return { shouldEscalate: false, reason: 'already retried once' }
98 }
99
100 const to = nextTierUp(input.tier)
101
102 return to === undefined
103 ? { shouldEscalate: false, reason: 'already on the top tier' }
104 : { shouldEscalate: true, to }
105}
106src/core/route.ts 148 lines1import type { TaskScore, TaskType } from './score.ts'
2import { TOP_RANK, nextTierUp, rankOf, rankOfModel, tierAt } from './tiers.ts'
3import type { Tier } from './tiers.ts'
4
5/** A complexity below `haikuMax` runs on Haiku, below `sonnetMax` on Sonnet, the rest on Opus. */
6export type Cutoffs = {
7 haikuMax: number
8 sonnetMax: number
9 /** The log sequence number these cutoffs were last tuned at; older outcomes are spent evidence. */
10 tunedAtSeq: number
11}
12
13export type Thresholds = Readonly<Record<TaskType, Cutoffs>>
14
15/**
16 * Starting cutoffs. Mechanical work gets generous ones; reasoning and
17 * unknown work get tight ones, so an unrecognised task defaults upward.
18 */
19export const DEFAULT_THRESHOLDS: Thresholds = {
20 search: { haikuMax: 45, sonnetMax: 75, tunedAtSeq: 0 },
21 'bulk-read': { haikuMax: 40, sonnetMax: 75, tunedAtSeq: 0 },
22 summarize: { haikuMax: 35, sonnetMax: 70, tunedAtSeq: 0 },
23 'repetitive-edit': { haikuMax: 30, sonnetMax: 65, tunedAtSeq: 0 },
24 edit: { haikuMax: 15, sonnetMax: 55, tunedAtSeq: 0 },
25 reasoning: { haikuMax: 5, sonnetMax: 35, tunedAtSeq: 0 },
26 general: { haikuMax: 10, sonnetMax: 50, tunedAtSeq: 0 },
27}
28
29export type RouteDecision = {
30 tier: Tier
31 /** The tier the cutoffs alone picked, before any upward adjustment. */
32 base: Tier
33 reasons: readonly string[]
34}
35
36const bump = (tier: Tier): Tier => nextTierUp(tier) ?? tier
37
38/**
39 * Picks a tier for a scored task. Every adjustment after the cutoffs moves
40 * up: a boundary tie, an unsure score, a boost and a risky task all raise.
41 */
42export const routeTier = (score: TaskScore, thresholds: Thresholds, boost = 0): RouteDecision => {
43 const cutoffs = thresholds[score.taskType]
44 const base: Tier =
45 score.complexity < cutoffs.haikuMax
46 ? 'haiku'
47 : score.complexity < cutoffs.sonnetMax
48 ? 'sonnet'
49 : 'opus'
50 const reasons: string[] = [`${score.taskType} at complexity ${String(score.complexity)}`]
51 let tier = base
52
53 if (score.isUnsure) {
54 tier = bump(tier)
55 reasons.push('unsure, so one tier up')
56 }
57
58 for (let step = 0; step < boost; step += 1) {
59 tier = bump(tier)
60 }
61
62 if (boost > 0) {
63 reasons.push('previous result rejected, so one tier up')
64 }
65
66 if (score.isRisky) {
67 tier = 'opus'
68 reasons.push(`risky (${score.risks.join(', ')}), so never below the top tier`)
69 }
70
71 return { tier, base, reasons }
72}
73
74export type SpawnContext = {
75 /** The parent's effective model id: what the sub-agent inherits when no model is set. */
76 parentModel: string
77 /** A model the caller asked for by name, if any. */
78 explicitModel?: string | undefined
79 /** A tier an escalation retry must run on. */
80 forcedTier?: Tier | undefined
81}
82
83export type SpawnRouting = {
84 /** The model to hand `agent.spawn`, or undefined to leave the spawn exactly as it came. */
85 model: string | undefined
86 /** The tier the sub-agent runs on, for the log and the dashboard; undefined when unknown. */
87 tier: Tier | undefined
88 note: string
89}
90
91/** A family above the top routable tier has no tier name of its own. */
92const tierForRank = (rank: number): Tier | undefined =>
93 rank > TOP_RANK ? undefined : tierAt(rank)
94
95/**
96 * Turns a tier decision into the model a spawn should run on.
97 *
98 * Invariants the safety tests hold this to: a risky task never runs below the
99 * parent's tier or below a model the caller named, and an unknown parent
100 * model is never rerouted.
101 */
102export const resolveSpawnModel = (
103 decision: RouteDecision,
104 score: TaskScore,
105 context: SpawnContext,
106): SpawnRouting => {
107 if (context.forcedTier !== undefined) {
108 return { model: context.forcedTier, tier: context.forcedTier, note: 'escalation retry' }
109 }
110
111 const parentRank = rankOfModel(context.parentModel)
112 const explicitRank = rankOfModel(context.explicitModel)
113
114 if (parentRank === undefined) {
115 return { model: undefined, tier: undefined, note: 'unknown parent model, left as is' }
116 }
117
118 if (context.explicitModel !== undefined && explicitRank === undefined) {
119 return { model: undefined, tier: undefined, note: 'caller named an unknown model, left as is' }
120 }
121
122 const decided = rankOf(decision.tier)
123 // The top decision means "the best available": under a parent above Opus that is the parent.
124 const wanted = explicitRank ?? (decision.tier === 'opus' ? Math.max(decided, parentRank) : decided)
125 const floor = score.isRisky ? Math.max(parentRank, explicitRank ?? 0) : 0
126 const final = Math.max(wanted, floor)
127 const current = explicitRank ?? parentRank
128
129 if (final === current) {
130 return {
131 model: undefined,
132 tier: tierForRank(final),
133 note: context.explicitModel === undefined ? 'inherits the parent model' : 'caller chose the model',
134 }
135 }
136
137 // Reaching the parent's rank from a lower named model: name the parent's own model.
138 if (final === parentRank) {
139 return {
140 model: context.parentModel,
141 tier: tierForRank(final),
142 note: 'risky, raised to the parent model',
143 }
144 }
145
146 return { model: tierAt(final), tier: tierAt(final), note: decision.reasons.join('; ') }
147}
148src/core/score.ts 200 lines1/**
2 * Deterministic task scoring. No model call: every signal is a regular
3 * expression or a count over the task text, so scoring costs zero tokens.
4 */
5
6export const TASK_TYPES = [
7 'search',
8 'bulk-read',
9 'repetitive-edit',
10 'summarize',
11 'edit',
12 'reasoning',
13 'general',
14] as const
15
16export type TaskType = (typeof TASK_TYPES)[number]
17
18export const RISK_KINDS = ['security', 'destructive', 'production'] as const
19
20export type RiskKind = (typeof RISK_KINDS)[number]
21
22/** Each feature is normalised to 0..1. */
23export type Features = {
24 scope: number
25 files: number
26 ambiguity: number
27 depth: number
28 risk: number
29}
30
31export type TaskScore = {
32 taskType: TaskType
33 features: Features
34 /** 0..100: how much reasoning the task needs, which is what picks the tier. */
35 complexity: number
36 isRisky: boolean
37 risks: readonly RiskKind[]
38 /** True when the heuristics cannot place the task with confidence. */
39 isUnsure: boolean
40 /** Files the task names, or the count it states. */
41 namedFiles: number
42}
43
44export type TaskInput = {
45 text: string
46 description?: string | undefined
47 subagentType?: string | undefined
48}
49
50const RISK_PATTERNS: Readonly<Record<RiskKind, RegExp>> = {
51 security:
52 /\b(auth(?:entication|orization|n|z)?|passwords?|passphrases?|secrets?|credentials?|(?:access|auth|api|bearer|refresh|session)[- ]tokens?|api[- ]?keys?|private keys?|encrypt\w*|decrypt\w*|oauth|jwt|csrf|xss|sql injection|vulnerab\w*|security|permissions?|sudo|chmod|chown|ssh|certificates?|iam|rbac|cve-\d+)\b|\.env\b/gi,
53 destructive:
54 /\b(rm -rf?|delete\w*|drop (?:table|database|schema|index)|truncate\w*|wipe\w*|purge\w*|destroy\w*|force[- ]push\w*|reset --hard|git clean|overwrit\w*|erase\w*|remove all|uninstall\w*|migrations?|rollback|format (?:the )?(?:disk|drive))\b|push (?:-f\b|--force)/gi,
55 production:
56 /\b(prod(?:uction)?|deploy\w*|releas\w*|live (?:site|system|server|database|db|traffic)|customer data|billing|payments?|invoic\w*|payroll|terraform|kubectl|kubernetes|infra(?:structure)?|ci\/cd|dns|rollout|hotfix)\b/gi,
57}
58
59/** Task types that can be told from the text. `general` is what is left. */
60type KnownType = Exclude<TaskType, 'general'>
61
62const TYPE_PATTERNS: Readonly<Record<KnownType, RegExp>> = {
63 summarize:
64 /\b(summari[sz]\w*|summary|tl;?dr|recap|digest|condense\w*|overview|gist)\b/gi,
65 search:
66 /\b(find|search\w*|grep|locate|look (?:for|up)|where (?:is|are|does|do)|which files?|list (?:all|every)|usages?|references? to|occurrences?|who calls|callers of)\b/gi,
67 'bulk-read':
68 /\b(read (?:all|every|each|through)|go through|scan\w*|inventory|catalogu?e?|survey|review (?:all|every|each)|collect (?:all|every)|gather\w*|crawl\w*)\b/gi,
69 'repetitive-edit':
70 /\b(renam\w*|replace (?:all|every|each)|(?:find|search) and replace|codemod|bulk|mass|in (?:all|every|each) files?|every (?:file|occurrence|instance)|across all files|update (?:all|every|each)|apply the same|convert (?:all|every|each)|reformat\w*|bump (?:all|the) versions?)\b/gi,
71 reasoning:
72 /\b(design\w*|architect\w*|why|root cause|debug\w*|diagnos\w*|race conditions?|deadlocks?|trade-?offs?|algorithms?|prove|optimi[sz]\w*|plan|strategy|investigat\w*|concurrency|refactor\w*|redesign\w*|evaluat\w*|compar\w*|decide|reason about|analy[sz]\w*|performance|memory leaks?)\b/gi,
73 edit: /\b(fix\w*|implement\w*|add|chang\w*|writ\w*|updat\w*|creat\w*|modif\w*|remov\w*|edit\w*|patch\w*|build|make)\b/gi,
74}
75
76/**
77 * Tie-break order: when two types match equally often, the one that routes
78 * to the higher tier wins, so a tie never sends work to a cheaper model.
79 */
80const TYPE_PRIORITY: readonly KnownType[] = [
81 'reasoning',
82 'edit',
83 'repetitive-edit',
84 'summarize',
85 'bulk-read',
86 'search',
87]
88
89const SCOPE_PATTERN =
90 /\b(all|every|each|entire|whole|across|everywhere|codebase|repo(?:sitory)?|project-wide|global(?:ly)?|monorepo|recursive(?:ly)?)\b|\*\*?\/|\*\.\w+/gi
91
92const PATH_PATTERN =
93 /(?:[\w.-]+\/)+[\w.-]+|\b[\w-]+\.(?:tsx?|jsx?|mjs|cjs|py|go|rs|java|rb|md|json|ya?ml|toml|s?css|html|sql|sh|c|cpp|h|cs|php|kt|swift)\b/gi
94
95const FILE_COUNT_PATTERN = /\b(\d{1,4}) files?\b/i
96
97const VAGUE_PATTERN =
98 /\b(somehow|maybe|perhaps|something|stuff|things?|etc|whatever|figure out|not sure|i guess|kind of|sort of|improve\w*|better|clean ?up|tidy|polish|as needed|if needed|appropriate(?:ly)?)\b/gi
99
100const IDENTIFIER_PATTERN = /`[^`]+`|\b[a-z]+[A-Z]\w*\b|\b\w+_\w+\b/g
101
102const STEP_PATTERN = /\b(then|after that|afterwards|finally|next,)\b|^\s*\d+[.)]\s/gim
103
104const clamp01 = (value: number): number => Math.min(1, Math.max(0, value))
105
106const matchesOf = (text: string, pattern: RegExp): readonly string[] =>
107 Array.from(text.matchAll(pattern), match => match[0].toLowerCase())
108
109const countOf = (text: string, pattern: RegExp): number => matchesOf(text, pattern).length
110
111const wordCountOf = (text: string): number => text.split(/\s+/).filter(word => word !== '').length
112
113const classify = (text: string): TaskType => {
114 const ranked = TYPE_PRIORITY.map((type, priority) => ({
115 type,
116 priority,
117 hits: countOf(text, TYPE_PATTERNS[type]),
118 })).sort((a, b) => b.hits - a.hits || a.priority - b.priority)
119
120 const best = ranked[0]
121
122 return best === undefined || best.hits === 0 ? 'general' : best.type
123}
124
125const risksOf = (text: string): readonly RiskKind[] =>
126 RISK_KINDS.filter(kind => countOf(text, RISK_PATTERNS[kind]) > 0)
127
128const namedFilesOf = (text: string): number => {
129 const paths = new Set(matchesOf(text, PATH_PATTERN)).size
130 const stated = Number(FILE_COUNT_PATTERN.exec(text)?.[1] ?? 0)
131
132 return Math.max(paths, stated)
133}
134
135const filesFeature = (count: number, scope: number): number => {
136
137 // A broad task that names no file touches an unknown number of them.
138 if (count === 0 && scope >= 0.34) {
139 return 0.5
140 }
141
142 return clamp01(count / 8)
143}
144
145const ambiguityFeature = (text: string, words: number): number => {
146 const anchors = countOf(text, IDENTIFIER_PATTERN) + new Set(matchesOf(text, PATH_PATTERN)).size
147 const vague = countOf(text, VAGUE_PATTERN) / 3
148 const isTerse = words < 6 && anchors === 0
149 const isConcrete = anchors >= 2
150
151 return clamp01(vague + (isTerse ? 0.4 : 0) - (isConcrete ? 0.2 : 0))
152}
153
154const depthFeature = (text: string, words: number): number => {
155 const reasoning = countOf(text, TYPE_PATTERNS.reasoning) / 3
156 const steps = Math.min(0.2, countOf(text, STEP_PATTERN) * 0.1)
157
158 return clamp01(reasoning + steps + (words > 150 ? 0.2 : 0))
159}
160
161/** Scores a task from its text alone. Pure and deterministic. */
162export const scoreTask = (input: TaskInput): TaskScore => {
163 const text = [input.description ?? '', input.text].join('\n').trim()
164 const words = wordCountOf(text)
165 const scope = clamp01(countOf(text, SCOPE_PATTERN) / 3)
166 const risks = risksOf(text)
167 const namedFiles = namedFilesOf(text)
168 const riskHits = RISK_KINDS.reduce((sum, kind) => sum + countOf(text, RISK_PATTERNS[kind]), 0)
169 const features: Features = {
170 scope,
171 files: filesFeature(namedFiles, scope),
172 ambiguity: ambiguityFeature(text, words),
173 depth: depthFeature(text, words),
174 risk: clamp01(riskHits / 2),
175 }
176 const taskType = classify(text)
177 const complexity = Math.round(
178 100 *
179 clamp01(
180 0.45 * features.depth +
181 0.25 * features.ambiguity +
182 0.15 * features.scope +
183 0.15 * features.files,
184 ),
185 )
186
187 return {
188 taskType,
189 features,
190 complexity,
191 isRisky: risks.length > 0,
192 risks,
193 isUnsure: features.ambiguity >= 0.6 || taskType === 'general' || words < 3,
194 namedFiles,
195 }
196}
197
198export const isTaskType = (value: unknown): value is TaskType =>
199 TASK_TYPES.some(type => type === value)
200