SLOPSHOPPER

tokensaver

Token-efficient delegation: routes sub-agents to the cheapest capable model tier, escalates on failure, learns from outcomes locally, and shows live savings.

newpanebandspinnerrowsguard
★ 2v0.2.1MITupdated 2026-10-09AGregDev/claude-tokensaver
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · tokensaver
│ ┃ tokensaver ✕ › fix the failing auth test and add an audit log call │ ┃ tokensaver ● on │ ┃ Context ████████████░░░░░░░░░░░░ 97.4k / 2 ⏺ Read(src/auth.ts) │ ┃ Session 7.9k tokens ⎿ Read 6 lines │ ┃ Saved ~0 est. by routing ⏺ Update(src/auth.ts) │ ┃ Kept out 0 of main context ⎿ Added 2 lines, removed 1 line │ ┃ Trend builds up over turns ⏺ Bash(bun test) │ ┃ Per tier ⎿ 3 pass, 1 fail │ ┃ haiku ░░░░░░░░░░░░░░░░░░░░░░░░ 0 │ ┃ sonnet ░░░░░░░░░░░░░░░░░░░░░░░░ 0 ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ opus ████████████████████████ 7.9k │ ┃ Delegation ✻ Worked for 42s · done 4:20 PM │ ┃ no sub-agents yet │ ┃ Learning no outcomes logged yet › /tokensaver │ ┃ [ Turn off ] [ Dry-run ] ⎿ tokensaver: tokensaver is on │ ⎿ tokensaver: 7.9k tokens this session (haiku 0, sonnet 0, opus 7. │ ⎿ tokensaver: ~0 saved by routing (estimate), 0 kept out of the ma │ ⎿ tokensaver: 0 logged outcomes, 100% clean, 0 escalations │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ tokensaver: ● tokensaver 7.9k tok · saved ~0 · H 0 S 0 O 7.9k

Draws

Pane · tokensaver
tokensaver ● on Context ████████████░░░░░░░░░░░░ 97.4k / 200.0k 49% Session 7.9k tokens Saved ~0 est. by routing Kept out 0 of main context Trend builds up over turns Per tier haiku ░░░░░░░░░░░░░░░░░░░░░░░░ 0 sonnet ░░░░░░░░░░░░░░░░░░░░░░░░ 0 opus ████████████████████████ 7.9k Delegation no sub-agents yet Learning no outcomes logged yet [ Turn off ] [ Dry-run ]
README

tokensaver

CI License: MIT

A Claude Code mod with one job: spend fewer tokens without lowering the quality of the result.

It scores each task with cheap, deterministic heuristics, sends sub-agents to the cheapest model tier that can do the work, retries one tier up when a result fails its checks, and learns from the outcomes. It works in the Claude Code CLI and in the Code tab of the desktop app.

The strip above the prompt while two sub-agents run

The tokensaver dashboard pane

The tokensaver status line in the CLI

<sub>These pictures are generated by npm run preview: the mod's real renderers, run over a sample session. The first is the animated strip exactly as the desktop app receives it; the other two are the dashboard and status line text painted as SVG. They are not screenshots of a live session.</sub>

Install

In a Claude Code session:

/plugin install tokensaver --marketplace AGregDev/claude-tokensaver

Or from your shell:

claude plugin install tokensaver --marketplace AGregDev/claude-tokensaver

Answer y to add the marketplace and pick a scope. The mod is active at once. Mods need Claude Code v2.1.287 or later; the shell form of the command needs v2.1.292.

A mod is code that runs with your permissions. Before you install any mod, you can list what it hooks and calls without running it: clone the repository and run claude plugin validate . in it. What it can reach has tokensaver's list.

What it does

Routes sub-agents by tier

When Claude starts a sub-agent, tokensaver scores the sub-agent's task and picks its model:

| Tier | Used for | | :- | :- | | Haiku | Simple, mechanical work: searches, bulk reads, repetitive edits | | Sonnet | Medium work, and anything Haiku-sized that the heuristics are unsure about | | Opus | Hard reasoning, architecture, and every risky task |

The score comes from five signals read off the task text with regular expressions and counts: scope (how broad), files touched, ambiguity, reasoning depth, and risk. No model is called to classify anything, so scoring costs no tokens.

Every adjustment after the first pick moves up, never down:

  • A score exactly on a cutoff goes to the higher tier.
  • An unsure score (vague wording, an unrecognised kind of task) goes up one tier.
  • A task that looks security-sensitive, destructive or production-touching always gets the top tier, and is never run below the model it would have had without tokensaver.
  • A model Claude named itself is respected; tokensaver only ever raises it, and only for risky work.

An agent type can name its own model in its definition, and a spawn does not say whether it does. So tokensaver leaves the first spawn of each agent type in a session untouched and watches which model Claude Code resolves for it:

  • If it resolves to the parent's model, the type inherits, and tokensaver routes that type from then on.
  • If it resolves to anything else, the type brings its own model, as a custom agent with model: in its definition does, and tokensaver never reroutes it.

Forks, teammates and workflow agents are left alone, and so is any session whose model family tokensaver does not recognise.

The main conversation's own model is not switched. Changing models mid-session invalidates the prompt cache, which costs more than routing would save.

Knows when delegating pays, and says so

You do not have to ask for sub-agents. For each prompt you type, tokensaver works out whether handing the reading to a sub-agent would cost less than Claude doing it inline, and when it would, it tells Claude to delegate.

The comparison uses measured numbers, not guesses:

  • Starting a sub-agent is expensive. Its own system prompt and tools cost 30k to 38k tokens before it reads anything, measured in live sessions.
  • Reading inline is expensive later. What the main loop reads stays in its context and is re-read on every following request.
  • A cheaper model does the reading. The sub-agent runs on the tier tokensaver routes it to, and the main context keeps only its report.

How much reading a task means is estimated from the files the prompt names and, for a broad task, from the size of the repository, which tokensaver samples once per session (file names only, never contents). So "find every usage across the repo" earns a suggestion in a 600-file repository and none in a 40-file one, where a single search is cheaper than any sub-agent.

When delegating wins by a clear margin, tokensaver adds one line to the prompt: about how many files, and to use one Explore sub-agent or up to three in parallel. Claude makes the final call, because it can see the code and tokensaver cannot. That line is the only thing tokensaver ever adds to the model's context for a prompt, and its cost is subtracted from the savings counter.

tokensaver never starts a sub-agent itself. A sub-agent that a mod starts has not been through the safety review Claude's own actions get, and Claude Code refuses it in auto mode. Delegation stays Claude's own action.

Escalates on failure

When a sub-agent's result fails its checks on a tier tokensaver picked, the same call is retried once on the next tier up, and Claude only sees the retry's result. The checks are:

  • the call errored
  • the answer is empty
  • the answer reports failing tests
  • the answer says it is unsure or did not finish

If your next prompt opens by rejecting the result ("no, that's wrong", "try again"), every sub-agent that turn is routed one tier up.

A denied call is never retried, and neither is a turn you interrupted.

What you see

tokensaver changes how a session looks while work is being delegated:

  • A strip above the prompt shows who is working for whom: the main loop, a line flowing out to each running sub-agent, and each sub-agent as a pill in its tier's colour. In the desktop app it is an animated vector; in the terminal it is a line of text with a spinner. It appears while a sub-agent runs and for a few seconds after an event (routed, escalated, finished), and is gone the rest of the time.
  • The spinner says how many sub-agents are running and on which tier.
  • A sub-agent's row in the transcript is marked with the tier tokensaver chose for it, and with where it came from if it was escalated.
  • A toast appears when a sub-agent is rerouted or escalated.
  • A status line under the prompt carries the session's tokens, the estimated savings, per-tier spend and a savings sparkline.
  • A dashboard pane has the context meter, per-tier spend, the delegation tree and the savings trend. It opens by itself in the desktop app; in the terminal, /tokensaver dashboard opens it. Close it and it stays closed until you ask for it again.

Animation runs only while something is happening, and the timer stops when things settle.

  • NO_COLOR (set to anything non-empty) removes every colour; state is still shown by glyphs and words.
  • The reducedMotion option, or /tokensaver motion off, stops all movement and shows final figures at once.

None of the drawing reaches the model. A /tokensaver command does print one short line into the transcript, which Claude can read on later turns; that is what makes a command visible on every surface.

What the two savings figures mean

Both are estimates, and they are kept apart on purpose because adding them would count the same tokens twice.

  • Saved (by routing): tokens a sub-agent used on a cheaper tier, weighted by relative cost (Haiku 1, Sonnet 3, Opus 5) against the model it would otherwise have run on. A failed attempt before a retry and tokensaver's own hint are subtracted, so the figure can go negative.
  • Kept out of main context: for delegations tokensaver suggested, the tokens the sub-agent read and wrote, less the summary it handed back.

The cost weights are ratios for estimating, not prices. The token meter counts new work (input, output and cache writes) and leaves out cache reads.

Commands

| Command | What it does | | :- | :- | | /tokensaver | Show status | | /tokensaver on / off | Master switch. Off, every event passes through untouched | | /tokensaver dry-run [on\|off] | Only show what would have been delegated or rerouted | | /tokensaver dashboard | Open the dashboard pane | | /tokensaver motion [on\|off] | Animations | | /tokensaver doctor | What tokensaver knows (model, repository size, agent types) and what the display reported. Run this first if nothing shows | | /tokensaver export [path] | Write the outcome log as JSON to a local file (default tokensaver-export.json) | | /tokensaver reset | Clear the outcome log and the learned thresholds | | /tokensaver help | List the commands |

Switches set with a command are remembered across sessions.

The most used ones are also listed in the / menu as /tokensaver:status, /tokensaver:dashboard, /tokensaver:doctor, /tokensaver:on, /tokensaver:off, /tokensaver:dry-run and /tokensaver:help. They do the same thing. They exist because /tokensaver itself is registered when a session starts running, so a brand-new session does not offer it in the menu until you have sent something; you can still type it in full.

Options

Set these when you install, or later with /plugin configure tokensaver@claude-tokensaver.

| Option | Default | What it does | | :- | :- | :- | | enabled | true | Master switch | | dryRun | false | Change nothing; show what would have happened | | contextHint | true | Add the one-line delegation hint to prompts where delegating pays | | statusLine | true | Show the status line under the prompt | | reducedMotion | false | Turn off the spinner and the counting animation |

How self-training works

Each finished delegation that tokensaver routed is appended to a local outcome log: the kind of task, the tier used, tokens, retries, whether it was escalated, and whether it passed its checks. The log holds counts and labels only. No prompt, file name or answer text is stored.

Routing is driven by two cutoffs per kind of task, and the log tunes them:

  • Tightening is immediate. One failure, escalation or rejection lowers the cutoff that let the task onto that tier by 10 points, so fewer tasks of that kind reach it next time.
  • Loosening is slow and conditional. A cutoff rises by 2 points only after at least 12 outcomes for that kind of task since its last change, every one of them a clean first-try success. A single failure in that window blocks it. Evidence is used once, and a cutoff can never drift more than 15 points above its default.
  • Risk is not tunable. No amount of clean history lets a risky task run below the top tier.

The log and the thresholds live in the mod's own store on your machine (Claude Code keeps it under your Claude Code configuration directory). The log is capped at 500 entries. A corrupted log is read entry by entry; whatever does not parse is dropped and the rest is kept.

The bar a delegation suggestion must clear is learned the same lopsided way. When a turn ends, the suggestion made in it is judged by what happened: if Claude declined and then read only a few files, or the sub-agent barely did more than start, the suggestion was not worth making and the bar rises by a quarter at once. If the sub-agent did real work, the bar drops by a few percent.

Nothing is sent anywhere. The mod makes no network call at all. /tokensaver export writes a file on your disk, and /tokensaver reset deletes the log and the thresholds.

Safety guarantees

Each guarantee below is enforced by a test, and the tests are part of CI.

| Guarantee | How it is held | | :- | :- | | Risky, destructive or production-touching work is never downgraded | A property test over every risky prompt, parent model and named model, under default thresholds and under the loosest thresholds learning could ever produce | | Escalation only moves up | A property test over every tier, check result, retry count and flag: the retry is exactly one tier higher or does not happen | | An agent type's own model is never overridden | A property test over every task and parent model, for a type known to bring its own model and for one not yet seen; a test inside Claude Code confirms it against the real agent.spawn event | | Permission decisions are never overridden | The mod has no tool.check hook, so it cannot approve anything; a test inside Claude Code confirms deny, ask and allow come back unchanged, and that a denied call is returned untouched and not retried | | It stacks with other mods | Every hook passes the event on with next, and every hook that could gate an event has a .catch that passes it on unchanged if the hook fails | | It degrades instead of breaking | A failed store, environment or display call never fails the hook that made it, and a test runs the mod with none of them available. When no per-request usage arrives, the meter falls back to turn totals |

One thing it cannot soften: Claude Code checks a mod's event names when it loads the mod. On a build that lacks one of the events above, Claude Code refuses to load tokensaver and says which event, and your session carries on without it. tokensaver is typed and tested against Claude Code v2.1.292.

tokensaver runs after Claude Code's built-in sec-default guard and after any mod your organization lists ahead of user mods. It does nothing to work around either.

What it can reach

This is the complete list, as claude plugin validate reads it from the source. The build fails if the list changes, so it cannot grow without a reviewed edit to scripts/build.mjs.

| It hooks | To | | :- | :- | | session.start | Load the log and settings, register /tokensaver, sample the repository's size | | session.attach | Redraw when an app connects to the session | | prompt.submit | Score the prompt; add the delegation line when it pays | | agent.spawn | Choose the sub-agent's model | | tool.call on Agent only | Read the result; retry once one tier up when it fails | | turn.step, turn.complete, session.measure | Read token usage for the meter, and count how much the main loop reads itself | | command.run for tokensaver | Answer its own command | | ui.render for its own pane, the strip above the prompt, the spinner, and tool rows | Draw the dashboard and the strip; add to the spinner; mark sub-agent rows. Every other tool row is passed on untouched | | ui.close for its own pane | Remember that you closed it |

| It calls | For | | :- | :- | | $.store.get, set, delete | The local log, thresholds and switches | | $.fs.write | /tokensaver export, to the path you give | | $.fs.list | Counting the repository's files once per session: names only, a few folders deep, capped | | $.env.get | Reading NO_COLOR, and nothing else | | $.ui.status, toast, log, open, resolve | Display | | $.state.get, set | Redrawing the pane | | $.clock.every | The animation timer | | $.command.register, $.session.surfaces, $.session.model | Its command; telling the CLI from the desktop app; knowing which model the session runs on |

It does not call a model, start a sub-agent or a process, make a network request, read the contents of your files, or change a permission.

Limits worth knowing

  • The heuristics read text. They can misjudge a task, which is why every unsure case goes up a tier and why risk keywords are deliberately broad. A prompt that mentions "delete" or "production" in passing is treated as risky and gets no cheaper model.
  • The first sub-agent of each agent type in a session is never rerouted, because that spawn is how tokensaver learns whether the type inherits its model. A session with a single sub-agent saves nothing by routing.
  • The delegation estimate is built from the prompt's wording and the repository's size, not from the code. Claude can still decline a suggestion, and when it does the cost of the line (about 80 tokens) shows as a small negative saving and the bar for the next suggestion rises.
  • Delegation only happens when Claude acts on the suggestion. tokensaver cannot make it happen, by design.
  • "Failing tests" and "unsure" are detected from the sub-agent's answer text, so a failure the sub-agent does not report is not caught.
  • A background sub-agent's result cannot be retried, because Claude has already moved on. Its outcome is still logged and still tightens the thresholds.
  • The savings figures are estimates built on the weights above, not a bill.

Development

git clone https://github.com/AGregDev/claude-tokensaver
cd claude-tokensaver
npm install
npm run types:sync
npm run check

npm run check runs the typecheck, lint, the unit, safety, integration and UI tests with an 80% coverage threshold, Claude Code's validator, the tests that run inside Claude Code, and the build. CONTRIBUTING.md has the layout and the details.

To try a working copy in a session:

claude --plugin-dir .

License

MIT

Source 19 files
hooks/register.tsx 628 lines
1/**
2 * tokensaver's hooks module: the thin layer between Claude Code's events and
3 * the pure session controller in ../src. Every hook here observes, routes or
4 * draws; none approves a tool call, and there is no `tool.check` hook at all,
5 * so the mod cannot override a permission rule. It never starts a sub-agent
6 * itself either: delegation is always Claude's own, reviewed, action.
7 */
8import { atom, read, update } from 'claude-code'
9import type { EngineInterface, Register, Timer, ToolCallResult, TurnUsage } from 'claude-code'
10
11import { isRecord } from '../src/core/log.ts'
12import {
13  hydrate,
14  isAnimating,
15  onAgentResult,
16  onCommand,
17  onDiagnostic,
18  onFacts,
19  onMeasure,
20  onPrompt,
21  onSpawn,
22  onSpawned,
23  onStepUsage,
24  onTick,
25  onTurnComplete,
26} from '../src/runtime/session.ts'
27import type { AgentResultInput, Effect, Session, Step, Usage } from '../src/runtime/session.ts'
28import { BAND_HEIGHT, bandLine, bandSvg, isBandVisible, routedBadge, spinnerSuffix } from '../src/ui/band.ts'
29import { renderDashboard, sparklineLine } from '../src/ui/dashboard.ts'
30import { sparklineSvg } from '../src/ui/format.ts'
31import { renderStatusLine } from '../src/ui/statusline.ts'
32import type { Line, Span, Tone } from '../src/ui/theme.ts'
33import { toView } from '../src/ui/view.ts'
34
35const NAME = 'tokensaver'
36const TICK_MS = 250
37const STORE_LOG = 'log'
38const STORE_THRESHOLDS = 'thresholds'
39const STORE_SETTINGS = 'settings'
40const STORE_DELEGATION = 'delegation'
41
42/** Tools that read the repository; counted to see how much the main loop reads for itself. */
43const READING_TOOLS = ['Read', 'Grep', 'Glob']
44
45/** Folders that hold no source of the project's own. */
46const SKIPPED_FOLDERS = [
47  'node_modules',
48  '.git',
49  'dist',
50  'build',
51  'out',
52  'coverage',
53  'target',
54  'vendor',
55  '.next',
56  '.venv',
57  '__pycache__',
58]
59const MAX_LISTINGS = 40
60const MAX_COUNTED = 2000
61
62/** About how wide one text cell is in the desktop app, to size a picture to its slot. */
63const CELL_PX = 7.4
64
65const revision = atom({ plugin: 'tokensaver', key: 'revision' } as const, 0)
66
67type TextStyle = { color?: Tone; bold?: true; dimColor?: true }
68
69const styleOf = (part: Span): TextStyle => ({
70  ...(part.tone === undefined ? {} : { color: part.tone }),
71  ...(part.isBold === true ? { bold: true } : {}),
72  ...(part.isDim === true ? { dimColor: true } : {}),
73})
74
75const usageOf = (usage: TurnUsage): Usage => ({
76  model: usage.model,
77  input: usage.input_tokens,
78  output: usage.output_tokens,
79  cacheWrite: usage.cache_creation_input_tokens,
80})
81
82/** The Agent tool's own token total. A denial, an error and a background launch carry none. */
83const totalTokensOf = (result: unknown): number | undefined => {
84  const total = isRecord(result) ? result['totalTokens'] : undefined
85
86  return typeof total === 'number' ? total : undefined
87}
88
89const resultOf = (toolUseId: string, result: ToolCallResult<'Agent'>): AgentResultInput => ({
90  toolUseId,
91  isDenied: result.deny !== undefined,
92  isError: result.isError === true,
93  text: result.text ?? '',
94  totalTokens: totalTokensOf(result.result),
95  now: Date.now(),
96})
97
98/** A call that falls back instead of failing its hook when it is not available. */
99const attempt = async <Value,>(call: () => Promise<Value>, fallback: Value): Promise<Value> => {
100  try {
101    return await call()
102  } catch {
103    return fallback
104  }
105}
106
107const reasonOf = (error: unknown): string => (error instanceof Error ? error.message : String(error))
108
109// Module state: a session's worth of bookkeeping. A reload starts it over, and
110// `session.start` then rebuilds it from the store.
111let session: Session = hydrate({
112  options: {},
113  storedLog: undefined,
114  storedThresholds: undefined,
115  storedSettings: undefined,
116  noColor: undefined,
117}).session
118let timer: Timer | undefined
119let shownStatus: string | undefined
120let isPaneOpen = false
121let wasPaneClosedByPerson = false
122
123/** Opens the dashboard pane and keeps what Claude Code answered, for `/tokensaver doctor`. */
124const openPane = async ($: EngineInterface, why: string): Promise<void> => {
125  try {
126    const opened = await $.ui.open({ id: NAME, title: NAME })
127
128    isPaneOpen = opened.isPlaced
129    session = onDiagnostic(
130      session,
131      opened.isPlaced ? `pane opened (${why})` : `pane waiting (${why}): ${opened.reason}`,
132    ).session
133  } catch (error) {
134    isPaneOpen = false
135    session = onDiagnostic(session, `pane did not open (${why}): ${reasonOf(error)}`).session
136  }
137}
138
139const run = async ($: EngineInterface, effect: Effect): Promise<void> => {
140  switch (effect.kind) {
141    case 'toast':
142      $.ui.toast(effect.text)
143      break
144    case 'log':
145      $.ui.log(effect.text)
146      break
147    case 'save-log':
148      await $.store.set(STORE_LOG, session.outcomes)
149      await $.store.set(STORE_THRESHOLDS, session.thresholds)
150      await $.store.set(STORE_DELEGATION, { minSaving: session.minSaving })
151      break
152    case 'save-settings':
153      await $.store.set(STORE_SETTINGS, session.overrides)
154      break
155    case 'clear-log':
156      await $.store.delete(STORE_LOG)
157      await $.store.delete(STORE_THRESHOLDS)
158      await $.store.delete(STORE_DELEGATION)
159      break
160    case 'open-dashboard':
161      wasPaneClosedByPerson = false
162      await openPane($, 'asked')
163      break
164    case 'export':
165      await $.fs.write(effect.path, effect.json)
166      break
167  }
168}
169
170/** Redraws the status line and everything the mod draws. Display only: nothing here reaches the model. */
171const refresh = ($: EngineInterface): void => {
172  const status = renderStatusLine(toView(session))
173
174  if (status !== shownStatus) {
175    shownStatus = status
176    $.ui.status(status)
177  }
178
179  void attempt(() => update($, revision, count => count + 1), undefined)
180
181  if (timer === undefined && isAnimating(session)) {
182    timer = $.clock.every(TICK_MS, () => {
183      session = onTick(session, Date.now()).session
184
185      if (!isAnimating(session)) {
186        timer?.cancel()
187        timer = undefined
188      }
189
190      refresh($)
191    })
192  }
193}
194
195/**
196 * Puts the display back in step with the session. What is sent while no
197 * surface is attached is dropped, so this runs again whenever one may have
198 * appeared: when an app attaches, and at the start of every turn.
199 */
200const sync = async ($: EngineInterface, why: string): Promise<void> => {
201  const surfaces = await attempt(() => $.session.surfaces(), [])
202
203  // Send the status line again even when its text has not changed.
204  shownStatus = undefined
205
206  // A surface with room beside the transcript gets the dashboard; the terminal gets the status line.
207  if (
208    session.config.isEnabled &&
209    !isPaneOpen &&
210    !wasPaneClosedByPerson &&
211    surfaces.some(surface => surface !== 'terminal')
212  ) {
213    await openPane($, why)
214  }
215
216  refresh($)
217}
218
219/** Applies a transition: keeps its session, carries out its effects, redraws. */
220const commit = async <Reply,>($: EngineInterface, step: Step<Reply>): Promise<Reply> => {
221  session = { ...step.session, now: Math.max(step.session.now, Date.now()) }
222
223  for (const effect of step.effects) {
224    // A display or storage effect that fails must never fail the hook that caused it.
225    await attempt(() => run($, effect), undefined)
226  }
227
228  refresh($)
229
230  return step.reply
231}
232
233/**
234 * Counts the project's files, a few folders deep, to judge how much reading a
235 * broad task means. A sample with a ceiling: it reads names only, never contents.
236 */
237const countFiles = async ($: EngineInterface): Promise<number | undefined> => {
238  const queue: { path: string; depth: number }[] = [{ path: '.', depth: 0 }]
239  let listings = 0
240  let files = 0
241
242  while (queue.length > 0 && listings < MAX_LISTINGS && files < MAX_COUNTED) {
243    const folder = queue.shift()
244
245    if (folder === undefined) {
246      break
247    }
248
249    const entries = await attempt(() => $.fs.list(folder.path), undefined)
250
251    if (entries === undefined) {
252      // The first listing failing means the count is unknown, not zero.
253      if (listings === 0) {
254        return undefined
255      }
256
257      continue
258    }
259
260    listings += 1
261
262    for (const entry of entries) {
263      if (entry.kind === 'file') {
264        files += 1
265      } else if (entry.kind === 'dir' && folder.depth < 3 && !SKIPPED_FOLDERS.includes(entry.name)) {
266        queue.push({
267          path: folder.path === '.' ? entry.name : `${folder.path}/${entry.name}`,
268          depth: folder.depth + 1,
269        })
270      }
271    }
272  }
273
274  return files
275}
276
277export const register: Register = (on, options) => {
278  session = hydrate({
279    options,
280    storedLog: undefined,
281    storedThresholds: undefined,
282    storedSettings: undefined,
283    noColor: undefined,
284  }).session
285  isPaneOpen = false
286  wasPaneClosedByPerson = false
287
288  on('session.start', async ($, e, next) => {
289    const [storedLog, storedThresholds, storedSettings, storedDelegation, noColor] =
290      await Promise.all([
291        attempt(() => $.store.get(STORE_LOG), undefined),
292        attempt(() => $.store.get(STORE_THRESHOLDS), undefined),
293        attempt(() => $.store.get(STORE_SETTINGS), undefined),
294        attempt(() => $.store.get(STORE_DELEGATION), undefined),
295        attempt(() => $.env.get('NO_COLOR'), undefined),
296      ])
297
298    await commit(
299      $,
300      hydrate({ options, storedLog, storedThresholds, storedSettings, storedDelegation, noColor }),
301    )
302    await attempt(
303      () =>
304        $.command.register({
305          name: 'tokensaver',
306          description: 'Token-saving delegation: status, on/off, dry-run, dashboard, doctor, export, reset',
307          argumentHint: '[on|off|dry-run|dashboard|doctor|motion|export|reset|help]',
308          immediate: true,
309        }),
310      undefined,
311    )
312
313    const [model, repoFiles] = await Promise.all([
314      attempt(() => $.session.model(), undefined),
315      countFiles($),
316    ])
317
318    await commit($, onFacts(session, { model, repoFiles }))
319    await sync($, 'session start')
320
321    return next(e)
322  })
323
324  on('session.attach', async ($, e, next) => {
325    const attached = await next(e)
326
327    session = onDiagnostic(session, `${e.surface} attached`).session
328    await sync($, `${e.surface} attached`)
329
330    return attached
331  })
332
333  on('ui.close', { id: 'tokensaver' }, async (_$, e, next) => {
334    const closed = await next(e)
335
336    isPaneOpen = false
337
338    // A pane the person closed stays closed until they ask for it again.
339    if (e.origin.kind === 'person') {
340      wasPaneClosedByPerson = true
341    }
342
343    return closed
344  }).catch((_$, e, next) => next(e))
345
346  on('command.run', { command: 'tokensaver' }, async ($, e) => {
347    const reply = await commit($, onCommand(session, e.args, Date.now()))
348
349    return { text: reply.text }
350  })
351
352  // The same commands, declared as files in commands/ so the menu lists them before this
353  // module has run: `/tokensaver:doctor` is `/tokensaver doctor`. Answered here, they
354  // never reach the model; their file text is only what Claude reads if this hook is absent.
355  on(
356    'command.run',
357    {
358      command: [
359        'tokensaver:status',
360        'tokensaver:dashboard',
361        'tokensaver:doctor',
362        'tokensaver:on',
363        'tokensaver:off',
364        'tokensaver:dry-run',
365        'tokensaver:help',
366      ],
367    },
368    async ($, e) => {
369      const word = e.command.slice('tokensaver:'.length)
370      const reply = await commit($, onCommand(session, `${word} ${e.args}`, Date.now()))
371
372      return { text: reply.text }
373    },
374  ).catch((_$, e, next) => next(e))
375
376  on('prompt.submit', async ($, e, next) => {
377    const kind = e.origin.kind
378    const model = await attempt(() => $.session.model(), undefined)
379
380    if (model !== undefined) {
381      session = onFacts(session, { model }).session
382    }
383
384    await sync($, 'turn start')
385
386    const reply = await commit(
387      $,
388      onPrompt(session, {
389        text: e.text,
390        isFromUser: kind === 'composer' || kind === 'bridge' || kind === 'sdk',
391        now: Date.now(),
392      }),
393    )
394
395    return reply.context === undefined
396      ? next(e)
397      : next({ ...e, context: [...(e.context ?? []), reply.context] })
398  }).catch((_$, e, next) => next(e))
399
400  let unnamed = 0
401
402  on('agent.spawn', async ($, e, next) => {
403    unnamed += 1
404
405    const toolUseId = e.tool_use_id === '' ? `spawn-${String(unnamed)}` : e.tool_use_id
406    const reply = await commit(
407      $,
408      onSpawn(session, {
409        toolUseId,
410        prompt: e.prompt,
411        description: e.description,
412        subagentType: e.subagentType,
413        explicitModel: e.model,
414        parentModel: e.parentModel,
415        parentAgentId: e.parentAgentId,
416        isBackground: e.background,
417        isPinned: e.fork || e.isTeammate === true || e.workflow !== undefined,
418        now: Date.now(),
419      }),
420    )
421    const started = await next(reply.model === undefined ? e : { ...e, model: reply.model })
422
423    await commit(
424      $,
425      onSpawned(session, {
426        toolUseId,
427        agentId: started.agentId,
428        isDenied: started.deny !== undefined,
429        resolvedModel: started.model,
430      }),
431    )
432
433    return started
434  }).catch((_$, e, next) => next(e))
435
436  on('tool.call', { tool: 'Agent' }, async ($, e, next) => {
437    const first = await next(e)
438    const reply = await commit($, onAgentResult(session, resultOf(e.tool_use_id, first)))
439
440    // A denied call, a passing result and a dry run all come back with nothing to retry.
441    if (reply.retryWith === undefined) {
442      return first
443    }
444
445    const second = await next({ ...e, model: reply.retryWith })
446
447    await commit($, onAgentResult(session, resultOf(e.tool_use_id, second)))
448
449    return second
450  }).catch((_$, e, next) => next(e))
451
452  on('turn.step', async function* ($, e, next) {
453    const result = yield* next(e)
454
455    if (result.usage !== null) {
456      await commit(
457        $,
458        onStepUsage(session, {
459          turnId: e.turnId,
460          agentId: e.agentId,
461          usage: usageOf(result.usage),
462          reads: result.toolUses.filter(use => READING_TOOLS.includes(use.name)).length,
463        }),
464      )
465    }
466
467    return result
468  })
469
470  on('turn.complete', async ($, e, next) => {
471    await commit(
472      $,
473      onTurnComplete(session, {
474        turnId: e.turnId,
475        agentId: e.agentId,
476        usage: e.usage === undefined ? undefined : usageOf(e.usage),
477        isFailed: e.reason === 'error' || e.reason === 'refusal',
478        isAborted: e.isAborted,
479        answer: e.answer,
480        now: Date.now(),
481      }),
482    )
483
484    return next(e)
485  }).catch((_$, e, next) => next(e))
486
487  on('session.measure', async ($, e, next) => {
488    await commit($, onMeasure(session, { tokens: e.context.tokens, window: e.context.window }))
489
490    return next(e)
491  }).catch((_$, e, next) => next(e))
492
493  // --- Drawing ---------------------------------------------------------------
494
495  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
496    // Reading the revision subscribes the strip to every later change.
497    await read($, revision)
498
499    const view = toView(session)
500
501    if (!isBandVisible(view) || e.props.hasSurvey) {
502      return next(e)
503    }
504
505    const { Box, Text } = $.ui.resolve(e)
506    // Whatever another mod draws in the strip stays, beneath tokensaver's line.
507    const beneath = await next(e)
508
509    if (e.surface === 'desktop') {
510      const width = Math.round(Math.max(40, e.props.bodyColumns) * CELL_PX)
511
512      return (
513        <Box flexDirection="column">
514          {h($.ui.resolve({ ...e, surface: 'desktop' as const }).Svg, {
515            source: bandSvg(view, width),
516            alt: bandLine(view, e.props.bodyColumns)
517              .map(part => part.text)
518              .join(''),
519            width,
520            height: BAND_HEIGHT,
521          })}
522          {beneath}
523        </Box>
524      )
525    }
526
527    return (
528      <Box flexDirection="column">
529        <Box>
530          {bandLine(view, e.props.bodyColumns).map(part => (
531            <Text {...styleOf(part)}>{part.text}</Text>
532          ))}
533        </Box>
534        {beneath}
535      </Box>
536    )
537  })
538
539  on('ui.render', { component: 'Spinner' }, async ($, e, next) => {
540    await read($, revision)
541
542    const suffix = spinnerSuffix(toView(session))
543
544    return suffix === undefined
545      ? next(e)
546      : next({ ...e, props: { ...e.props, suffix: `${e.props.suffix}${suffix}` } })
547  })
548
549  on('ui.render', { component: 'ToolUse' }, async ($, e, next) => {
550    if (e.props.tool !== 'Agent') {
551      return next(e)
552    }
553
554    // Only a running call is kept in step: its model is decided just after its row appears.
555    if (e.props.isRunning) {
556      await read($, revision)
557    }
558
559    const badge = routedBadge(toView(session), e.props.tool_use_id)
560    const row = await next(e)
561
562    if (badge === undefined) {
563      return row
564    }
565
566    const { Box, Text } = $.ui.resolve(e)
567
568    return (
569      <Box flexDirection="column">
570        <Box>
571          {badge.map(part => (
572            <Text {...styleOf(part)}>{part.text}</Text>
573          ))}
574        </Box>
575        {row}
576      </Box>
577    )
578  })
579
580  on('ui.render', { component: 'Pane', requestId: 'tokensaver' }, async ($, e) => {
581    await read($, revision)
582    isPaneOpen = true
583
584    const { Box, Text, Button } = $.ui.resolve(e)
585    const view = toView(session)
586    const columns = e.props.bodyColumns
587    const dashboard = renderDashboard(view, columns)
588    const row = (line: Line) => (
589      <Box>
590        {line.map(part => (
591          <Text {...styleOf(part)}>{part.text}</Text>
592        ))}
593      </Box>
594    )
595    const command = (args: string) => () => commit($, onCommand(session, args, Date.now()))
596
597    return (
598      <Box flexDirection="column">
599        {dashboard.top.map(row)}
600        {e.surface === 'desktop' && view.history.length >= 2 ? (
601          <Box>
602            <Text>{'Trend    '}</Text>
603            {h($.ui.resolve({ ...e, surface: 'desktop' as const }).Svg, {
604              source: sparklineSvg(view.history),
605              alt: `Estimated savings over the last ${String(view.history.length)} turns`,
606            })}
607          </Box>
608        ) : (
609          row(sparklineLine(view, columns))
610        )}
611        {dashboard.bottom.map(row)}
612        <Box gap={2}>
613          <Button
614            key="power"
615            label={view.mode === 'off' ? 'Turn on' : 'Turn off'}
616            onPress={command(view.mode === 'off' ? 'on' : 'off')}
617          />
618          <Button
619            key="dry-run"
620            label={view.mode === 'dry-run' ? 'Leave dry-run' : 'Dry-run'}
621            onPress={command(view.mode === 'dry-run' ? 'dry-run off' : 'dry-run on')}
622          />
623        </Box>
624      </Box>
625    )
626  })
627}
628
src/core/log.ts 177 lines
1import { isCheckSignal } from './escalate.ts'
2import type { CheckSignal } from './escalate.ts'
3import { DEFAULT_THRESHOLDS } from './route.ts'
4import type { Cutoffs, Thresholds } from './route.ts'
5import { TASK_TYPES, isTaskType } from './score.ts'
6import type { TaskType } from './score.ts'
7import { TIERS, isTier } from './tiers.ts'
8import type { Tier } from './tiers.ts'
9
10/** One finished delegation, as the local outcome log keeps it. No prompt text is stored. */
11export type Outcome = {
12  seq: number
13  at: number
14  taskType: TaskType
15  tier: Tier
16  tokens: number
17  retries: number
18  wasEscalated: boolean
19  isSuccess: boolean
20  signal: CheckSignal
21}
22
23/** The log is capped so it stays far below the store's size limit. */
24export const MAX_OUTCOMES = 500
25
26export const isRecord = (value: unknown): value is Readonly<Record<string, unknown>> =>
27  typeof value === 'object' && value !== null && !Array.isArray(value)
28
29const isCount = (value: unknown): value is number =>
30  typeof value === 'number' && Number.isFinite(value) && value >= 0
31
32/** Reads one stored entry, or undefined when it is not a well-formed outcome. */
33export const parseOutcome = (value: unknown): Outcome | undefined => {
34  if (!isRecord(value)) {
35    return undefined
36  }
37
38  const { seq, at, taskType, tier, tokens, retries, wasEscalated, isSuccess, signal } = value
39
40  if (
41    !isCount(seq) ||
42    !isCount(at) ||
43    !isTaskType(taskType) ||
44    !isTier(tier) ||
45    !isCount(tokens) ||
46    !isCount(retries) ||
47    typeof wasEscalated !== 'boolean' ||
48    typeof isSuccess !== 'boolean' ||
49    !isCheckSignal(signal)
50  ) {
51    return undefined
52  }
53
54  return { seq, at, taskType, tier, tokens, retries, wasEscalated, isSuccess, signal }
55}
56
57export type ParsedLog = {
58  outcomes: readonly Outcome[]
59  /** Entries that were present but unreadable, including a log that was not a list at all. */
60  dropped: number
61}
62
63/**
64 * Reads the stored log defensively. A missing log is empty; a corrupted one
65 * keeps every entry that still parses and counts the rest as dropped.
66 */
67export const parseLog = (value: unknown): ParsedLog => {
68  if (value === undefined || value === null) {
69    return { outcomes: [], dropped: 0 }
70  }
71
72  if (!Array.isArray(value)) {
73    return { outcomes: [], dropped: 1 }
74  }
75
76  const entries: readonly unknown[] = value
77  const outcomes = entries
78    .map(parseOutcome)
79    .filter(outcome => outcome !== undefined)
80    .slice(-MAX_OUTCOMES)
81
82  return { outcomes, dropped: entries.length - outcomes.length }
83}
84
85export const nextSeq = (outcomes: readonly Outcome[]): number =>
86  outcomes.reduce((highest, outcome) => Math.max(highest, outcome.seq), 0) + 1
87
88export const appendOutcome = (outcomes: readonly Outcome[], outcome: Outcome): readonly Outcome[] =>
89  [...outcomes, outcome].slice(-MAX_OUTCOMES)
90
91/** A clean outcome is the only evidence that counts toward loosening a rule. */
92export const isCleanSuccess = (outcome: Outcome): boolean =>
93  outcome.isSuccess && !outcome.wasEscalated && outcome.retries === 0
94
95const clampCutoff = (value: unknown, fallback: number): number =>
96  typeof value === 'number' && Number.isFinite(value)
97    ? Math.min(100, Math.max(0, Math.round(value)))
98    : fallback
99
100const parseCutoffs = (value: unknown, fallback: Cutoffs): Cutoffs => {
101  if (!isRecord(value)) {
102    return fallback
103  }
104
105  const haikuMax = clampCutoff(value['haikuMax'], fallback.haikuMax)
106  const sonnetMax = Math.max(haikuMax, clampCutoff(value['sonnetMax'], fallback.sonnetMax))
107  const tunedAtSeq = isCount(value['tunedAtSeq']) ? value['tunedAtSeq'] : 0
108
109  return { haikuMax, sonnetMax, tunedAtSeq }
110}
111
112/** Reads stored thresholds, falling back to the default for anything missing or malformed. */
113export const parseThresholds = (value: unknown): Thresholds => {
114  const stored = isRecord(value) ? value : {}
115
116  return {
117    search: parseCutoffs(stored['search'], DEFAULT_THRESHOLDS.search),
118    'bulk-read': parseCutoffs(stored['bulk-read'], DEFAULT_THRESHOLDS['bulk-read']),
119    'repetitive-edit': parseCutoffs(
120      stored['repetitive-edit'],
121      DEFAULT_THRESHOLDS['repetitive-edit'],
122    ),
123    summarize: parseCutoffs(stored['summarize'], DEFAULT_THRESHOLDS.summarize),
124    edit: parseCutoffs(stored['edit'], DEFAULT_THRESHOLDS.edit),
125    reasoning: parseCutoffs(stored['reasoning'], DEFAULT_THRESHOLDS.reasoning),
126    general: parseCutoffs(stored['general'], DEFAULT_THRESHOLDS.general),
127  }
128}
129
130export type LogSummary = {
131  count: number
132  cleanRate: number
133  escalations: number
134  failures: number
135  byTier: Readonly<Record<Tier, number>>
136}
137
138export const summarizeLog = (outcomes: readonly Outcome[]): LogSummary => {
139  const clean = outcomes.filter(isCleanSuccess).length
140  const countAt = (tier: Tier): number => outcomes.filter(outcome => outcome.tier === tier).length
141
142  return {
143    count: outcomes.length,
144    cleanRate: outcomes.length === 0 ? 1 : clean / outcomes.length,
145    escalations: outcomes.filter(outcome => outcome.wasEscalated).length,
146    failures: outcomes.filter(outcome => !outcome.isSuccess).length,
147    byTier: { haiku: countAt('haiku'), sonnet: countAt('sonnet'), opus: countAt('opus') },
148  }
149}
150
151export type ExportBundle = {
152  format: 'tokensaver-export'
153  version: 1
154  exportedAt: number
155  taskTypes: readonly TaskType[]
156  tiers: readonly Tier[]
157  thresholds: Thresholds
158  summary: LogSummary
159  outcomes: readonly Outcome[]
160}
161
162/** Everything `/tokensaver export` writes: the log, the tuned thresholds and a summary. */
163export const buildExport = (
164  outcomes: readonly Outcome[],
165  thresholds: Thresholds,
166  exportedAt: number,
167): ExportBundle => ({
168  format: 'tokensaver-export',
169  version: 1,
170  exportedAt,
171  taskTypes: TASK_TYPES,
172  tiers: TIERS,
173  thresholds,
174  summary: summarizeLog(outcomes),
175  outcomes,
176})
177
src/runtime/session.ts 1002 lines
1/**
2 * The session controller: every mod event becomes one pure transition,
3 * `(session, input) -> { session, effects, reply }`. Nothing here touches
4 * Claude Code; hooks/register.tsx feeds events in and carries effects out.
5 */
6import { checkResult, isUserRejection, planEscalation } from '../core/escalate.ts'
7import type { CheckSignal } from '../core/escalate.ts'
8import {
9  appendOutcome,
10  buildExport,
11  nextSeq,
12  parseLog,
13  parseThresholds,
14  summarizeLog,
15} from '../core/log.ts'
16import type { Outcome } from '../core/log.ts'
17import {
18  DEFAULT_MIN_SAVING,
19  decideDelegation,
20  judgeHint,
21  parseMinSaving,
22  tuneMinSaving,
23} from '../core/delegate.ts'
24import { DEFAULT_THRESHOLDS, resolveSpawnModel, routeTier } from '../core/route.ts'
25import type { Thresholds } from '../core/route.ts'
26import { contextSavings, estimateTokens, routingSavings } from '../core/savings.ts'
27import { scoreTask } from '../core/score.ts'
28import type { TaskType } from '../core/score.ts'
29import { rankOf, rankOfModel, tierOfModel } from '../core/tiers.ts'
30import type { Tier } from '../core/tiers.ts'
31import { learnFromOutcome, tighten, tuneFromLog } from '../core/tune.ts'
32import { DEFAULT_EXPORT_PATH, HELP_TEXT, parseCommand } from './commands.ts'
33import { applyOverrides, isNoColor, parseOptions, parseOverrides } from './config.ts'
34import type { Config, Overrides } from './config.ts'
35import { formatTokens } from '../ui/format.ts'
36
37export type AgentStatus = 'running' | 'retrying' | 'done' | 'failed' | 'planned'
38
39/**
40 * How an agent type gets its model when nobody names one: from its parent, or
41 * from its own definition. Only the first kind is tokensaver's to route.
42 */
43export type AgentTypeModel = 'inherits' | 'own-model'
44
45/** One sub-agent in the delegation tree. */
46export type AgentNode = {
47  /** The Agent tool call it belongs to. */
48  key: string
49  subagentType: string
50  parentModel: string
51  /** True for the first, untouched spawn of an agent type: it shows how the type picks its model. */
52  isProbe: boolean
53  agentId: string | undefined
54  parentAgentId: string | undefined
55  label: string
56  taskType: TaskType
57  /** The tier it runs on; undefined when it inherits a model tokensaver cannot name. */
58  tier: Tier | undefined
59  /** True when tokensaver chose its model. */
60  isRouted: boolean
61  escalatedFrom: Tier | undefined
62  /** Set between a failed attempt and its retry: the tier the retry must use. */
63  pendingTier: Tier | undefined
64  status: AgentStatus
65  tokens: number
66  /** Tokens a failed attempt spent before its retry. */
67  wastedTokens: number
68  answerTokens: number
69  attempts: number
70  isBackground: boolean
71  isDryRun: boolean
72  /** True when tokensaver's hint asked for this delegation. */
73  isHinted: boolean
74  isRecorded: boolean
75  /** The rank of the model it would have run on without tokensaver. */
76  baselineRank: number | undefined
77}
78
79export type Spend = Readonly<Record<Tier | 'other', number>>
80
81export type Session = {
82  config: Config
83  overrides: Overrides
84  hasNoColor: boolean
85  thresholds: Thresholds
86  outcomes: readonly Outcome[]
87  spend: Spend
88  context: { tokens: number; window: number }
89  /**
90   * Estimated saving from running work on cheaper tiers, in tokens of the
91   * model it would otherwise have run on, net of failed attempts and of
92   * tokensaver's own hint. Negative when the overhead outweighs the gain.
93   */
94  saved: number
95  /** Tokens of delegated work the main context never had to hold: the work less its summary. */
96  keptOut: number
97  savedHistory: readonly number[]
98  agents: readonly AgentNode[]
99  /** What this session has seen of each agent type; a type not listed has not been observed yet. */
100  agentTypes: Readonly<Record<string, AgentTypeModel>>
101  /** Delegations of the turn in progress, by key; the next prompt reads them to judge that turn. */
102  turnKeys: readonly string[]
103  /** Tiers to add to every routing decision this turn (after a user rejection). */
104  boost: number
105  isHintActive: boolean
106  /** Turns whose usage arrived step by step, so their totals are not counted twice. */
107  steppedTurns: readonly string[]
108  /** The bar a delegation suggestion must clear, tuned from how suggestions turn out. */
109  minSaving: number
110  /** What the session knows about where it runs. */
111  facts: SessionFacts
112  /** Files the main loop read or searched itself in the turn in progress. */
113  mainReads: number
114  /** The newest event worth showing above the prompt, for a few seconds. */
115  flash: Flash | undefined
116  /** The clock as of the last event or animation frame. */
117  now: number
118  /** What the display layer reported, for `/tokensaver doctor`. */
119  diagnostics: readonly string[]
120  notes: readonly string[]
121  frame: number
122  shownSaved: number
123}
124
125export type SessionFacts = {
126  /** Files under the working directory, sampled once; undefined until counted. */
127  repoFiles: number | undefined
128  /** The session's model id; undefined until Claude Code has said. */
129  model: string | undefined
130}
131
132export type FlashTone = 'route' | 'escalate' | 'done' | 'fail' | 'info'
133
134export type Flash = {
135  text: string
136  tone: FlashTone
137  at: number
138}
139
140/** How long an event stays above the prompt after it happens. */
141export const FLASH_MS = 6000
142
143export type Effect =
144  | { kind: 'toast'; text: string }
145  | { kind: 'log'; text: string }
146  | { kind: 'save-log' }
147  | { kind: 'save-settings' }
148  | { kind: 'clear-log' }
149  | { kind: 'open-dashboard' }
150  | { kind: 'export'; path: string; json: string }
151
152export type Step<Reply> = {
153  session: Session
154  effects: readonly Effect[]
155  reply: Reply
156}
157
158const MAX_AGENTS = 24
159const MAX_HISTORY = 32
160const MAX_NOTES = 4
161const MAX_STEPPED_TURNS = 64
162const MAX_DIAGNOSTICS = 12
163const HINT_NOUNS: Readonly<Record<TaskType, string>> = {
164  search: 'search',
165  'bulk-read': 'read-through',
166  'repetitive-edit': 'repetitive edit',
167  summarize: 'summarising job',
168  edit: 'edit',
169  reasoning: 'problem',
170  general: 'task',
171}
172
173const withFlash = (session: Session, text: string, tone: FlashTone, at: number): Session => ({
174  ...session,
175  flash: { text, tone, at },
176  now: Math.max(session.now, at),
177})
178
179const still = <Reply>(session: Session, reply: Reply): Step<Reply> => ({
180  session,
181  effects: [],
182  reply,
183})
184
185const shorten = (text: string, limit = 40): string => {
186  const line = text.replace(/\s+/g, ' ').trim()
187
188  return line.length <= limit ? line : `${line.slice(0, limit - 1)}…`
189}
190
191const withNote = (session: Session, note: string): Session => ({
192  ...session,
193  notes: [...session.notes, note].slice(-MAX_NOTES),
194})
195
196const replaceAgent = (session: Session, node: AgentNode): Session => ({
197  ...session,
198  agents: session.agents.map(agent => (agent.key === node.key ? node : agent)),
199})
200
201const signalText = (signal: CheckSignal): string => signal.replace('-', ' ')
202
203export type HydrateInput = {
204  options: Readonly<Record<string, unknown>>
205  storedLog: unknown
206  storedThresholds: unknown
207  storedSettings: unknown
208  storedDelegation?: unknown
209  noColor: string | undefined
210}
211
212/** Builds the session from the plugin's options and whatever the local store holds. */
213export const hydrate = (input: HydrateInput): Step<undefined> => {
214  const log = parseLog(input.storedLog)
215  const overrides = parseOverrides(input.storedSettings)
216  const session: Session = {
217    config: applyOverrides(parseOptions(input.options), overrides),
218    overrides,
219    hasNoColor: isNoColor(input.noColor),
220    thresholds: tuneFromLog(parseThresholds(input.storedThresholds), log.outcomes),
221    outcomes: log.outcomes,
222    spend: { haiku: 0, sonnet: 0, opus: 0, other: 0 },
223    context: { tokens: 0, window: 0 },
224    saved: 0,
225    keptOut: 0,
226    savedHistory: [],
227    agents: [],
228    agentTypes: {},
229    turnKeys: [],
230    boost: 0,
231    isHintActive: false,
232    steppedTurns: [],
233    minSaving: parseMinSaving(input.storedDelegation),
234    facts: { repoFiles: undefined, model: undefined },
235    mainReads: 0,
236    flash: undefined,
237    now: 0,
238    diagnostics: [],
239    notes: [],
240    frame: 0,
241    shownSaved: 0,
242  }
243
244  if (log.dropped === 0) {
245    return still(session, undefined)
246  }
247
248  return {
249    session,
250    effects: [
251      {
252        kind: 'log',
253        text: `tokensaver: ${String(log.dropped)} unreadable entr${log.dropped === 1 ? 'y' : 'ies'} dropped from the outcome log`,
254      },
255      { kind: 'save-log' },
256    ],
257    reply: undefined,
258  }
259}
260
261/** What the display layer learned about the session: its model and how big the repository is. */
262export const onFacts = (session: Session, facts: Partial<SessionFacts>): Step<undefined> =>
263  still({ ...session, facts: { ...session.facts, ...facts } }, undefined)
264
265/** Keeps a line for `/tokensaver doctor`: what was asked of the display and what came back. */
266export const onDiagnostic = (session: Session, line: string): Step<undefined> =>
267  still(
268    { ...session, diagnostics: [...session.diagnostics, line].slice(-MAX_DIAGNOSTICS) },
269    undefined,
270  )
271
272export type PromptInput = {
273  text: string
274  /** True for a prompt the person typed; notifications and peers are not scored. */
275  isFromUser: boolean
276  now: number
277}
278
279/** A user rejection counts against every model tokensaver picked in the turn it rejects. */
280const recordRejections = (session: Session, now: number): Session =>
281  session.agents
282    .filter(agent => session.turnKeys.includes(agent.key) && agent.isRouted && !agent.isDryRun)
283    .reduce((current, agent) => {
284      if (agent.tier === undefined) {
285        return current
286      }
287
288      const outcome: Outcome = {
289        seq: nextSeq(current.outcomes),
290        at: now,
291        taskType: agent.taskType,
292        tier: agent.tier,
293        tokens: agent.tokens,
294        retries: agent.attempts,
295        wasEscalated: agent.escalatedFrom !== undefined,
296        isSuccess: false,
297        signal: 'user-rejection',
298      }
299
300      return {
301        ...current,
302        outcomes: appendOutcome(current.outcomes, outcome),
303        thresholds: tighten(current.thresholds, agent.taskType, agent.tier, outcome.seq),
304      }
305    }, session)
306
307/**
308 * A submitted prompt: judges the previous turn if the user rejects it, then
309 * scores the new task and decides whether delegating would save tokens.
310 */
311export const onPrompt = (
312  session: Session,
313  input: PromptInput,
314): Step<{ context: string | undefined }> => {
315  if (!session.config.isEnabled || !input.isFromUser) {
316    return still(session, { context: undefined })
317  }
318
319  const effects: Effect[] = []
320  const hadRouted = session.agents.some(
321    agent => session.turnKeys.includes(agent.key) && agent.isRouted && !agent.isDryRun,
322  )
323  const isRejected = hadRouted && isUserRejection(input.text)
324  let next: Session = isRejected ? recordRejections(session, input.now) : session
325
326  if (isRejected) {
327    effects.push(
328      { kind: 'toast', text: 'tokensaver: result rejected, routing one tier up this turn' },
329      { kind: 'save-log' },
330    )
331    next = withNote(next, 'rejected: one tier up this turn')
332  }
333
334  const score = scoreTask({ text: input.text })
335  const delegation = decideDelegation(score, {
336    repoFiles: session.facts.repoFiles,
337    parentRank: rankOfModel(session.facts.model),
338    minSaving: session.minSaving,
339  })
340  const shouldHint =
341    delegation.shouldDelegate && session.config.hasContextHint && !session.config.isDryRun
342  const helpers =
343    delegation.agents > 1
344      ? `up to ${String(delegation.agents)} Explore sub-agents in parallel, one per independent part,`
345      : 'one Explore sub-agent'
346  // Directive, because Claude only delegates on a clear instruction; with a way out, because
347  // this estimate is made from the wording and Claude can see the code.
348  const hint = `[tokensaver] This ${HINT_NOUNS[score.taskType]} likely means reading about ${String(delegation.estimatedFiles)} files. Before reading them yourself, hand that reading to ${helpers} with the Agent tool, and work from what they report. Skip this only if one or two targeted lookups would answer it.`
349
350  if (delegation.shouldDelegate && session.config.isDryRun) {
351    const text = `tokensaver dry-run: would suggest delegating this ${HINT_NOUNS[score.taskType]} (about ${String(delegation.estimatedFiles)} files, ~${formatTokens(delegation.estimatedSaving)} tokens saved)`
352    effects.push({ kind: 'toast', text })
353    next = withNote(next, text.replace('tokensaver ', ''))
354  }
355
356  if (shouldHint) {
357    next = withFlash(
358      next,
359      `suggested delegating about ${String(delegation.estimatedFiles)} files of reading`,
360      'info',
361      input.now,
362    )
363  }
364
365  return {
366    session: {
367      ...next,
368      turnKeys: [],
369      mainReads: 0,
370      boost: isRejected ? 1 : 0,
371      isHintActive: shouldHint,
372      // The hint is the one thing tokensaver adds to the model's context: count it as a cost.
373      saved: shouldHint ? next.saved - estimateTokens(hint) : next.saved,
374    },
375    effects,
376    reply: { context: shouldHint ? hint : undefined },
377  }
378}
379
380export type SpawnInput = {
381  toolUseId: string
382  prompt: string
383  description: string
384  subagentType: string
385  explicitModel: string | undefined
386  parentModel: string
387  parentAgentId: string | undefined
388  isBackground: boolean
389  /** True for a fork, a teammate or a workflow agent: their model is not tokensaver's to set. */
390  isPinned: boolean
391  now?: number | undefined
392}
393
394/**
395 * A sub-agent is about to start: scores its task and answers with the model
396 * it should run on, or with nothing to leave the spawn untouched.
397 */
398export const onSpawn = (
399  session: Session,
400  input: SpawnInput,
401): Step<{ model: string | undefined }> => {
402  if (!session.config.isEnabled) {
403    return still(session, { model: undefined })
404  }
405
406  const existing = session.agents.find(agent => agent.key === input.toolUseId)
407  const label = shorten(input.description === '' ? input.prompt : input.description)
408  const score = scoreTask({ text: input.prompt, description: input.description })
409  const decision = routeTier(score, session.thresholds, session.boost)
410  const known = session.agentTypes[input.subagentType]
411  // An agent type may name its own model, and a spawn does not say whether it does. So
412  // tokensaver sets a model only where that cannot override the type's choice: on a retry of
413  // its own, on a model the caller named (which it only ever raises), or on a type it has
414  // watched inherit its parent's model.
415  const canRoute =
416    !input.isPinned &&
417    (existing?.pendingTier !== undefined || input.explicitModel !== undefined || known === 'inherits')
418  const routing = canRoute
419    ? resolveSpawnModel(decision, score, {
420        parentModel: input.parentModel,
421        explicitModel: input.explicitModel,
422        forcedTier: existing?.pendingTier,
423      })
424    : { model: undefined, tier: undefined, note: 'the agent type picks its own model' }
425  const isRouted = routing.model !== undefined
426  const isDryRun = session.config.isDryRun
427  const node: AgentNode = {
428    key: input.toolUseId,
429    subagentType: input.subagentType,
430    parentModel: input.parentModel,
431    isProbe:
432      existing === undefined &&
433      known === undefined &&
434      !input.isPinned &&
435      input.explicitModel === undefined,
436    agentId: undefined,
437    parentAgentId: input.parentAgentId,
438    label,
439    taskType: score.taskType,
440    tier: routing.tier,
441    isRouted,
442    escalatedFrom: existing?.escalatedFrom,
443    pendingTier: undefined,
444    status: isDryRun ? 'planned' : 'running',
445    tokens: existing?.tokens ?? 0,
446    wastedTokens: existing?.wastedTokens ?? 0,
447    answerTokens: 0,
448    attempts: existing?.attempts ?? 0,
449    isBackground: input.isBackground,
450    isDryRun,
451    isHinted: existing?.isHinted ?? session.isHintActive,
452    isRecorded: false,
453    // A retry keeps the baseline of the first attempt: what would have run without tokensaver.
454    baselineRank: existing?.baselineRank ?? rankOfModel(input.explicitModel ?? input.parentModel),
455  }
456  const agents =
457    existing === undefined
458      ? [...session.agents, node].slice(-MAX_AGENTS)
459      : session.agents.map(agent => (agent.key === node.key ? node : agent))
460  const tracked: Session = {
461    ...session,
462    agents,
463    turnKeys: session.turnKeys.includes(node.key) ? session.turnKeys : [...session.turnKeys, node.key],
464  }
465
466  if (!isRouted) {
467    return still(tracked, { model: undefined })
468  }
469
470  const target = routing.tier ?? 'the parent model'
471
472  if (isDryRun) {
473    const text = `tokensaver dry-run: would run "${label}" on ${target} (${score.taskType})`
474
475    return {
476      session: withNote(tracked, text.replace('tokensaver ', '')),
477      effects: [{ kind: 'toast', text }],
478      reply: { model: undefined },
479    }
480  }
481
482  const text =
483    existing?.escalatedFrom === undefined
484      ? `tokensaver: "${label}" → ${target} (${score.taskType})`
485      : `tokensaver: escalated "${label}" ${existing.escalatedFrom} → ${target}`
486
487  return {
488    session: withFlash(
489      withNote(tracked, text.replace('tokensaver: ', '')),
490      existing?.escalatedFrom === undefined
491        ? `${label} → ${target}`
492        : `${label} escalated ${existing.escalatedFrom} → ${target}`,
493      existing?.escalatedFrom === undefined ? 'route' : 'escalate',
494      input.now ?? session.now,
495    ),
496    effects: [{ kind: 'toast', text }],
497    reply: { model: routing.model },
498  }
499}
500
501export type SpawnedInput = {
502  toolUseId: string
503  agentId: string | undefined
504  isDenied: boolean
505  /** The model Claude Code resolved for the sub-agent. */
506  resolvedModel?: string | undefined
507}
508
509const isSameModel = (a: string, b: string): boolean => {
510  const rank = rankOfModel(a)
511
512  return a === b || (rank !== undefined && rank === rankOfModel(b))
513}
514
515/**
516 * The spawn settled: remembers the agent's id, and from an untouched first
517 * spawn learns whether its agent type inherits the parent's model. A refused
518 * spawn is dropped.
519 */
520export const onSpawned = (session: Session, input: SpawnedInput): Step<undefined> => {
521  const node = session.agents.find(agent => agent.key === input.toolUseId)
522
523  if (node === undefined) {
524    return still(session, undefined)
525  }
526
527  if (input.isDenied) {
528    return still(replaceAgent(session, { ...node, status: 'failed', isRecorded: true }), undefined)
529  }
530
531  const started: AgentNode = {
532    ...node,
533    agentId: input.agentId,
534    tier: node.isRouted ? node.tier : (tierOfModel(input.resolvedModel) ?? node.tier),
535  }
536
537  if (!node.isProbe || input.resolvedModel === undefined) {
538    return still(replaceAgent(session, started), undefined)
539  }
540
541  const learned: AgentTypeModel = isSameModel(input.resolvedModel, node.parentModel)
542    ? 'inherits'
543    : 'own-model'
544  const note =
545    learned === 'inherits'
546      ? `"${node.subagentType}" agents inherit their model: routing them from now on`
547      : `"${node.subagentType}" agents choose their own model: left alone`
548
549  return still(
550    withNote(
551      {
552        ...replaceAgent(session, started),
553        agentTypes: { ...session.agentTypes, [node.subagentType]: learned },
554      },
555      note,
556    ),
557    undefined,
558  )
559}
560
561/** Books a finished delegation: its outcome, what it teaches, and what it saved. */
562const finalize = (session: Session, node: AgentNode, signal: CheckSignal, now: number): Step<undefined> => {
563  const isSuccess = signal === 'ok'
564  const usefulTokens = Math.max(0, node.tokens - node.wastedTokens)
565  // A dry run changed nothing, so it saved nothing.
566  const routed =
567    node.isRouted && !node.isDryRun && node.tier !== undefined && node.baselineRank !== undefined
568      ? routingSavings(usefulTokens, rankOf(node.tier), node.baselineRank)
569      : 0
570  const kept =
571    node.isHinted && !node.isDryRun && isSuccess
572      ? contextSavings(usefulTokens, node.answerTokens)
573      : 0
574  const done: AgentNode = { ...node, status: isSuccess ? 'done' : 'failed', isRecorded: true }
575  const settled: Session = {
576    ...withFlash(
577      replaceAgent(session, done),
578      `${node.label} ${isSuccess ? 'done' : `failed (${signalText(signal)})`} · ${formatTokens(node.tokens)}${node.tier === undefined ? '' : ` on ${node.tier}`}`,
579      isSuccess ? 'done' : 'fail',
580      now,
581    ),
582    // Two different savings, kept apart: adding them would count the same tokens twice.
583    saved: session.saved + routed - node.wastedTokens,
584    keptOut: session.keptOut + kept,
585  }
586
587  // The log records tokensaver's own decisions; a model it did not pick teaches it nothing.
588  if (!node.isRouted || node.tier === undefined || node.isDryRun) {
589    return still(settled, undefined)
590  }
591
592  const outcome: Outcome = {
593    seq: nextSeq(session.outcomes),
594    at: now,
595    taskType: node.taskType,
596    tier: node.tier,
597    tokens: node.tokens,
598    retries: node.attempts,
599    wasEscalated: node.escalatedFrom !== undefined,
600    isSuccess,
601    signal,
602  }
603  const outcomes = appendOutcome(session.outcomes, outcome)
604
605  return {
606    session: {
607      ...settled,
608      outcomes,
609      thresholds: learnFromOutcome(session.thresholds, outcomes, outcome),
610    },
611    effects: [{ kind: 'save-log' }],
612    reply: undefined,
613  }
614}
615
616export type AgentResultInput = {
617  toolUseId: string
618  isDenied: boolean
619  isError: boolean
620  text: string
621  /** The Agent tool's own token total, used when no step usage was seen. */
622  totalTokens: number | undefined
623  now: number
624}
625
626/**
627 * An Agent call returned. A result that fails its checks on a tier
628 * tokensaver picked is retried once, one tier up; anything else is final.
629 */
630export const onAgentResult = (
631  session: Session,
632  input: AgentResultInput,
633): Step<{ retryWith: Tier | undefined }> => {
634  const found = session.agents.find(agent => agent.key === input.toolUseId)
635
636  if (!session.config.isEnabled || found === undefined || found.isRecorded) {
637    return still(session, { retryWith: undefined })
638  }
639
640  // A background agent's call returns as it starts; its turn.complete settles it.
641  if (found.isBackground && !input.isDenied && !input.isError) {
642    return still(session, { retryWith: undefined })
643  }
644
645  const node: AgentNode =
646    found.tokens === 0 && input.totalTokens !== undefined
647      ? { ...found, tokens: input.totalTokens }
648      : found
649
650  // A denial is a permission decision, not a quality signal: no retry, no outcome.
651  if (input.isDenied) {
652    return still(replaceAgent(session, { ...node, status: 'failed', isRecorded: true }), {
653      retryWith: undefined,
654    })
655  }
656
657  const signal = checkResult({ isError: input.isError, text: input.text })
658  const plan = planEscalation({
659    signal,
660    tier: node.isRouted ? node.tier : undefined,
661    attempts: node.attempts,
662    isDenied: input.isDenied,
663    isDryRun: node.isDryRun,
664  })
665
666  if (!plan.shouldEscalate || node.tier === undefined) {
667    return { ...finalize(session, node, signal, input.now), reply: { retryWith: undefined } }
668  }
669
670  const retrying: AgentNode = {
671    ...node,
672    status: 'retrying',
673    escalatedFrom: node.tier,
674    pendingTier: plan.to,
675    attempts: node.attempts + 1,
676    wastedTokens: node.tokens,
677  }
678  const text = `tokensaver: "${node.label}" failed on ${node.tier} (${signalText(signal)}), retrying on ${plan.to}`
679
680  return {
681    session: withNote(
682      {
683        ...withFlash(
684          replaceAgent(session, retrying),
685          `${node.label} failed on ${node.tier}, retrying on ${plan.to}`,
686          'escalate',
687          input.now,
688        ),
689        // Tighten now, not when the retry finishes: the cheaper tier has already failed.
690        thresholds: tighten(session.thresholds, node.taskType, node.tier, nextSeq(session.outcomes)),
691      },
692      text.replace('tokensaver: ', ''),
693    ),
694    effects: [{ kind: 'toast', text }, { kind: 'save-log' }],
695    reply: { retryWith: plan.to },
696  }
697}
698
699export type Usage = {
700  model: string
701  input: number
702  output: number
703  cacheWrite: number
704}
705
706const applyUsage = (session: Session, agentId: string | undefined, usage: Usage): Session => {
707  // Cache reads are left out: they are re-read context, not new work.
708  const tokens = usage.input + usage.output + usage.cacheWrite
709  const bucket = tierOfModel(usage.model) ?? 'other'
710  const spend = { ...session.spend, [bucket]: session.spend[bucket] + tokens }
711
712  return {
713    ...session,
714    spend,
715    agents:
716      agentId === undefined
717        ? session.agents
718        : session.agents.map(agent =>
719            agent.agentId === agentId ? { ...agent, tokens: agent.tokens + tokens } : agent,
720          ),
721  }
722}
723
724export type StepUsageInput = {
725  turnId: string
726  agentId: string | undefined
727  usage: Usage
728  /** Read, Grep and Glob calls the response asked for. */
729  reads?: number | undefined
730}
731
732/** One model request finished: adds its tokens to the meter and to its agent. */
733export const onStepUsage = (session: Session, input: StepUsageInput): Step<undefined> => {
734  if (!session.config.isEnabled) {
735    return still(session, undefined)
736  }
737
738  const stepKey = `${input.turnId}/${input.agentId ?? 'main'}`
739
740  return still(
741    {
742      ...applyUsage(session, input.agentId, input.usage),
743      mainReads: session.mainReads + (input.agentId === undefined ? (input.reads ?? 0) : 0),
744      steppedTurns: session.steppedTurns.includes(stepKey)
745        ? session.steppedTurns
746        : [...session.steppedTurns, stepKey].slice(-MAX_STEPPED_TURNS),
747    },
748    undefined,
749  )
750}
751
752export type TurnCompleteInput = {
753  turnId: string
754  agentId: string | undefined
755  /** The turn's totals; used only when no step of this turn reported usage. */
756  usage: Usage | undefined
757  isFailed: boolean
758  /** True when the user interrupted the turn. */
759  isAborted: boolean
760  answer: string
761  now: number
762}
763
764/** A turn ended: the main loop's closes the books on the turn, a sub-agent's on itself. */
765export const onTurnComplete = (session: Session, input: TurnCompleteInput): Step<undefined> => {
766  if (!session.config.isEnabled) {
767    return still(session, undefined)
768  }
769
770  const stepKey = `${input.turnId}/${input.agentId ?? 'main'}`
771  const counted =
772    input.usage === undefined || session.steppedTurns.includes(stepKey)
773      ? session
774      : applyUsage(session, input.agentId, input.usage)
775
776  if (input.agentId === undefined) {
777    const closed: Session = {
778      ...counted,
779      savedHistory: [...counted.savedHistory, counted.saved].slice(-MAX_HISTORY),
780    }
781
782    if (!closed.isHintActive || input.isAborted) {
783      return still(closed, undefined)
784    }
785
786    // The suggestion is judged by what the turn then did, and the bar moves with the verdict.
787    const verdict = judgeHint({
788      largestAgentTokens: Math.max(
789        0,
790        ...closed.agents.filter(agent => closed.turnKeys.includes(agent.key)).map(agent => agent.tokens),
791      ),
792      mainReads: closed.mainReads,
793    })
794    const minSaving = tuneMinSaving(closed.minSaving, verdict)
795    const judged: Session = { ...closed, isHintActive: false, minSaving }
796
797    if (minSaving === closed.minSaving) {
798      return still(judged, undefined)
799    }
800
801    return {
802      session: withNote(
803        judged,
804        verdict === 'over-fired'
805          ? 'delegation suggestion was not needed: raising the bar'
806          : 'delegation suggestion paid off',
807      ),
808      effects: [{ kind: 'save-log' }],
809      reply: undefined,
810    }
811  }
812
813  const node = counted.agents.find(agent => agent.agentId === input.agentId)
814
815  if (node === undefined || node.isRecorded) {
816    return still(counted, undefined)
817  }
818
819  // An interrupt is the user's choice, not a quality signal: no outcome, and never a retry.
820  if (input.isAborted) {
821    return still(replaceAgent(counted, { ...node, status: 'failed', isRecorded: true }), undefined)
822  }
823
824  const answered: AgentNode = { ...node, answerTokens: estimateTokens(input.answer) }
825
826  // A foreground agent is settled by its Agent call, which can still retry it.
827  if (!node.isBackground) {
828    return still(replaceAgent(counted, answered), undefined)
829  }
830
831  const signal = checkResult({ isError: input.isFailed, text: input.answer })
832
833  return finalize(replaceAgent(counted, answered), answered, signal, input.now)
834}
835
836/** The context window moved: feeds the live meter. */
837export const onMeasure = (
838  session: Session,
839  input: { tokens: number | undefined; window: number },
840): Step<undefined> =>
841  still({ ...session, context: { tokens: input.tokens ?? session.context.tokens, window: input.window } }, undefined)
842
843export const activeCount = (session: Session): number =>
844  session.agents.filter(agent => agent.status === 'running' || agent.status === 'retrying').length
845
846/** True while an event is recent enough to still be shown above the prompt. */
847export const hasFreshFlash = (session: Session): boolean =>
848  session.flash !== undefined && session.now - session.flash.at < FLASH_MS
849
850/**
851 * True while the animation clock is needed. A fresh event keeps it running
852 * even under reduced motion, because the clock is also what takes the event
853 * down again; reduced motion stops only the movement.
854 */
855export const isAnimating = (session: Session): boolean =>
856  session.config.isEnabled &&
857  (hasFreshFlash(session) ||
858    (!session.config.hasReducedMotion &&
859      (activeCount(session) > 0 || session.shownSaved !== session.saved)))
860
861/** One animation frame: turns the spinner and eases the saved counter toward its value. */
862export const onTick = (session: Session, now = session.now): Step<undefined> => {
863  if (session.config.hasReducedMotion) {
864    return still({ ...session, now, shownSaved: session.saved }, undefined)
865  }
866
867  const gap = session.saved - session.shownSaved
868  const move = Math.abs(gap) <= 3 ? gap : Math.trunc(gap / 3)
869
870  return still(
871    { ...session, now, frame: session.frame + 1, shownSaved: session.shownSaved + move },
872    undefined,
873  )
874}
875
876const modeOf = (config: Config): string => {
877  if (!config.isEnabled) {
878    return 'off'
879  }
880
881  return config.isDryRun ? 'dry-run' : 'on'
882}
883
884const statusReport = (session: Session): string => {
885  const summary = summarizeLog(session.outcomes)
886  const total = session.spend.haiku + session.spend.sonnet + session.spend.opus + session.spend.other
887
888  return [
889    `tokensaver is ${modeOf(session.config)}`,
890    `${formatTokens(total)} tokens this session (haiku ${formatTokens(session.spend.haiku)}, sonnet ${formatTokens(session.spend.sonnet)}, opus ${formatTokens(session.spend.opus)})`,
891    `~${formatTokens(session.saved)} saved by routing (estimate), ${formatTokens(session.keptOut)} kept out of the main context`,
892    `${String(summary.count)} logged outcomes, ${String(Math.round(summary.cleanRate * 100))}% clean, ${String(summary.escalations)} escalations`,
893  ].join('\n')
894}
895
896/** What tokensaver knows and what the display told it: the first thing to read when nothing shows. */
897const doctorReport = (session: Session): string => {
898  const types = Object.entries(session.agentTypes)
899    .map(([name, kind]) => `${name} ${kind === 'inherits' ? 'inherits' : 'has its own model'}`)
900    .join(', ')
901
902  return [
903    `tokensaver is ${modeOf(session.config)}`,
904    `session model: ${session.facts.model ?? 'not reported yet'}`,
905    `repository: ${session.facts.repoFiles === undefined ? 'not counted yet' : `about ${String(session.facts.repoFiles)} files`}`,
906    `agent types seen: ${types === '' ? 'none yet' : types}`,
907    `a delegation suggestion must save at least ${formatTokens(session.minSaving)} tokens`,
908    `colour ${session.hasNoColor ? 'off (NO_COLOR)' : 'on'}, motion ${session.config.hasReducedMotion ? 'off' : 'on'}`,
909    'display:',
910    ...(session.diagnostics.length === 0
911      ? ['  nothing reported yet']
912      : session.diagnostics.map(line => `  ${line}`)),
913  ].join('\n')
914}
915
916export type CommandReply = {
917  /** What the command prints. A command's output is a transcript row, so it is kept short. */
918  text: string
919}
920
921const answer = (session: Session, text: string, effects: readonly Effect[] = []): Step<CommandReply> => ({
922  session,
923  effects,
924  reply: { text },
925})
926
927const withOverride = (session: Session, overrides: Overrides, text: string): Step<CommandReply> => {
928  const merged = { ...session.overrides, ...overrides }
929
930  return answer(
931    { ...session, overrides: merged, config: applyOverrides(session.config, merged) },
932    text,
933    [{ kind: 'save-settings' }],
934  )
935}
936
937/**
938 * `/tokensaver <args>`. Every command answers with a line of text, so it is
939 * visible on every surface even where a pane or a toast is not.
940 */
941export const onCommand = (session: Session, args: string, now: number): Step<CommandReply> => {
942  const command = parseCommand(args)
943
944  switch (command.kind) {
945    case 'status':
946      return answer(session, statusReport(session))
947    case 'help':
948      return answer(session, HELP_TEXT)
949    case 'doctor':
950      return answer(session, doctorReport(session))
951    case 'dashboard':
952      return answer(session, 'tokensaver dashboard opened', [{ kind: 'open-dashboard' }])
953    case 'enable':
954      return withOverride(
955        session,
956        { isEnabled: command.value },
957        `tokensaver is ${command.value ? 'on' : 'off'}`,
958      )
959    case 'dry-run': {
960      const isDryRun = command.value ?? !session.config.isDryRun
961
962      return withOverride(session, { isDryRun }, `tokensaver dry-run is ${isDryRun ? 'on' : 'off'}`)
963    }
964    case 'motion': {
965      const hasMotion = command.value ?? session.config.hasReducedMotion
966
967      return withOverride(
968        { ...session, shownSaved: session.saved },
969        { hasReducedMotion: !hasMotion },
970        `tokensaver animations are ${hasMotion ? 'on' : 'off'}`,
971      )
972    }
973    case 'reset':
974      return answer(
975        {
976          ...session,
977          outcomes: [],
978          thresholds: DEFAULT_THRESHOLDS,
979          minSaving: DEFAULT_MIN_SAVING,
980          saved: 0,
981          keptOut: 0,
982          shownSaved: 0,
983          savedHistory: [],
984        },
985        'tokensaver: outcome log and everything learned from it cleared',
986        [{ kind: 'clear-log' }],
987      )
988    case 'export': {
989      const path = command.path ?? DEFAULT_EXPORT_PATH
990      const json = JSON.stringify(buildExport(session.outcomes, session.thresholds, now), null, 2)
991
992      return answer(
993        session,
994        `tokensaver: wrote ${String(session.outcomes.length)} outcomes to ${path} (local file, nothing was sent anywhere)`,
995        [{ kind: 'export', path, json }],
996      )
997    }
998    case 'unknown':
999      return answer(session, `tokensaver: unknown command "${command.word}"\n${HELP_TEXT}`)
1000  }
1001}
1002
src/ui/band.ts 226 lines
1/**
2 * The strip above the prompt: who is working for whom, right now. It shows
3 * while a sub-agent runs and for a few seconds after an event, and is gone
4 * the rest of the time. Text for the terminal, an animated vector for
5 * surfaces that draw one.
6 */
7import type { Tier } from '../core/tiers.ts'
8import type { FlashTone } from '../runtime/session.ts'
9import { fit, formatTokens, spinner } from './format.ts'
10import { span } from './theme.ts'
11import type { Line, Tone } from './theme.ts'
12import { themeOf } from './view.ts'
13import type { View, ViewAgent } from './view.ts'
14
15const TIER_TONES: Readonly<Record<Tier, Tone>> = {
16  haiku: 'success',
17  sonnet: 'suggestion',
18  opus: 'claude',
19}
20
21const FLASH_TONES: Readonly<Record<FlashTone, Tone>> = {
22  route: 'success',
23  escalate: 'warning',
24  done: 'success',
25  fail: 'error',
26  info: 'suggestion',
27}
28
29const FLASH_GLYPHS: Readonly<Record<FlashTone, string>> = {
30  route: '→',
31  escalate: '↑',
32  done: '✓',
33  fail: '✗',
34  info: '·',
35}
36
37/** Colours that read on a light and on a dark background alike. */
38const HEX: Readonly<Record<Tone, string>> = {
39  success: '#3fa55b',
40  suggestion: '#7b83eb',
41  claude: '#d77757',
42  warning: '#d99a00',
43  error: '#e5484d',
44  inactive: '#8b8b8b',
45}
46
47const MAX_SHOWN = 3
48const LABEL_CHARS = 22
49
50/** True when the strip has something to say. */
51export const isBandVisible = (view: View): boolean =>
52  view.mode !== 'off' && (view.band.agents.length > 0 || view.band.flash !== undefined)
53
54const tierName = (agent: ViewAgent): string => agent.tier ?? 'inherit'
55
56/** The strip as one line of text. */
57export const bandLine = (view: View, columns: number): Line => {
58  const theme = themeOf(view)
59  const { agents, flash } = view.band
60  const isActive = agents.length > 0
61  const parts = [
62    span(theme, `${isActive ? spinner(view.frame, theme.hasMotion) : '●'} `, { tone: 'claude' }),
63    span(theme, 'tokensaver ', { isBold: true }),
64  ]
65
66  if (isActive) {
67    parts.push(span(theme, 'main', { isDim: true }))
68
69    for (const agent of agents.slice(0, MAX_SHOWN)) {
70      parts.push(
71        span(theme, ' ─▸ ', { isDim: true }),
72        span(theme, tierName(agent), agent.tier === undefined ? {} : { tone: TIER_TONES[agent.tier] }),
73        span(theme, ` ${fit(agent.label, LABEL_CHARS)}`),
74      )
75    }
76
77    if (agents.length > MAX_SHOWN) {
78      parts.push(span(theme, `  +${String(agents.length - MAX_SHOWN)}`, { isDim: true }))
79    }
80  } else if (flash !== undefined) {
81    parts.push(
82      span(theme, `${FLASH_GLYPHS[flash.tone]} `, { tone: FLASH_TONES[flash.tone] }),
83      span(theme, fit(flash.text, Math.max(20, columns - 16))),
84    )
85  }
86
87  return parts
88}
89
90const escapeXml = (text: string): string =>
91  text.replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/>/g, '&gt;').replace(/"/g, '&quot;')
92
93const STYLE =
94  '<style>' +
95  'text{font-family:ui-sans-serif,system-ui,-apple-system,"Segoe UI",sans-serif;font-size:11.5px}' +
96  '.ink{fill:#2b2a27}.dim{fill:#77756f}.rail{stroke:#77756f}' +
97  '@media (prefers-color-scheme: dark){.ink{fill:#ecebe6}.dim{fill:#9c9a92}.rail{stroke:#9c9a92}}' +
98  '</style>'
99
100/** Approximate width of proportional text at the strip's font size. */
101const textWidth = (text: string): number => Math.ceil(text.length * 6.1)
102
103const HEIGHT = 30
104const MID = HEIGHT / 2
105
106/**
107 * The strip as an animated SVG picture: the main loop on the left, a dashed
108 * line flowing out to each running sub-agent, each agent a pill in its
109 * tier's colour with a pulsing dot. Under reduced motion nothing moves, and
110 * under NO_COLOR everything is drawn in the text colour.
111 */
112export const bandSvg = (view: View, width: number): string => {
113  const theme = themeOf(view)
114  const { agents, flash } = view.band
115  const colour = (tone: Tone): string => (theme.hasColor ? `fill="${HEX[tone]}"` : 'class="ink"')
116  const stroke = (tone: Tone): string => (theme.hasColor ? `stroke="${HEX[tone]}"` : 'class="rail"')
117  const moving = (markup: string): string => (theme.hasMotion ? markup : '')
118  const parts: string[] = [STYLE]
119  const isActive = agents.length > 0
120
121  // The main loop.
122  parts.push(
123    `<circle cx="12" cy="${String(MID)}" r="9" ${colour('claude')} opacity="0.18">${moving(
124      isActive
125        ? '<animate attributeName="r" values="6;11;6" dur="1.6s" repeatCount="indefinite"/><animate attributeName="opacity" values="0.3;0.05;0.3" dur="1.6s" repeatCount="indefinite"/>'
126        : '',
127    )}</circle>`,
128    `<circle cx="12" cy="${String(MID)}" r="4.5" ${colour('claude')}/>`,
129    `<text x="24" y="${String(MID + 4)}" class="ink" font-weight="600">tokensaver</text>`,
130  )
131
132  let x = 24 + textWidth('tokensaver') + 14
133
134  if (isActive) {
135    agents.slice(0, MAX_SHOWN).forEach((agent, index) => {
136      const tone: Tone = agent.status === 'retrying' ? 'warning' : agent.tier === undefined ? 'inactive' : TIER_TONES[agent.tier]
137      const label = `${tierName(agent)} · ${fit(agent.label, LABEL_CHARS)}`
138      const pill = textWidth(label) + 30
139      const rail = 26
140      const delay = (index * 0.25).toFixed(2)
141
142      // The line the work flows along.
143      parts.push(
144        `<line x1="${String(x)}" y1="${String(MID)}" x2="${String(x + rail)}" y2="${String(MID)}" ${stroke(tone)} stroke-width="1.6" stroke-dasharray="4 4" stroke-linecap="round" fill="none">${moving(
145          `<animate attributeName="stroke-dashoffset" from="8" to="0" dur="0.6s" begin="${delay}s" repeatCount="indefinite"/>`,
146        )}</line>`,
147      )
148      x += rail + 4
149
150      // The sub-agent.
151      parts.push(
152        `<rect x="${String(x)}" y="5" width="${String(pill)}" height="20" rx="10" ${colour(tone)} opacity="0.16"/>`,
153        `<rect x="${String(x)}" y="5" width="${String(pill)}" height="20" rx="10" fill="none" ${stroke(tone)} stroke-width="1"/>`,
154        `<circle cx="${String(x + 12)}" cy="${String(MID)}" r="3.5" ${colour(tone)}>${moving(
155          `<animate attributeName="opacity" values="1;0.25;1" dur="1.1s" begin="${delay}s" repeatCount="indefinite"/>`,
156        )}</circle>`,
157        `<text x="${String(x + 21)}" y="${String(MID + 4)}" class="ink">${escapeXml(label)}</text>`,
158      )
159      x += pill + 8
160    })
161
162    if (agents.length > MAX_SHOWN) {
163      parts.push(
164        `<text x="${String(x + 2)}" y="${String(MID + 4)}" class="dim">+${String(agents.length - MAX_SHOWN)}</text>`,
165      )
166      x += 30
167    }
168  } else if (flash !== undefined) {
169    const tone = FLASH_TONES[flash.tone]
170    const text = `${FLASH_GLYPHS[flash.tone]} ${fit(flash.text, 70)}`
171
172    parts.push(
173      `<text x="${String(x)}" y="${String(MID + 4)}" ${colour(tone)}>${escapeXml(text)}${moving(
174        '<animate attributeName="opacity" values="0;1" dur="0.35s" fill="freeze"/>',
175      )}</text>`,
176    )
177    x += textWidth(text) + 12
178  }
179
180  // The running total, where there is room for it.
181  const total = `saved ~${formatTokens(view.hasReducedMotion ? view.saved : view.shownSaved)}`
182
183  if (x + textWidth(total) + 8 <= width) {
184    parts.push(
185      `<text x="${String(width - 6)}" y="${String(MID + 4)}" text-anchor="end" class="dim">${escapeXml(total)}</text>`,
186    )
187  }
188
189  return `<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 ${String(width)} ${String(HEIGHT)}" width="${String(width)}" height="${String(HEIGHT)}">${parts.join('')}</svg>`
190}
191
192export const BAND_HEIGHT = HEIGHT
193
194/** What to add after the spinner's word while sub-agents run, or undefined to leave it alone. */
195export const spinnerSuffix = (view: View): string | undefined => {
196  const { agents } = view.band
197
198  if (view.mode === 'off' || agents.length === 0) {
199    return undefined
200  }
201
202  const tiers = [...new Set(agents.map(tierName))].join(', ')
203
204  return ` · ${String(agents.length)} sub-agent${agents.length === 1 ? '' : 's'} on ${tiers}`
205}
206
207/** A one-line badge for an Agent call tokensaver routed, or undefined when it did not. */
208export const routedBadge = (view: View, toolUseId: string): Line | undefined => {
209  const agent = view.routed.find(entry => entry.key === toolUseId)
210
211  if (agent?.tier === undefined) {
212    return undefined
213  }
214
215  const theme = themeOf(view)
216
217  return [
218    span(theme, '◆ tokensaver ', { isDim: true }),
219    span(theme, agent.escalatedFrom === undefined ? '→ ' : `${agent.escalatedFrom} ↑ `, {
220      tone: agent.escalatedFrom === undefined ? TIER_TONES[agent.tier] : 'warning',
221    }),
222    span(theme, agent.tier, { tone: TIER_TONES[agent.tier], isBold: true }),
223    span(theme, ` (${agent.taskType})`, { isDim: true }),
224  ]
225}
226
src/ui/dashboard.ts 204 lines
1import type { Tier } from '../core/tiers.ts'
2import type { AgentStatus } from '../runtime/session.ts'
3import { bar, fit, formatTokens, padEnd, padStart, sparkline, spinner } from './format.ts'
4import { savedOnScreen } from './statusline.ts'
5import { span } from './theme.ts'
6import type { Line, Theme, Tone } from './theme.ts'
7import { themeOf } from './view.ts'
8import type { View, ViewAgent } from './view.ts'
9
10const MIN_COLUMNS = 34
11const MAX_COLUMNS = 72
12const LABEL_WIDTH = 9
13
14const TIER_TONES: Readonly<Record<Tier, Tone>> = {
15  haiku: 'success',
16  sonnet: 'suggestion',
17  opus: 'claude',
18}
19
20const STATUS_TONES: Readonly<Record<AgentStatus, Tone>> = {
21  running: 'claude',
22  retrying: 'warning',
23  done: 'success',
24  failed: 'error',
25  planned: 'inactive',
26}
27
28const STATUS_GLYPHS: Readonly<Record<Exclude<AgentStatus, 'running'>, string>> = {
29  retrying: '↑',
30  done: '✓',
31  failed: '✗',
32  planned: '○',
33}
34
35const MODE_TEXT: Readonly<Record<View['mode'], { glyph: string; word: string; tone: Tone }>> = {
36  on: { glyph: '●', word: 'on', tone: 'success' },
37  off: { glyph: '○', word: 'off', tone: 'inactive' },
38  'dry-run': { glyph: '◌', word: 'dry-run', tone: 'warning' },
39}
40
41const widthOf = (columns: number): number =>
42  Math.min(MAX_COLUMNS, Math.max(MIN_COLUMNS, Math.floor(columns)))
43
44const heading = (theme: Theme, text: string): Line => [span(theme, text, { isBold: true })]
45
46const headerLine = (view: View, theme: Theme): Line => {
47  const mode = MODE_TEXT[view.mode]
48
49  return [
50    span(theme, 'tokensaver ', { isBold: true }),
51    span(theme, `${mode.glyph} ${mode.word}`, { tone: mode.tone }),
52  ]
53}
54
55const contextLine = (view: View, theme: Theme, barWidth: number): Line => {
56  const { tokens, window } = view.context
57  const fraction = window > 0 ? tokens / window : 0
58  const tone: Tone = fraction >= 0.85 ? 'error' : fraction >= 0.6 ? 'warning' : 'success'
59  const figure =
60    window > 0
61      ? `${formatTokens(tokens)} / ${formatTokens(window)}  ${String(Math.round(fraction * 100))}%`
62      : 'not measured yet'
63
64  return [
65    span(theme, padEnd('Context', LABEL_WIDTH)),
66    span(theme, bar(fraction, barWidth), { tone }),
67    span(theme, `  ${figure}`, { isDim: true }),
68  ]
69}
70
71const sessionLine = (view: View, theme: Theme): Line => [
72  span(theme, padEnd('Session', LABEL_WIDTH)),
73  span(theme, `${formatTokens(view.total)} tokens`),
74]
75
76const savedLine = (view: View, theme: Theme): Line => [
77  span(theme, padEnd('Saved', LABEL_WIDTH)),
78  span(theme, `~${formatTokens(savedOnScreen(view))}`, { tone: 'success', isBold: true }),
79  span(theme, ' est. by routing', { isDim: true }),
80]
81
82const keptOutLine = (view: View, theme: Theme): Line => [
83  span(theme, padEnd('Kept out', LABEL_WIDTH)),
84  span(theme, formatTokens(view.keptOut), { tone: 'success' }),
85  span(theme, ' of main context', { isDim: true }),
86]
87
88/** The savings history as one text row: the terminal's sparkline. */
89export const sparklineLine = (view: View, columns: number): Line => {
90  const theme = themeOf(view)
91  const width = widthOf(columns) - LABEL_WIDTH
92
93  return [
94    span(theme, padEnd('Trend', LABEL_WIDTH)),
95    view.history.length < 2
96      ? span(theme, 'builds up over turns', { isDim: true })
97      : span(theme, sparkline(view.history, width), { tone: 'success' }),
98  ]
99}
100
101const tierLines = (view: View, theme: Theme, barWidth: number): readonly Line[] => {
102  const rows: readonly (readonly [string, number, Tone | undefined])[] = [
103    ['haiku', view.spend.haiku, TIER_TONES.haiku],
104    ['sonnet', view.spend.sonnet, TIER_TONES.sonnet],
105    ['opus', view.spend.opus, TIER_TONES.opus],
106    ...(view.spend.other > 0 ? [['other', view.spend.other, undefined] as const] : []),
107  ]
108  const peak = Math.max(1, ...rows.map(([, tokens]) => tokens))
109
110  return rows.map(([name, tokens, tone]) => [
111    span(theme, `  ${padEnd(name, LABEL_WIDTH - 2)}`),
112    span(theme, bar(tokens / peak, barWidth), tone === undefined ? {} : { tone }),
113    span(theme, `  ${padStart(formatTokens(tokens), 6)}`, { isDim: true }),
114  ])
115}
116
117const agentLine = (agent: ViewAgent, view: View, theme: Theme, width: number): Line => {
118  const rails = agent.guides.map(hasGuide => (hasGuide ? '│  ' : '   ')).join('')
119  const indent = `  ${rails}${agent.isLast ? '└─ ' : '├─ '}`
120  const glyph =
121    agent.status === 'running'
122      ? spinner(view.frame, theme.hasMotion)
123      : STATUS_GLYPHS[agent.status]
124  const tier = padEnd(agent.tier ?? 'inherit', 7)
125  const tokens = padStart(formatTokens(agent.tokens), 6)
126  const room = Math.max(8, width - indent.length - 2 - tier.length - 1 - tokens.length - 2)
127  const wasOn = agent.escalatedFrom === undefined ? '' : ` (was ${agent.escalatedFrom})`
128  // The escalation mark is added after the label is cut, so a long label never hides it;
129  // where the row is too narrow for the words, an arrow stands for them.
130  const escalation = wasOn === '' || room - wasOn.length >= 10 ? wasOn : ' ↑'
131
132  return [
133    span(theme, indent, { isDim: true }),
134    span(theme, `${glyph} `, { tone: STATUS_TONES[agent.status] }),
135    span(theme, `${tier} `, agent.tier === undefined ? { isDim: true } : { tone: TIER_TONES[agent.tier] }),
136    span(theme, padEnd(`${fit(agent.label, room - escalation.length)}${escalation}`, room)),
137    span(theme, `  ${tokens}`, { isDim: true }),
138  ]
139}
140
141const treeLines = (view: View, theme: Theme, width: number): readonly Line[] =>
142  view.agents.length === 0
143    ? [[span(theme, '  no sub-agents yet', { isDim: true })]]
144    : [
145        [span(theme, '  main', { isDim: true })],
146        ...view.agents.map(agent => agentLine(agent, view, theme, width)),
147      ]
148
149const learningLine = (view: View, theme: Theme, width: number): Line => [
150  span(theme, padEnd('Learning', LABEL_WIDTH)),
151  span(
152    theme,
153    view.learning.count === 0
154      ? 'no outcomes logged yet'
155      : fit(
156          `${String(view.learning.count)} runs · ${String(Math.round(view.learning.cleanRate * 100))}% clean · ${String(view.learning.escalations)} escalated`,
157          width - LABEL_WIDTH,
158        ),
159    { isDim: true },
160  ),
161]
162
163export type Dashboard = {
164  /** Header, meters and the saved counter, drawn above the savings sparkline. */
165  top: readonly Line[]
166  /** Per-tier spend, the delegation tree, learning and recent events, drawn below it. */
167  bottom: readonly Line[]
168}
169
170/**
171 * The dashboard as styled lines. The savings sparkline sits between the two
172 * halves so each surface can draw it its own way: `sparklineLine` as text on
173 * the terminal, `sparklineSvg` as a vector on the desktop.
174 */
175export const renderDashboard = (view: View, columns: number): Dashboard => {
176  const theme = themeOf(view)
177  const width = widthOf(columns)
178  // The widest row is the context meter: label, bar, then up to 21 cells of figures.
179  const barWidth = Math.max(4, Math.min(24, width - LABEL_WIDTH - 21))
180
181  return {
182    top: [
183      headerLine(view, theme),
184      contextLine(view, theme, barWidth),
185      sessionLine(view, theme),
186      savedLine(view, theme),
187      keptOutLine(view, theme),
188    ],
189    bottom: [
190      heading(theme, 'Per tier'),
191      ...tierLines(view, theme, barWidth),
192      heading(theme, 'Delegation'),
193      ...treeLines(view, theme, width),
194      learningLine(view, theme, width),
195      ...(view.notes.length === 0
196        ? []
197        : [
198            heading(theme, 'Recent'),
199            ...view.notes.map((note): Line => [span(theme, `  ${fit(note, width - 2)}`, { isDim: true })]),
200          ]),
201    ],
202  }
203}
204
src/ui/format.ts 85 lines
1const BLOCKS = ['▁', '▂', '▃', '▄', '▅', '▆', '▇', '█'] as const
2const SPINNER = ['◐', '◓', '◑', '◒'] as const
3
4/** 1234 -> "1.2k", 950 -> "950". A negative count keeps its sign. */
5export const formatTokens = (tokens: number): string => {
6  const size = Math.abs(tokens)
7  const sign = tokens < 0 ? '-' : ''
8
9  if (size >= 1_000_000) {
10    return `${sign}${(size / 1_000_000).toFixed(1)}M`
11  }
12
13  return size >= 1000 ? `${sign}${(size / 1000).toFixed(1)}k` : `${sign}${String(Math.round(size))}`
14}
15
16/** A fixed-width meter: `fraction` of `width` cells filled. */
17export const bar = (fraction: number, width: number): string => {
18  const cells = Math.max(1, Math.round(width))
19  const ratio = Number.isFinite(fraction) ? Math.min(1, Math.max(0, fraction)) : 0
20  const filled = Math.round(ratio * cells)
21
22  return '█'.repeat(filled) + '░'.repeat(cells - filled)
23}
24
25/** The last `width` values as block characters, scaled between their own low and high. */
26export const sparkline = (values: readonly number[], width: number): string => {
27  const shown = values.slice(-Math.max(1, width))
28
29  if (shown.length === 0) {
30    return ''
31  }
32
33  const low = Math.min(...shown)
34  const range = Math.max(...shown) - low
35
36  return shown
37    .map(value => {
38      const level = range === 0 ? 0 : Math.round(((value - low) / range) * (BLOCKS.length - 1))
39
40      return BLOCKS[level] ?? BLOCKS[0]
41    })
42    .join('')
43}
44
45/** The spinner glyph for a frame; one fixed glyph when motion is off. */
46export const spinner = (frame: number, hasMotion: boolean): string =>
47  hasMotion ? (SPINNER[Math.abs(frame) % SPINNER.length] ?? SPINNER[0]) : '◆'
48
49export const padEnd = (text: string, width: number): string =>
50  text.length >= width ? text : text + ' '.repeat(width - text.length)
51
52export const padStart = (text: string, width: number): string =>
53  text.length >= width ? text : ' '.repeat(width - text.length) + text
54
55/** Cuts text to a width, marking the cut. */
56export const fit = (text: string, width: number): string => {
57  if (text.length <= width) {
58    return text
59  }
60
61  return width <= 1 ? text.slice(0, Math.max(0, width)) : `${text.slice(0, width - 1)}…`
62}
63
64/**
65 * The savings history as a standalone SVG polyline, for surfaces that draw
66 * vectors. `currentColor` keeps it readable in any theme and under NO_COLOR.
67 */
68export const sparklineSvg = (values: readonly number[], width = 240, height = 36): string => {
69  // Fewer than two readings draw as a flat line rather than nothing.
70  const points = values.length < 2 ? [values[0] ?? 0, values[0] ?? 0] : values
71  const low = Math.min(...points)
72  const range = Math.max(...points) - low
73  const stepX = width / (points.length - 1)
74  const path = points
75    .map((value, index) => {
76      const x = (index * stepX).toFixed(1)
77      const y = (height - 2 - (range === 0 ? 0 : ((value - low) / range) * (height - 4))).toFixed(1)
78
79      return `${x},${y}`
80    })
81    .join(' ')
82
83  return `<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 ${String(width)} ${String(height)}" width="${String(width)}" height="${String(height)}"><polyline fill="none" stroke="currentColor" stroke-width="2" stroke-linejoin="round" stroke-linecap="round" points="${path}"/></svg>`
84}
85
src/ui/statusline.ts 39 lines
1import { formatTokens, sparkline, spinner } from './format.ts'
2import { themeOf } from './view.ts'
3import type { View } from './view.ts'
4
5const SPARK_WIDTH = 8
6
7/** The saved figure on screen: mid-ease while animating, final under reduced motion. */
8export const savedOnScreen = (view: View): number =>
9  view.hasReducedMotion ? view.saved : view.shownSaved
10
11/**
12 * The compact status line for the CLI, or undefined to clear it. Plain text:
13 * it carries no colour codes, so NO_COLOR needs nothing more here.
14 */
15export const renderStatusLine = (view: View): string | undefined => {
16  if (view.mode === 'off' || !view.hasStatusLine) {
17    return undefined
18  }
19
20  const theme = themeOf(view)
21  const glyph = view.activeCount > 0 ? spinner(view.frame, theme.hasMotion) : '●'
22  const label = view.mode === 'dry-run' ? 'tokensaver [dry-run]' : 'tokensaver'
23  const parts = [
24    `${glyph} ${label} ${formatTokens(view.total)} tok`,
25    `saved ~${formatTokens(savedOnScreen(view))}`,
26    `H ${formatTokens(view.spend.haiku)} S ${formatTokens(view.spend.sonnet)} O ${formatTokens(view.spend.opus)}`,
27  ]
28
29  if (view.history.length >= 2) {
30    parts.push(sparkline(view.history, SPARK_WIDTH))
31  }
32
33  if (view.activeCount > 0) {
34    parts.push(`${String(view.activeCount)} agent${view.activeCount === 1 ? '' : 's'}`)
35  }
36
37  return parts.join(' · ')
38}
39
src/ui/theme.ts 36 lines
1/** Colours are named by meaning, as Claude Code's theme keys are, so they follow the user's theme. */
2export type Tone = 'success' | 'warning' | 'error' | 'claude' | 'suggestion' | 'inactive'
3
4/** One run of text with one style. The renderers build these; a surface paints them. */
5export type Span = {
6  text: string
7  tone?: Tone
8  isBold?: boolean
9  isDim?: boolean
10}
11
12export type Line = readonly Span[]
13
14export type Theme = {
15  /** False under NO_COLOR: no span carries a tone, and meaning rests on glyphs and words. */
16  hasColor: boolean
17  /** False under reduced motion: one static frame, and counters show their final value. */
18  hasMotion: boolean
19}
20
21export type SpanStyle = {
22  tone?: Tone
23  isBold?: boolean
24  isDim?: boolean
25}
26
27/** Builds a span for the theme: under NO_COLOR the tone is left out entirely. */
28export const span = (theme: Theme, text: string, style: SpanStyle = {}): Span => ({
29  text,
30  ...(theme.hasColor && style.tone !== undefined ? { tone: style.tone } : {}),
31  ...(style.isBold === true ? { isBold: true } : {}),
32  ...(style.isDim === true ? { isDim: true } : {}),
33})
34
35export const plain = (line: Line): string => line.map(part => part.text).join('')
36
src/ui/view.ts 131 lines
1import { summarizeLog } from '../core/log.ts'
2import type { TaskType } from '../core/score.ts'
3import type { Tier } from '../core/tiers.ts'
4import { activeCount, hasFreshFlash } from '../runtime/session.ts'
5import type { AgentNode, AgentStatus, FlashTone, Session, Spend } from '../runtime/session.ts'
6import type { Theme } from './theme.ts'
7
8export type ViewAgent = {
9  key: string
10  label: string
11  tier: Tier | undefined
12  taskType: TaskType
13  status: AgentStatus
14  tokens: number
15  escalatedFrom: Tier | undefined
16  /** One entry per ancestor level: true where a guide line still runs down beside this row. */
17  guides: readonly boolean[]
18  isLast: boolean
19}
20
21/** Everything the status line and the dashboard draw, as plain data. */
22export type View = {
23  mode: 'on' | 'off' | 'dry-run'
24  hasNoColor: boolean
25  hasReducedMotion: boolean
26  hasStatusLine: boolean
27  context: { tokens: number; window: number }
28  spend: Spend
29  total: number
30  saved: number
31  shownSaved: number
32  keptOut: number
33  history: readonly number[]
34  agents: readonly ViewAgent[]
35  activeCount: number
36  learning: { count: number; cleanRate: number; escalations: number }
37  notes: readonly string[]
38  frame: number
39  /** The strip above the prompt: the sub-agents running now and the newest event, while fresh. */
40  band: {
41    agents: readonly ViewAgent[]
42    flash: { text: string; tone: FlashTone } | undefined
43  }
44  /** Every sub-agent whose model tokensaver chose, for marking its row in the transcript. */
45  routed: readonly ViewAgent[]
46}
47
48const MAX_DEPTH = 4
49
50/** Orders agents as a tree walk: each agent, then the agents it spawned. */
51const walk = (
52  nodes: readonly AgentNode[],
53  parentAgentId: string | undefined,
54  guides: readonly boolean[],
55): readonly ViewAgent[] => {
56  const known = new Set(nodes.flatMap(node => (node.agentId === undefined ? [] : [node.agentId])))
57  const children = nodes.filter(node =>
58    parentAgentId === undefined
59      ? node.parentAgentId === undefined || !known.has(node.parentAgentId)
60      : node.parentAgentId === parentAgentId,
61  )
62
63  return children.flatMap((node, index) => [
64    {
65      key: node.key,
66      label: node.label,
67      tier: node.tier,
68      taskType: node.taskType,
69      status: node.status,
70      tokens: node.tokens,
71      escalatedFrom: node.escalatedFrom,
72      guides,
73      isLast: index === children.length - 1,
74    },
75    ...(node.agentId === undefined || guides.length >= MAX_DEPTH
76      ? []
77      : walk(nodes, node.agentId, [...guides, index < children.length - 1])),
78  ])
79}
80
81const modeOf = (session: Session): View['mode'] => {
82  if (!session.config.isEnabled) {
83    return 'off'
84  }
85
86  return session.config.isDryRun ? 'dry-run' : 'on'
87}
88
89export const toView = (session: Session): View => {
90  const summary = summarizeLog(session.outcomes)
91  const agents = walk(session.agents, undefined, [])
92
93  return {
94    mode: modeOf(session),
95    hasNoColor: session.hasNoColor,
96    hasReducedMotion: session.config.hasReducedMotion,
97    hasStatusLine: session.config.hasStatusLine,
98    context: session.context,
99    spend: session.spend,
100    total: session.spend.haiku + session.spend.sonnet + session.spend.opus + session.spend.other,
101    saved: session.saved,
102    shownSaved: session.shownSaved,
103    keptOut: session.keptOut,
104    history: session.savedHistory,
105    agents,
106    activeCount: activeCount(session),
107    learning: {
108      count: summary.count,
109      cleanRate: summary.cleanRate,
110      escalations: summary.escalations,
111    },
112    notes: session.notes,
113    frame: session.frame,
114    band: {
115      agents: agents.filter(agent => agent.status === 'running' || agent.status === 'retrying'),
116      flash:
117        session.flash !== undefined && hasFreshFlash(session)
118          ? { text: session.flash.text, tone: session.flash.tone }
119          : undefined,
120    },
121    routed: agents.filter(agent =>
122      session.agents.some(node => node.key === agent.key && node.isRouted && !node.isDryRun),
123    ),
124  }
125}
126
127export const themeOf = (view: View): Theme => ({
128  hasColor: !view.hasNoColor,
129  hasMotion: !view.hasReducedMotion,
130})
131
src/core/escalate.ts 106 lines
1import { nextTierUp } from './tiers.ts'
2import type { Tier } from './tiers.ts'
3
4export const CHECK_SIGNALS = [
5  'ok',
6  'error',
7  'failing-tests',
8  'low-confidence',
9  'empty',
10  'refusal',
11  'user-rejection',
12] as const
13
14/** What the checks made of a sub-agent's result. Anything but `ok` is a failure. */
15export type CheckSignal = (typeof CHECK_SIGNALS)[number]
16
17export const isCheckSignal = (value: unknown): value is CheckSignal =>
18  CHECK_SIGNALS.some(signal => signal === value)
19
20const FAILING_TESTS_PATTERN =
21  /\b\d+\s+(?:tests?\s+)?fail(?:ed|ing|ures?)\b|\btests?\s+(?:are\s+)?(?:still\s+)?fail(?:ed|ing)\b|\bAssertionError\b|^\s*FAIL\b/im
22
23/**
24 * Explicit statements of doubt or of not finishing. "Found nothing" is not
25 * here on purpose: an empty search result is a valid answer, not a failure.
26 */
27const LOW_CONFIDENCE_PATTERN =
28  /\b(?:i(?:'m| am) not (?:sure|certain|confident)|not confident|i (?:may|might) be wrong|(?:unable|was not able|wasn't able) to (?:complete|finish|determine)|could(?: not|n't) (?:complete|finish|determine)|ran out of (?:time|context|turns)|gave up|incomplete|best guess)\b/i
29
30const REJECTION_PATTERN =
31  /^\s*(?:no\b|nope\b|wrong\b|incorrect\b|that(?:'s| is| was)? (?:wrong|not right|incorrect|not what)|not what i|this is wrong|undo\b|revert\b|try again\b|redo\b|(?:it|that|this) (?:didn'?t|doesn'?t|did not|does not) work|still (?:broken|failing|wrong|not working))/i
32
33export type ResultCheck = {
34  isError: boolean
35  text: string
36}
37
38/** Runs the deterministic checks over a sub-agent's result. */
39export const checkResult = ({ isError, text }: ResultCheck): CheckSignal => {
40  if (isError) {
41    return 'error'
42  }
43
44  if (text.trim() === '') {
45    return 'empty'
46  }
47
48  if (FAILING_TESTS_PATTERN.test(text)) {
49    return 'failing-tests'
50  }
51
52  return LOW_CONFIDENCE_PATTERN.test(text) ? 'low-confidence' : 'ok'
53}
54
55/** True when the user's next prompt opens by rejecting the previous result. */
56export const isUserRejection = (prompt: string): boolean => REJECTION_PATTERN.test(prompt)
57
58export type EscalationInput = {
59  signal: CheckSignal
60  /** The tier the failed attempt ran on; undefined when tokensaver did not pick it. */
61  tier: Tier | undefined
62  /** Retries already spent on this task. */
63  attempts: number
64  isDenied: boolean
65  isDryRun: boolean
66}
67
68export type EscalationPlan =
69  | { shouldEscalate: true; to: Tier }
70  | { shouldEscalate: false; reason: string }
71
72/** The most retries one task gets: escalation is one step up, once. */
73export const MAX_RETRIES = 1
74
75/**
76 * Decides whether a failed result is retried one tier up. The answer is
77 * either a strictly higher tier or no retry at all: it never moves down.
78 */
79export const planEscalation = (input: EscalationInput): EscalationPlan => {
80  if (input.isDenied) {
81    return { shouldEscalate: false, reason: 'the call was denied, and a denial is final' }
82  }
83
84  if (input.signal === 'ok') {
85    return { shouldEscalate: false, reason: 'the result passed its checks' }
86  }
87
88  if (input.isDryRun) {
89    return { shouldEscalate: false, reason: 'dry run' }
90  }
91
92  if (input.tier === undefined) {
93    return { shouldEscalate: false, reason: 'tokensaver did not pick this model' }
94  }
95
96  if (input.attempts >= MAX_RETRIES) {
97    return { shouldEscalate: false, reason: 'already retried once' }
98  }
99
100  const to = nextTierUp(input.tier)
101
102  return to === undefined
103    ? { shouldEscalate: false, reason: 'already on the top tier' }
104    : { shouldEscalate: true, to }
105}
106
src/core/route.ts 148 lines
1import type { TaskScore, TaskType } from './score.ts'
2import { TOP_RANK, nextTierUp, rankOf, rankOfModel, tierAt } from './tiers.ts'
3import type { Tier } from './tiers.ts'
4
5/** A complexity below `haikuMax` runs on Haiku, below `sonnetMax` on Sonnet, the rest on Opus. */
6export type Cutoffs = {
7  haikuMax: number
8  sonnetMax: number
9  /** The log sequence number these cutoffs were last tuned at; older outcomes are spent evidence. */
10  tunedAtSeq: number
11}
12
13export type Thresholds = Readonly<Record<TaskType, Cutoffs>>
14
15/**
16 * Starting cutoffs. Mechanical work gets generous ones; reasoning and
17 * unknown work get tight ones, so an unrecognised task defaults upward.
18 */
19export const DEFAULT_THRESHOLDS: Thresholds = {
20  search: { haikuMax: 45, sonnetMax: 75, tunedAtSeq: 0 },
21  'bulk-read': { haikuMax: 40, sonnetMax: 75, tunedAtSeq: 0 },
22  summarize: { haikuMax: 35, sonnetMax: 70, tunedAtSeq: 0 },
23  'repetitive-edit': { haikuMax: 30, sonnetMax: 65, tunedAtSeq: 0 },
24  edit: { haikuMax: 15, sonnetMax: 55, tunedAtSeq: 0 },
25  reasoning: { haikuMax: 5, sonnetMax: 35, tunedAtSeq: 0 },
26  general: { haikuMax: 10, sonnetMax: 50, tunedAtSeq: 0 },
27}
28
29export type RouteDecision = {
30  tier: Tier
31  /** The tier the cutoffs alone picked, before any upward adjustment. */
32  base: Tier
33  reasons: readonly string[]
34}
35
36const bump = (tier: Tier): Tier => nextTierUp(tier) ?? tier
37
38/**
39 * Picks a tier for a scored task. Every adjustment after the cutoffs moves
40 * up: a boundary tie, an unsure score, a boost and a risky task all raise.
41 */
42export const routeTier = (score: TaskScore, thresholds: Thresholds, boost = 0): RouteDecision => {
43  const cutoffs = thresholds[score.taskType]
44  const base: Tier =
45    score.complexity < cutoffs.haikuMax
46      ? 'haiku'
47      : score.complexity < cutoffs.sonnetMax
48        ? 'sonnet'
49        : 'opus'
50  const reasons: string[] = [`${score.taskType} at complexity ${String(score.complexity)}`]
51  let tier = base
52
53  if (score.isUnsure) {
54    tier = bump(tier)
55    reasons.push('unsure, so one tier up')
56  }
57
58  for (let step = 0; step < boost; step += 1) {
59    tier = bump(tier)
60  }
61
62  if (boost > 0) {
63    reasons.push('previous result rejected, so one tier up')
64  }
65
66  if (score.isRisky) {
67    tier = 'opus'
68    reasons.push(`risky (${score.risks.join(', ')}), so never below the top tier`)
69  }
70
71  return { tier, base, reasons }
72}
73
74export type SpawnContext = {
75  /** The parent's effective model id: what the sub-agent inherits when no model is set. */
76  parentModel: string
77  /** A model the caller asked for by name, if any. */
78  explicitModel?: string | undefined
79  /** A tier an escalation retry must run on. */
80  forcedTier?: Tier | undefined
81}
82
83export type SpawnRouting = {
84  /** The model to hand `agent.spawn`, or undefined to leave the spawn exactly as it came. */
85  model: string | undefined
86  /** The tier the sub-agent runs on, for the log and the dashboard; undefined when unknown. */
87  tier: Tier | undefined
88  note: string
89}
90
91/** A family above the top routable tier has no tier name of its own. */
92const tierForRank = (rank: number): Tier | undefined =>
93  rank > TOP_RANK ? undefined : tierAt(rank)
94
95/**
96 * Turns a tier decision into the model a spawn should run on.
97 *
98 * Invariants the safety tests hold this to: a risky task never runs below the
99 * parent's tier or below a model the caller named, and an unknown parent
100 * model is never rerouted.
101 */
102export const resolveSpawnModel = (
103  decision: RouteDecision,
104  score: TaskScore,
105  context: SpawnContext,
106): SpawnRouting => {
107  if (context.forcedTier !== undefined) {
108    return { model: context.forcedTier, tier: context.forcedTier, note: 'escalation retry' }
109  }
110
111  const parentRank = rankOfModel(context.parentModel)
112  const explicitRank = rankOfModel(context.explicitModel)
113
114  if (parentRank === undefined) {
115    return { model: undefined, tier: undefined, note: 'unknown parent model, left as is' }
116  }
117
118  if (context.explicitModel !== undefined && explicitRank === undefined) {
119    return { model: undefined, tier: undefined, note: 'caller named an unknown model, left as is' }
120  }
121
122  const decided = rankOf(decision.tier)
123  // The top decision means "the best available": under a parent above Opus that is the parent.
124  const wanted = explicitRank ?? (decision.tier === 'opus' ? Math.max(decided, parentRank) : decided)
125  const floor = score.isRisky ? Math.max(parentRank, explicitRank ?? 0) : 0
126  const final = Math.max(wanted, floor)
127  const current = explicitRank ?? parentRank
128
129  if (final === current) {
130    return {
131      model: undefined,
132      tier: tierForRank(final),
133      note: context.explicitModel === undefined ? 'inherits the parent model' : 'caller chose the model',
134    }
135  }
136
137  // Reaching the parent's rank from a lower named model: name the parent's own model.
138  if (final === parentRank) {
139    return {
140      model: context.parentModel,
141      tier: tierForRank(final),
142      note: 'risky, raised to the parent model',
143    }
144  }
145
146  return { model: tierAt(final), tier: tierAt(final), note: decision.reasons.join('; ') }
147}
148
src/core/score.ts 200 lines
1/**
2 * Deterministic task scoring. No model call: every signal is a regular
3 * expression or a count over the task text, so scoring costs zero tokens.
4 */
5
6export const TASK_TYPES = [
7  'search',
8  'bulk-read',
9  'repetitive-edit',
10  'summarize',
11  'edit',
12  'reasoning',
13  'general',
14] as const
15
16export type TaskType = (typeof TASK_TYPES)[number]
17
18export const RISK_KINDS = ['security', 'destructive', 'production'] as const
19
20export type RiskKind = (typeof RISK_KINDS)[number]
21
22/** Each feature is normalised to 0..1. */
23export type Features = {
24  scope: number
25  files: number
26  ambiguity: number
27  depth: number
28  risk: number
29}
30
31export type TaskScore = {
32  taskType: TaskType
33  features: Features
34  /** 0..100: how much reasoning the task needs, which is what picks the tier. */
35  complexity: number
36  isRisky: boolean
37  risks: readonly RiskKind[]
38  /** True when the heuristics cannot place the task with confidence. */
39  isUnsure: boolean
40  /** Files the task names, or the count it states. */
41  namedFiles: number
42}
43
44export type TaskInput = {
45  text: string
46  description?: string | undefined
47  subagentType?: string | undefined
48}
49
50const RISK_PATTERNS: Readonly<Record<RiskKind, RegExp>> = {
51  security:
52    /\b(auth(?:entication|orization|n|z)?|passwords?|passphrases?|secrets?|credentials?|(?:access|auth|api|bearer|refresh|session)[- ]tokens?|api[- ]?keys?|private keys?|encrypt\w*|decrypt\w*|oauth|jwt|csrf|xss|sql injection|vulnerab\w*|security|permissions?|sudo|chmod|chown|ssh|certificates?|iam|rbac|cve-\d+)\b|\.env\b/gi,
53  destructive:
54    /\b(rm -rf?|delete\w*|drop (?:table|database|schema|index)|truncate\w*|wipe\w*|purge\w*|destroy\w*|force[- ]push\w*|reset --hard|git clean|overwrit\w*|erase\w*|remove all|uninstall\w*|migrations?|rollback|format (?:the )?(?:disk|drive))\b|push (?:-f\b|--force)/gi,
55  production:
56    /\b(prod(?:uction)?|deploy\w*|releas\w*|live (?:site|system|server|database|db|traffic)|customer data|billing|payments?|invoic\w*|payroll|terraform|kubectl|kubernetes|infra(?:structure)?|ci\/cd|dns|rollout|hotfix)\b/gi,
57}
58
59/** Task types that can be told from the text. `general` is what is left. */
60type KnownType = Exclude<TaskType, 'general'>
61
62const TYPE_PATTERNS: Readonly<Record<KnownType, RegExp>> = {
63  summarize:
64    /\b(summari[sz]\w*|summary|tl;?dr|recap|digest|condense\w*|overview|gist)\b/gi,
65  search:
66    /\b(find|search\w*|grep|locate|look (?:for|up)|where (?:is|are|does|do)|which files?|list (?:all|every)|usages?|references? to|occurrences?|who calls|callers of)\b/gi,
67  'bulk-read':
68    /\b(read (?:all|every|each|through)|go through|scan\w*|inventory|catalogu?e?|survey|review (?:all|every|each)|collect (?:all|every)|gather\w*|crawl\w*)\b/gi,
69  'repetitive-edit':
70    /\b(renam\w*|replace (?:all|every|each)|(?:find|search) and replace|codemod|bulk|mass|in (?:all|every|each) files?|every (?:file|occurrence|instance)|across all files|update (?:all|every|each)|apply the same|convert (?:all|every|each)|reformat\w*|bump (?:all|the) versions?)\b/gi,
71  reasoning:
72    /\b(design\w*|architect\w*|why|root cause|debug\w*|diagnos\w*|race conditions?|deadlocks?|trade-?offs?|algorithms?|prove|optimi[sz]\w*|plan|strategy|investigat\w*|concurrency|refactor\w*|redesign\w*|evaluat\w*|compar\w*|decide|reason about|analy[sz]\w*|performance|memory leaks?)\b/gi,
73  edit: /\b(fix\w*|implement\w*|add|chang\w*|writ\w*|updat\w*|creat\w*|modif\w*|remov\w*|edit\w*|patch\w*|build|make)\b/gi,
74}
75
76/**
77 * Tie-break order: when two types match equally often, the one that routes
78 * to the higher tier wins, so a tie never sends work to a cheaper model.
79 */
80const TYPE_PRIORITY: readonly KnownType[] = [
81  'reasoning',
82  'edit',
83  'repetitive-edit',
84  'summarize',
85  'bulk-read',
86  'search',
87]
88
89const SCOPE_PATTERN =
90  /\b(all|every|each|entire|whole|across|everywhere|codebase|repo(?:sitory)?|project-wide|global(?:ly)?|monorepo|recursive(?:ly)?)\b|\*\*?\/|\*\.\w+/gi
91
92const PATH_PATTERN =
93  /(?:[\w.-]+\/)+[\w.-]+|\b[\w-]+\.(?:tsx?|jsx?|mjs|cjs|py|go|rs|java|rb|md|json|ya?ml|toml|s?css|html|sql|sh|c|cpp|h|cs|php|kt|swift)\b/gi
94
95const FILE_COUNT_PATTERN = /\b(\d{1,4}) files?\b/i
96
97const VAGUE_PATTERN =
98  /\b(somehow|maybe|perhaps|something|stuff|things?|etc|whatever|figure out|not sure|i guess|kind of|sort of|improve\w*|better|clean ?up|tidy|polish|as needed|if needed|appropriate(?:ly)?)\b/gi
99
100const IDENTIFIER_PATTERN = /`[^`]+`|\b[a-z]+[A-Z]\w*\b|\b\w+_\w+\b/g
101
102const STEP_PATTERN = /\b(then|after that|afterwards|finally|next,)\b|^\s*\d+[.)]\s/gim
103
104const clamp01 = (value: number): number => Math.min(1, Math.max(0, value))
105
106const matchesOf = (text: string, pattern: RegExp): readonly string[] =>
107  Array.from(text.matchAll(pattern), match => match[0].toLowerCase())
108
109const countOf = (text: string, pattern: RegExp): number => matchesOf(text, pattern).length
110
111const wordCountOf = (text: string): number => text.split(/\s+/).filter(word => word !== '').length
112
113const classify = (text: string): TaskType => {
114  const ranked = TYPE_PRIORITY.map((type, priority) => ({
115    type,
116    priority,
117    hits: countOf(text, TYPE_PATTERNS[type]),
118  })).sort((a, b) => b.hits - a.hits || a.priority - b.priority)
119
120  const best = ranked[0]
121
122  return best === undefined || best.hits === 0 ? 'general' : best.type
123}
124
125const risksOf = (text: string): readonly RiskKind[] =>
126  RISK_KINDS.filter(kind => countOf(text, RISK_PATTERNS[kind]) > 0)
127
128const namedFilesOf = (text: string): number => {
129  const paths = new Set(matchesOf(text, PATH_PATTERN)).size
130  const stated = Number(FILE_COUNT_PATTERN.exec(text)?.[1] ?? 0)
131
132  return Math.max(paths, stated)
133}
134
135const filesFeature = (count: number, scope: number): number => {
136
137  // A broad task that names no file touches an unknown number of them.
138  if (count === 0 && scope >= 0.34) {
139    return 0.5
140  }
141
142  return clamp01(count / 8)
143}
144
145const ambiguityFeature = (text: string, words: number): number => {
146  const anchors = countOf(text, IDENTIFIER_PATTERN) + new Set(matchesOf(text, PATH_PATTERN)).size
147  const vague = countOf(text, VAGUE_PATTERN) / 3
148  const isTerse = words < 6 && anchors === 0
149  const isConcrete = anchors >= 2
150
151  return clamp01(vague + (isTerse ? 0.4 : 0) - (isConcrete ? 0.2 : 0))
152}
153
154const depthFeature = (text: string, words: number): number => {
155  const reasoning = countOf(text, TYPE_PATTERNS.reasoning) / 3
156  const steps = Math.min(0.2, countOf(text, STEP_PATTERN) * 0.1)
157
158  return clamp01(reasoning + steps + (words > 150 ? 0.2 : 0))
159}
160
161/** Scores a task from its text alone. Pure and deterministic. */
162export const scoreTask = (input: TaskInput): TaskScore => {
163  const text = [input.description ?? '', input.text].join('\n').trim()
164  const words = wordCountOf(text)
165  const scope = clamp01(countOf(text, SCOPE_PATTERN) / 3)
166  const risks = risksOf(text)
167  const namedFiles = namedFilesOf(text)
168  const riskHits = RISK_KINDS.reduce((sum, kind) => sum + countOf(text, RISK_PATTERNS[kind]), 0)
169  const features: Features = {
170    scope,
171    files: filesFeature(namedFiles, scope),
172    ambiguity: ambiguityFeature(text, words),
173    depth: depthFeature(text, words),
174    risk: clamp01(riskHits / 2),
175  }
176  const taskType = classify(text)
177  const complexity = Math.round(
178    100 *
179      clamp01(
180        0.45 * features.depth +
181          0.25 * features.ambiguity +
182          0.15 * features.scope +
183          0.15 * features.files,
184      ),
185  )
186
187  return {
188    taskType,
189    features,
190    complexity,
191    isRisky: risks.length > 0,
192    risks,
193    isUnsure: features.ambiguity >= 0.6 || taskType === 'general' || words < 3,
194    namedFiles,
195  }
196}
197
198export const isTaskType = (value: unknown): value is TaskType =>
199  TASK_TYPES.some(type => type === value)
200