SLOPSHOPPER

effort-router

Make your Claude Code usage go up to twice as far. Picks the right reasoning effort for every prompt and every subagent, so easy work stops burning your limits…

newbandspinnerguardcommandprompt
v0.19.0MITupdated 2026-10-07tommy5dollar/effort-router/effort-router
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · effort-router
› fix the failing auth test and add an audit log call ● effort-router: effort-router: first sighting with 1 prompts already in the session ● effort-router: effort-router: rules from built-in defaults (base) ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /effort-router ● effort-router: effort-router: assessment settled in 30040 ms (after a prompt) ● effort-router: effort-router: fork assessment said (after a prompt) OK Effort router: unlocked. Your effort setting applies. Locks after 3 more prompts. 1: Hide 2: Lock 3: Turn off 4: Assess ⟨Claude Code's own drawing⟩ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⟨Claude Code's own drawing⟩ 🔓 …

Draws

Band
Effort router: unlocked. Your effort setting applies. Locks after 3 more prompts. 1: Hide 2: Lock 3: Turn off 4: Assess ⟨Claude Code's own drawing⟩
README

effort-router

A Claude Code plugin that saves time and money by running each task at the lowest reasoning effort that does it well. Your session's own model assesses each of your first five prompts and moves the level to whichever gets the work done fastest and cheapest, up or down. Then the level locks for the rest of the session. Each subagent gets its own level, chosen by the agent that launches it.

<img src="https://raw.githubusercontent.com/tommy5dollar/effort-router/main/docs/launch.gif" width="720" alt="Claude Code is overthinking your renames. Everyday prompts all run at high and usage drains fast. effort-router picks the effort for every prompt, easy ones drop to low, and usage lasts far longer. Up to 2x the usage, up to 3.5x faster, same results">

It works with Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 5.5, in the terminal and in the Desktop app's Code tab. It requires Claude Code 2.1.287 or later in the terminal, or the Desktop app with Claude Code 2.1.286 or later. Haiku 5.5 needs Claude Code 2.1.293 or later.

In the Desktop app, first add the marketplace: go to Customize, then Plugins, Add marketplace, Add from a repository, and enter tommy5dollar/effort-router. The app then opens the new marketplace: click Add on effort-router (if you've left that screen, it's under effort-router at the bottom of Plugins, Discover). In the terminal:

claude plugin install effort-router --marketplace tommy5dollar/effort-router

Before Claude Code 2.1.292 that's two commands: claude plugin marketplace add tommy5dollar/effort-router, then claude plugin install effort-router@effort-router.

Then start a new session. In the Desktop app the footer appears once you've sent the first message.

Contents

Why

Claude Code runs every request at the level you picked. Anthropic's Using Claude Code: Spending your effort (Thariq Shihipar, 25 September 2026) found that effort buys verification and edge-case testing, not a better approach, so the right level depends on the task. The router gives your session's model those principles and sourced notes on what each level can do on that model, then lets it judge. It never maps a kind of task to a fixed level, because level names mean different things on Opus, Sonnet and Fable.

Does it save money and time? That's what it's for. When two levels would both do the work, it picks the cheaper one. The higher you run, the more it saves. In our test, Fable 5.1 on xhigh was given three small chores in a payments repo to hand to subagents. The router moved it to medium and its Opus subagents to medium or low. It finished in about 2 minutes for $1.10 to $1.65, against 6 to 8.5 minutes and $3.20 to $3.75 left on xhigh, and every test passed both ways. Those figures include the router's own assessments. On Opus 5.5's default of medium there's less to step down from, so it mostly picks off the small tasks. It still steps up when the work clearly needs it, because a hard task done right first time costs less than the rework, in tokens and in your own time. /er report shows what ran at each level, so you can see what it did to your own work.

What routing costs. Assessments run on your own login, so they come out of your normal usage, with nothing extra to pay or sign up for. The figures here are at API prices. Each of the first five prompts waits about 1.5 seconds for an assessment. The first is a separate call that can't use the prompt cache: about 5 cents on Opus 5.5 or 13 cents on Fable 5.1. The other four read your conversation from the session's cache, about 3 cents each on Opus. So a session costs about 20 cents to route on Opus, then nothing more. A subagent's assessment is about 2 cents.

What it reads, sends and stores

This plugin runs inside Claude Code without a sandbox, so here's exactly what it does:

  • Sends: assessments go to your session's own model through Claude Code, on your existing login. If your organisation has set up Claude Code's OpenTelemetry, it adds a few attributes to Claude Code's own records for that collector (see Telemetry). Nothing else leaves your machine.
  • Reads: your conversation, your CLAUDE.md files, rules and memory, your agent definitions and its own rules files (see How it assesses for what goes into each assessment).
  • Writes: one small JSON file per session in ~/.claude/effort-router/spend/, and nothing else.
  • Never: runs a process, changes your saved effort setting or changes the model.

Common questions

  • Does it change my model? No. It only changes effort, and stands aside on models it doesn't support.
  • Can I ask for more or less effort in my prompt? Yes, while the level is unlocked: name a level ("use low effort") or say "think really hard about this" or "quick one", and the assessment follows it. Asking for max gets xhigh unless highestLevel is max. Once the level has locked, prompts aren't assessed, so use the band or /er assess. Without the router, words like that only nudge how much the model thinks within the level you set.
  • What if I disagree with it? Change the effort picker and routing turns off for that session, with your level in force. Or use the band: Lock keeps the current level, Unlock lets it choose again and Assess has it look again.
  • What if an assessment fails or is slow? The prompt runs at the level it already had. Nothing waits longer than 30 seconds.
  • Does it work in VS Code or with -p? It routes there too, but there's no footer or band. Use /er instead.
  • Can it keep assessing instead of locking? Set promptsToAssess higher (say 50). Each extra assessment costs about 3 cents and 1.5 seconds on Opus 5.5.
  • Why does Claude Code still say "with medium effort"? That line shows your setting, not the level the request was sent at. See Known limits.

How it works

flowchart LR
    P["Prompts 1 to 5"] --> R["Your model picks<br>the level"]
    R --> C{"Different from<br>the level running?"}
    C -- "yes" --> M["Step up<br>or down"]
    C -- "no" --> K["Stay"]
    M --> L["After prompt 5,<br>lock"]
    K --> L
  1. It starts from your effort setting. Nothing changes until an assessment picks another level.
  2. Each of your first five prompts is assessed before its turn runs. Your session's model is asked one question: which level gets this session's work done in the least time and total inference cost, counting the rework that too little effort causes? It's told that people use the router to spend less, so when two levels would both do the work it picks the cheaper one, and that your own words about effort ("think really hard about this", "quick one") are your call. It answers with a level and a reason, or says no task has been stated yet.
  3. The session goes to the level it picked. Switching costs you nothing (no approval, no review), so the router doesn't second-guess the answer. If it picks the level already running, nothing changes.
  4. After the fifth assessment it locks whatever level is running. A move never locks early.
  5. Only an assessment changes the level, or a button whose label names the level.

Locking after a fixed number of prompts is deliberate. A fixed window always ends, still catches a task that grows over the first few prompts, and has one number to tune (promptsToAssess).

The footer

The footer sits beside the native model and effort pickers. It shows the router's status, the level running and, while unlocked, how much of the window is used.

<img src="https://raw.githubusercontent.com/tommy5dollar/effort-router/main/docs/footer.gif" width="720" alt="The footer in the Desktop app. A rename is assessed and moves from medium to low, a production bug moves from low to high while the effort picker still says Medium, and clicking the footer opens the band">

FooterWhat it means
🔓 MEDIUM ○Unlocked, nothing assessed yet. Your setting (medium) runs
🔓 HIGH ◔Unlocked with one prompt assessed, which moved the level to high
🔓 HIGH ◑Two or three assessed
🔓 HIGH ◕Four assessed. The next prompt is the last one assessed
🔓 …An assessment is running (the turn starts when it's done), or a new session hasn't shown your level yet
🔒 HIGHLocked. Every request on the main thread runs at high
⏸️ MEDIUMOff. Your own effort setting applies

The circle never fills. When the window ends the padlock closes instead. The level is in capitals, to tell it apart from the picker's own label, which shows your setting.

Levels have no colours, because a scale from green to red would suggest that low effort is good.

The band

Clicking the footer opens the band above the prompt. /er does the same. It never opens by itself.

Effort router: unlocked. High (chosen by the router). Locks after 1 more prompt.
Last assessment: high (bug fix touching three services), so it moved from medium.
Subagents get their own level: 4 routed this session.
1 Hide   2 Lock at high   3 Turn off   4 Assess

The first line is the footer in words: the status, the level and where it came from, and what happens next. The second is the last assessment. The third appears once a subagent has been routed.

The four buttons always sit in the same slots, so the digit keys are learnable. They run from doing nothing to taking action:

SlotUnlockedLockedOff
1HideHideHide
2Lock at highUnlockTurn on, locked at high
3Turn offTurn offTurn on, unlocked
4Assess (greyed out)AssessTurn on and assess

The band in the Desktop app after you locked it at high, with the last assessment and why

  • Lock ends assessing early when the level is plainly right.
  • Unlock keeps the locked level running and assesses your next five prompts from there. Use it when you're about to steer the work somewhere new.
  • Assess while locked runs one fresh assessment now. If it moves the level, the new level stays locked. This is the "the work changed, look again" case the article recommends. While unlocked it's greyed out, because your next prompt is assessed anyway and a re-roll invites fishing for an answer.
  • Turn off stops routing for this session. Your effort setting applies and subagents aren't routed.
  • Turn on, locked at high brings back the router's last level in one press. It's greyed out when the router never had a level of its own.
  • Turn on, unlocked starts a fresh window from your own setting.

A greyed-out button stays in its slot. Pressing it says why it's greyed out.

The band in the Desktop app after a move to high, with its four buttons

Commands

/effort-router sits beside /effort in the typeahead, and /er is its short form. Each verb does what the band's matching button does, so the commands are also the controls where there's no band (VS Code and -p).

CommandBandWhat it does
/erClicking the footerOpens the band. Where there's no band it prints the band's lines
/er lockSlot 2Locks at the level running. When off, turns on locked at the router's last level
/er unlockSlot 2Unlocks. When off, turns on unlocked
/er off, /er onSlot 3Turns routing off, or on and unlocked
/er assess [hint]Slot 4Assesses now. While unlocked it keeps the hint for your next prompt's assessment instead (/er assess this is a security review)
`/er report [session\week\month\all]`Where the effort went
/er statusThe band's lines, then the details for troubleshooting: the last reply in full and what it was judged against, assessments used, the last error and the routed subagents
/er rulesThe rules in force, where each layer came from, and the notes for your model

The verbs are explicit rather than toggles, so repeating one is safe ("Already locked at high"). Anything else is refused with the list above, so a typo never runs an assessment.

Messages in the conversation

Each change adds one dim line to the conversation, labelled effort-router by Claude Code. These lines are for you and are never sent to the model. An assessment that stays put adds nothing.

What changedMessage
An assessment moved the levelAssessed, medium to high (bug fix touching three services).
It locked after the last promptLocked at high.
You pressed LockYou locked it at high.
You pressed UnlockUnlocked. Assessing again from your next prompt.
You changed the effort pickerYou changed the effort to xhigh, so routing is off.
You turned it offOff. Your effort (medium) applies.
You turned it onOn, locked at high. Or: On, unlocked.

Changing the level yourself

Changing the effort picker turns routing off, whether it was locked or unlocked, and your new level applies from that request on. The router never sets your effort setting itself: it sets the level on each request instead. So the picker keeps showing your own level while the footer shows the one in use.

In the Desktop app the picker keeps showing your setting while the router runs another level. Picking the level it already shows changes nothing there, so use Turn off instead.

How it assesses

  • Before your prompt runs. Each of the first five prompts waits for one assessment, so the turn's first request already carries the level. If the assessment takes longer than 30 seconds or fails, the turn runs at the level it had, and /er status says why. A failed assessment still uses up its prompt, so the window ends when the footer says it will.
  • On your session's own model. The model you chose to work in judges the task, because it judges better than a small model and the savings from getting the level right scale with it. When the conversation has a request to fork, an assessment is a fork of it: the session's own request (system prompt, tools, CLAUDE.md, memory and the whole conversation) with one question added, served from the prompt cache. Measured on Opus 5.5 with a 72k-token conversation: 1.6 seconds, about 2.8k fresh input tokens and 40 output tokens.
  • The first prompt is a separate call. Before the session has sent anything there is no request to fork, and a plugin can't build one with Claude Code's system prompt and tools. So the first assessment is one call to the same model with your CLAUDE.md files, rules and memory (up to 80,000 characters, about 20k tokens) and your prompt. The same happens after /clear, or after a resume that starts afresh.
  • What a separate call reads. Your earlier prompts and answers (up to 4,000 characters each), Claude's replies shortened and tool calls as names only, up to 24,000 characters in all. Tool results, file contents and thinking never go in. Over the cap it keeps your first prompt (the original task), then the newest lines. The prompt being assessed always goes in whole, outside the cap, because a long dictated brief is the prompt that matters most.
  • Told the level running. That's the router's own level after a move, or your effort setting. The router learns your setting from the first request (nothing else shows it), so the first assessment's level is applied when that request arrives.
  • Up to xhigh. Assessments are offered levels up to highestLevel (xhigh by default): on all three models max rarely beats xhigh and can overthink. If your own setting is higher (max, say), they're offered levels up to yours, so a session you set to max can stay there.
  • Your answers count too. Answers to Claude's multiple-choice questions on the main thread are a human turn as well. They're assessed before they go back to Claude, by a fork that carries them. They count toward the window like a prompt.
  • No clear task yet. The model answers that only for opening filler: greetings, housekeeping such as "pull the latest code", or questions asked before any work. Once you've stated a real task it picks the level that task most likely needs, even while the details are open. The assessment still uses up its prompt.
  • The latest exchange counts most. A later clarification overrides an earlier ask, and a short reply is read against the question it answers.
  • A prompt sent while a turn is running isn't assessed. The next one is.
  • Every assessment is kept. Each one's level, reason, the level the session was on and what it did go into the session's ledger, for calibrating the model notes.

Sessions that started before the router

The first time the router sees a session, it counts the prompts already in it. Those count toward the window. A session that already has five or more prompts starts off and the band says "This session started before the router". Its subagents are still routed, because each brief is a fresh, whole task.

The router keeps each session's status in that session's ledger, so claude --resume picks up where it was.

Subagents

Without the router every subagent runs at the session's level, unless its agent definition sets one or you ask Claude for one. Since Claude Code 2.1.292 you can ask (run the reviewer at high), but Claude is told never to pick a level on its own judgement.

  • Its parent decides. When Claude launches a subagent, the launch waits for one fork of the parent's conversation, asked which level the subagent needs, with its brief. The parent knows the task and why it's delegating this part, which a brief alone often doesn't say. Then the subagent starts, and every request it makes carries that level. Measured on Opus 5.5: 2.3 to 3.3 seconds, the parent's conversation read from cache, about 2 cents. Before the parent's first reply there's nothing to fork, so it's a separate call that reads the brief alone.
  • On its own model. The assessment is told which model the subagent runs on (the Agent call's model, else its definition's, else the parent's) and gets that model's notes. Your rules and your organisation's apply here too.
  • Haiku 5.5 agents are judged on Haiku. A subagent on Haiku 5.5 gets a separate call on Haiku that reads its brief alone, not a fork of the parent. Haiku jobs are short and self-contained, and a fork on a bigger parent model could cost about what it saves and slow down the helper chosen for speed. It's offered up to high, because each level above that buys little on Haiku for many more steps. Without the router, a Haiku subagent runs at its parent's level, so an Opus session on xhigh runs its Haiku helpers on xhigh too.
  • Other models are left alone. A subagent on Haiku 4.5, or on another model the router doesn't support, isn't assessed.
  • A level you ask for wins. If the Agent call sets an effort because you, a CLAUDE.md or a skill asked for one, the router leaves that subagent alone.
  • An agent's own effort: wins. If the agent's definition sets an effort, the router leaves its requests alone and the engine applies that level. The router finds the definition by its name: in the project's .claude/agents/*.md, then your ~/.claude/agents/*.md, and in the agents key of policy, project and user settings. The first definition with that name decides, as it does for the engine.
  • Forks and failures take the parent's level. A fork shares its parent's context, so it isn't assessed. If an assessment fails, times out or gives no level, the subagent takes its parent's level too.
  • Turning the router off sends subagents back to your effort setting. Turning it on brings their routed levels back.
  • Seeing it. The band counts the subagents routed this session. /er status lists the last ten, newest first, with each one's level and why. Every routed subagent is also kept in the session's ledger, with the level it would have inherited from its parent, the level it got and why.

Set routeSubagents to false to leave subagents at the session's level.

Models

The router supports Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 5.5. Level names don't mean the same amount of thinking on each, and each responds to effort differently. In Claude Code, Opus 5.5, Sonnet 5.5 and Haiku 5.5 default to medium and Fable 5.1 to high. Opus 5.5 gains most from low to medium and little above high, while Sonnet 5.5 gains a lot at every step. Routing one like another would be a mistake.

Haiku 5.5 is the first Haiku with effort levels, and in practice a subagent model. The router picks up to high on it, below the highestLevel option, unless your own setting in a Haiku session is higher. Its price goes up 5 times on every token of a request whose prompt passes 100,000 tokens, so its notes tell the check to step up only when the task clearly needs it.

Each has a notes file in rules/models/ on how its levels behave, in the same shape for every model: how effort pays on it, Anthropic's advice for it, then each level with its cost and time against medium and how it behaves. They are heuristics. Benchmark scores are left out, because a few points on a hard benchmark means a few more of the hardest tasks solved, not every task done better. Every assessment carries the notes for the model it's about, after the routing rules. Lines about max are left out unless max is on offer. /er rules prints them. The evidence behind each line, with sources, is in rules/models/research-2026-10.md and its addendum. No eval results are in the notes, ours or anyone's. The router's routing eval (eval/routing.ts) checks 87 prompts against an approved level for each, and any change to the prompt, rules or notes has to pass it.

The notes guide the level instead of fixed rules because of an eval on 4 October 2026. With rules that tied kinds of task to levels, all three models gave almost the same answers and ignored their notes. Without those rules, each model's answers moved the way its evidence predicts (TESTING.md, "Prompt variants").

On any other model the router stands aside: the footer shows ⏸️ with your level, the band names the models it works with, and nothing is assessed. Its state is kept, so switching back with /model picks up where it was. A new model needs a new version of the router.

Where the effort went

/er report shows what your requests spent at each level over the last 7 days. /er report session, month or all cover other spans. For example:

Effort for the last 7 days (since 2026-09-28): 412 requests in 9 sessions, 610k output tokens.
By level:
- low: 120 requests, 31k out
Source 2 files
hooks/register.ts 1303 lines
1import type { AgentSpawnInput, EngineInterface, On, PluginOptions } from 'claude-code'
2
3import {
4  telemetryAttributes,
5  type AgentDefinition,
6  type BandAction,
7  type CheckKind,
8  type ComposedRules,
9  type LastAssessment,
10  type Level,
11  type ModelNotes,
12  type Proposal,
13  type ReadDiagnostics,
14  type RoutedAgent,
15  type RouterState,
16  type SpendLedger,
17  type SpendPeriod,
18  type SpendUsage,
19  type SubagentStatus,
20  type TranscriptMessage,
21  type View,
22  DEFAULT_HIGHEST,
23  DEFAULT_TRIM,
24  QUESTION_TOOL,
25  ROUTE_USAGE,
26  SUPPORTED_NAMES,
27  agentFileDefinition,
28  appliedLevel,
29  bandActions,
30  bandHeadline,
31  clampLevel,
32  classifierPrompt,
33  classifierSystem,
34  composeRules,
35  dayOf,
36  definitionFor,
37  emptyLedger,
38  firstSighting,
39  footerLabel,
40  forkPrompt,
41  freshState,
42  humanPromptCount,
43  isLevel,
44  lastAssessmentLine,
45  lockedByYou,
46  message,
47  modelName,
48  offeredLevels,
49  onModel,
50  parentLevel,
51  parseDecision,
52  parseLedger,
53  parseRoute,
54  parseSubagentReply,
55  rank,
56  renderTranscript,
57  restored,
58  routeReport,
59  routesSubagents,
60  ruleLayers,
61  runningLevel,
62  savedOf,
63  settingsAgentDefinitions,
64  settingsRulesOf,
65  settle,
66  spendReport,
67  subagentForkPrompt,
68  subagentLine,
69  subagentPrompt,
70  subagentSystem,
71  supportedModel,
72  turnedOff,
73  turnedOnLocked,
74  turnedOnUnlocked,
75  unlocked,
76  wantsAssessment,
77  withQuestionAnswer,
78  withRead,
79  withSpend,
80  withVerdictOutcome,
81  withSubagentRow,
82  withVerdictRow,
83} from './policy'
84
85/**
86 * effort-router 0.18. On each of a session's first prompts (`promptsToAssess`,
87 * 5 by default) the router assesses the conversation before the turn runs:
88 * a fork of the conversation on the session's own model (`$.model.fork`,
89 * served from its prompt cache), or, when there is nothing to fork yet, one
90 * separate call carrying the session's instructions (CLAUDE.md, rules,
91 * memory). It asks which level gets the work done in the least time and total
92 * inference cost, and the session goes to that level. After the last prompt
93 * of the window it locks whatever is running. A move never locks.
94 *
95 * Statuses: unlocked (it may still move the level), locked, off. The footer
96 * shows the status glyph, the level running and the share of the window
97 * used; clicking it opens the band, which says the same in words and holds
98 * four fixed slots: Hide, the lock, on/off, Assess. `/effort-router` (alias
99 * `/er`) is the band without a mouse.
100 *
101 * Changing the effort picker yourself turns routing off. Only an assessment,
102 * or a button whose label names a level, changes the level.
103 *
104 * Subagents are routed apart: each spawn waits for one read, a fork of its
105 * parent plus the subagent's brief, and its requests carry that level (forks,
106 * and failed reads, take the parent's level). A Haiku 5.5 subagent is read by
107 * Haiku from its brief alone, and offered up to high. An agent whose definition
108 * sets an effort is left to it, and so is one on a model the router doesn't
109 * support.
110 *
111 * The session's state lives in its ledger (`~/.claude/effort-router/spend/
112 * <session>.json`) beside its requests and assessments. Fail open everywhere:
113 * any error leaves the request at the effort it arrived with.
114 */
115
116type Settings = {
117  /** Prompts assessed before the router locks. */
118  promptsToAssess: number
119  /** The highest level the router picks, unless your own setting is higher. */
120  highestLevel: Level
121  /** Route each subagent from its own brief at spawn. */
122  routeSubagents: boolean
123}
124
125/** A session's ledger, where it is saved, and whether it changed since. */
126type Spend = { path?: string; ledger: SpendLedger; dirty: boolean; writing: Promise<void> }
127
128/** Per-session runtime facts. Only the state is saved (in the ledger). */
129type Session = {
130  state: RouterState
131  spend: Spend
132  /** An assessment is in flight. */
133  reading: boolean
134  /**
135   * Your effort setting: `e.effort` as the last main-thread request arrived,
136   * before the router's rewrite. Until a request shows it, a guess from the
137   * settings file's `effortLevel`, shown but never judged against.
138   */
139  picker?: string | number
140  pickerSeen?: boolean
141  /** The band was opened from the footer or `/er`. */
142  bandOpen: boolean
143  /** A line the band shows under its own (why a slot is greyed out, what an action did). */
144  note?: string
145  /** An assessment is running: the footer shows an ellipsis where the progress circle goes. */
146  assessing: boolean
147  /** Assessments this session, for `/er status`. */
148  calls: number
149  verdict?: ReadDiagnostics['verdict']
150  error?: ReadDiagnostics['error']
151  lastReadMs?: number
152  sent?: ReadDiagnostics['sent']
153  /** Routed subagents by agentId, oldest first (memory only: a resume starts empty). */
154  agents: Map<string, RoutedAgent>
155  /** The user's and project's agent definitions, scanned once per session. */
156  definitions?: Promise<AgentDefinition[]>
157  /** The main loop's model, as its last request named it (or `$.session.model()`). */
158  model?: string
159  /** The session's instructions block (CLAUDE.md files, rules, memory), as `prompt.context` carried it. */
160  instructions?: string
161  /** A main-thread turn is running. */
162  busy: boolean
163}
164
165/** What one assessment sees beyond the stored transcript. */
166type ReadInput = {
167  /** The prompt being submitted (not yet in the transcript at prompt.submit). */
168  current?: string
169  /** An AskUserQuestion call just answered (not yet in the transcript at tool.call). */
170  answer?: { toolUseId?: string; input: unknown; text: string }
171  hint?: string
172  /** What prompted it, for `/er status`. */
173  trigger: string
174}
175
176/** What one assessment found: its level (undefined = no clear task, or failed), how it was made, when it answered. */
177type Check = { proposal?: Proposal; kind: CheckKind; model: string; withInstructions?: boolean; checkedAt?: number; failed?: string }
178
179const SESSIONS = new Map<string, Session>()
180const LOADING = new Map<string, Promise<Session>>()
181/** Whether this session draws a band: an interactive terminal or Desktop. VS Code and `-p` don't. */
182let hasBand = true
183/** An assessment that takes longer is abandoned and the prompt runs at the level it has (fail open). */
184// 30 s: on Fable 5.1 at xhigh a fork took 7 s and 15 s (2026-10-05). A shorter limit made the prompt wait and then
185// threw away an answer already paid for.
186const TIMEOUT_MS = 30_000
187/** The most of the conversation a separate call reads, before the prompt being assessed (about 6k tokens). */
188const MAX_CHARS = DEFAULT_TRIM.totalChars
189/** The most instruction text a first assessment sends (about 20k tokens). */
190const MAX_INSTRUCTIONS_CHARS = 80_000
191/** An assessment may think at the model's default effort: room for that and the reply. */
192const SESSION_CHECK_MAX_TOKENS = 4000
193/** Routed subagents kept per session, for turn.step and `/er status`. */
194const MAX_ROUTED_AGENTS = 200
195/** The tool that launches subagents. Since 2.1.292 its call can carry an `effort`, set only when someone asked for one. */
196const AGENT_TOOL = 'Agent'
197/** Efforts asked for in Agent calls, by tool_use_id, from the call until its spawn. */
198const ASKED_EFFORT = new Map<string, Level>()
199/** The reason recorded for a subagent whose Agent call set its effort. */
200const ASKED_REASON = 'asked for when it was launched'
201
202const FALLBACK_RULES =
203  'Effort buys verification, edge-case testing and independent judgement, not a better approach. Weigh how much is ' +
204  'hidden (what a careful engineer could miss: edge cases, existing code, concurrency, security), whether the user ' +
205  'is in the loop, ' +
206  'and how well specified and how big the task is. Pick the level that does the work well on this model without ' +
207  "paying for thinking it won't use."
208
209const HUMAN_ORIGINS = new Set(['composer', 'bridge', 'sdk'])
210/** Sent with each telemetry record, so a collector can tell versions apart. Keep in step with plugin.json. */
211const VERSION = '0.19.0'
212const COMMANDS = ['effort-router', 'er']
213
214function settingsOf(options: PluginOptions): Settings {
215  const n = Number(options.promptsToAssess)
216  return {
217    promptsToAssess: Number.isFinite(n) && n >= 1 ? Math.floor(n) : 5,
218    highestLevel: isLevel(options.highestLevel) && options.highestLevel !== 'low' ? options.highestLevel : DEFAULT_HIGHEST,
219    routeSubagents: options.routeSubagents !== false && options.routeSubagents !== 'false',
220  }
221}
222
223// --- session state ---------------------------------------------------------------
224
225/** Where ledgers are kept: ~/.claude/effort-router/spend (undefined with no home directory). */
226async function spendDir($: EngineInterface): Promise<{ dir?: string; sep: string }> {
227  const { sep, homeDir } = await homeOf($)
228  return { dir: homeDir ? `${homeDir}${sep}.claude${sep}effort-router${sep}spend` : undefined, sep }
229}
230
231/** The session's repository, as a short name: its owner/name when known, else its root folder's name. */
232async function repoName($: EngineInterface): Promise<string> {
233  const repo = await $.session.repo().catch(() => null)
234  if (repo?.name) return repo.name
235  const root = repo?.root ?? (await $.session.root().catch(() => undefined))
236  return root?.split(/[\\/]/).filter(Boolean).pop() ?? 'unknown'
237}
238
239/** This session's ledger, from its file so a resume or a reload carries on from it. */
240async function loadSpend($: EngineInterface, id: string): Promise<Spend> {
241  const { dir, sep } = await spendDir($)
242  const path = dir ? `${dir}${sep}${id}.json` : undefined
243  const text = path ? await readText($, path) : undefined
244  const saved = text === undefined ? undefined : parseLedger(text)
245  return { path, ledger: saved?.session === id ? saved : emptyLedger(id, await repoName($)), dirty: false, writing: Promise.resolve() }
246}
247
248async function loadSession($: EngineInterface, id: string): Promise<Session> {
249  const spend = await loadSpend($, id)
250  let state = restored(spend.ledger.state)
251  const session: Session = { state: state ?? freshState(), spend, reading: false, bandOpen: false, assessing: false, calls: 0, agents: new Map(), busy: false }
252  // Your setting as the last request before a reload, restart or resume showed it, so a change made since is seen.
253  if (spend.ledger.setting) {
254    session.picker = spend.ledger.setting
255    session.pickerSeen = true
256  }
257  if (!state) {
258    // First sighting: prompts already in the session count toward the window. A session with the window already
259    // used up started before the router, and is left off.
260    const messages = (await $.session.messages().catch(() => [])) as TranscriptMessage[]
261    const prior = humanPromptCount(messages)
262    state = firstSighting(prior, currentSettings.promptsToAssess)
263    session.state = state
264    if (prior > 0) {
265      $.ui.log(`effort-router: first sighting with ${prior} prompts already in the session${state.status === 'off' ? ', left off' : ''}`, { to: 'debug' })
266      spend.ledger = { ...spend.ledger, state: savedOf(state) }
267      spend.dirty = true
268    }
269  }
270  return session
271}
272
273async function sessionOf($: EngineInterface): Promise<{ id: string; session: Session }> {
274  const id = await $.session.id()
275  let session = SESSIONS.get(id)
276  if (!session) {
277    let loading = LOADING.get(id)
278    if (!loading) {
279      loading = loadSession($, id)
280      LOADING.set(id, loading)
281    }
282    try {
283      session = await loading
284    } finally {
285      LOADING.delete(id)
286    }
287    if (!SESSIONS.has(id)) SESSIONS.set(id, session)
288    session = SESSIONS.get(id) as Session
289  }
290  return { id, session }
291}
292
293/** The main loop's model now (it can change with /model), remembered for drawing. */
294async function modelOf($: EngineInterface, session: Session): Promise<string | undefined> {
295  const model = await $.session.model().catch(() => undefined)
296  if (model) session.model = model
297  return session.model
298}
299
300/** The session's state as it applies on its model: on an unsupported one the router stands aside. */
301const stateOf = (session: Session): RouterState => onModel(session.state, session.model)
302
303/** Your effort setting as shown: the picker's level, or the settings file's guess before a request shows it. */
304const shownSetting = (session: Session): Level | undefined => (isLevel(session.picker) ? session.picker : undefined)
305
306/** Your effort setting as judged against: only once a main-thread request has shown it. */
307const seenSetting = (session: Session): Level | undefined => (session.pickerSeen ? shownSetting(session) : undefined)
308
309/** The level an assessment is judged against: the router's own, else your setting once seen. */
310const levelInForce = (session: Session): Level | undefined => appliedLevel(session.state) ?? seenSetting(session)
311
312/** The highest level the router picks on a model: the `highestLevel` option, or the model's own lower cap. */
313const highestOn = (settings: Settings, model: string | undefined): Level => {
314  const cap = supportedModel(model)?.highest
315  return cap && rank(cap) < rank(settings.highestLevel) ? cap : settings.highestLevel
316}
317
318/** The levels an assessment is offered. */
319const levelsFor = (settings: Settings, session: Session): readonly Level[] => offeredLevels(highestOn(settings, session.model), shownSetting(session))
320
321/** The last assessment, from the ledger, so the band survives a resume. */
322function lastOf(session: Session): LastAssessment | undefined {
323  const row = session.spend.ledger.verdicts?.at(-1)
324  if (!row) return undefined
325  return {
326    at: row.at,
327    ...(row.level ? { level: row.level } : {}),
328    ...(row.against ? { against: row.against } : {}),
329    ...(row.reason ? { reason: row.reason } : {}),
330    ...(row.why ? { why: row.why } : {}),
331    outcome: row.outcome,
332  }
333}
334
335function viewOf(session: Session, settings: Settings): View {
336  const routed = [...session.agents.values()].filter(agent => !agent.byDefinition).length
337  return {
338    setting: shownSetting(session),
339    last: lastOf(session),
340    limit: settings.promptsToAssess,
341    offered: levelsFor(settings, session),
342    assessing: session.assessing,
343    subagents: routed,
344  }
345}
346
347/**
348 * Loads the session the process is in now and draws it: at start, and after a resume or /clear moves the process to
349 * another session.
350 */
351async function prime($: EngineInterface): Promise<void> {
352  const { session } = await sessionOf($)
353  await modelOf($, session)
354  // A guess at your setting until a request shows it, for drawing only. The settings file's effortLevel
355  // doesn't apply to Opus 5.5 (Claude Code's model-config docs), so it isn't used there.
356  if (session.picker === undefined && supportedModel(session.model)?.id !== 'claude-opus-5-5') {
357    const configured = (await $.settings.read().catch(() => ({}))) as { effortLevel?: unknown }
358    if (isLevel(configured.effortLevel)) session.picker = configured.effortLevel
359  }
360  await saveSpend($, session) // a first sighting's state
361  show($)
362}
363
364// --- showing, saving and saying ---------------------------------------------------------
365
366function show($: EngineInterface): void {
367  try {
368    $.ui.invalidate('ui.render')
369  } catch {
370    // no UI (-p): nothing to draw
371  }
372}
373
374/** One dim line in the conversation (never sent to the model). */
375function say($: EngineInterface, text: string): void {
376  try {
377    $.ui.log(text)
378  } catch {
379    // headless
380  }
381}
382
383/** Writes the ledger when it changed, one write at a time, each with the newest rows and state. */
384async function saveSpend($: EngineInterface, session: Session): Promise<void> {
385  const spend = session.spend
386  if (!spend.path || !spend.dirty) return
387  const path = spend.path
388  spend.dirty = false
389  spend.writing = spend.writing
390    .then(() => $.fs.write(path, JSON.stringify(spend.ledger)))
391    .catch((error: unknown) => {
392      spend.dirty = true
393      $.ui.log(`effort-router: ledger not saved: ${String(error)}`, { to: 'debug' })
394    })
395  await spend.writing
396}
397
398/** Sets a new state, redraws, and saves it in the session's ledger. */
399async function commit($: EngineInterface, session: Session, state: RouterState): Promise<void> {
400  session.state = state
401  session.spend.ledger = { ...session.spend.ledger, state: savedOf(state) }
402  session.spend.dirty = true
403  show($)
404  await saveSpend($, session)
405}
406
407/** Adds a change to the ledger (in memory; it is written at the next commit or when a turn ends). */
408function record(session: Session, change: (ledger: SpendLedger) => SpendLedger): void {
409  session.spend.ledger = change(session.spend.ledger)
410  session.spend.dirty = true
411}
412
413async function today($: EngineInterface): Promise<string> {
414  return dayOf(await $.clock.now().catch(() => Date.now()))
415}
416
417// --- rules ---------------------------------------------------------------------------
418
419async function readText($: EngineInterface, path: string): Promise<string | undefined> {
420  try {
421    return await $.fs.read(path)
422  } catch {
423    return undefined
424  }
425}
426
427/** The path separator, and the home directory (undefined when neither variable is set). */
428async function homeOf($: EngineInterface): Promise<{ sep: string; homeDir?: string }> {
429  const sep = $.plugin.root.includes('\\') ? '\\' : '/'
430  const [profile, home] = await Promise.all([$.env.get('USERPROFILE').catch(() => undefined), $.env.get('HOME').catch(() => undefined)])
431  return { sep, homeDir: sep === '\\' ? (profile ?? home) : (home ?? profile) }
432}
433
434async function rulePaths($: EngineInterface): Promise<{ defaults: string; user?: string; project?: string }> {
435  const [{ sep, homeDir }, root] = await Promise.all([homeOf($), $.session.root().catch(() => undefined)])
436  return {
437    defaults: `${$.plugin.root}${sep}rules${sep}default.md`,
438    user: homeDir ? `${homeDir}${sep}.claude${sep}effort-router.md` : undefined,
439    project: root ? `${root}${sep}.claude${sep}effort-router.md` : undefined,
440  }
441}
442
443async function settingsSource($: EngineInterface, source: 'user' | 'project' | 'policy'): Promise<unknown> {
444  try {
445    return await $.settings.read({ source })
446  } catch {
447    return undefined
448  }
449}
450
451/**
452 * Shipped defaults, then the organisation's (managed settings), yours, the
453 * project's: re-read on every call so edits apply without a reload. Each
454 * source fails open to absent.
455 */
456async function loadRules($: EngineInterface): Promise<{ composed: ComposedRules; defaults: string }> {
457  const paths = await rulePaths($)
458  const [defaults, userText, projectText, policy, user, project] = await Promise.all([
459    readText($, paths.defaults),
460    paths.user ? readText($, paths.user) : Promise.resolve(undefined),
461    paths.project ? readText($, paths.project) : Promise.resolve(undefined),
462    settingsSource($, 'policy'),
463    settingsSource($, 'user'),
464    settingsSource($, 'project'),
465  ])
466  const layers = ruleLayers({
467    defaults: defaults ?? FALLBACK_RULES,
468    org: settingsRulesOf(policy, $.plugin.name),
469    orgSource: 'managed settings',
470    userFile: paths.user ? { path: paths.user, text: userText } : undefined,
471    userSettings: settingsRulesOf(user, $.plugin.name),
472    projectFile: paths.project ? { path: paths.project, text: projectText } : undefined,
473    projectSettings: settingsRulesOf(project, $.plugin.name),
474  })
475  const composed = composeRules(layers)
476  $.ui.log(`effort-router: rules from ${composed.contributors.map(c => `${c.source} (${c.how})`).join(' → ')}`, { to: 'debug' })
477  return { composed, defaults: defaults ?? FALLBACK_RULES }
478}
479
480/** The notes on what effort means on a model (`rules/models/<model>.md`); undefined on an unsupported model or with no file. */
481async function modelNotes($: EngineInterface, model: string | undefined): Promise<(ModelNotes & { path: string }) | undefined> {
482  const known = supportedModel(model)
483  if (!known) return undefined
484  const { sep } = await homeOf($)
485  const path = `${$.plugin.root}${sep}rules${sep}models${sep}${known.notesFile}`
486  const notes = (await readText($, path))?.replace(/<!--[\s\S]*?-->/g, '').trim()
487  return notes ? { name: known.name, notes, path } : undefined
488}
489
490// --- assessing ----------------------------------------------------------------------
491
492/** The last assistant reply's text, capped: a fork replays the request before it, so it is sent along. */
493function lastReplyOf(messages: readonly TranscriptMessage[]): string | undefined {
494  for (let i = messages.length - 1; i >= 0; i--) {
495    const m = messages[i]
496    if (m?.role === 'assistant' && m.text.trim() !== '') return m.text.trim().slice(-DEFAULT_TRIM.lastAssistantChars)
497    if (m?.role === 'user' && m.text.trim() !== '') return undefined
498  }
499  return undefined
500}
501
502/**
503 * Waits for `work` at most `ms`: `{ ok: false }` on a timeout. The work goes on
504 * in the background; its late answer is ignored.
505 */
506async function timed<T>($: EngineInterface, ms: number, work: Promise<T>): Promise<{ ok: true; value: T } | { ok: false }> {
507  work.catch(() => undefined) // a late failure is not unhandled
508  const abort = typeof AbortController === 'function' ? new AbortController() : undefined
509  const timeout = $.clock.sleep(ms, abort ? { signal: abort.signal } : undefined).then(
510    () => ({ ok: false as const }),
511    () => ({ ok: false as const }),
512  )
513  try {
514    return await Promise.race([work.then(value => ({ ok: true as const, value })), timeout])
515  } finally {
516    abort?.abort()
517  }
518}
519
520/**
521 * One assessment of the whole conversation on the session's model: a fork
522 * when the conversation has a request to fork, else (a new session's first
523 * prompt, or after /clear or some resumes) a separate call with the session's
524 * instructions and the trimmed transcript. A spread is judged against the
525 * level in force; with none known yet it is judged at the next request.
526 */
527async function classifyNow($: EngineInterface, settings: Settings, session: Session, input: ReadInput): Promise<Check | undefined> {
528  const now = async () => $.clock.now().catch(() => Date.now())
529  try {
530    const [stored, rules, model] = await Promise.all([
531      $.session.messages().catch(() => [] as TranscriptMessage[]),
532      loadRules($),
533      modelOf($, session),
534    ])
535    const known = supportedModel(model)
536    if (!known) return undefined
537    const messages = input.answer ? withQuestionAnswer(stored as TranscriptMessage[], input.answer) : (stored as TranscriptMessage[])
538    const rendered = renderTranscript(messages, input.current, DEFAULT_TRIM)
539    if (rendered.text.trim() === '' && !input.hint) return undefined
540    const notes = await modelNotes($, model)
541    const levels = levelsFor(settings, session)
542    const inForce = levelInForce(session)
543    session.calls += 1
544    let check: Check = { kind: 'fork', model: known.id }
545    let reply = await $.model.fork({
546      prompt: forkPrompt({ rules: rules.composed.text, model: notes, current: input.current, lastReply: lastReplyOf(messages), hint: input.hint, answered: input.answer?.text, levels, inForce }),
547    })
548    if (!reply.isAnswered && reply.reason === 'nothing-to-fork') {
549      const instructions = session.instructions ? session.instructions.slice(0, MAX_INSTRUCTIONS_CHARS) : undefined
550      session.sent = { sentChars: rendered.sentChars, fullChars: rendered.fullChars, maxChars: MAX_CHARS, omitted: rendered.omitted }
551      reply = await $.model.complete({
552        model: known.id,
553        system: classifierSystem(rules.composed.text, notes, levels),
554        prompt: classifierPrompt(rendered.text, input.hint, instructions, inForce),
555        maxTokens: SESSION_CHECK_MAX_TOKENS,
556        timeoutMs: TIMEOUT_MS,
557      })
558      check = { kind: 'first', model: known.id, withInstructions: instructions !== undefined }
559    } else {
560      session.sent = undefined
561    }
562    if ('usage' in reply && reply.usage) {
563      const day = await today($)
564      const usage = reply.usage as SpendUsage
565      record(session, ledger => withRead(ledger, day, usage, check.kind))
566    }
567    const checkedAt = await now()
568    if (!reply.isAnswered) {
569      session.error = { at: checkedAt, text: `the assessment got no answer (${reply.reason})` }
570      return { ...check, checkedAt, failed: 'no answer' }
571    }
572    session.verdict = { at: checkedAt, trigger: input.trigger, raw: reply.text, kind: check.kind }
573    $.ui.log(`effort-router: ${check.kind} assessment said (${input.trigger}) ${reply.text.trim().slice(0, 200)}`, { to: 'debug' })
574    const decision = parseDecision(reply.text)
575    if (decision.decision !== 'lock') return { ...check, checkedAt }
576    const { decision: _, ...found } = decision
577    return { ...check, checkedAt, proposal: { ...found, level: clampLevel(found.level, levels), ...(inForce ? { against: inForce } : {}), checkedAt } }
578  } catch (error) {
579    session.error = { at: await now(), text: String(error) }
580    throw error
581  }
582}
583
584/** An assessment's row in the ledger. `prompt` is the prompt of the window it assessed (a manual one: the last counted). */
585function recordVerdict(session: Session, check: Check, outcome: string, manual: boolean, prompt: number): void {
586  const { proposal } = check
587  const at = check.checkedAt ?? Date.now()
588  record(session, ledger =>
589    withVerdictRow(ledger, {
590      at,
591      kind: check.kind,
592      model: check.model,
593      prompt,
594      ...(proposal ? { level: proposal.level, reason: proposal.reason } : {}),
595      ...(proposal?.why ? { why: proposal.why } : {}),
596      ...(proposal?.against ? { against: proposal.against } : {}),
597      outcome,
598      ...(check.withInstructions !== undefined ? { withInstructions: check.withInstructions } : {}),
599      ...(manual ? { manual: true as const } : {}),
600    }),
601  )
602}
603
604/** Says what an assessment did: a move, then a lock. */
605function sayChanges($: EngineInterface, settled: ReturnType<typeof settle>, reason: string | undefined): string[] {
606  const said: string[] = []
607  if (settled.moved) said.push(message.moved(settled.moved.from, settled.moved.to, reason ?? 'router'))
608  if (settled.locked) said.push(message.locked(settled.locked))
609  for (const line of said) say($, line)
610  return said
611}
612
613/**
614 * One assessment, applied. `counted` uses up one prompt of the window (the
615 * automatic ones, and Turn on and assess); a manual one while locked moves the
616 * locked level and stays locked. A level picked with no level in force known
617 * yet is applied at the next main-thread request (judgeWaiting).
618 */
619async function assess($: EngineInterface, id: string, session: Session, settings: Settings, input: ReadInput, counted: boolean): Promise<string> {
620  if (session.reading) return 'Already assessing. Try again in a moment.'
621  session.reading = true
622  session.assessing = true
623  show($)
624  const started = await $.clock.now().catch(() => Date.now())
625  let summary = 'The assessment failed, so nothing changed.'
626  try {
627    // A failed assessment (an error, a timeout, no answer) still uses up its prompt: the window stays the length
628    // the band promises.
629    const result = await timed($, TIMEOUT_MS, classifyNow($, settings, session, input).catch((): undefined => undefined))
630    const prompt = counted ? Math.min(settings.promptsToAssess, session.state.assessed + 1) : session.state.assessed
631    const check = result.ok ? result.value : undefined
632    if (!result.ok) session.error = { at: await $.clock.now().catch(() => Date.now()), text: `the assessment timed out after ${TIMEOUT_MS / 1000} s, so the prompt ran at the level it had` }
633    const proposal = check?.proposal
634    if (check && proposal && !proposal.against) {
635      // No level in force known yet: applied at the next request, which shows it.
636      const state = { ...session.state, pending: proposal, hint: undefined, assessed: prompt }
637      recordVerdict(session, check, 'judged at the first request', !counted, prompt)
638      await commit($, session, state)
639      summary = 'Assessed. It is judged against your level when the next request shows it.'
640      return summary
641    }
642    const failed = !check || check.failed !== undefined
643    const settled = settle(session.state, proposal, { limit: settings.promptsToAssess, running: levelInForce(session), counted })
644    if (check) recordVerdict(session, check, failed ? 'failed' : settled.outcome, !counted, prompt)
645    await commit($, session, settled.state)
646    const said = sayChanges($, settled, proposal?.reason)
647    summary = said.length > 0 ? said.join(' ') : failed ? summary : (lastAssessmentLine(lastOf(session), session.state.status === 'locked') ?? 'Nothing changed.')
648    return summary
649  } catch (error) {
650    $.ui.log(`effort-router: assessment failed: ${String(error)}`, { to: 'debug' })
651    return summary
652  } finally {
653    session.reading = false
654    session.assessing = false
655    show($)
656    session.lastReadMs = (await $.clock.now().catch(() => Date.now())) - started
657    $.ui.log(`effort-router: assessment settled in ${session.lastReadMs} ms (${input.trigger})`, { to: 'debug' })
658  }
659}
660
661/** A first assessment made before the level in force was known: judged now, against the level this request shows. */
662async function judgeWaiting($: EngineInterface, session: Session, settings: Settings, setting: Level): Promise<void> {
663  const waiting = session.state.pending
664  if (!waiting) return
665  const judged: Proposal = { ...waiting, against: setting }
666  const settled = settle(session.state, judged, { limit: settings.promptsToAssess, running: setting, counted: false })
667  record(session, ledger => withVerdictOutcome(ledger, settled.outcome, waiting.checkedAt, judged))
668  await commit($, session, settled.state)
669  sayChanges($, settled, judged.reason)
670}
671
672/**
673 * A human turn of the conversation (a prompt, or answers to the model's
674 * questions): assessed while unlocked and within the window, before the turn
675 * goes on, so its first request carries the level.
676 */
677async function humanTurn($: EngineInterface, settings: Settings, input: ReadInput): Promise<void> {
678  const { id, session } = await sessionOf($)
679  if (!supportedModel(await modelOf($, session))) return
680  // A prompt sent while a turn runs isn't assessed (not tested live); answers mid-turn are, by a fork carrying them.
681  if (session.busy && !input.answer) return
682  if (!wantsAssessment(session.state, settings.promptsToAssess)) return
683  await assess($, id, session, settings, { ...input, hint: session.state.hint }, true)
684}
685
686// --- actions (the band's slots, and /er) ---------------------------------------------------
687
688/** What an action did: a message for the conversation, or a reply that changes nothing. */
689type Done = { said?: string; reply?: string }
690
691async function doLock($: EngineInterface, session: Session): Promise<Done> {
692  const state = stateOf(session)
693  if (state.unsupported) return { reply: unsupportedText(session.model) }
694  if (state.status === 'locked') return { reply: `Already locked at ${state.level}.` }
695  if (state.status === 'off') return doTurnOnLocked($, session)
696  // The band's label names the level it locks at (your setting as shown, a guess from the settings file before a
697  // request shows it), so lock exactly that.
698  const running = runningLevel(state, shownSetting(session))
699  if (!running) return { reply: 'Nothing to lock yet: the level shows with the first request.' }
700  await commit($, session, lockedByYou(session.state, running))
701  return { said: message.lockedByYou(running) }
702}
703
704async function doUnlock($: EngineInterface, session: Session): Promise<Done> {
705  const state = stateOf(session)
706  if (state.unsupported) return { reply: unsupportedText(session.model) }
707  if (state.status === 'off') return doTurnOn($, session)
708  if (state.status === 'unlocked') return { reply: `Already unlocked. Assessed ${state.assessed} of ${currentSettings.promptsToAssess} prompts.` }
709  await commit($, session, unlocked(session.state))
710  return { said: message.unlocked() }
711}
712
713async function doTurnOn($: EngineInterface, session: Session): Promise<Done> {
714  const state = stateOf(session)
715  if (state.unsupported) return { reply: unsupportedText(session.model) }
716  if (state.status !== 'off') return { reply: state.status === 'locked' ? `Already on, locked at ${state.level}.` : 'Already on, unlocked.' }
717  await commit($, session, turnedOnUnlocked(session.state))
718  return { said: message.onUnlocked() }
719}
720
721async function doTurnOnLocked($: EngineInterface, session: Session): Promise<Done> {
722  if (session.state.status !== 'off') return doLock($, session)
723  if (!session.state.lastLevel) return { reply: 'The router has no level of its own to lock at yet. /er on turns it on, unlocked.' }
724  await commit($, session, turnedOnLocked(session.state))
725  return { said: message.onLocked(session.state.lastLevel as Level) }
726}
727
728async function doTurnOff($: EngineInterface, session: Session): Promise<Done> {
729  if (session.state.status === 'off') return { reply: 'Already off.' }
730  await commit($, session, turnedOff(session.state, 'you'))
731  return { said: message.off(seenSetting(session)) }
732}
733
734async function doAssess($: EngineInterface, id: string, session: Session, settings: Settings, hint: string | undefined): Promise<Done> {
735  if (!supportedModel(await modelOf($, session))) return { reply: unsupportedText(session.model) }
736  if (session.state.status === 'unlocked') {
737    if (hint) {
738      await commit($, session, { ...session.state, hint })
739      return { reply: 'Your hint is used when your next prompt is assessed.' }
740    }
741    return { reply: 'It assesses before your next prompt anyway.' }
742  }
743  if (session.state.status === 'off') {
744    await commit($, session, turnedOnUnlocked(session.state))
745    say($, message.onUnlocked())
746    return { reply: await assess($, id, session, settings, { hint, trigger: 'Turn on and assess' }, true) }
747  }
748  return { reply: await assess($, id, session, settings, { hint, trigger: hint ? 'Assess, with a hint' : 'Assess' }, false) }
749}
750
751const unsupportedText = (model: string | undefined): string =>
752  `The router doesn't support ${modelName(model)}, so your effort setting applies. It works with ${SUPPORTED_NAMES}.`
753
754// --- the band ------------------------------------------------------------------------
755
756function closeBand($: EngineInterface, session: Session): void {
757  session.bandOpen = false
758  session.note = undefined
759  show($)
760}
761
762/** The footer button: opens the band, or closes it when it is showing. */
763function toggleBand($: EngineInterface, session: Session): void {
764  if (session.bandOpen) closeBand($, session)
765  else {
766    session.bandOpen = true
767    session.note = undefined
768    show($)
769  }
770}
771
772/** A band slot: runs its action and keeps the band open on the new state (Hide closes it). */
773async function bandAction($: EngineInterface, id: string, session: Session, settings: Settings, action: BandAction): Promise<void> {
774  if (action.value === 'hide') return closeBand($, session)
775  if (action.disabled) {
776    session.note = action.disabled
777    show($)
778    return
779  }
780  session.note = undefined
781  try {
782    const done =
783      action.value === 'lock' ? await doLock($, session)
784      : action.value === 'unlock' ? await doUnlock($, session)
785      : action.value === 'off' ? await doTurnOff($, session)
786      : action.value === 'on-unlocked' ? await doTurnOn($, session)
787      : action.value === 'on-locked' ? await doTurnOnLocked($, session)
788      : await doAssess($, id, session, settings, undefined)
789    if (done.said) say($, done.said)
790    else if (done.reply && action.value !== 'assess' && action.value !== 'on-assess') session.note = done.reply
791  } catch (error) {
792    $.ui.log(`effort-router: band action ${action.value} failed: ${String(error)}`, { to: 'debug' })
793    session.note = 'That failed. Nothing changed.'
794  }
795  show($)
796}
797
798// --- /er -------------------------------------------------------------------------------
799
800async function route($: EngineInterface, args: string, settings: Settings): Promise<{ text?: string }> {
801  const { id, session } = await sessionOf($)
802  await modelOf($, session)
803  const command = parseRoute(args)
804  const reply = (done: Done): { text: string } => ({ text: done.said ?? done.reply ?? '' })
805  switch (command.kind) {
806    case 'band': {
807      if (hasBand) {
808        session.bandOpen = true
809        session.note = undefined
810        show($)
811        return {}
812      }
813      const view = viewOf(session, settings)
814      const state = stateOf(session)
815      const last = lastAssessmentLine(view.last, state.status === 'locked')
816      return { text: [bandHeadline(state, view), last, ROUTE_USAGE].filter(Boolean).join('\n') }
817    }
818    case 'lock':
819      return reply(await doLock($, session))
820    case 'unlock':
821      return reply(await doUnlock($, session))
822    case 'on':
823      return reply(await doTurnOn($, session))
824    case 'off':
825      return reply(await doTurnOff($, session))
826    case 'assess':
827      return reply(await doAssess($, id, session, settings, command.hint))
828    case 'report':
829      return { text: await spendReportFor($, id, session, command.period) }
830    case 'status':
831      return {
832        text: routeReport(stateOf(session), viewOf(session, settings), {
833          now: await $.clock.now().catch(() => Date.now()),
834          calls: Math.max(session.calls, session.spend.ledger.reads.reduce((n, r) => n + r.calls, 0)),
835          verdict: session.verdict,
836          error: session.error,
837          lastReadMs: session.lastReadMs,
838          sent: session.sent,
839          subagents: { routing: subagentRouting(settings, session), agents: [...session.agents.values()] },
840        }),
841      }
842    case 'rules': {
843      const { composed } = await loadRules($)
844      const notes = await modelNotes($, session.model)
845      const from = composed.contributors
846        .map(c => (c.how === 'base' ? `  ${c.source}` : c.how === 'spliced' ? `  + ${c.source}` : `  ${c.source} (replaces the rules above)`))
847        .join('\n')
848      const onModel = notes ? `\n  + notes on ${notes.name} (${notes.path})` : ''
849      const notesText = notes ? `\n\nOn ${notes.name}:\n${notes.notes}` : ''
850      return { text: `Routing rules in use:\n${from}${onModel}\n\n${composed.text}${notesText}` }
851    }
852    case 'unknown':
853      return { text: `Unknown: ${command.text}. ${ROUTE_USAGE}` }
854  }
855}
856
857// --- subagents -----------------------------------------------------------------------
858
859/** The agent definition files in one `.claude/agents` folder; a missing folder or unreadable file is skipped. */
860async function agentFiles($: EngineInterface, dir: string, sep: string): Promise<AgentDefinition[]> {
861  const entries = await $.fs.list(dir).catch(() => [])
862  const files = entries.filter(entry => entry.kind !== 'dir' && /\.md$/i.test(entry.name)).map(entry => entry.name).sort()
863  const found = await Promise.all(
864    files.map(async name => {
865      const path = `${dir}${sep}${name}`
866      const text = await readText($, path)
867      return text === undefined ? undefined : agentFileDefinition(text, name, path)
868    }),
869  )
870  return found.filter((definition): definition is AgentDefinition => definition !== undefined)
871}
872
873/**
874 * Every agent definition the user and project hold, highest precedence first:
875 * policy settings' `agents`, the project's `.claude/agents/*.md` (the session's
876 * directory, then the project root), project settings' `agents`, the user's
877 * `~/.claude/agents/*.md`, then user settings' `agents`. Plugins' agents are
878 * not here (see definitionFor).
879 */
880async function loadDefinitions($: EngineInterface): Promise<AgentDefinition[]> {
881  const [{ sep, homeDir }, cwd, root, policy, project, user] = await Promise.all([
882    homeOf($),
883    $.session.cwd().catch(() => undefined),
884    $.session.root().catch(() => undefined),
885    settingsSource($, 'policy'),
886    settingsSource($, 'project'),
887    settingsSource($, 'user'),
888  ])
889  const agentsDir = (base: string) => `${base}${sep}.claude${sep}agents`
890  const projectDirs = [...new Set([cwd, root].filter((dir): dir is string => typeof dir === 'string' && dir !== ''))].map(agentsDir)
891  const [projectFiles, userFiles] = await Promise.all([
892    Promise.all(projectDirs.map(dir => agentFiles($, dir, sep))).then(lists => lists.flat()),
893    homeDir ? agentFiles($, agentsDir(homeDir), sep) : Promise.resolve([]),
894  ])
895  const all = [
896    ...settingsAgentDefinitions(policy, 'policy settings'),
897    ...projectFiles,
898    ...settingsAgentDefinitions(project, 'project settings'),
899    ...userFiles,
900    ...settingsAgentDefinitions(user, 'user settings'),
901  ]
902  $.ui.log(`effort-router: ${all.length} agent definitions found, ${all.filter(d => d.effort !== undefined).length} with their own effort`, { to: 'debug' })
903  return all
904}
905
906function definitionsOf($: EngineInterface, session: Session): Promise<AgentDefinition[]> {
907  session.definitions ??= loadDefinitions($).catch(() => [])
908  return session.definitions
909}
910
911/** Whether subagents are routed now, and if not, why. */
912function subagentRouting(settings: Settings, session: Session): SubagentStatus['routing'] {
913  if (!settings.routeSubagents) return 'setting'
914  return routesSubagents(session.state) ? 'on' : 'user-off'
915}
916
917/**
918 * The level for a spawn, decided before it starts. One whose Agent call asked
919 * for an effort keeps it, untouched like a definition's. A fork takes the parent's
920 * level. An agent whose definition sets an effort keeps it: no read, and
921 * `byDefinition` so its requests are left to the engine. Anything else waits
922 * (at most 30 s) for one read: a fork of the parent plus the brief, else, with
923 * nothing to fork, a separate call on the brief alone. A failed, late or
924 * unusable read takes the parent's level. Undefined leaves the subagent's
925 * requests as they would have been.
926 */
927async function routeSpawn($: EngineInterface, settings: Settings, session: Session, e: AgentSpawnInput): Promise<Pick<RoutedAgent, 'level' | 'reason' | 'byDefinition'> | undefined> {
928  const inherited = parentLevel(session.state, session.agents, e.parentAgentId)
929  const fallback = (why: string): Proposal | undefined => (inherited ? { level: inherited, reason: `same as its parent: ${why}` } : undefined)
930  const asked = ASKED_EFFORT.get(e.tool_use_id)
931  if (asked) return { level: asked, reason: ASKED_REASON, byDefinition: true }
932  if (e.fork) return fallback("it's a fork")
933  const definition = definitionFor(e.subagentType, await definitionsOf($, session))
934  if (definition?.effort !== undefined) return { level: definition.effort, reason: `from ${definition.source}`, byDefinition: true }
935  // The model it runs on: the Agent call's, else its definition's, else the parent's. A model the router doesn't
936  // support is left alone. (An alias that resolves to an older model, such as haiku on a cloud provider, gets no
937  // effort at turn.step, which only rewrites a request that carries one.)
938  const runsOn = e.model ?? definition?.model ?? e.parentModel
939  const known = supportedModel(runsOn)
940  if (!known) {
941    $.ui.log(`effort-router: subagent (${e.subagentType}: ${e.description}) runs on ${modelName(runsOn)}, left alone`, { to: 'debug' })
942    return undefined
943  }
944  const rules = await loadRules($)
945  const brief = { subagentType: e.subagentType, description: e.description, prompt: e.prompt }
946  const notes = await modelNotes($, known.id)
947  // Up to the parent's own setting at most, and never above this model's cap.
948  const levels = levelsFor(settings, session).filter(level => rank(level) <= rank(highestOn(settings, known.id)))
949  const alone = () =>
950    $.model.complete({
951      model: known.id,
952      system: subagentSystem(rules.composed.text, notes, levels),
953      prompt: subagentPrompt(brief, MAX_CHARS),
954      maxTokens: SESSION_CHECK_MAX_TOKENS,
955      timeoutMs: TIMEOUT_MS,
956    })
957  const read = async () => {
958    // A small model's subagents get short, self-contained briefs: judging one on the parent's model could cost what
959    // it saves, and slow the helper chosen for speed. Read the brief on the model itself.
960    if (known.checksOwnBrief) return alone()
961    // The parent knows the task and why it delegates this part: ask a fork of it (cached, a few seconds).
962    const forked = await $.model.fork({ prompt: subagentForkPrompt({ rules: rules.composed.text, brief, runsOn: known.name, model: notes, maxChars: MAX_CHARS, levels }) })
963    if (forked.isAnswered || forked.reason !== 'nothing-to-fork') return forked
964    return alone()
965  }
966  const result = await timed($, TIMEOUT_MS, read()).catch((error: unknown) => {
967    $.ui.log(`effort-router: subagent read failed: ${String(error)}`, { to: 'debug' })
968    return undefined
969  })
970  if (!result) return fallback('the assessment failed')
971  if (!result.ok) return fallback('the assessment timed out')
972  if ('usage' in result.value && result.value.usage) {
973    const day = await today($)
974    const usage = result.value.usage as SpendUsage
975    record(session, ledger => withRead(ledger, day, usage, 'subagent'))
976  }
977  if (!result.value.isAnswered) return fallback('the assessment got no answer')
978  $.ui.log(`effort-router: subagent assessment said ${result.value.text.trim().slice(0, 200)}`, { to: 'debug' })
979  const found = parseSubagentReply(result.value.text)
980  return found ? { ...found, level: clampLevel(found.level, levels) } : fallback('the assessment gave no level')
981}
982
983/** Keeps a routed subagent, dropping the oldest past `MAX_ROUTED_AGENTS`. */
984function remember(session: Session, agentId: string, agent: RoutedAgent): void {
985  session.agents.delete(agentId)
986  session.agents.set(agentId, agent)
987  for (const oldest of session.agents.keys()) {
988    if (session.agents.size <= MAX_ROUTED_AGENTS) break
989    session.agents.delete(oldest)
990  }
991}
992
993// --- the report ----------------------------------------------------------------------
994
995/** `/er report`: this session's ledger (in memory, the newest) and, beyond it, every saved ledger touched in the period. */
996async function spendReportFor($: EngineInterface, id: string, session: Session, period: SpendPeriod): Promise<string> {
997  const now = await $.clock.now().catch(() => Date.now())
998  const others: SpendLedger[] = []
999  const { dir, sep } = await spendDir($)
1000  if (period !== 'session' && dir) {
1001    const oldest = period === 'all' ? 0 : now - (period === 'week' ? 8 : 31) * 86_400_000
1002    const entries = await $.fs.list(dir).catch(() => [])
1003    const files = entries.filter(f => f.kind === 'file' && /\.json$/i.test(f.name) && f.name !== `${id}.json` && (f.mtimeMs === 0 || f.mtimeMs >= oldest))
1004    const texts = await Promise.all(files.map(f => readText($, `${dir}${sep}${f.name}`)))
1005    for (const text of texts) {
1006      const ledger = text === undefined ? undefined : parseLedger(text)
1007      if (ledger) others.push(ledger)
1008    }
1009  }
1010  return spendReport([session.spend.ledger, ...others], period, { today: dayOf(now), session: id })
1011}
1012
1013// --- hooks ---------------------------------------------------------------------------
1014
1015/** The settings in force, for code that runs outside a hook's own closure (a first sighting). */
1016let currentSettings: Settings = settingsOf({})
1017
1018export function register(on: On, options: PluginOptions): void {
1019  const settings = settingsOf(options)
1020  currentSettings = settings
1021
1022  // Claude Code's api_request records go to the collector your organisation configured, never anywhere else. They already
1023  // carry the level each request went out at. This adds your own setting and the router's status, so the collector can
1024  // see what the router changed.
1025  on('telemetry.log', { to: 'collector' }, async ($, e, next) => {
1026    if (e.to !== 'collector' || e.event !== 'api_request') return next(e)
1027    try {
1028      const { session } = await sessionOf($)
1029      const extra = telemetryAttributes(stateOf(session), seenSetting(session), VERSION)
1030      return next({ ...e, attributes: { ...e.attributes, ...extra } })
1031    } catch {
1032      return next(e)
1033    }
1034  })
1035
1036  on('session.start', async ($, e, next) => {
1037    hasBand = e.isInteractive && (e.surface === 'terminal' || e.surface === 'desktop')
1038    try {
1039      for (const name of COMMANDS) {
1040        await $.command.register({
1041          name,
1042          description: 'Effort router: open the band, or lock, unlock, on, off, assess [hint], report, status, rules',
1043          argumentHint: '[lock|unlock|on|off|assess|report|status|rules]',
1044          immediate: true,
1045        }).catch((error: unknown) => $.ui.log(`effort-router: /${name} not registered: ${String(error)}`, { to: 'debug' }))
1046      }
1047      await prime($)
1048    } catch (error) {
1049      $.ui.log(`effort-router: start failed: ${String(error)}`, { to: 'debug' })
1050    }
1051    return next(e)
1052  })
1053
1054  // A resume or /clear inside a running process goes on under another session id, and no session.start fires for
1055  // it. Without this the footer kept drawing the session that ended until something else redrew it.
1056  on('session.end', async ($, e, next) => {
1057    const result = await next(e)
1058    if (e.reason === 'resume' || e.reason === 'clear') {
1059      const again = () => void prime($).catch((error: unknown) => $.ui.log(`effort-router: resume failed: ${String(error)}`, { to: 'debug' }))
1060      try {
1061        $.clock.after(50, again)
1062        $.clock.after(1000, again)
1063      } catch {
1064        // no timers: the next redraw picks the new session up
1065      }
1066    }
1067    return result
1068  })
1069
1070  for (const name of COMMANDS) {
1071    on('command.run', { command: name }, async ($, e) => {
1072      try {
1073        return await route($, e.args ?? '', settings)
1074      } catch (error) {
1075        return { text: `/${name} failed: ${String(error)}` }
1076      }
1077    })
1078  }
1079
1080  // The context the conversation's first message carries: keep the instructions block (CLAUDE.md files, rules,
1081  // memory) for the first assessment, which cannot fork a conversation that has sent nothing yet. Unchanged.
1082  on('prompt.context', async ($, e, next) => {
1083    try {
1084      const text = e.blocks.find(block => block.name === 'claudeMd')?.text
1085      if (text) (await sessionOf($)).session.instructions = text
1086    } catch {
1087      // the first assessment goes without them
1088    }
1089    return next(e)
1090  })
1091
1092  // After each human prompt while unlocked and within the window: assess the whole conversation BEFORE the turn
1093  // runs, so its first request carries the level. At most 30 s; on a timeout or error the turn goes ahead.
1094  on('prompt.submit', async ($, e, next) => {
1095    try {
1096      if (HUMAN_ORIGINS.has(e.origin.kind) && !e.text.trimStart().startsWith('/')) {
1097        await humanTurn($, settings, { current: e.text, trigger: 'after a prompt' })
1098      }
1099    } catch (error) {
1100      $.ui.log(`effort-router: prompt.submit failed: ${String(error)}`, { to: 'debug' })
1101    }
1102    return next(e)
1103  })
1104
1105  // The model asked the user multiple-choice questions on the main thread and got answers: that is a human turn
1106  // too (platforms, scope, "keep it simple"), assessed before the answers go back to the model.
1107  on('tool.call', { tool: QUESTION_TOOL }, async ($, e, next) => {
1108    const result = await next(e)
1109    try {
1110      const answered = !('deny' in result && result.deny) && !result.isError && typeof result.text === 'string' && result.text.trim() !== ''
1111      if (e.agentId === undefined && answered) {
1112        await humanTurn($, settings, {
1113          answer: { toolUseId: e.tool_use_id, input: { questions: e.questions }, text: result.text as string },
1114          trigger: 'after answered questions',
1115        })
1116      }
1117    } catch (error) {
1118      $.ui.log(`effort-router: tool.call failed: ${String(error)}`, { to: 'debug' })
1119    }
1120    return result
1121  })
1122
1123  // An Agent call that asks for an effort (you, CLAUDE.md or a skill asked for one): its subagent keeps it.
1124  on('tool.call', { tool: AGENT_TOOL }, async ($, e, next) => {
1125    const effort = (e as { effort?: unknown }).effort
1126    if (!isLevel(effort)) return next(e)
1127    ASKED_EFFORT.set(e.tool_use_id, effort)
1128    try {
1129      return await next(e)
1130    } finally {
1131      ASKED_EFFORT.delete(e.tool_use_id)
1132    }
1133  })
1134
1135  // A subagent is about to start: read its brief (a fork takes its parent's level) BEFORE it starts, and key the
1136  // level to its agentId. next(e) resolves with the id before the agent's first turn.step (verified live), so its
1137  // first request already carries it.
1138  on('agent.spawn', async ($, e, next) => {
1139    let routed: { session: Session; proposal: Pick<RoutedAgent, 'level' | 'reason' | 'byDefinition'>; took: number; at: number; parent?: Level } | undefined
1140    try {
1141      const { session } = await sessionOf($)
1142      if (subagentRouting(settings, session) === 'on') {
1143        const started = await $.clock.now().catch(() => Date.now())
1144        const parent = parentLevel(session.state, session.agents, e.parentAgentId) ?? seenSetting(session)
1145        const proposal = await routeSpawn($, settings, session, e)
1146        if (proposal) routed = { session, proposal, took: (await $.clock.now().catch(() => Date.now())) - started, at: started, parent }
1147      }
1148    } catch (error) {
1149      $.ui.log(`effort-router: agent.spawn failed: ${String(error)}`, { to: 'debug' })
1150    }
1151    const result = await next(e)
1152    try {
1153      if (routed && result.agentId !== undefined) {
1154        const { session, proposal, took, at, parent } = routed
1155        remember(session, result.agentId, { ...proposal, subagentType: e.subagentType, description: e.description })
1156        record(session, ledger =>
1157          withSubagentRow(ledger, {
1158            at,
1159            model: e.model ?? e.parentModel ?? 'unknown',
1160            subagentType: e.subagentType,
1161            description: e.description,
1162            ...(parent ? { parent } : {}),
1163            level: proposal.level,
1164            reason: proposal.reason,
1165            ...(proposal.byDefinition ? { byDefinition: true as const } : {}),
1166            ms: took,
1167          }),
1168        )
1169        $.ui.log(
1170          `effort-router: subagent ${result.agentId} (${e.subagentType}${e.fork ? ', fork' : ''}: ${e.description}) -> ${proposal.level}${proposal.byDefinition ? ', left alone' : ''} (${proposal.reason}) in ${took} ms`,
1171          { to: 'debug' },
1172        )
1173        show($)
1174      }
1175    } catch {
1176      // best effort: an unrouted subagent runs as it would have
1177    }
1178    return result
1179  })
1180
1181  // Every model request. On the main thread `e.effort` as it arrives is your effort setting (the router has not
1182  // rewritten it): a change of it turns routing off, a first assessment waiting for it is judged against it, and
1183  // the window's end locks. Then the level: a routed subagent's own; a subagent whose definition sets its effort,
1184  // untouched; otherwise the router's level, or untouched.
1185  on('turn.step', async function* ($, e, next) {
1186    let effort = e.effort
1187    let byDefinition = false
1188    try {
1189      const { session } = await sessionOf($)
1190      if (e.agentId === undefined) {
1191        session.busy = true
1192        if (session.model !== e.model) {
1193          session.model = e.model
1194          show($)
1195        }
1196        if (isLevel(e.effort)) {
1197          const was = session.pickerSeen ? session.picker : undefined
1198          session.picker = e.effort
1199          if (session.spend.ledger.setting !== e.effort) record(session, ledger => ({ ...ledger, setting: e.effort as Level }))
1200          if (!session.pickerSeen) {
hooks/policy.ts 1600 lines
1/**
2 * The pure half of effort-router: the levels, the classifier prompt, the
3 * transcript trimming, parsing the classifier's reply, the /er grammar and
4 * the text the mod shows. No `$`, no engine: `bun test` runs it directly.
5 *
6 * Policy source: Anthropic, "Using Claude Code: Spending your effort",
7 * Thariq Shihipar, 2026-09-25 (Anthropic's blog).
8 */
9
10export type Level = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
11
12export const LEVELS: readonly Level[] = ['low', 'medium', 'high', 'xhigh', 'max']
13
14export const isLevel = (value: unknown): value is Level =>
15  typeof value === 'string' && (LEVELS as readonly string[]).includes(value)
16
17export const rank = (level: Level): number => LEVELS.indexOf(level)
18
19// --- the classifier prompt -----------------------------------------------------
20
21/** The highest level the router picks unless the `highestLevel` option says otherwise. Max rarely beats xhigh. */
22export const DEFAULT_HIGHEST: Level = 'xhigh'
23
24/** The levels a check may pick: low up to `highest`. */
25export const levelsUpTo = (highest: Level = DEFAULT_HIGHEST): readonly Level[] => LEVELS.slice(0, LEVELS.indexOf(highest) + 1)
26
27/** A level held inside `levels` (a run of adjacent levels): the nearest end when it falls outside. */
28export function clampLevel(level: Level, levels: readonly Level[]): Level {
29  const lowest = levels[0]
30  const highest = levels[levels.length - 1]
31  if (lowest === undefined || highest === undefined) return level
32  return rank(level) < rank(lowest) ? lowest : rank(level) > rank(highest) ? highest : level
33}
34
35/**
36 * The fixed frame around the routing rules: the job, what to optimise and
37 * when there is nothing to judge yet. It never ties a kind of task to a
38 * level: what a level can do differs by model, and the model notes say it.
39 * (An eval on 2026-10-04 showed a frame whose examples named levels overrode
40 * the notes.) It has no worked examples: the checks run on Opus 5.5 and
41 * Fable 5.1, which read a conversation without being shown how. The rules
42 * (`rules/default.md` and the user's files) hold the principles for choosing.
43 */
44export const classifierFrame = (levels: readonly Level[] = levelsUpTo()): string => `You pick the reasoning-effort level for this Claude Code session. Levels you may pick, lowest to highest: ${levels.join(', ')}.
45
46Pick the level that gets the work from here done in the least time and total inference cost. People turn this router on to spend less, so when two levels would both get the work done, pick the cheaper one, and go higher only when the work clearly needs it. Too little effort is not cheaper when it leads to mistakes, rework or a second attempt. Too much pays for thinking the work won't use. The session switches to your pick straight away and switching costs nothing, so the level it is on now has no special weight.
47
48If the user says how hard to think or how quickly to go ("think really hard about this", "quick one"), that is their call: pick the level that matches it.
49
50If no task has been stated yet (a greeting, setup such as "pull the latest code", a question asked before any work), answer undecided. Once there is a task, pick a level for it even if details are still unclear: you are asked again after each of the user's next few messages.
51
52Principles for choosing. Later rules override earlier ones where they conflict:`
53
54/** The frame at the default highest level. */
55export const CLASSIFIER_FRAME = classifierFrame()
56
57export const classifierContract = (levels: readonly Level[] = levelsUpTo()): string => `Reply with one JSON object and nothing else:
58{"level":"<undecided|${levels.join('|')}>","reason":"<what the task is, 3-8 words>","why":"<one or two sentences: why this level, and what would change your pick>"}`
59
60export const CLASSIFIER_CONTRACT = classifierContract()
61
62/** What effort means on a model: its name and the notes in `rules/models/`. */
63export type ModelNotes = { name: string; notes: string }
64
65/**
66 * A model's notes cut to the levels on offer: a line that starts with a level
67 * the check can't pick (a `| max |` table row or a `- max:` bullet) is left out, so
68 * the notes don't argue about a level nobody is choosing.
69 */
70export function notesFor(notes: string, levels: readonly Level[]): string {
71  const absent = LEVELS.filter(level => !levels.includes(level))
72  if (absent.length === 0) return notes.trim()
73  const starts = new RegExp(`^(?:- |\\|\\s*)?(?:${absent.join('|')})\\b`)
74  return notes
75    .split('\n')
76    .filter(line => !starts.test(line.trim()))
77    .join('\n')
78    .trim()
79}
80
81const modelBlock = (model: ModelNotes | undefined, who: string, levels: readonly Level[]): string =>
82  model && model.notes.trim() !== ''
83    ? `\n\n${who} ${model.name}. Level names buy different amounts of thinking on different models. On this one:\n<model_notes>\n${notesFor(model.notes, levels)}\n</model_notes>`
84    : ''
85
86/** The classifier's whole system prompt around the composed rules, with the session model's notes when there are any. */
87export const classifierSystem = (rules: string, model?: ModelNotes, levels: readonly Level[] = levelsUpTo()): string =>
88  `${classifierFrame(levels)}\n\n<rules>\n${rules.trim()}\n</rules>${modelBlock(model, 'The session runs on', levels)}\n\n${classifierContract(levels)}`
89
90/**
91 * The one message a check sends into a fork of the session (`$.model.fork`):
92 * the whole conversation as the session's model last saw it, its own system
93 * prompt, CLAUDE.md and memory included, then this. A fork replays the last
94 * request, which does not hold the reply it produced, so that reply comes
95 * along here, as do the prompt being submitted and, mid-turn, the answers the
96 * user just gave to the model's questions.
97 */
98/** The question that ends every session check. */
99export const QUESTION = "Which effort level gets this session's work done in the least time and total inference cost? JSON only."
100
101export function forkPrompt(input: { rules: string; model?: ModelNotes; current?: string; lastReply?: string; hint?: string; answered?: string; levels?: readonly Level[]; inForce?: Level }): string {
102  const parts = [
103    'Pause the task for a moment. Do not use any tools and do not carry on with the work: answer only the question below.',
104    classifierSystem(input.rules, input.model, input.levels),
105  ]
106  const lastReply = input.lastReply?.trim()
107  if (lastReply) parts.push(`Your last reply in this conversation, which is not shown above:\n<last_reply>\n${lastReply}\n</last_reply>`)
108  const answered = input.answered?.trim()
109  if (answered) parts.push(`You asked the user questions, and they have just answered:\n<answers>\n${answered}\n</answers>`)
110  const current = input.current?.trim()
111  if (current) parts.push(`The user has just sent this new message, and the work goes on from it:\n<new_message>\n${current}\n</new_message>`)
112  const hint = input.hint?.trim()
113  if (hint) parts.push(`<user_hint>\n${hint}\n</user_hint>\nThe user asked for this routing explicitly and gave this hint; weigh it strongly.`)
114  if (input.inForce) parts.push(`The session is at ${input.inForce} effort now.`)
115  parts.push(QUESTION)
116  return parts.join('\n\n')
117}
118
119// --- rule files and their composition ---------------------------------------------
120
121/** The line that splices in the layer beneath (shipped defaults, then the user's). */
122export const DEFAULTS_MARKER = '$defaults'
123
124export type RuleLayer = {
125  /** Where it came from, for `/er rules` (a path or `defaults`). */
126  source: string
127  /** The file's text; undefined when the file is absent or unreadable. */
128  text: string | undefined
129}
130
131export type ComposedRules = {
132  text: string
133  /** Each layer that contributed, bottom first, and how. */
134  contributors: { source: string; how: 'base' | 'spliced' | 'replaced' }[]
135}
136
137const stripComments = (text: string): string => text.replace(/<!--[\s\S]*?-->/g, '')
138
139/**
140 * Composes rule layers bottom-up. The first layer is the base (the shipped
141 * defaults). Each later layer that exists and has content either splices the
142 * result so far wherever it has a line that is exactly `$defaults`, or, with
143 * no such line, replaces it. An absent, unreadable or empty layer changes
144 * nothing. HTML comments are dropped (the starter file's example lives in one).
145 */
146export function composeRules(layers: readonly RuleLayer[]): ComposedRules {
147  let text = ''
148  const contributors: ComposedRules['contributors'] = []
149
150  layers.forEach((layer, index) => {
151    if (layer.text === undefined) return
152    const body = stripComments(layer.text)
153    if (index === 0) {
154      text = body.trim()
155      contributors.push({ source: layer.source, how: 'base' })
156      return
157    }
158    const lines = body.split(/\r?\n/)
159    const hasMarker = lines.some(line => line.trim() === DEFAULTS_MARKER)
160    const meaningful = lines.some(line => line.trim() !== '' && line.trim() !== DEFAULTS_MARKER)
161    if (!meaningful && !hasMarker) return // empty file: nothing to say
162    if (!meaningful && hasMarker) return // only `$defaults`: identity
163    text = hasMarker
164      ? lines.map(line => (line.trim() === DEFAULTS_MARKER ? text : line)).join('\n').trim()
165      : body.trim()
166    contributors.push({ source: layer.source, how: hasMarker ? 'spliced' : 'replaced' })
167  })
168
169  return { text, contributors }
170}
171
172// --- transcript trimming ---------------------------------------------------------
173
174/** The shape `$.session.messages()` returns, as far as trimming needs it. */
175export type TranscriptToolUse = {
176  tool: string
177  tool_use_id?: string
178  input?: unknown
179  /** The result as the model read it; absent while the call is in flight. */
180  text?: string
181}
182
183export type TranscriptMessage = {
184  role: 'user' | 'assistant'
185  text: string
186  toolUses?: readonly TranscriptToolUse[]
187  toolResults?: readonly unknown[]
188}
189
190/** The tool the model asks the user multiple-choice questions with; its answers are kept. */
191export const QUESTION_TOOL = 'AskUserQuestion'
192
193/** `Which platforms? [options: Xero | QuickBooks]; Where? [options: UK | EU]` from AskUserQuestion's input. */
194export function questionText(input: unknown): string {
195  const questions = (input as { questions?: unknown } | null | undefined)?.questions
196  if (!Array.isArray(questions)) return ''
197  return questions
198    .map(q => {
199      const record = (q ?? {}) as { question?: unknown; options?: unknown }
200      const question = typeof record.question === 'string' ? record.question.trim() : ''
201      const labels = Array.isArray(record.options)
202        ? record.options.map(o => (typeof o === 'string' ? o : (o as { label?: unknown } | null)?.label)).filter((l): l is string => typeof l === 'string')
203        : []
204      return `${question}${labels.length ? ` [options: ${labels.join(' | ')}]` : ''}`
205    })
206    .filter(Boolean)
207    .join('; ')
208}
209
210/**
211 * The transcript with an AskUserQuestion call's answer filled in: at
212 * `tool.call` the answer is known before the transcript holds it. Patches the
213 * matching tool use, or appends one when the transcript has not got it yet.
214 */
215export function withQuestionAnswer(
216  messages: readonly TranscriptMessage[],
217  answer: { toolUseId?: string; input: unknown; text: string },
218): TranscriptMessage[] {
219  const out = messages.map(m => ({ ...m }))
220  for (const message of out) {
221    const uses = message.toolUses ?? []
222    const at = uses.findIndex(u => u.tool === QUESTION_TOOL && answer.toolUseId !== undefined && u.tool_use_id === answer.toolUseId)
223    if (at >= 0) {
224      if (uses[at]?.text) return out
225      message.toolUses = uses.map((u, i) => (i === at ? { ...u, text: answer.text } : u))
226      return out
227    }
228  }
229  out.push({ role: 'assistant', text: '', toolUses: [{ tool: QUESTION_TOOL, tool_use_id: answer.toolUseId, input: answer.input, text: answer.text }] })
230  return out
231}
232
233export type TrimLimits = {
234  /** Cap per human prompt; human prompts are kept whole up to this. */
235  userChars: number
236  /** Cap per assistant message's text. */
237  assistantChars: number
238  /** Cap for the last assistant message, often the question a short reply answers. */
239  lastAssistantChars: number
240  /** Cap on the rendered conversation before the prompt being assessed: the first prompt and the newest lines are kept, human lines before assistant text. */
241  totalChars: number
242}
243
244export const DEFAULT_TRIM: TrimLimits = { userChars: 4000, assistantChars: 300, lastAssistantChars: 2000, totalChars: 24000 }
245
246const COMMAND_MESSAGE = /^\s*<(command-name|command-message|local-command-stdout|local-command-stderr)>/
247
248const cut = (text: string, max: number): string =>
249  text.length <= max ? text : `${text.slice(0, max)}… [${text.length - max} more chars]`
250
251/** Tool names with repeat counts, in first-use order: `Read×3, Edit, Bash`. AskUserQuestion is left out: it is rendered in full. */
252export function toolNames(uses: readonly { tool: string }[] | undefined): string {
253  uses = uses?.filter(use => use.tool !== QUESTION_TOOL)
254  if (!uses || uses.length === 0) return ''
255  const counts = new Map<string, number>()
256  for (const use of uses) counts.set(use.tool, (counts.get(use.tool) ?? 0) + 1)
257  return [...counts].map(([tool, n]) => (n > 1 ? `${tool}×${n}` : tool)).join(', ')
258}
259
260/**
261 * Renders the transcript for the classifier: human prompts in full (capped),
262 * assistant text truncated, tool uses as names only, tool results and slash
263 * command echoes dropped. AskUserQuestion is the exception: its questions
264 * (`ASSISTANT asked:`) and the user's answers (`USER answered:`) are kept,
265 * because they are the user's words about the task. The last assistant message keeps more of its text
266 * (`lastAssistantChars`): it is often the question that a short reply such as
267 * "2" answers. `current` is the prompt being submitted, which
268 * `$.session.messages()` does not hold yet at `prompt.submit`.
269 *
270 * When over `totalChars`, the first human prompt and the most recent lines
271 * are kept and the middle is replaced by a marker.
272 */
273export function trimTranscript(
274  messages: readonly TranscriptMessage[],
275  current?: string,
276  limits: TrimLimits = DEFAULT_TRIM,
277): string {
278  return renderTranscript(messages, current, limits).text
279}
280
281/** The conversation's lines as the classifier reads them, without the prompt being submitted. */
282function transcriptLines(messages: readonly TranscriptMessage[], limits: TrimLimits): string[] {
283  const lines: string[] = []
284  let lastAssistant = -1
285  messages.forEach((message, index) => {
286    if (message.role === 'assistant' && (message.text ?? '').trim() !== '') lastAssistant = index
287  })
288  for (const [index, message] of messages.entries()) {
289    const text = (message.text ?? '').trim()
290    if (message.role === 'user') {
291      if (text === '' || COMMAND_MESSAGE.test(text)) continue // tool results, /commands
292      lines.push(`USER: ${cut(text, limits.userChars)}`)
293    } else {
294      const tools = toolNames(message.toolUses)
295      const cap = index === lastAssistant ? limits.lastAssistantChars : limits.assistantChars
296      const said = text === '' ? '' : cut(text.replace(/\s+/g, ' '), cap)
297      if (said !== '' || tools !== '') lines.push(`ASSISTANT: ${said}${said && tools ? ' ' : ''}${tools ? `[tools: ${tools}]` : ''}`)
298      for (const use of message.toolUses ?? []) {
299        if (use.tool !== QUESTION_TOOL) continue
300        const asked = questionText(use.input)
301        if (asked) lines.push(`ASSISTANT asked: ${cut(asked, limits.lastAssistantChars)}`)
302        const answered = typeof use.text === 'string' ? use.text.replace(/\s+/g, ' ').trim() : ''
303        if (answered) lines.push(`USER answered: ${cut(answered, limits.userChars)}`)
304      }
305    }
306  }
307  return lines
308}
309
310/** What a capped transcript kept, for `/er status`. */
311export type CapStats = { text: string; fullChars: number; sentChars: number; omitted: number }
312
313/** The transcript as `trimTranscript` renders it, with what the cap dropped. */
314export function renderTranscript(messages: readonly TranscriptMessage[], current?: string, limits: TrimLimits = DEFAULT_TRIM): CapStats {
315  const lines = transcriptLines(messages, limits)
316  const now = (current ?? '').trim()
317  // The prompt being assessed goes in whole, outside the cap: a long dictated brief is the prompt that matters most.
318  const currentLine = now !== '' && !COMMAND_MESSAGE.test(now) ? `USER: ${now}` : undefined
319  const capped = capLines(lines, limits.totalChars)
320  const text = [capped.text, currentLine].filter(part => part !== undefined && part !== '').join('\n')
321  const fullChars = [...lines, ...(currentLine ? [currentLine] : [])].join('\n').length
322  return { text, sentChars: text.length, omitted: capped.omitted, fullChars }
323}
324
325const isHumanLine = (line: string): boolean => line.startsWith('USER') || line.startsWith('ASSISTANT asked:')
326
327/**
328 * Fits rendered lines into `max` characters. Always keeps the first human
329 * prompt (the original task). Then, newest first, the human side (prompts,
330 * AskUserQuestion questions and answers) and the last assistant line (often
331 * the question a short reply answers); then, newest first, other assistant
332 * text with what is left. Order is kept; each gap becomes one marker line.
333 */
334export function capLines(lines: readonly string[], max: number): { text: string; sentChars: number; omitted: number } {
335  const size = (line: string) => line.length + 1
336  const whole = lines.reduce((n, line) => n + size(line), 0)
337  if (whole <= max) {
338    const text = lines.join('\n')
339    return { text, sentChars: text.length, omitted: 0 }
340  }
341  // Budget the output as it will be rendered: kept lines, plus one marker per
342  // run of dropped lines (adding a line can split a run, shrink it or close it).
343  const n = lines.length
344  const keep = new Set<number>()
345  const MARKER = 30 // `[… 123 messages omitted …]` and its newline
346  let used = n > 0 ? MARKER : 0
347  const dropped = (i: number) => i >= 0 && i < n && !keep.has(i)
348  const add = (i: number): boolean => {
349    const left = dropped(i - 1)
350    const right = dropped(i + 1)
351    const markers = left && right ? 1 : !left && !right ? -1 : 0
352    const cost = size(lines[i] as string) + markers * MARKER
353    if (used + cost > max) return false
354    keep.add(i)
355    used += cost
356    return true
357  }
358  const firstUser = lines.findIndex(line => line.startsWith('USER: '))
359  if (firstUser >= 0) add(firstUser)
360  let lastAssistant = -1
361  lines.forEach((line, i) => {
362    if (line.startsWith('ASSISTANT: ')) lastAssistant = i
363  })
364  for (let i = lines.length - 1; i >= 0; i--) {
365    if (keep.has(i)) continue
366    if (isHumanLine(lines[i] as string) || i === lastAssistant) {
367      if (!add(i)) break
368    }
369  }
370  for (let i = lines.length - 1; i >= 0; i--) {
371    if (keep.has(i)) continue
372    if (!add(i)) break
373  }
374  const out: string[] = []
375  let gap = 0
376  lines.forEach((line, i) => {
377    if (keep.has(i)) {
378      if (gap > 0) out.push(`[… ${gap} messages omitted …]`)
379      gap = 0
380      out.push(line)
381    } else gap++
382  })
383  if (gap > 0) out.push(`[… ${gap} messages omitted …]`)
384  const text = out.join('\n')
385  return { text, sentChars: text.length, omitted: lines.length - keep.size }
386}
387
388/** Human prompts in a stored transcript: the ones a person typed, not tool results, /commands or interruptions. */
389export function humanPromptCount(messages: readonly TranscriptMessage[]): number {
390  return messages.filter(message => {
391    if (message.role !== 'user') return false
392    const text = (message.text ?? '').trim()
393    return text !== '' && !COMMAND_MESSAGE.test(text) && !text.startsWith('[Request interrupted')
394  }).length
395}
396
397/**
398 * The user message sent to the classifier. The session's instructions
399 * (CLAUDE.md files, rules, memory) come first when known; a manual
400 * `/er assess <hint>` adds the hint after the transcript.
401 */
402/** The line that tells a check which level the session is on now. */
403const inForceLine = (inForce: Level | undefined): string => (inForce ? `\n\nThe session is at ${inForce} effort now.` : '')
404
405export const classifierPrompt = (transcript: string, hint?: string, instructions?: string, inForce?: Level): string => {
406  const said = hint?.trim()
407  const hintBlock = said ? `\n\n<user_hint>\n${said}\n</user_hint>\nThe user asked for this routing explicitly and gave this hint; weigh it strongly.` : ''
408  const given = instructions?.trim()
409  const instructionsBlock = given ? `The session's instructions (CLAUDE.md files, rules and memory), as its model sees them:\n<instructions>\n${given}\n</instructions>\n\n` : ''
410  return `${instructionsBlock}Transcript so far (oldest first):\n<transcript>\n${transcript}\n</transcript>${hintBlock}${inForceLine(inForce)}\n\n${QUESTION}`
411}
412
413// --- parsing the classifier's reply ---------------------------------------------
414
415export type Decision =
416  | { decision: 'undecided' }
417  | { decision: 'lock'; level: Level; reason: string; why?: string; against?: Level }
418
419/**
420 * Reads the classifier's reply: the first `{...}` in it, so a reply fenced
421 * in a json code block or wrapped in prose still parses. A reply with a level
422 * and no `decision` is read as that level. Anything unparseable or an unknown
423 * level is `undecided`: the router never moves on a reply it cannot read
424 * (fail open).
425 */
426export function parseDecision(reply: string | undefined | null): Decision {
427  const record = levelInDecision(jsonObjectOf(reply))
428  const decided = record?.decision === 'level' || record?.decision === 'lock' || record?.decision === 'suggest' || (record?.decision === undefined && record?.level !== undefined)
429  if (!record || !decided) return { decision: 'undecided' }
430  const proposal = proposalOf(record)
431  return proposal ? { decision: 'lock', ...proposal } : { decision: 'undecided' }
432}
433
434/**
435 * A reply that names its level as the decision (`{"decision":"medium"}`, seen
436 * from Sonnet 5.5 on a subagent check) is read as that level.
437 */
438function levelInDecision(record: Record<string, unknown> | undefined): Record<string, unknown> | undefined {
439  const named = typeof record?.decision === 'string' ? record.decision.trim().toLowerCase() : undefined
440  return record && isLevel(named) && record.level === undefined ? { ...record, decision: 'level', level: named } : record
441}
442
443/**
444 * The first JSON object in a reply, parsed; undefined when there is none.
445 * Reads from the first `{` to the brace that closes it, so text after the
446 * object is ignored, and closes braces left open at the end of the reply
447 * (Fable 5.1 sometimes stops before its last `}`).
448 */
449function jsonObjectOf(reply: string | undefined | null): Record<string, unknown> | undefined {
450  if (typeof reply !== 'string') return undefined
451  const start = reply.indexOf('{')
452  if (start < 0) return undefined
453  let depth = 0
454  let inString = false
455  let end = reply.length
456  for (let i = start; i < reply.length; i++) {
457    const c = reply[i]
458    if (inString) {
459      if (c === '\\') i++
460      else if (c === '"') inString = false
461    } else if (c === '"') inString = true
462    else if (c === '{') depth++
463    else if (c === '}' && --depth === 0) {
464      end = i + 1
465      break
466    }
467  }
468  const text = reply.slice(start, end).trimEnd() + (end === reply.length && depth > 0 && !inString ? '}'.repeat(depth) : '')
469  try {
470    const data: unknown = JSON.parse(text)
471    return typeof data === 'object' && data !== null ? (data as Record<string, unknown>) : undefined
472  } catch {
473    return undefined
474  }
475}
476
477/** A reply's level (case-insensitive) and reason (capped); undefined for an unknown level. */
478function proposalOf(record: Record<string, unknown>): Proposal | undefined {
479  const level = typeof record.level === 'string' ? record.level.trim().toLowerCase() : undefined
480  if (!isLevel(level)) return undefined
481  const reason = typeof record.reason === 'string' ? record.reason.replace(/\s+/g, ' ').trim() : ''
482  const why = typeof record.why === 'string' ? record.why.replace(/\s+/g, ' ').trim().slice(0, 400) : ''
483  return { level, reason: reason === '' ? 'classifier' : cut(reason, 60).replace(/… \[\d+ more chars\]$/, '…'), ...(why ? { why } : {}) }
484}
485
486
487// --- supported models ---------------------------------------------------------------
488
489/**
490 * The models the router supports: the current generation, each with a notes
491 * file in `rules/models/` on what effort means there. On any other model the
492 * router stands aside, because its rules and notes were written for these
493 * levels. A new model needs a new version of the plugin.
494 */
495export type SupportedModel = {
496  id: string
497  alias: string
498  name: string
499  notesFile: string
500  /** The highest level the router picks on this model, below the highestLevel option. */
501  highest?: Level
502  /** Its subagents are checked by a call on this model with their brief alone, not a fork of the parent. */
503  checksOwnBrief?: boolean
504}
505
506export const SUPPORTED_MODELS: readonly SupportedModel[] = [
507  { id: 'claude-fable-5-1', alias: 'fable', name: 'Fable 5.1', notesFile: 'fable-5-1.md' },
508  { id: 'claude-opus-5-5', alias: 'opus', name: 'Opus 5.5', notesFile: 'opus-5-5.md' },
509  { id: 'claude-sonnet-5-5', alias: 'sonnet', name: 'Sonnet 5.5', notesFile: 'sonnet-5-5.md' },
510  // A subagent model in practice: short, scoped jobs that a fork of a bigger parent would cost as much to judge as they
511  // save, and where each step above high buys little for many more steps (Haiku 5.5 needs Claude Code 2.1.293).
512  { id: 'claude-haiku-5-5', alias: 'haiku', name: 'Haiku 5.5', notesFile: 'haiku-5-5.md', highest: 'high', checksOwnBrief: true },
513]
514
515/** `Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 5.5`. */
516export const SUPPORTED_NAMES = SUPPORTED_MODELS.map(m => m.name).join(', ').replace(/, ([^,]*)$/, ' and $1')
517
518/** The supported model a model id or alias names (`claude-opus-5-5`, `claude-opus-5-5[1m]`, `opus`, a cloud provider's id), or undefined. */
519export function supportedModel(model: string | undefined): SupportedModel | undefined {
520  if (!model) return undefined
521  const id = model.toLowerCase().replace(/\[[^\]]*\]$/, '').trim()
522  return SUPPORTED_MODELS.find(m => id === m.alias || id.includes(m.id))
523}
524
525/** A model's name as people say it: `Opus 5.5`, `Haiku 4.5`; the id itself when it is not a Claude id. */
526export function modelName(model: string | undefined): string {
527  if (!model) return 'this model'
528  const known = supportedModel(model)
529  if (known) return known.name
530  const match = model.toLowerCase().match(/claude-([a-z]+)-(\d+)(?:-(\d{1,2})(?!\d))?/)
531  if (!match?.[1] || !match[2]) return model
532  return `${match[1].charAt(0).toUpperCase()}${match[1].slice(1)} ${match[2]}${match[3] ? `.${match[3]}` : ''}`
533}
534
535// --- subagents --------------------------------------------------------------------
536
537/**
538 * The frame for a subagent's read. A subagent is routed once, at spawn. Like
539 * the session frame it ties no kind of task to a level: the notes on the
540 * model the subagent runs on say what each level can do. The same rules sit
541 * inside it, so a user's or organisation's rules ("payments code is never
542 * below high") still apply to subagents.
543 */
544export const subagentFrame = (levels: readonly Level[] = levelsUpTo()): string => `You pick the reasoning-effort level for one Claude Code subagent. Levels you may pick, lowest to highest: ${levels.join(', ')}.
545
546No user is in the loop: the subagent works alone from its brief until it reports back. Pick the level that gets its work done in the least time and total inference cost. People turn this router on to spend less, so when two levels would both get the work done, pick the cheaper one. If the brief says how hard to think, follow it. Too little effort is not cheaper when its work is wrong or has to be redone. Too much pays for thinking the work won't use.
547
548Principles for choosing. They were written for whole sessions, so read them for a subagent. Later rules override earlier ones where they conflict:`
549
550export const SUBAGENT_FRAME = subagentFrame()
551
552export const subagentContract = (levels: readonly Level[] = levelsUpTo()): string => `Reply with one JSON object and nothing else:
553{"level":"<${levels.join('|')}>","reason":"<what the subagent's task is, 3-8 words>"}`
554
555export const SUBAGENT_CONTRACT = subagentContract()
556
557/** The subagent read's whole system prompt around the composed rules, with the notes on the model it runs on. */
558export const subagentSystem = (rules: string, model?: ModelNotes, levels: readonly Level[] = levelsUpTo()): string =>
559  `${subagentFrame(levels)}\n\n<rules>\n${rules.trim()}\n</rules>${modelBlock(model, 'The subagent runs on', levels)}\n\n${subagentContract(levels)}`
560
561/** What `agent.spawn` says about the subagent, as far as its read needs it. */
562export type SubagentBrief = { subagentType: string; description: string; prompt: string }
563
564/**
565 * Fits a brief into `max` characters: the head (the task is usually stated
566 * first) and the tail (often what to report back), with a marker between.
567 */
568export function capBrief(text: string, max: number): string {
569  if (text.length <= max) return text
570  const marker = (n: number) => `\n[… ${n} chars omitted …]\n`
571  const room = Math.max(0, max - marker(text.length).length)
572  const head = Math.ceil(room * 0.75)
573  const tail = room - head
574  return `${text.slice(0, head)}${marker(text.length - head - tail)}${tail > 0 ? text.slice(-tail) : ''}`
575}
576
577/** The user message for a subagent's read on another model: its type, description and brief (capped at `maxChars`). */
578export const subagentPrompt = (brief: SubagentBrief, maxChars: number = DEFAULT_TRIM.totalChars): string =>
579  `Agent type: ${brief.subagentType || 'unknown'}\nDescription: ${brief.description.trim() || '(none)'}\n<brief>\n${capBrief(brief.prompt.trim(), maxChars)}\n</brief>\n\nPick the effort level this subagent should run at, from its brief alone: decide from it and do not assume context it does not state. JSON only.`
580
581/**
582 * The message a subagent's read sends into a fork of its parent at spawn:
583 * the parent knows the task and why it is delegating this part, which the
584 * brief alone often doesn't say.
585 */
586export function subagentForkPrompt(input: { rules: string; brief: SubagentBrief; runsOn: string; model?: ModelNotes; maxChars?: number; levels?: readonly Level[] }): string {
587  const { brief } = input
588  return [
589    `Pause the task for a moment. Do not use any tools and do not start the subagent yourself: answer only the question below. You are about to start a ${brief.subagentType || 'general-purpose'} subagent on ${input.runsOn}${brief.description.trim() ? ` ("${brief.description.trim()}")` : ''} with the brief below. You know the task and why you are delegating this part of it: use that.`,
590    subagentSystem(input.rules, input.model, input.levels),
591    `<brief>\n${capBrief(brief.prompt.trim(), input.maxChars ?? DEFAULT_TRIM.totalChars)}\n</brief>`,
592    'Pick the effort level this subagent should run at. JSON only.',
593  ].join('\n\n')
594}
595
596/**
597 * Reads a subagent read's reply: a level and reason, from the first `{...}`.
598 * The `decision` field may be left out; an explicit undecided, an unknown
599 * level or anything unparseable is undefined, and the caller falls back.
600 */
601export function parseSubagentReply(reply: string | undefined | null): Proposal | undefined {
602  const record = levelInDecision(jsonObjectOf(reply))
603  if (!record) return undefined
604  if (record.decision !== undefined && record.decision !== 'level' && record.decision !== 'lock' && record.decision !== 'suggest') return undefined
605  return proposalOf(record)
606}
607
608/**
609 * A subagent the router routed: the level its requests carry, and why. With
610 * `byDefinition`, its agent definition sets the level (the engine applies it;
611 * the router leaves its requests alone and only records it), which may be a
612 * number.
613 */
614export type RoutedAgent = { level: Level | number; reason: string; subagentType: string; description: string; byDefinition?: boolean }
615
616// --- agent definitions that set their own effort ------------------------------------
617
618/** One agent definition, as far as the router needs it: its name, and its effort when it sets one. */
619export type AgentDefinition = { name: string; effort?: Level | number; model?: string; source: string }
620
621/** A definition's model, unless it inherits the parent's. */
622const definitionModel = (value: unknown): string | undefined => {
623  const text = typeof value === 'string' ? value.trim().replace(/^(['"])(.*)\1$/, '$2').trim() : ''
624  return text === '' || text.toLowerCase() === 'inherit' ? undefined : text
625}
626
627/** A definition's effort when it is one the engine takes: a level, or a positive number. */
628export function definitionEffort(value: unknown): Level | number | undefined {
629  if (typeof value === 'number') return Number.isFinite(value) && value > 0 ? value : undefined
630  if (typeof value !== 'string') return undefined
631  const text = value.trim().replace(/^(['"])(.*)\1$/, '$2').trim().toLowerCase()
632  if (isLevel(text)) return text
633  return /^\d+(\.\d+)?$/.test(text) && Number(text) > 0 ? Number(text) : undefined
634}
635
636/**
637 * The top-level `key: value` pairs of a markdown file's YAML frontmatter
638 * (between `---` lines at the very start); undefined without one. Values are
639 * unquoted; nested and list values are left out (the router reads only
640 * `name` and `effort`).
641 */
642export function frontmatterOf(text: string): Record<string, string> | undefined {
643  const match = text.replace(/^\uFEFF/, '').match(/^---\r?\n([\s\S]*?)\r?\n---[ \t]*(\r?\n|$)/)
644  if (!match) return undefined
645  const fields: Record<string, string> = {}
646  for (const line of (match[1] as string).split(/\r?\n/)) {
647    const field = line.match(/^([A-Za-z_][\w-]*)[ \t]*:[ \t]*(.*?)[ \t]*$/)
648    if (!field) continue
649    const value = (field[2] as string).replace(/[ \t]+#.*$/, '').replace(/^(['"])(.*)\1$/, '$2')
650    if (value !== '' && !(field[1] as string in fields)) fields[field[1] as string] = value
651  }
652  return fields
653}
654
655/**
656 * An agent definition file (`.claude/agents/*.md`): named by its frontmatter
657 * `name:`, else by its file name. Undefined for a file with no frontmatter,
658 * which is not an agent definition.
659 */
660export function agentFileDefinition(text: string, fileName: string, source: string): AgentDefinition | undefined {
661  const fields = frontmatterOf(text)
662  if (!fields) return undefined
663  const name = fields.name?.trim() || fileName.replace(/\.md$/i, '')
664  const effort = definitionEffort(fields.effort)
665  const model = definitionModel(fields.model)
666  return { name, ...(effort === undefined ? {} : { effort }), ...(model === undefined ? {} : { model }), source }
667}
668
669/**
670 * The agent definitions in a settings source's `agents` key: an object keyed
671 * by agent name (as `--agents` takes them), or a list of `{ name, ... }`.
672 * Anything malformed is skipped.
673 */
674export function settingsAgentDefinitions(settings: unknown, source: string): AgentDefinition[] {
675  const agents = (settings as { agents?: unknown } | null | undefined)?.agents
676  if (typeof agents !== 'object' || agents === null) return []
677  const entries: [unknown, unknown][] = Array.isArray(agents)
678    ? agents.map(agent => [(agent as { name?: unknown } | null)?.name, agent])
679    : Object.entries(agents)
680  const out: AgentDefinition[] = []
681  for (const [name, spec] of entries) {
682    if (typeof name !== 'string' || name.trim() === '' || typeof spec !== 'object' || spec === null) continue
683    const effort = definitionEffort((spec as { effort?: unknown }).effort)
684    const model = definitionModel((spec as { model?: unknown }).model)
685    out.push({ name: name.trim(), ...(effort === undefined ? {} : { effort }), ...(model === undefined ? {} : { model }), source })
686  }
687  return out
688}
689
690/**
691 * The definition a spawn of `subagentType` runs under: the first one with that
692 * name, highest precedence first, whether or not it sets an effort (a project
693 * definition without one still overrides a user definition with one). A
694 * plugin's agent (`<plugin>:<name>`) is never looked up here.
695 */
696export function definitionFor(subagentType: string, definitions: readonly AgentDefinition[]): AgentDefinition | undefined {
697  if (subagentType.includes(':')) return undefined
698  return definitions.find(definition => definition.name === subagentType)
699}
700
701/**
702 * Whether subagents are routed in this state. They are unless you turned the
703 * router off (Turn off, `/er off`, or changing the effort picker). A session
704 * that started before the router still routes them: each brief is a new,
705 * whole task.
706 */
707export const routesSubagents = (state: RouterState): boolean => state.status !== 'off' || state.offReason === 'mid-flow'
708
709/**
710 * The level a spawn inherits, for a fork or when its read fails: the parent
711 * subagent's level for a nested spawn (routed, or a level its definition
712 * set), else the main thread's level in use; undefined when neither has one
713 * (the request is left alone).
714 */
715export function parentLevel(state: RouterState, agents: ReadonlyMap<string, RoutedAgent>, parentAgentId?: string): Level | undefined {
716  const parent = parentAgentId !== undefined ? agents.get(parentAgentId)?.level : undefined
717  return isLevel(parent) ? parent : appliedLevel(state)
718}
719
720/** Why subagents are or are not routed, for `/er status`. */
721export type SubagentStatus = { routing: 'on' | 'setting' | 'user-off'; agents: readonly RoutedAgent[] }
722
723/** `/er status`'s subagent lines: whether they are routed, then the newest `shown`, newest first. */
724export function subagentReport(status: SubagentStatus, shown = 10): string[] {
725  const why = {
726    on: 'each gets its own level from its task',
727    setting: 'not routed (routeSubagents is off), so they use the session level',
728    'user-off': 'not routed while the router is off, so they use your effort setting',
729  }[status.routing]
730  const lines = [`Subagents: ${why}.`]
731  if (status.agents.length === 0) return lines
732  const recent = status.agents.slice(-shown).reverse()
733  lines.push(`Recent subagents (${status.agents.length}${status.agents.length > recent.length ? `, newest ${recent.length} shown` : ''}):`)
734  for (const agent of recent) {
735    const description = cut(agent.description.replace(/\s+/g, ' ').trim() || agent.subagentType || 'a subagent', 60).replace(/… \[\d+ more chars\]$/, '…')
736    // A bullet, not an indent: the Desktop app drops leading spaces in command output.
737    lines.push(`- ${agent.level}: ${description} (${agent.byDefinition && agent.reason.startsWith('from ') ? 'set by its agent definition' : agent.reason})`)
738  }
739  return lines
740}
741
742// --- the spend ledger -------------------------------------------------------------
743
744/**
745 * What the router records about each model request, so `/er report` can
746 * say where the effort went: the level the request arrived at (the picker's,
747 * or the level a subagent would have inherited), the level it went out at,
748 * and what it cost as the API reported it. Requests are summed into rows per
749 * UTC day and pair of levels. One file per session, written by that session.
750 */
751export type SpendRow = {
752  /** UTC day, YYYY-MM-DD. */
753  day: string
754  caller: 'main' | 'subagent'
755  /** The level the request arrived at: what it would have run at without the router. `none` for a model without effort. */
756  from: string
757  /** The level it went out at. */
758  to: string
759  /** A subagent whose own definition set its level (the router left it alone). */
760  byDefinition?: true
761  requests: number
762  /** Output tokens: thinking and the answer, the part effort changes most. */
763  output: number
764  /** Input tokens, cached and uncached. */
765  input: number
766}
767
768/**
769 * How a check was made: `first`, a separate call on the session's model
770 * before the conversation has a request to fork; `fork`, a fork of the
771 * conversation; `separate`, a call on the check model setting's model;
772 * `subagent`, a subagent's brief.
773 */
774export type CheckKind = 'first' | 'fork' | 'separate' | 'subagent'
775
776export const isCheckKind = (value: unknown): value is CheckKind =>
777  value === 'first' || value === 'fork' || value === 'separate' || value === 'subagent'
778
779/** The router's own reads, per UTC day and kind (no kind: recorded before 0.10). */
780export type ReadRow = { day: string; kind?: CheckKind; calls: number; output: number; input: number }
781
782/**
783 * One assessment and what came of it, kept for calibration: its level, the
784 * level the session was on, and `outcome` (`stayed`, `moved to high`, `no
785 * clear task`, `judged at the first request`). Rows before 0.18 also hold the
786 * check's spread and confidence, and older ones the ask-era outcomes.
787 */
788export type VerdictRow = {
789  at: number
790  kind: CheckKind
791  model: string
792  prompt: number
793  level?: Level
794  confidence?: number
795  /** The check's short task summary and its why, for calibration. */
796  reason?: string
797  why?: string
798  /** Before 0.18: the check's probability for each level. */
799  spread?: Spread
800  /** The level the session was on when the check ran. */
801  against?: Level
802  outcome: string
803  /** A first check that carried the session's instructions (CLAUDE.md, rules, memory), to learn whether they help. */
804  withInstructions?: boolean
805  /** Asked for by you (the band's Assess, `/er assess`), not counted toward the window. */
806  manual?: true
807}
808
809/** Verdicts kept per session. */
810export const MAX_VERDICTS = 200
811
812/** A subagent's level as set at its spawn, against the level it would have inherited, for calibration. */
813export type SubagentRow = {
814  at: number
815  /** The model it runs on: the Agent call's, else its parent's. */
816  model: string
817  subagentType: string
818  description: string
819  /** The level its parent was running at, which it would otherwise have run at. */
820  parent?: Level
821  level: Level | number
822  reason: string
823  /** Its definition set its effort, so the router left it alone. */
824  byDefinition?: true
825  /** How long the spawn waited for its level, in ms. */
826  ms: number
827}
828
829/** One session's ledger: its requests, the router's reads, its assessments and, since 0.17, its state. */
830export type SpendLedger = {
831  version: 1; session: string; repo: string; rows: SpendRow[]; reads: ReadRow[]; verdicts?: VerdictRow[]; state?: SavedState
832  /** Each routed subagent's level and why (since 0.17.3). */
833  subagents?: SubagentRow[]
834  /** Your effort setting as the last main-thread request showed it, so a picker change is still seen after a reload, restart or resume. */
835  setting?: Level
836}
837
838/** A request's usage, in the API's spelling. */
839export type SpendUsage = { input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number }
840
841export type SpendPeriod = 'session' | 'week' | 'month' | 'all'
842
843export const isSpendPeriod = (value: unknown): value is SpendPeriod =>
844  value === 'session' || value === 'week' || value === 'month' || value === 'all'
845
846export const emptyLedger = (session: string, repo: string): SpendLedger => ({ version: 1, session, repo, rows: [], reads: [] })
847
848/** The UTC day of a time in epoch ms, YYYY-MM-DD. */
849export const dayOf = (ms: number): string => new Date(ms).toISOString().slice(0, 10)
850
851const inputOf = (usage: SpendUsage): number =>
852  usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
853
854const levelName = (value: unknown): string => (value === undefined || value === null ? 'none' : String(value))
855
856export type SpendEntry = { day: string; caller: 'main' | 'subagent'; from: unknown; to: unknown; byDefinition?: boolean; usage: SpendUsage }
857
858/** Adds one request to its row. */
859export function withSpend(ledger: SpendLedger, entry: SpendEntry): SpendLedger {
860  const from = levelName(entry.from)
861  const to = levelName(entry.to)
862  const byDefinition = entry.byDefinition ? (true as const) : undefined
863  const at = ledger.rows.findIndex(
864    r => r.day === entry.day && r.caller === entry.caller && r.from === from && r.to === to && r.byDefinition === byDefinition,
865  )
866  const old: SpendRow = ledger.rows[at] ?? { day: entry.day, caller: entry.caller, from, to, ...(byDefinition ? { byDefinition } : {}), requests: 0, output: 0, input: 0 }
867  const row = { ...old, requests: old.requests + 1, output: old.output + entry.usage.output_tokens, input: old.input + inputOf(entry.usage) }
868  return { ...ledger, rows: at >= 0 ? ledger.rows.map((r, i) => (i === at ? row : r)) : [...ledger.rows, row] }
869}
870
871/** Adds one of the router's own reads. */
872export function withRead(ledger: SpendLedger, day: string, usage: SpendUsage, kind?: CheckKind): SpendLedger {
873  const at = ledger.reads.findIndex(r => r.day === day && r.kind === kind)
874  const old: ReadRow = ledger.reads[at] ?? { day, ...(kind ? { kind } : {}), calls: 0, output: 0, input: 0 }
875  const row = { ...old, calls: old.calls + 1, output: old.output + usage.output_tokens, input: old.input + inputOf(usage) }
876  return { ...ledger, reads: at >= 0 ? ledger.reads.map((r, i) => (i === at ? row : r)) : [...ledger.reads, row] }
877}
878
879/** Adds one check's verdict, keeping the newest `MAX_VERDICTS`. */
880export function withVerdictRow(ledger: SpendLedger, row: VerdictRow): SpendLedger {
881  return { ...ledger, verdicts: [...(ledger.verdicts ?? []), row].slice(-MAX_VERDICTS) }
882}
883
884export function withSubagentRow(ledger: SpendLedger, row: SubagentRow): SpendLedger {
885  return { ...ledger, subagents: [...(ledger.subagents ?? []), row].slice(-MAX_VERDICTS) }
886}
887
888/**
889 * Sets what came of a verdict (the answer to its question): the one checked at `at`, else the newest. A check
890 * judged later (at the first request) also sets the level the session was on then.
891 */
892export function withVerdictOutcome(ledger: SpendLedger, outcome: string, at?: number, judged?: { level: Level; against?: Level }): SpendLedger {
893  const verdicts = ledger.verdicts ?? []
894  let index = at === undefined ? -1 : verdicts.findLastIndex(v => v.at === at)
895  if (index < 0) index = verdicts.length - 1
896  const row = verdicts[index]
897  const update = {
898    outcome,
899    ...(judged ? { level: judged.level } : {}),
900    ...(judged?.against ? { against: judged.against } : {}),
901  }
902  return row ? { ...ledger, verdicts: verdicts.map((v, i) => (i === index ? { ...row, ...update } : v)) } : ledger
903}
904
905const isCount = (value: unknown): value is number => typeof value === 'number' && Number.isFinite(value) && value >= 0
906
907/** A ledger file's text, checked; rows that do not fit the shape are dropped. Undefined when it is not a ledger. */
908export function parseLedger(text: string): SpendLedger | undefined {
909  let raw: unknown
910  try {
911    raw = JSON.parse(text)
912  } catch {
913    return undefined
914  }
915  if (typeof raw !== 'object' || raw === null) return undefined
916  const value = raw as Record<string, unknown>
917  if (value.version !== 1 || typeof value.session !== 'string') return undefined
918  const rows = (Array.isArray(value.rows) ? value.rows : []).filter(
919    (r): r is SpendRow =>
920      typeof r === 'object' && r !== null &&
921      typeof r.day === 'string' && (r.caller === 'main' || r.caller === 'subagent') &&
922      typeof r.from === 'string' && typeof r.to === 'string' &&
923      (r.byDefinition === undefined || r.byDefinition === true) &&
924      isCount(r.requests) && isCount(r.output) && isCount(r.input),
925  )
926  const reads = (Array.isArray(value.reads) ? value.reads : []).filter(
927    (r): r is ReadRow =>
928      typeof r === 'object' && r !== null && typeof r.day === 'string' && (r.kind === undefined || isCheckKind(r.kind)) &&
929      isCount(r.calls) && isCount(r.output) && isCount(r.input),
930  )
931  const verdicts = (Array.isArray(value.verdicts) ? value.verdicts : []).filter(
932    (r): r is VerdictRow =>
933      typeof r === 'object' && r !== null && isCount(r.at) && isCheckKind(r.kind) && typeof r.model === 'string' && isCount(r.prompt) &&
934      (r.level === undefined || isLevel(r.level)) && (r.confidence === undefined || isCount(r.confidence)) && typeof r.outcome === 'string' &&
935      (r.withInstructions === undefined || typeof r.withInstructions === 'boolean'),
936  )
937  const subagents = (Array.isArray(value.subagents) ? value.subagents : []).filter(
938    (r): r is SubagentRow =>
939      typeof r === 'object' && r !== null && isCount(r.at) && typeof r.model === 'string' && typeof r.subagentType === 'string' &&
940      typeof r.description === 'string' && (isLevel(r.level) || isCount(r.level)) && typeof r.reason === 'string' && isCount(r.ms) &&
941      (r.parent === undefined || isLevel(r.parent)),
942  )
943  const state = restored(value.state)
944  return {
945    version: 1, session: value.session, repo: typeof value.repo === 'string' ? value.repo : 'unknown', rows, reads,
946    ...(verdicts.length > 0 ? { verdicts } : {}),
947    ...(subagents.length > 0 ? { subagents } : {}),
948    ...(state ? { state: savedOf(state) } : {}),
949    ...(isLevel(value.setting) ? { setting: value.setting } : {}),
950  }
951}
952
953/** A token count in a few characters: 950, 12.3k, 450k, 1.23M. */
954export function tokens(n: number): string {
955  if (n < 1000) return String(Math.round(n))
956  if (n < 10_000) return `${(n / 1000).toFixed(1)}k`
957  if (n < 1_000_000) return `${Math.round(n / 1000)}k`
958  return `${(n / 1_000_000).toFixed(2)}M`
959}
960
961const PERIOD_DAYS: Record<Exclude<SpendPeriod, 'session' | 'all'>, number> = { week: 7, month: 30 }
962
963const periodLabel = (period: SpendPeriod, since?: string): string =>
964  period === 'session' ? 'this session'
965  : period === 'all' ? 'all recorded sessions'
966  : `the last ${PERIOD_DAYS[period]} days (since ${since})`
967
968/** Levels first in their order, then anything else (numbers, none). */
969const byLevelOrder = (a: string, b: string): number => {
970  const rankOf = (s: string) => (isLevel(s) ? rank(s) : LEVELS.length)
971  return rankOf(a) - rankOf(b) || a.localeCompare(b)
972}
973
974const plural = (n: number, word: string): string => `${n} ${word}${n === 1 ? '' : 's'}`
975
976/**
977 * `/er report`: where the effort went over a period, measured. Requests and
978 * output tokens per level; the requests the router moved off the level they
979 * arrived at, with the average size of requests left at that level beside
980 * them; agent definitions' own levels; the router's own reads; and, beyond
981 * one session, the split by repo. No "saved" figure (see the last line).
982 */
983export function spendReport(ledgers: readonly SpendLedger[], period: SpendPeriod, at: { today: string; session: string }): string {
984  const since = period === 'week' || period === 'month'
985    ? dayOf(Date.parse(`${at.today}T00:00:00Z`) - (PERIOD_DAYS[period] - 1) * 86_400_000)
986    : undefined
987  const inPeriod = (day: string) => since === undefined || day >= since
988  const chosen = (period === 'session' ? ledgers.filter(l => l.session === at.session) : ledgers).map(l => ({
989    ...l,
990    rows: l.rows.filter(r => inPeriod(r.day)),
991    reads: l.reads.filter(r => inPeriod(r.day)),
992  }))
993  const rows = chosen.flatMap(l => l.rows)
994  const label = periodLabel(period, since)
995  if (rows.length === 0) {
996    return `Nothing recorded for ${label}. Recording started with version 0.9.0.`
997  }
998  const sum = (list: readonly SpendRow[]) => list.reduce((t, r) => ({ requests: t.requests + r.requests, output: t.output + r.output }), { requests: 0, output: 0 })
999  const avg = (t: { requests: number; output: number }) => tokens(t.output / Math.max(1, t.requests))
1000  const group = <K extends string>(list: readonly SpendRow[], key: (r: SpendRow) => K): Map<K, SpendRow[]> => {
1001    const out = new Map<K, SpendRow[]>()
1002    for (const r of list) out.set(key(r), [...(out.get(key(r)) ?? []), r])
1003    return out
1004  }
1005  const all = sum(rows)
1006  const sessions = chosen.filter(l => l.rows.length > 0).length
1007  const lines = [`Effort for ${label}: ${plural(all.requests, 'request')}${period === 'session' ? '' : ` in ${plural(sessions, 'session')}`}, ${tokens(all.output)} output tokens.`]
1008
1009  lines.push('By level:')
1010  for (const [level, list] of [...group(rows, r => r.to)].sort(([a], [b]) => byLevelOrder(a, b))) {
1011    const t = sum(list)
1012    lines.push(`- ${level}: ${plural(t.requests, 'request')}, ${tokens(t.output)} output tokens (avg ${avg(t)})`)
1013  }
1014
1015  const unmoved = group(rows.filter(r => !r.byDefinition && r.from === r.to), r => r.to)
1016  const moved = rows.filter(r => !r.byDefinition && r.from !== r.to)
1017  if (moved.length === 0) {
1018    lines.push('Changed by the router: none.')
1019  } else {
1020    lines.push(`Changed by the router: ${plural(sum(moved).requests, 'request')}`)
1021    const groups = [...group(moved, r => `${r.caller}|${r.from}|${r.to}`)].map(([key, list]) => {
1022      const [caller, from, to] = key.split('|') as [string, string, string]
1023      return { caller, from, to, t: sum(list) }
1024    })
1025    for (const { caller, from, to, t } of groups.sort((a, b) => b.t.requests - a.t.requests)) {
1026      const left = unmoved.get(from)
1027      const beside = left ? `, vs ${avg(sum(left))} for those left at ${from}` : ''
1028      lines.push(`- ${caller === 'main' ? 'main conversation' : 'subagents'}, ${from} → ${to}: ${plural(t.requests, 'request')}, ${tokens(t.output)} output tokens (avg ${avg(t)}${beside})`)
1029    }
1030  }
1031
1032  const defined = rows.filter(r => r.byDefinition)
1033  if (defined.length > 0) {
1034    const levels = [...group(defined, r => r.to)].sort(([a], [b]) => byLevelOrder(a, b)).map(([level, list]) => `${level} ${sum(list).requests}`)
1035    lines.push(`Set by you or an agent definition: ${plural(sum(defined).requests, 'request')} (${levels.join(', ')}).`)
1036  }
1037
1038  const reads = chosen.flatMap(l => l.reads)
1039  if (reads.length > 0) {
1040    const r = reads.reduce((t, x) => ({ calls: t.calls + x.calls, output: t.output + x.output, input: t.input + x.input }), { calls: 0, output: 0, input: 0 })
1041    const count = (kinds: readonly (CheckKind | undefined)[]) => reads.filter(x => kinds.includes(x.kind)).reduce((n, x) => n + x.calls, 0)
1042    const split = [
1043      [count(['first']), 'of a first prompt'],
1044      [count(['fork', 'separate', undefined]), 'of a conversation'],
1045      [count(['subagent']), 'for subagents'],
1046    ].filter(([n]) => (n as number) > 0).map(([n, what]) => `${n} ${what}`)
1047    const by = split.length > 1 ? ` (${split.join(', ')})` : ''
1048    lines.push(`The router's own assessments: ${r.calls}${by}, using ${tokens(r.output)} output and ${tokens(r.input)} input tokens.`)
1049  }
1050
1051  if (period !== 'session') {
1052    const byRepo = new Map<string, number>()
1053    for (const l of chosen) if (l.rows.length > 0) byRepo.set(l.repo, (byRepo.get(l.repo) ?? 0) + sum(l.rows).output)
1054    const repos = [...byRepo].sort(([, a], [, b]) => b - a)
1055    if (repos.length > 1) {
1056      const shown = repos.slice(0, 6).map(([repo, output]) => `${repo} ${tokens(output)}`)
1057      lines.push(`By repo (output tokens): ${shown.join(', ')}${repos.length > 6 ? `, ${repos.length - 6} more` : ''}.`)
1058    }
1059  }
1060
1061  lines.push(
1062    'No "saved" figure: the router lowers easy tasks and raises hard ones, so these averages can\'t show what a changed request would have cost.',
1063  )
1064  return lines.join('\n')
1065}
1066
1067// --- state and what the mod shows -------------------------------------------------
1068
1069/** A check's probability for each level, normalised to sum to 1. */
1070export type Spread = Partial<Record<Level, number>>
1071
1072/**
1073 * An assessment's level. `against`: the level the session was on when it ran.
1074 * `checkedAt` (when it ran, never
1075 * saved) ties a later judgement back to its ledger row.
1076 */
1077export type Proposal = { level: Level; reason: string; why?: string; against?: Level; checkedAt?: number }
1078
1079/**
1080 * The router's three statuses, as the footer's glyph shows them: `unlocked`
1081 * (it may still move the level, assessing each of the first prompts),
1082 * `locked` (the level holds), `off` (your effort setting applies).
1083 */
1084export type Status = 'unlocked' | 'locked' | 'off'
1085
1086/** Why the router is off: you turned it off, you changed the effort picker, or the session started before the router. */
1087export type OffReason = 'you' | 'picker' | 'mid-flow'
1088
1089/**
1090 * Everything the router remembers about one session.
1091 *
1092 * `level` is the router's own level: after a move while unlocked, or the
1093 * locked level. Undefined while unlocked means your effort setting runs.
1094 * Only an assessment moves it, or a button whose label names a level.
1095 */
1096export type RouterState = {
1097  status: Status
1098  level?: Level
1099  /** Prompts assessed in this window (an earlier session's prompts count on a first sighting). */
1100  assessed: number
1101  /** Who locked it: the router after the last prompt of the window, or you. */
1102  lockedBy?: 'router' | 'you'
1103  /** Prompts assessed when it locked, for the band. */
1104  lockedAfter?: number
1105  offReason?: OffReason
1106  /** The router's last level of its own, for `Turn on, locked at <level>`. */
1107  lastLevel?: Level
1108  /** A hint from `/er assess <hint>` while unlocked, used by the next prompt's assessment. */
1109  hint?: string
1110  /**
1111   * A first assessment made before any request showed the level in force: its
1112   * level is compared with the one the next main-thread request shows. Never saved.
1113   */
1114  pending?: Proposal
1115  /** Shown only, never saved: the session's model, which the router does not support, so it stands aside. */
1116  unsupported?: string
1117}
1118
1119export const freshState = (): RouterState => ({ status: 'unlocked', assessed: 0 })
1120
1121/** The state as it applies on a model: on one the router does not support, it stands aside without forgetting its state. */
1122export function onModel(state: RouterState, model: string | undefined): RouterState {
1123  return model === undefined || supportedModel(model) ? state : { ...state, unsupported: modelName(model) }
1124}
1125
1126/** The level `turn.step` applies to the main thread, or undefined to leave the request at your effort setting. */
1127export function appliedLevel(state: RouterState): Level | undefined {
1128  return state.status !== 'off' && !state.unsupported ? state.level : undefined
1129}
1130
1131/**
1132 * What the router adds to Claude Code's own `api_request` records for an organisation's telemetry collector. The record
1133 * already carries `effort`, the level the request went out at, so these say what it would have been and why it wasn't.
1134 * The values are the session's: the engine gives a mod no way to tie a record to one request, so a subagent's request
1135 * carries its session's values too.
1136 */
1137export function telemetryAttributes(state: RouterState, setting: Level | undefined, version: string): Record<string, string> {
1138  const status = state.unsupported ? 'standing aside' : state.status
1139  const attributes: Record<string, string> = { 'effort_router.version': version, 'effort_router.status': status }
1140  if (setting) attributes['effort_router.setting'] = setting
1141  const level = appliedLevel(state)
1142  if (level) attributes['effort_router.level'] = level
1143  if (state.status === 'off' && state.offReason) attributes['effort_router.off_reason'] = state.offReason
1144  return attributes
1145}
1146
1147/** Whether a human prompt should be assessed now. */
1148export function wantsAssessment(state: RouterState, limit: number): boolean {
1149  return state.status === 'unlocked' && !state.unsupported && state.assessed < limit
1150}
1151
1152/**
1153 * The state for a session the router first sees with prompts already in it.
1154 * Those prompts count toward the window; with the window already used up, the
1155 * session started before the router and it is left off.
1156 */
1157export function firstSighting(prior: number, limit: number): RouterState {
1158  if (prior < limit) return { ...freshState(), assessed: prior }
1159  return { status: 'off', assessed: 0, offReason: 'mid-flow' }
1160}
1161
1162/**
1163 * The levels an assessment is offered: low up to `highestLevel`, or up to your
1164 * own setting when that is higher (a session at max would otherwise always
1165 * read max as too high, since the assessment could never vote for it).
1166 */
1167export function offeredLevels(highest: Level, setting?: Level): readonly Level[] {
1168  return levelsUpTo(setting && rank(setting) > rank(highest) ? setting : highest)
1169}
1170
1171/** What one assessment did, for the messages and the ledger. */
1172export type Settled = { state: RouterState; moved?: { from?: Level; to: Level }; locked?: Level; outcome: string }
1173
1174/**
1175 * Applies an assessment: the session goes to the level it picked. Switching
1176 * costs the user nothing (no approval, no review),
1177 * so the router does what the check says rather than second-guessing it with
1178 * a confidence bar in code. The check is asked for the level that gets the
1179 * work done in the least time and total cost, and weighs the risk itself.
1180 * Counted assessments use up the window, and the last one locks whatever is
1181 * running. A manual assessment while locked moves the locked level and stays
1182 * locked. `running` is the level in force (the router's own, else your
1183 * setting); undefined when no request has shown it yet.
1184 */
1185export function settle(
1186  state: RouterState,
1187  judged: Proposal | undefined,
1188  options: { limit: number; running?: Level; counted: boolean },
1189): Settled {
1190  let next: RouterState = { ...state, pending: undefined, hint: undefined }
1191  if (options.counted) next.assessed = Math.min(options.limit, state.assessed + 1)
1192  let moved: Settled['moved']
1193  let outcome = judged ? 'stayed' : 'no clear task'
1194  if (judged && options.running !== undefined && judged.level !== options.running) {
1195    moved = { from: options.running, to: judged.level }
1196    next = { ...next, level: judged.level, lastLevel: judged.level }
1197    outcome = `moved to ${judged.level}`
1198  }
1199  let locked: Level | undefined
1200  const running = next.level ?? options.running