SLOPSHOPPER

model-router

Haiku sorts every subagent task into simple, standard, hard or long run and picks the model; a pane shows what each agent cost.

newpaneguardcommandstatusmodel
v0.4.1no licenseupdated 2026-10-09aott33/model-router
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · model-router
│ ┃ Model router ✕ › fix the failing auth test and add an audit log call │ ┃ $0 spent · $0 unrouted │ ┃ No model calls yet. · 0 routed · classifier ⏺ Read(src/auth.ts) │ ┃ ⎿ Read 6 lines │ ┃ agent T model cost unr ⏺ Update(src/auth.ts) │ ┃ ⎿ Added 2 lines, removed 1 line │ ┃ T: S simple→Haiku 5.5 at low effort, M stand ⏺ Bash(bun test) │ ┃ at medium effort at most, H hard→Opus, L lon ⎿ 3 pass, 1 fail │ ┃ (background only). ! risky, sent to Opus at │ ┃ one tier down for rate limits. Dim rows are ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ routed. API list prices; plans billed by sub │ ┃ see rate-limit use, not dollars. ✻ Worked for 42s · done 4:20 PM │ │ › /router │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · Model router
$0 spent · $0 unrouted No model calls yet. · 0 routed · classifier $0 agent T model cost unrouted T: S simple→Haiku 5.5 at low effort, M standard→Sonnet at medium effort at most, H hard→Opus, L long→Fable (background only). ! risky, sent to Opus at least; ↓ one tier down for rate limits. Dim rows are not routed. API list prices; plans billed by subscription see rate-limit use, not dollars.
README

model-router

A Claude Code mod that picks the model for each subagent before it starts, and shows what each one cost.

What it does

  1. Four roles. Adds model-router:architect (plans, no edits), model-router:builder (writes code), model-router:runner (runs tests and commands, no edits) and model-router:fixer (fixes what failed).
  2. Haiku reads every task first. On each subagent spawn, Haiku 4.5 sorts the task into simple, standard, hard or long. The classifier stays on Haiku 4.5 because it answers without thinking, so an 8-token reply cap holds. It gets 3 seconds; if it fails or names no tier, the engine's built-in classifier gets another 3 seconds before the role default is used.
  3. The mod routes it.
TierModelEffortTypical work
simpleHaiku 5.5lowrename, search, run tests
standardSonnet 5.5medium at mostnormal feature work
hardOpus 5.5unchangedhard bugs, architecture
longFable 5.1unchangedmulti-hour background runs

Effort is only ever lowered, never raised, and only for agents the router picked a model for. An Agent call that sets its own effort keeps it.

Risk floor. Haiku also flags a task as risky when carrying it out could do costly or hard-to-reverse harm: deploying to production, deleting data, a migration or destructive command on shared data, a forced push. It judges the act, not the subject, so writing code that deals with payments or databases is not risky. A risky task runs on Opus at least, past the runner's cap and past rate-limit pressure. If no classifier answers, a narrow pattern check on destructive acts stands in. The pane marks these rows !.

Rate limits. When the fullest rate-limit window (5-hour or weekly) is at 80% or more (setting), every routed task goes one tier lower and never to Fable. Role floors still hold. The status line shows the window, and the pane marks these rows ↓. Off a subscription there are no windows, so nothing changes.

Guard rails: runner never goes above Sonnet; architect and fixer never go below Sonnet; long (Fable, 2.5 times Opus) only for background agents, so a foreground task the parent waits on tops out at Opus. If neither classifier gives a tier, the role default is used (architect: Opus, builder and fixer: Sonnet, runner: Haiku). Forks, workflow agents and agent-team teammates are not routed (the engine ignores a model change for the first two; a teammate is long-lived, so one classification of its first message is not a good guide). A model Claude already named in the Agent call is kept (setting), and the pane shows it at that model's tier.

  1. Modes. /router off leaves every subagent on the model it would have had; /router haiku, /router sonnet or /router opus sends every routed subagent to that model without asking the classifier (a model Claude named is still kept, per the setting below); /router on goes back to classifying. The mode is kept across sessions and shows in the status line; /router status names it.
  2. Bill pane. /router opens a pane: each agent, its tier, the model that ran it, what it cost, and what the same tokens would have cost without the router. The main conversation and unrouted agents are dim rows, and the Haiku classifier calls are counted in the total. A status line under the prompt shows the running total. /router reset clears it.

Settings (/config)

  • Compare against: unrouted (default) prices each routed agent on its parent's model, the one it would have inherited, and every unrouted row on its own model, so only routing shows as saved. fable, opus or sonnet price everything, the main conversation included, on that model.
  • Route every subagent: on routes built-in subagents (Explore, general-purpose) too; off routes only the four router agents.
  • Keep a model Claude asked for: on by default.
  • Lower effort for easy tasks: on by default; off leaves every agent's effort alone.
  • Send risky tasks to Opus: on by default.
  • Route cheaper near rate limits: 80 (default), 70, 90 or off; the percentage of the fullest rate-limit window at which routing goes one tier lower.

Install

/plugin install model-router --marketplace aott33/model-router

Or for development: claude --plugin-dir ./model-router

Needs Claude Code 2.1.287 or later (mods).

Limits

  • The risk flag is Haiku's judgment of a short task description; the pattern check without it only catches plainly named acts. Neither replaces permission prompts.
  • The built-in classifier's calls are not in the bill: the engine does not report their tokens.
  • A routed subagent's first request may wait up to 2 seconds for its spawn to be recorded, so it gets its effort from the start.
  • Only subagents are routed. The main conversation stays on your session model, so total savings depend on how much work goes through subagents.
  • Costs use API list prices (Oct 2026) and treat cache writes as 5-minute writes. Haiku 5.5 costs five times as much above a 100K-token prompt; the actual cost is priced per request, the baseline on each row's summed tokens at the lower rate. On a Pro or Max plan the dollar figures are API-equivalent; what you actually save is rate-limit headroom.
  • The baseline assumes the same token counts on the baseline model, and that an agent would have inherited its parent's model (a built-in agent with its own model, such as Explore, may not have). Different models use different token counts for the same task, so the "saved" figure is an estimate.
  • Model ids live in hooks/lib/pricing.ts. Update them and the price table when models change.

Develop

claude plugin validate .
claude plugin test .

Releases follow RELEASING.md; changes are listed in CHANGELOG.md.

Source 4 files
hooks/register.tsx 484 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { AgentRow, RouterLedger, Tier, Tokens } from '../types'
5import { ROLES } from './lib/agents'
6import {
7  CLASSIFIER_MODEL,
8  CLASSIFIER_SYSTEM,
9  EFFORT_CAP,
10  FORCED_TIER,
11  MODEL_FOR_TIER,
12  ORDER,
13  ZERO,
14  addTokens,
15  builtinClassifierText,
16  capEffort,
17  classifierPrompt,
18  costOf,
19  fallbackTier,
20  familyOf,
21  fullestWindow,
22  looksRisky,
23  parseMode,
24  parseRisky,
25  parseTier,
26  pressureThreshold,
27  roleOf,
28  routeTier,
29  tierOfModel,
30  tokensFromUsage,
31  usd,
32  type Family,
33  type Effort,
34  type Mode,
35} from './lib/pricing'
36
37const PANE = 'model-router'
38const EMPTY: RouterLedger = { rows: {}, classifierCost: 0, routed: 0 }
39const ledger = atom({ plugin: 'model-router', key: 'ledger' } as const, EMPTY)
40
41const TIER_TAG: Record<Tier, string> = { simple: 'S', standard: 'M', hard: 'H', long: 'L' }
42
43/** The classifier's budget. A spawn waits on it, so a slow answer costs more than a wrong tier. */
44const CLASSIFY_MS = 3000
45/** How long a subagent's first request waits for its spawn to be recorded, so it gets its effort. */
46const SPAWN_WAIT_MS = 2000
47const MODE_KEY = 'mode'
48
49type Baseline = 'unrouted' | Exclude<Family, 'haiku4' | 'haiku5'>
50
51function blankRow(id: string, label: string, agentType: string, now: number): AgentRow {
52  return { id, label, agentType, tokens: ZERO, cost: 0, startedAt: now }
53}
54
55function shortModel(model: string | undefined): string {
56  if (!model) return '?'
57  return { haiku4: 'Haiku4', haiku5: 'Haiku', sonnet: 'Sonnet', opus: 'Opus', fable: 'Fable' }[familyOf(model)]
58}
59
60/**
61 * The same tokens on the baseline. Priced on the row's summed tokens, so a fixed baseline
62 * or a later change of the setting reprices the whole bill.
63 */
64function baselineCost(r: AgentRow, baseline: Baseline): number {
65  const family = baseline === 'unrouted' ? familyOf(r.baselineModel ?? r.model) : baseline
66  return costOf(r.tokens, family)
67}
68
69function totals(l: RouterLedger, baseline: Baseline) {
70  let cost = l.classifierCost
71  let base = 0
72  for (const r of Object.values(l.rows)) {
73    cost += r.cost
74    base += baselineCost(r, baseline)
75  }
76  return { cost, base, saved: base > 0 ? 1 - cost / base : 0 }
77}
78
79/** Writes a spawned agent's row; its steps may have landed first, so it merges into them. */
80async function record($: EngineInterface, id: string, now: number, fields: Partial<AgentRow>, routed: boolean) {
81  await update($, ledger, l => {
82    const row = l.rows[id] ?? blankRow(id, 'subagent', '?', now)
83    return { ...l, routed: l.routed + (routed ? 1 : 0), rows: { ...l.rows, [id]: { ...row, ...fields } } }
84  })
85}
86
87/** Tier tag for the pane: `!` when the risk floor raised it, `↓` when rate limits lowered it. */
88function tierTag(r: AgentRow): string {
89  if (!r.tier) return '- '
90  return (TIER_TAG[r.tier] + (r.risky ? '!' : r.pressure !== undefined ? '↓' : '')).padEnd(2)
91}
92
93/** The fullest rate-limit window in percent; 0 when the engine has none or the call fails. */
94async function windowPercent($: EngineInterface): Promise<number> {
95  try {
96    return fullestWindow((await $.session.usage()).rateLimits)
97  } catch {
98    return 0
99  }
100}
101
102/** What a mode does, as `/router` and `/router status` say it. */
103function modeText(mode: Mode): string {
104  if (mode === 'on') return 'on: Haiku picks the model for each subagent'
105  if (mode === 'off') return 'off: subagents run on the model they would have without the router'
106  return `sending every routed subagent to ${shortModel(MODEL_FOR_TIER[FORCED_TIER[mode]])}`
107}
108
109function modeLabel(mode: Mode): string | undefined {
110  if (mode === 'on') return undefined
111  if (mode === 'off') return 'router off'
112  return `router: all ${shortModel(MODEL_FOR_TIER[FORCED_TIER[mode]])}`
113}
114
115async function refreshStatus(
116  $: EngineInterface,
117  baseline: Baseline,
118  baselineName: string,
119  mode: Mode,
120  pressureAt: number | undefined,
121) {
122  const t = totals(await read($, ledger), baseline)
123  const parts = [modeLabel(mode)]
124  if (mode === 'on' && pressureAt !== undefined) {
125    const pct = await windowPercent($)
126    if (pct >= pressureAt) parts.push(`limits ${Math.round(pct)}%: one tier down`)
127  }
128  if (t.base > 0) parts.push(`${usd(t.cost)} vs ${usd(t.base)} ${baselineName} (${Math.round(t.saved * 100)}% saved)`)
129  const text = parts.filter(Boolean).join(' · ')
130  $.ui.status(text || undefined)
131}
132
133/** Resolves undefined after `ms`, so a call raced against it can't hold a spawn up. */
134function timeout($: EngineInterface, ms: number): Promise<undefined> {
135  return $.clock.sleep(ms).then(() => undefined)
136}
137
138/** The `/router` mode saved by an earlier session, if any. */
139async function storedMode($: EngineInterface): Promise<Mode | undefined> {
140  try {
141    return parseMode(await $.store.get(MODE_KEY))
142  } catch {
143    return undefined
144  }
145}
146
147/**
148 * The effort cap of a routed subagent. Its first step can beat its spawn's answer, so it
149 * waits briefly for spawns in flight. The caps are kept in memory because a hook reads
150 * `$.state` as of the moment it started: a row written during the wait is not seen. The
151 * ledger is the fallback after a hot reload has emptied memory.
152 */
153async function effortCapOf(
154  $: EngineInterface,
155  agentId: string,
156  caps: ReadonlyMap<string, Effort | undefined>,
157  starting: ReadonlySet<Promise<unknown>>,
158) {
159  if (!caps.has(agentId) && starting.size > 0) {
160    await Promise.race([Promise.allSettled([...starting]), timeout($, SPAWN_WAIT_MS)])
161  }
162  return caps.has(agentId) ? caps.get(agentId) : (await read($, ledger)).rows[agentId]?.effort
163}
164
165export const register: Register = (on, options) => {
166  const baseline = (['unrouted', 'fable', 'opus', 'sonnet'].includes(String(options.baseline))
167    ? options.baseline
168    : 'unrouted') as Baseline
169  const routeAll = options.routeAll !== false
170  const respectExplicit = options.respectExplicitModel !== false
171  const baselineName = baseline === 'unrouted' ? 'unrouted' : `on ${shortModel(baseline)}`
172  const effortByTier = options.effortByTier !== false
173  const riskFloor = options.riskFloor !== false
174  const pressureAt = options.limitPressure === 'off' ? undefined : pressureThreshold(options.limitPressure)
175
176  // Module state starts over on a hot reload; the mode is read back from the store.
177  let mode: Mode = 'on'
178  let modeLoaded = false
179  /** Agent calls that set their own effort, by tool_use_id, until their spawn reads it. */
180  const callEffort = new Set<string>()
181  /** Spawns between the call to `next` and its answer, which a first step may overtake. */
182  const starting = new Set<Promise<unknown>>()
183  /** Each routed subagent's effort cap, by agent id, as soon as it has started. */
184  const effortCaps = new Map<string, Effort | undefined>()
185
186  on('session.start', async ($, e, next) => {
187    if (!modeLoaded) {
188      modeLoaded = true
189      mode = (await storedMode($)) ?? mode
190    }
191    if (mode !== 'on') $.ui.status(modeLabel(mode))
192    for (const role of ROLES) {
193      try {
194        await $.agent.register({ ...role, model: 'sonnet' })
195      } catch (err) {
196        $.ui.log(`could not add agent ${role.name}: ${String(err)}`)
197      }
198    }
199    try {
200      await $.command.register({
201        name: 'router',
202        description: 'Open the model router bill pane; on, off, haiku, sonnet or opus sets the mode; reset clears the bill',
203        argumentHint: '[on|off|haiku|sonnet|opus|status|reset]',
204        immediate: true,
205      })
206    } catch (err) {
207      $.ui.log(`could not add /router: ${String(err)}`)
208    }
209    // Opened unasked, so it only seats in a wide terminal; /router opens it anywhere.
210    void $.ui.open({ id: PANE, title: 'Model router' })
211    return next(e)
212  })
213
214  on('command.run', { command: 'router' }, async ($, e) => {
215    const arg = e.args.trim().toLowerCase()
216    if (!modeLoaded) {
217      modeLoaded = true
218      mode = (await storedMode($)) ?? mode
219    }
220    if (arg === 'reset') {
221      await update($, ledger, () => EMPTY)
222      $.ui.status(modeLabel(mode))
223      return { text: 'Model router bill cleared.' }
224    }
225    if (arg === 'status') return { text: `Model router ${modeText(mode)}.` }
226    const picked = parseMode(arg)
227    if (picked) {
228      mode = picked
229      modeLoaded = true
230      try {
231        await $.store.set(MODE_KEY, mode)
232      } catch (err) {
233        $.ui.log(`could not save the router mode: ${String(err)}`)
234      }
235      await refreshStatus($, baseline, baselineName, mode, pressureAt)
236      const back = mode === 'on' ? '' : '; /router on to go back'
237      return { text: `Model router ${modeText(mode)}. Kept across sessions${back}.` }
238    }
239    if (arg) return { text: `Unknown option "${arg}". Use on, off, haiku, sonnet, opus, status or reset.` }
240    await $.ui.open({ id: PANE, title: 'Model router', focus: true, closeOnEscape: true })
241    return {}
242  })
243
244  // An Agent call that sets its own effort was asked for that effort: remember it so
245  // the router leaves that agent's effort alone.
246  on('tool.call', { tool: 'Agent' }, async ($, e, next) => {
247    if (e.tool === 'Agent' && e.effort && e.tool_use_id) {
248      // Read and removed by the spawn; a call that never spawns leaves one id behind.
249      if (callEffort.size > 200) callEffort.clear()
250      callEffort.add(e.tool_use_id)
251    }
252    return next(e)
253  })
254
255  // 1. Haiku reads the task and sorts it. 2. The router sets the model.
256  on('agent.spawn', async ($, e, next) => {
257    if (!modeLoaded) {
258      modeLoaded = true
259      mode = (await storedMode($)) ?? mode
260    }
261    const callSetEffort = callEffort.delete(e.tool_use_id)
262    const role = roleOf(e.subagentType)
263    const now = await $.clock.now()
264
265    // Not routed: forks always inherit, a workflow agent's model cannot be rewritten,
266    // a teammate is long-lived and always background (so it would qualify for Fable on
267    // its first message alone), and with routing narrowed only the router's own agents
268    // are touched. `/router off` leaves everything alone.
269    if (mode === 'off' || e.fork || e.workflow || e.isTeammate || (!routeAll && !role)) {
270      const result = await next(e)
271      if (result.deny === undefined && result.agentId) {
272        await record($, result.agentId, now, {
273          label: e.description || e.name || e.subagentType,
274          agentType: e.subagentType,
275          via: 'unrouted',
276          model: result.model,
277          baselineModel: result.model,
278        }, false)
279      }
280      return result
281    }
282
283    let tier: Tier
284    let via: AgentRow['via']
285    let model: string | undefined
286    let risky = false
287    let pressure: number | undefined
288
289    if (respectExplicit && e.model) {
290      via = 'explicit'
291      model = e.model
292      tier = tierOfModel(e.model)
293    } else if (mode !== 'on') {
294      via = 'forced'
295      tier = FORCED_TIER[mode]
296      model = MODEL_FOR_TIER[tier]
297    } else {
298      let picked: Tier | undefined
299      try {
300        const r = await $.model.complete(
301          {
302            model: CLASSIFIER_MODEL,
303            system: CLASSIFIER_SYSTEM,
304            prompt: classifierPrompt({
305              agentType: e.subagentType,
306              description: e.description,
307              prompt: e.prompt,
308              background: e.background,
309            }),
310            maxTokens: 8,
311            timeoutMs: CLASSIFY_MS,
312          },
313          { signal: next.signal },
314        )
315        const c = costOf(tokensFromUsage(r.usage), familyOf(CLASSIFIER_MODEL))
316        await update($, ledger, l => ({ ...l, classifierCost: l.classifierCost + c }))
317        if (r.isAnswered) {
318          picked = parseTier(r.text)
319          risky = parseRisky(r.text)
320        }
321      } catch {
322        // try the built-in classifier
323      }
324      via = picked ? 'haiku' : undefined
325      if (!picked) {
326        // The engine's own small model, with no rubric: worse than Haiku 4.5 with one, better
327        // than a fixed default, and it keeps working if the pinned classifier model goes away.
328        try {
329          const label = await Promise.race([
330            $.model.classify(builtinClassifierText({
331              agentType: e.subagentType,
332              description: e.description,
333              prompt: e.prompt,
334            }), ORDER),
335            timeout($, CLASSIFY_MS),
336          ])
337          picked = parseTier(label ?? '')
338          if (picked) via = 'builtin'
339        } catch {
340          // fall through to the role's default
341        }
342      }
343      via ??= 'fallback'
344      // Without Haiku's judgment, only a plainly destructive act counts as risky.
345      if (via !== 'haiku') risky ||= looksRisky(`${e.description}\n${e.prompt}`)
346      risky &&= riskFloor
347      const asked = picked ?? fallbackTier(role)
348      tier = routeTier(asked, role, e.background, { risky })
349      if (pressureAt !== undefined) {
350        const pct = await windowPercent($)
351        const lower = pct >= pressureAt ? routeTier(asked, role, e.background, { risky, underPressure: true }) : tier
352        // Marked only when the window actually moved the agent down.
353        if (lower !== tier) {
354          tier = lower
355          pressure = pct
356        }
357      }
358      model = MODEL_FOR_TIER[tier]
359    }
360
361    // Effort is set per request in turn.step; here the router only decides the cap.
362    const effort =
363      effortByTier && via !== 'explicit' && via !== 'forced' && !callSetEffort
364        ? EFFORT_CAP[tier]
365        : undefined
366
367    // Held in `starting` until the cap is known: the agent's first request can arrive
368    // before `next` answers and must wait for its effort cap.
369    const spawned = (async () => {
370      const result = await next({ ...e, model })
371      if (result.deny !== undefined || !result.agentId) return result
372      if (effortCaps.size > 500) effortCaps.clear()
373      effortCaps.set(result.agentId, effort)
374      await record($, result.agentId, now, {
375        label: e.description,
376        agentType: e.subagentType,
377        tier,
378        via,
379        model: result.model,
380        effort,
381        risky: risky || undefined,
382        pressure,
383        // A model Claude named would have run anyway; otherwise the agent inherits its parent's.
384        baselineModel: via === 'explicit' ? result.model : e.parentModel,
385      }, true)
386      return result
387    })()
388    starting.add(spawned)
389    try {
390      return await spawned
391    } finally {
392      starting.delete(spawned)
393    }
394  }).catch(($, e, next) => next(e)) // routing is an optimisation: on failure, spawn as asked
395
396  // 3. Effort, lowered for easy tiers. 4. The bill: every model request, priced on the model that answered it.
397  on('turn.step', async function* ($, e, next) {
398    let request = e
399    if (effortByTier && e.agentId && e.effort !== undefined) {
400      try {
401        const effort = capEffort(e.effort, await effortCapOf($, e.agentId, effortCaps, starting))
402        if (effort !== e.effort) request = { ...e, effort }
403      } catch {
404        // leave the request as it is
405      }
406    }
407    const result = yield* next(request)
408    if (!result.usage) return result
409    const usage = result.usage
410    const t: Tokens = tokensFromUsage(usage)
411    const id = e.agentId ?? 'main'
412    const actual = costOf(t, familyOf(usage.model))
413    const now = await $.clock.now()
414    await update($, ledger, l => {
415      const row =
416        l.rows[id] ??
417        blankRow(id, id === 'main' ? 'main conversation' : 'subagent', id === 'main' ? 'main' : '?', now)
418      return {
419        ...l,
420        rows: {
421          ...l.rows,
422          [id]: {
423            ...row,
424            model: row.model ?? usage.model,
425            via: row.via ?? (id === 'main' ? 'unrouted' : row.via),
426            baselineModel: row.baselineModel ?? (id === 'main' ? usage.model : undefined),
427            tokens: addTokens(row.tokens, t),
428            cost: row.cost + actual,
429          },
430        },
431      }
432    })
433    return result
434  })
435
436  on('turn.complete', async ($, e, next) => {
437    await refreshStatus($, baseline, baselineName, mode, pressureAt)
438    return next(e)
439  })
440
441  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
442    const { Box, Text } = $.ui.resolve(e)
443    const l = await read($, ledger)
444    const t = totals(l, baseline)
445    const width = Math.max(30, (e.props.bodyColumns ?? 60) - 2)
446    const rows = Object.values(l.rows).sort((a, b) => b.startedAt - a.startedAt)
447    const room = Math.max(3, (e.viewport?.rows ?? 30) - 10)
448    const labelWidth = Math.max(8, width - 33)
449    const baseHead = baseline === 'unrouted' ? 'unrouted' : shortModel(baseline)
450
451    return (
452      <Box flexDirection="column" width={width}>
453        <Text bold>
454          {usd(t.cost)} spent · {usd(t.base)} {baselineName}
455        </Text>
456        <Text color={t.saved > 0 ? 'green' : undefined}>
457          {t.base > 0 ? `${Math.round(t.saved * 100)}% saved` : 'No model calls yet.'} · {l.routed} routed ·
458          classifier {usd(l.classifierCost)}
459        </Text>
460        {mode !== 'on' && <Text color="yellow">{modeLabel(mode)} (/router on to go back)</Text>}
461        <Text> </Text>
462        <Text dimColor wrap="truncate">
463          {'agent'.padEnd(labelWidth)} T  model   {'cost'.padStart(7)} {baseHead.padStart(8)}
464        </Text>
465        {rows.slice(0, room).map(r => (
466          <Text wrap="truncate" dimColor={r.via === 'unrouted'}>
467            {(r.label.length > labelWidth - 1 ? r.label.slice(0, labelWidth - 2) + '…' : r.label).padEnd(labelWidth)}
468            {tierTag(r)} {shortModel(r.model).padEnd(7)} {usd(r.cost).padStart(7)}{' '}
469            {usd(baselineCost(r, baseline)).padStart(8)}
470          </Text>
471        ))}
472        {rows.length > room && <Text dimColor>…{rows.length - room} more</Text>}
473        <Text> </Text>
474        <Text dimColor wrap="wrap">
475          T: S simple→Haiku 5.5 at low effort, M standard→Sonnet at medium effort at most, H hard→Opus,
476          L long→Fable (background only). ! risky, sent to Opus at least; ↓ one tier down for rate limits.
477          Dim rows are not routed. API list prices; plans billed by
478          subscription see rate-limit use, not dollars.
479        </Text>
480      </Box>
481    )
482  })
483}
484
hooks/lib/agents.ts 36 lines
1/** The four roles the router adds. Their model is decided per task at spawn time. */
2export const ROLES = [
3  {
4    name: 'architect',
5    description:
6      'Plans before code is written: reads the codebase, weighs designs, and returns a step-by-step plan with the files to touch. Use for new features, refactors and anything with design choices. Does not edit files.',
7    tools: ['Read', 'Grep', 'Glob', 'WebFetch', 'WebSearch'],
8    prompt: `You are the architect. Read the relevant code, then return a plan another agent can follow without re-deriving it.
9Output: the goal in one line; the files to change and why; ordered steps; risks and how to test the result.
10Do not edit files. Keep the plan as short as the task allows.`,
11  },
12  {
13    name: 'builder',
14    description:
15      'Writes and edits code to carry out a defined task or an architect plan. Use for implementing features, writing tests, and refactors with a clear target.',
16    prompt: `You are the builder. Make the change you are given, following the existing style of the codebase.
17Keep the diff focused. When done, list the files you changed and anything you could not finish.`,
18  },
19  {
20    name: 'runner',
21    description:
22      'Runs commands and reports results: tests, builds, linters, type checks, searches. Use when the job is to execute and summarize, not to change code.',
23    tools: ['Bash', 'Read', 'Grep', 'Glob'],
24    prompt: `You are the runner. Run what you are asked to run and report the result.
25Report: the command, pass or fail, and for each failure the test or file, the error, and the line. Quote errors exactly; do not paraphrase.
26Do not change code.`,
27  },
28  {
29    name: 'fixer',
30    description:
31      'Fixes what failed: takes a failing test, build error or bug report, finds the root cause, and changes the code until it passes. Use after a runner reports failures.',
32    prompt: `You are the fixer. Find the root cause of the failure you are given before changing anything.
33Fix the cause, not the symptom. Re-run the failing check to confirm. Report the cause in one or two lines and the files you changed.`,
34  },
35] as const
36
hooks/lib/pricing.ts 268 lines
1import type { Tier, Tokens } from '../../types'
2
3/** USD per million tokens, from platform.claude.com/docs/en/about-claude/pricing (Oct 2026). */
4export type Price = { input: number; output: number; cacheWrite: number; cacheRead: number }
5
6export const PRICES: Record<'haiku4' | 'haiku5' | 'sonnet' | 'opus' | 'fable', Price> = {
7  haiku4: { input: 1, output: 5, cacheWrite: 1.25, cacheRead: 0.1 },
8  // Prompts up to 100,000 tokens; HAIKU5_LONG above that.
9  haiku5: { input: 0.1, output: 0.5, cacheWrite: 0.125, cacheRead: 0.01 },
10  sonnet: { input: 2, output: 10, cacheWrite: 2.5, cacheRead: 0.1 },
11  opus: { input: 4, output: 20, cacheWrite: 5, cacheRead: 0.2 },
12  fable: { input: 10, output: 50, cacheWrite: 12.5, cacheRead: 0.25 },
13}
14
15/** Haiku 5.5 for a prompt (input, cache reads and cache writes together) over 100,000 tokens. */
16export const HAIKU5_LONG: Price = { input: 0.5, output: 2.5, cacheWrite: 0.625, cacheRead: 0.05 }
17export const HAIKU5_LONG_FROM = 100_000
18
19export type Family = keyof typeof PRICES
20
21/** Full model ids the router spawns on. */
22export const MODEL_FOR_TIER: Record<Tier, string> = {
23  simple: 'claude-haiku-5-5',
24  standard: 'claude-sonnet-5-5',
25  hard: 'claude-opus-5-5',
26  long: 'claude-fable-5-1',
27}
28
29/**
30 * The classifier stays on Haiku 4.5: it answers without thinking, so an 8-token cap holds.
31 * Haiku 5.5 thinks by default and would often come back empty at that cap.
32 */
33export const CLASSIFIER_MODEL = 'claude-haiku-4-5-20251001'
34
35export const ORDER: Tier[] = ['simple', 'standard', 'hard', 'long']
36
37export type Effort = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
38const EFFORTS: Effort[] = ['low', 'medium', 'high', 'xhigh', 'max']
39
40/**
41 * The most reasoning effort a routed agent's requests may ask for, by tier. Easy work
42 * thinks less; hard and long work keep whatever the session would have used.
43 */
44export const EFFORT_CAP: Record<Tier, Effort | undefined> = {
45  simple: 'low',
46  standard: 'medium',
47  hard: undefined,
48  long: undefined,
49}
50
51/**
52 * The lower of a request's effort and the cap. An integer budget or a level this
53 * table does not know is left alone, as is a request without effort.
54 */
55export function capEffort<E>(effort: E, cap: Effort | undefined): E | Effort {
56  if (!cap || typeof effort !== 'string') return effort
57  const i = EFFORTS.indexOf(effort as Effort)
58  return i < 0 || i <= EFFORTS.indexOf(cap) ? effort : cap
59}
60
61/** `/router` modes: `on` classifies, `off` leaves every spawn alone, a model name sends every routed spawn there. */
62export type Mode = 'on' | 'off' | 'haiku' | 'sonnet' | 'opus'
63export const MODES: Mode[] = ['on', 'off', 'haiku', 'sonnet', 'opus']
64
65export function parseMode(text: unknown): Mode | undefined {
66  const m = String(text ?? '').trim().toLowerCase()
67  return (MODES as string[]).includes(m) ? (m as Mode) : undefined
68}
69
70export const FORCED_TIER: Record<Exclude<Mode, 'on' | 'off'>, Tier> = {
71  haiku: 'simple',
72  sonnet: 'standard',
73  opus: 'hard',
74}
75
76/**
77 * Which price row a model id or alias falls in. Mythos bills as Fable.
78 * A bare `haiku` alias is the current Haiku (5.5). Unknown ids price as Opus.
79 */
80export function familyOf(model: string | undefined): Family {
81  const m = (model ?? '').toLowerCase()
82  if (m.includes('haiku')) return /haiku-[34]/.test(m) ? 'haiku4' : 'haiku5'
83  if (m.includes('sonnet')) return 'sonnet'
84  if (m.includes('fable') || m.includes('mythos')) return 'fable'
85  return 'opus'
86}
87
88const FAMILY_TIER: Record<Family, Tier> = {
89  haiku4: 'simple',
90  haiku5: 'simple',
91  sonnet: 'standard',
92  opus: 'hard',
93  fable: 'long',
94}
95
96/** The tier a model belongs to, for a model Claude named itself. */
97export function tierOfModel(model: string): Tier {
98  return FAMILY_TIER[familyOf(model)]
99}
100
101/**
102 * Cost of one request's tokens on a family. Haiku 5.5 switches price above a 100K-token prompt,
103 * so call it per request; on summed tokens it prices everything at the lower rate.
104 * Cache writes are priced as 5-minute writes; 1-hour writes cost more, so a session that uses
105 * them is under-counted here.
106 */
107export function costOf(t: Tokens, family: Family): number {
108  const prompt = t.input + t.cacheRead + t.cacheWrite
109  const p = family === 'haiku5' && prompt > HAIKU5_LONG_FROM ? HAIKU5_LONG : PRICES[family]
110  return (t.input * p.input + t.output * p.output + t.cacheWrite * p.cacheWrite + t.cacheRead * p.cacheRead) / 1e6
111}
112
113export function tokensFromUsage(u: {
114  input_tokens: number
115  output_tokens: number
116  cache_read_input_tokens: number
117  cache_creation_input_tokens: number
118}): Tokens {
119  return {
120    input: u.input_tokens ?? 0,
121    output: u.output_tokens ?? 0,
122    cacheRead: u.cache_read_input_tokens ?? 0,
123    cacheWrite: u.cache_creation_input_tokens ?? 0,
124  }
125}
126
127export function addTokens(a: Tokens, b: Tokens): Tokens {
128  return {
129    input: a.input + b.input,
130    output: a.output + b.output,
131    cacheRead: a.cacheRead + b.cacheRead,
132    cacheWrite: a.cacheWrite + b.cacheWrite,
133  }
134}
135
136export const ZERO: Tokens = { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }
137
138/**
139 * Role guard rails, applied after Haiku's answer.
140 * runner never goes above standard (running tests is not Opus work).
141 * architect and fixer never go below standard (a bad plan or a bad fix costs more than the tokens).
142 */
143export const ROLE_LIMITS: Record<string, { min?: Tier; max?: Tier; fallback: Tier }> = {
144  architect: { min: 'standard', fallback: 'hard' },
145  builder: { fallback: 'standard' },
146  runner: { max: 'standard', fallback: 'simple' },
147  fixer: { min: 'standard', fallback: 'standard' },
148}
149
150export function roleOf(agentType: string): string | undefined {
151  const m = /^model-router:(\w+)$/.exec(agentType)
152  return m ? m[1] : undefined
153}
154
155/**
156 * The role's limits, then `long` only for background work: a foreground task the
157 * parent waits on goes to Opus, never to Fable at 2.5 times the price.
158 */
159export function clampTier(tier: Tier, role: string | undefined, background = true): Tier {
160  let i = ORDER.indexOf(tier)
161  const lim = role ? ROLE_LIMITS[role] : undefined
162  if (lim?.min) i = Math.max(i, ORDER.indexOf(lim.min))
163  if (lim?.max) i = Math.min(i, ORDER.indexOf(lim.max))
164  if (!background) i = Math.min(i, ORDER.indexOf('hard'))
165  return ORDER[i]!
166}
167
168/**
169 * The tier after the risk floor and rate-limit pressure.
170 *
171 * Under pressure (a rate-limit window past the threshold) the tier drops one step and
172 * `long` is off the table; the role's limits still hold. A risky task then goes to at
173 * least `hard`, past a role's cap too: getting a destructive act wrong costs more than
174 * the tokens, so the risk floor wins over pressure.
175 */
176export function routeTier(
177  tier: Tier,
178  role: string | undefined,
179  background: boolean,
180  o: { risky?: boolean; underPressure?: boolean } = {},
181): Tier {
182  const i = ORDER.indexOf(tier)
183  let t = clampTier(o.underPressure ? ORDER[Math.max(0, i - 1)]! : tier, role, background && !o.underPressure)
184  if (o.risky && ORDER.indexOf(t) < ORDER.indexOf('hard')) t = 'hard'
185  return t
186}
187
188/**
189 * Destructive acts, for when no classifier answered. Narrow on purpose: it names the act
190 * (deploying to production, dropping a table, a forced push), not the subject, so code
191 * that merely deals with payments or databases does not match.
192 */
193const RISKY_ACT =
194  /\b(deploy|ship|release|roll\s?out)\w*\b[^.\n]{0,40}\b(to|on|in)\s+(prod|production|live)\b|\bdrop\s+(table|database|schema)\b|\btruncate\s+table\b|\brm\s+-rf\s+\/|\bgit\s+push\s+(-f|--force)\b|\bforce[- ]push\b|\b(delete|wipe|purge)\b[^.\n]{0,30}\b(prod|production)\s+(data|database|db|bucket|users?)\b/i
195
196export function looksRisky(text: string): boolean {
197  return RISKY_ACT.test(text)
198}
199
200/** Reads the classifier's risk flag: the word `risky` after the tier. */
201export function parseRisky(text: string): boolean {
202  return /\brisky\b/i.test(text)
203}
204
205/** The fullest rate-limit window, in percent; 0 off a subscription or before the first reading. */
206export function fullestWindow(limits: readonly { percentUsed: number }[] | undefined): number {
207  return (limits ?? []).reduce((m, l) => Math.max(m, l.percentUsed), 0)
208}
209
210/** The `limitPressure` setting as a percentage, or undefined when off. */
211export function pressureThreshold(setting: unknown): number | undefined {
212  const n = Number(setting ?? 80)
213  return Number.isFinite(n) && n > 0 && n <= 100 ? n : undefined
214}
215
216export function fallbackTier(role: string | undefined): Tier {
217  return (role && ROLE_LIMITS[role]?.fallback) || 'standard'
218}
219
220/** Reads the classifier's reply: the first tier word it names. */
221export function parseTier(text: string): Tier | undefined {
222  const m = /\b(simple|standard|hard|long)\b/i.exec(text)
223  return m ? (m[1]!.toLowerCase() as Tier) : undefined
224}
225
226export const CLASSIFIER_SYSTEM = `You route coding subagent tasks to a model. Read the task and reply with exactly one word.
227
228simple   - mechanical, little judgment: rename, find/search/grep, list files, run an existing test or build command and report, format, small single-file edits with an obvious answer.
229standard - normal feature work: implement a function or component, write tests, a refactor within a few files, a bug with a clear repro.
230hard     - needs deep reasoning: architecture or design decisions, an intermittent or cross-cutting bug, concurrency, security, performance work, a migration with subtle risk.
231long     - a multi-hour autonomous run: a large multi-step build or migration across many files, explicitly long-running or background work that must keep going unattended.
232
233Pick the cheapest tier that will succeed. If unsure between two, pick the higher.
234
235Then add the word risky if carrying out the task itself could do costly or hard-to-reverse harm: deploying to production, running a migration or a destructive command against production or shared data, deleting data, rotating or exposing credentials, moving money, force-pushing over shared history. Judge the act, not the subject: writing, refactoring or testing code that deals with payments, databases or credentials is not risky.
236
237Reply with the tier word alone, or the tier word and risky: for example "standard" or "hard risky".`
238
239/** What the engine's built-in classifier reads when Haiku 4.5 gives no tier: no rubric, so keep it short. */
240export function builtinClassifierText(input: { agentType: string; description: string; prompt: string }): string {
241  const body = input.prompt.length > 2000 ? input.prompt.slice(0, 2000) + ' [...]' : input.prompt
242  return `How hard is this coding task for an AI agent? Agent: ${input.agentType}. ${input.description}. ${body}`
243}
244
245export function classifierPrompt(input: {
246  agentType: string
247  description: string
248  prompt: string
249  background: boolean
250}): string {
251  const body = input.prompt.length > 6000 ? input.prompt.slice(0, 6000) + '\n[...truncated]' : input.prompt
252  return [
253    `Agent type: ${input.agentType}`,
254    `Runs in background: ${input.background ? 'yes' : 'no'}`,
255    `Short description: ${input.description}`,
256    'Task:',
257    '<task>',
258    body,
259    '</task>',
260  ].join('\n')
261}
262
263export function usd(n: number): string {
264  if (n === 0) return '$0'
265  if (n < 0.01) return '<$0.01'
266  return '$' + n.toFixed(2)
267}
268
types/index.d.ts 52 lines
1export type Tier = 'simple' | 'standard' | 'hard' | 'long'
2
3export type Effort = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
4
5export type Tokens = {
6  input: number
7  output: number
8  cacheRead: number
9  cacheWrite: number
10}
11
12/** One subagent (or the main loop, id "main") as the bill pane shows it. */
13export type AgentRow = {
14  id: string
15  label: string
16  agentType: string
17  tier?: Tier
18  /**
19   * How the model was decided: the Haiku classifier, the engine's built-in classifier
20   * when Haiku gave no tier, the role fallback, a model Claude named, a `/router` mode
21   * that sends everything to one model, or not routed (main, forks, workflow agents,
22   * teammates, built-ins when routing is narrowed, everything under `/router off`).
23   */
24  via?: 'haiku' | 'builtin' | 'fallback' | 'explicit' | 'forced' | 'unrouted'
25  model?: string
26  /** True when the task was judged risky and sent to at least the hard tier. */
27  risky?: boolean
28  /** The fullest rate-limit window, in percent, when it pushed this agent a tier down. */
29  pressure?: number
30  /** The most effort this agent's requests may ask for; absent when the router leaves effort alone. */
31  effort?: Effort
32  /** What this agent would have run on without the router: the parent's model, or its own when unrouted. */
33  baselineModel?: string
34  tokens: Tokens
35  /** USD at the price of the model that actually answered each step. */
36  cost: number
37  startedAt: number
38}
39
40export type RouterLedger = {
41  rows: Record<string, AgentRow>
42  /** What the Haiku classifier calls themselves cost. */
43  classifierCost: number
44  routed: number
45}
46
47declare module 'claude-code' {
48  interface PluginState {
49    'model-router': { ledger: RouterLedger }
50  }
51}
52