Haiku sorts every subagent task into simple, standard, hard or long run and picks the model; a pane shows what each agent cost.

A Claude Code mod that picks the model for each subagent before it starts, and shows what each one cost.
model-router:architect (plans, no edits), model-router:builder (writes code), model-router:runner (runs tests and commands, no edits) and model-router:fixer (fixes what failed).simple, standard, hard or long. The classifier stays on Haiku 4.5 because it answers without thinking, so an 8-token reply cap holds. It gets 3 seconds; if it fails or names no tier, the engine's built-in classifier gets another 3 seconds before the role default is used.| Tier | Model | Effort | Typical work |
|---|---|---|---|
| simple | Haiku 5.5 | low | rename, search, run tests |
| standard | Sonnet 5.5 | medium at most | normal feature work |
| hard | Opus 5.5 | unchanged | hard bugs, architecture |
| long | Fable 5.1 | unchanged | multi-hour background runs |
Effort is only ever lowered, never raised, and only for agents the router picked a model for. An Agent call that sets its own effort keeps it.
Risk floor. Haiku also flags a task as risky when carrying it out could do costly or hard-to-reverse harm: deploying to production, deleting data, a migration or destructive command on shared data, a forced push. It judges the act, not the subject, so writing code that deals with payments or databases is not risky. A risky task runs on Opus at least, past the runner's cap and past rate-limit pressure. If no classifier answers, a narrow pattern check on destructive acts stands in. The pane marks these rows !.
Rate limits. When the fullest rate-limit window (5-hour or weekly) is at 80% or more (setting), every routed task goes one tier lower and never to Fable. Role floors still hold. The status line shows the window, and the pane marks these rows ↓. Off a subscription there are no windows, so nothing changes.
Guard rails: runner never goes above Sonnet; architect and fixer never go below Sonnet; long (Fable, 2.5 times Opus) only for background agents, so a foreground task the parent waits on tops out at Opus. If neither classifier gives a tier, the role default is used (architect: Opus, builder and fixer: Sonnet, runner: Haiku). Forks, workflow agents and agent-team teammates are not routed (the engine ignores a model change for the first two; a teammate is long-lived, so one classification of its first message is not a good guide). A model Claude already named in the Agent call is kept (setting), and the pane shows it at that model's tier.
/router off leaves every subagent on the model it would have had; /router haiku, /router sonnet or /router opus sends every routed subagent to that model without asking the classifier (a model Claude named is still kept, per the setting below); /router on goes back to classifying. The mode is kept across sessions and shows in the status line; /router status names it./router opens a pane: each agent, its tier, the model that ran it, what it cost, and what the same tokens would have cost without the router. The main conversation and unrouted agents are dim rows, and the Haiku classifier calls are counted in the total. A status line under the prompt shows the running total. /router reset clears it./config)unrouted (default) prices each routed agent on its parent's model, the one it would have inherited, and every unrouted row on its own model, so only routing shows as saved. fable, opus or sonnet price everything, the main conversation included, on that model.80 (default), 70, 90 or off; the percentage of the fullest rate-limit window at which routing goes one tier lower./plugin install model-router --marketplace aott33/model-router
Or for development: claude --plugin-dir ./model-router
Needs Claude Code 2.1.287 or later (mods).
hooks/lib/pricing.ts. Update them and the price table when models change.claude plugin validate .
claude plugin test .
Releases follow RELEASING.md; changes are listed in CHANGELOG.md.
hooks/register.tsx 484 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { AgentRow, RouterLedger, Tier, Tokens } from '../types'
5import { ROLES } from './lib/agents'
6import {
7 CLASSIFIER_MODEL,
8 CLASSIFIER_SYSTEM,
9 EFFORT_CAP,
10 FORCED_TIER,
11 MODEL_FOR_TIER,
12 ORDER,
13 ZERO,
14 addTokens,
15 builtinClassifierText,
16 capEffort,
17 classifierPrompt,
18 costOf,
19 fallbackTier,
20 familyOf,
21 fullestWindow,
22 looksRisky,
23 parseMode,
24 parseRisky,
25 parseTier,
26 pressureThreshold,
27 roleOf,
28 routeTier,
29 tierOfModel,
30 tokensFromUsage,
31 usd,
32 type Family,
33 type Effort,
34 type Mode,
35} from './lib/pricing'
36
37const PANE = 'model-router'
38const EMPTY: RouterLedger = { rows: {}, classifierCost: 0, routed: 0 }
39const ledger = atom({ plugin: 'model-router', key: 'ledger' } as const, EMPTY)
40
41const TIER_TAG: Record<Tier, string> = { simple: 'S', standard: 'M', hard: 'H', long: 'L' }
42
43/** The classifier's budget. A spawn waits on it, so a slow answer costs more than a wrong tier. */
44const CLASSIFY_MS = 3000
45/** How long a subagent's first request waits for its spawn to be recorded, so it gets its effort. */
46const SPAWN_WAIT_MS = 2000
47const MODE_KEY = 'mode'
48
49type Baseline = 'unrouted' | Exclude<Family, 'haiku4' | 'haiku5'>
50
51function blankRow(id: string, label: string, agentType: string, now: number): AgentRow {
52 return { id, label, agentType, tokens: ZERO, cost: 0, startedAt: now }
53}
54
55function shortModel(model: string | undefined): string {
56 if (!model) return '?'
57 return { haiku4: 'Haiku4', haiku5: 'Haiku', sonnet: 'Sonnet', opus: 'Opus', fable: 'Fable' }[familyOf(model)]
58}
59
60/**
61 * The same tokens on the baseline. Priced on the row's summed tokens, so a fixed baseline
62 * or a later change of the setting reprices the whole bill.
63 */
64function baselineCost(r: AgentRow, baseline: Baseline): number {
65 const family = baseline === 'unrouted' ? familyOf(r.baselineModel ?? r.model) : baseline
66 return costOf(r.tokens, family)
67}
68
69function totals(l: RouterLedger, baseline: Baseline) {
70 let cost = l.classifierCost
71 let base = 0
72 for (const r of Object.values(l.rows)) {
73 cost += r.cost
74 base += baselineCost(r, baseline)
75 }
76 return { cost, base, saved: base > 0 ? 1 - cost / base : 0 }
77}
78
79/** Writes a spawned agent's row; its steps may have landed first, so it merges into them. */
80async function record($: EngineInterface, id: string, now: number, fields: Partial<AgentRow>, routed: boolean) {
81 await update($, ledger, l => {
82 const row = l.rows[id] ?? blankRow(id, 'subagent', '?', now)
83 return { ...l, routed: l.routed + (routed ? 1 : 0), rows: { ...l.rows, [id]: { ...row, ...fields } } }
84 })
85}
86
87/** Tier tag for the pane: `!` when the risk floor raised it, `↓` when rate limits lowered it. */
88function tierTag(r: AgentRow): string {
89 if (!r.tier) return '- '
90 return (TIER_TAG[r.tier] + (r.risky ? '!' : r.pressure !== undefined ? '↓' : '')).padEnd(2)
91}
92
93/** The fullest rate-limit window in percent; 0 when the engine has none or the call fails. */
94async function windowPercent($: EngineInterface): Promise<number> {
95 try {
96 return fullestWindow((await $.session.usage()).rateLimits)
97 } catch {
98 return 0
99 }
100}
101
102/** What a mode does, as `/router` and `/router status` say it. */
103function modeText(mode: Mode): string {
104 if (mode === 'on') return 'on: Haiku picks the model for each subagent'
105 if (mode === 'off') return 'off: subagents run on the model they would have without the router'
106 return `sending every routed subagent to ${shortModel(MODEL_FOR_TIER[FORCED_TIER[mode]])}`
107}
108
109function modeLabel(mode: Mode): string | undefined {
110 if (mode === 'on') return undefined
111 if (mode === 'off') return 'router off'
112 return `router: all ${shortModel(MODEL_FOR_TIER[FORCED_TIER[mode]])}`
113}
114
115async function refreshStatus(
116 $: EngineInterface,
117 baseline: Baseline,
118 baselineName: string,
119 mode: Mode,
120 pressureAt: number | undefined,
121) {
122 const t = totals(await read($, ledger), baseline)
123 const parts = [modeLabel(mode)]
124 if (mode === 'on' && pressureAt !== undefined) {
125 const pct = await windowPercent($)
126 if (pct >= pressureAt) parts.push(`limits ${Math.round(pct)}%: one tier down`)
127 }
128 if (t.base > 0) parts.push(`${usd(t.cost)} vs ${usd(t.base)} ${baselineName} (${Math.round(t.saved * 100)}% saved)`)
129 const text = parts.filter(Boolean).join(' · ')
130 $.ui.status(text || undefined)
131}
132
133/** Resolves undefined after `ms`, so a call raced against it can't hold a spawn up. */
134function timeout($: EngineInterface, ms: number): Promise<undefined> {
135 return $.clock.sleep(ms).then(() => undefined)
136}
137
138/** The `/router` mode saved by an earlier session, if any. */
139async function storedMode($: EngineInterface): Promise<Mode | undefined> {
140 try {
141 return parseMode(await $.store.get(MODE_KEY))
142 } catch {
143 return undefined
144 }
145}
146
147/**
148 * The effort cap of a routed subagent. Its first step can beat its spawn's answer, so it
149 * waits briefly for spawns in flight. The caps are kept in memory because a hook reads
150 * `$.state` as of the moment it started: a row written during the wait is not seen. The
151 * ledger is the fallback after a hot reload has emptied memory.
152 */
153async function effortCapOf(
154 $: EngineInterface,
155 agentId: string,
156 caps: ReadonlyMap<string, Effort | undefined>,
157 starting: ReadonlySet<Promise<unknown>>,
158) {
159 if (!caps.has(agentId) && starting.size > 0) {
160 await Promise.race([Promise.allSettled([...starting]), timeout($, SPAWN_WAIT_MS)])
161 }
162 return caps.has(agentId) ? caps.get(agentId) : (await read($, ledger)).rows[agentId]?.effort
163}
164
165export const register: Register = (on, options) => {
166 const baseline = (['unrouted', 'fable', 'opus', 'sonnet'].includes(String(options.baseline))
167 ? options.baseline
168 : 'unrouted') as Baseline
169 const routeAll = options.routeAll !== false
170 const respectExplicit = options.respectExplicitModel !== false
171 const baselineName = baseline === 'unrouted' ? 'unrouted' : `on ${shortModel(baseline)}`
172 const effortByTier = options.effortByTier !== false
173 const riskFloor = options.riskFloor !== false
174 const pressureAt = options.limitPressure === 'off' ? undefined : pressureThreshold(options.limitPressure)
175
176 // Module state starts over on a hot reload; the mode is read back from the store.
177 let mode: Mode = 'on'
178 let modeLoaded = false
179 /** Agent calls that set their own effort, by tool_use_id, until their spawn reads it. */
180 const callEffort = new Set<string>()
181 /** Spawns between the call to `next` and its answer, which a first step may overtake. */
182 const starting = new Set<Promise<unknown>>()
183 /** Each routed subagent's effort cap, by agent id, as soon as it has started. */
184 const effortCaps = new Map<string, Effort | undefined>()
185
186 on('session.start', async ($, e, next) => {
187 if (!modeLoaded) {
188 modeLoaded = true
189 mode = (await storedMode($)) ?? mode
190 }
191 if (mode !== 'on') $.ui.status(modeLabel(mode))
192 for (const role of ROLES) {
193 try {
194 await $.agent.register({ ...role, model: 'sonnet' })
195 } catch (err) {
196 $.ui.log(`could not add agent ${role.name}: ${String(err)}`)
197 }
198 }
199 try {
200 await $.command.register({
201 name: 'router',
202 description: 'Open the model router bill pane; on, off, haiku, sonnet or opus sets the mode; reset clears the bill',
203 argumentHint: '[on|off|haiku|sonnet|opus|status|reset]',
204 immediate: true,
205 })
206 } catch (err) {
207 $.ui.log(`could not add /router: ${String(err)}`)
208 }
209 // Opened unasked, so it only seats in a wide terminal; /router opens it anywhere.
210 void $.ui.open({ id: PANE, title: 'Model router' })
211 return next(e)
212 })
213
214 on('command.run', { command: 'router' }, async ($, e) => {
215 const arg = e.args.trim().toLowerCase()
216 if (!modeLoaded) {
217 modeLoaded = true
218 mode = (await storedMode($)) ?? mode
219 }
220 if (arg === 'reset') {
221 await update($, ledger, () => EMPTY)
222 $.ui.status(modeLabel(mode))
223 return { text: 'Model router bill cleared.' }
224 }
225 if (arg === 'status') return { text: `Model router ${modeText(mode)}.` }
226 const picked = parseMode(arg)
227 if (picked) {
228 mode = picked
229 modeLoaded = true
230 try {
231 await $.store.set(MODE_KEY, mode)
232 } catch (err) {
233 $.ui.log(`could not save the router mode: ${String(err)}`)
234 }
235 await refreshStatus($, baseline, baselineName, mode, pressureAt)
236 const back = mode === 'on' ? '' : '; /router on to go back'
237 return { text: `Model router ${modeText(mode)}. Kept across sessions${back}.` }
238 }
239 if (arg) return { text: `Unknown option "${arg}". Use on, off, haiku, sonnet, opus, status or reset.` }
240 await $.ui.open({ id: PANE, title: 'Model router', focus: true, closeOnEscape: true })
241 return {}
242 })
243
244 // An Agent call that sets its own effort was asked for that effort: remember it so
245 // the router leaves that agent's effort alone.
246 on('tool.call', { tool: 'Agent' }, async ($, e, next) => {
247 if (e.tool === 'Agent' && e.effort && e.tool_use_id) {
248 // Read and removed by the spawn; a call that never spawns leaves one id behind.
249 if (callEffort.size > 200) callEffort.clear()
250 callEffort.add(e.tool_use_id)
251 }
252 return next(e)
253 })
254
255 // 1. Haiku reads the task and sorts it. 2. The router sets the model.
256 on('agent.spawn', async ($, e, next) => {
257 if (!modeLoaded) {
258 modeLoaded = true
259 mode = (await storedMode($)) ?? mode
260 }
261 const callSetEffort = callEffort.delete(e.tool_use_id)
262 const role = roleOf(e.subagentType)
263 const now = await $.clock.now()
264
265 // Not routed: forks always inherit, a workflow agent's model cannot be rewritten,
266 // a teammate is long-lived and always background (so it would qualify for Fable on
267 // its first message alone), and with routing narrowed only the router's own agents
268 // are touched. `/router off` leaves everything alone.
269 if (mode === 'off' || e.fork || e.workflow || e.isTeammate || (!routeAll && !role)) {
270 const result = await next(e)
271 if (result.deny === undefined && result.agentId) {
272 await record($, result.agentId, now, {
273 label: e.description || e.name || e.subagentType,
274 agentType: e.subagentType,
275 via: 'unrouted',
276 model: result.model,
277 baselineModel: result.model,
278 }, false)
279 }
280 return result
281 }
282
283 let tier: Tier
284 let via: AgentRow['via']
285 let model: string | undefined
286 let risky = false
287 let pressure: number | undefined
288
289 if (respectExplicit && e.model) {
290 via = 'explicit'
291 model = e.model
292 tier = tierOfModel(e.model)
293 } else if (mode !== 'on') {
294 via = 'forced'
295 tier = FORCED_TIER[mode]
296 model = MODEL_FOR_TIER[tier]
297 } else {
298 let picked: Tier | undefined
299 try {
300 const r = await $.model.complete(
301 {
302 model: CLASSIFIER_MODEL,
303 system: CLASSIFIER_SYSTEM,
304 prompt: classifierPrompt({
305 agentType: e.subagentType,
306 description: e.description,
307 prompt: e.prompt,
308 background: e.background,
309 }),
310 maxTokens: 8,
311 timeoutMs: CLASSIFY_MS,
312 },
313 { signal: next.signal },
314 )
315 const c = costOf(tokensFromUsage(r.usage), familyOf(CLASSIFIER_MODEL))
316 await update($, ledger, l => ({ ...l, classifierCost: l.classifierCost + c }))
317 if (r.isAnswered) {
318 picked = parseTier(r.text)
319 risky = parseRisky(r.text)
320 }
321 } catch {
322 // try the built-in classifier
323 }
324 via = picked ? 'haiku' : undefined
325 if (!picked) {
326 // The engine's own small model, with no rubric: worse than Haiku 4.5 with one, better
327 // than a fixed default, and it keeps working if the pinned classifier model goes away.
328 try {
329 const label = await Promise.race([
330 $.model.classify(builtinClassifierText({
331 agentType: e.subagentType,
332 description: e.description,
333 prompt: e.prompt,
334 }), ORDER),
335 timeout($, CLASSIFY_MS),
336 ])
337 picked = parseTier(label ?? '')
338 if (picked) via = 'builtin'
339 } catch {
340 // fall through to the role's default
341 }
342 }
343 via ??= 'fallback'
344 // Without Haiku's judgment, only a plainly destructive act counts as risky.
345 if (via !== 'haiku') risky ||= looksRisky(`${e.description}\n${e.prompt}`)
346 risky &&= riskFloor
347 const asked = picked ?? fallbackTier(role)
348 tier = routeTier(asked, role, e.background, { risky })
349 if (pressureAt !== undefined) {
350 const pct = await windowPercent($)
351 const lower = pct >= pressureAt ? routeTier(asked, role, e.background, { risky, underPressure: true }) : tier
352 // Marked only when the window actually moved the agent down.
353 if (lower !== tier) {
354 tier = lower
355 pressure = pct
356 }
357 }
358 model = MODEL_FOR_TIER[tier]
359 }
360
361 // Effort is set per request in turn.step; here the router only decides the cap.
362 const effort =
363 effortByTier && via !== 'explicit' && via !== 'forced' && !callSetEffort
364 ? EFFORT_CAP[tier]
365 : undefined
366
367 // Held in `starting` until the cap is known: the agent's first request can arrive
368 // before `next` answers and must wait for its effort cap.
369 const spawned = (async () => {
370 const result = await next({ ...e, model })
371 if (result.deny !== undefined || !result.agentId) return result
372 if (effortCaps.size > 500) effortCaps.clear()
373 effortCaps.set(result.agentId, effort)
374 await record($, result.agentId, now, {
375 label: e.description,
376 agentType: e.subagentType,
377 tier,
378 via,
379 model: result.model,
380 effort,
381 risky: risky || undefined,
382 pressure,
383 // A model Claude named would have run anyway; otherwise the agent inherits its parent's.
384 baselineModel: via === 'explicit' ? result.model : e.parentModel,
385 }, true)
386 return result
387 })()
388 starting.add(spawned)
389 try {
390 return await spawned
391 } finally {
392 starting.delete(spawned)
393 }
394 }).catch(($, e, next) => next(e)) // routing is an optimisation: on failure, spawn as asked
395
396 // 3. Effort, lowered for easy tiers. 4. The bill: every model request, priced on the model that answered it.
397 on('turn.step', async function* ($, e, next) {
398 let request = e
399 if (effortByTier && e.agentId && e.effort !== undefined) {
400 try {
401 const effort = capEffort(e.effort, await effortCapOf($, e.agentId, effortCaps, starting))
402 if (effort !== e.effort) request = { ...e, effort }
403 } catch {
404 // leave the request as it is
405 }
406 }
407 const result = yield* next(request)
408 if (!result.usage) return result
409 const usage = result.usage
410 const t: Tokens = tokensFromUsage(usage)
411 const id = e.agentId ?? 'main'
412 const actual = costOf(t, familyOf(usage.model))
413 const now = await $.clock.now()
414 await update($, ledger, l => {
415 const row =
416 l.rows[id] ??
417 blankRow(id, id === 'main' ? 'main conversation' : 'subagent', id === 'main' ? 'main' : '?', now)
418 return {
419 ...l,
420 rows: {
421 ...l.rows,
422 [id]: {
423 ...row,
424 model: row.model ?? usage.model,
425 via: row.via ?? (id === 'main' ? 'unrouted' : row.via),
426 baselineModel: row.baselineModel ?? (id === 'main' ? usage.model : undefined),
427 tokens: addTokens(row.tokens, t),
428 cost: row.cost + actual,
429 },
430 },
431 }
432 })
433 return result
434 })
435
436 on('turn.complete', async ($, e, next) => {
437 await refreshStatus($, baseline, baselineName, mode, pressureAt)
438 return next(e)
439 })
440
441 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
442 const { Box, Text } = $.ui.resolve(e)
443 const l = await read($, ledger)
444 const t = totals(l, baseline)
445 const width = Math.max(30, (e.props.bodyColumns ?? 60) - 2)
446 const rows = Object.values(l.rows).sort((a, b) => b.startedAt - a.startedAt)
447 const room = Math.max(3, (e.viewport?.rows ?? 30) - 10)
448 const labelWidth = Math.max(8, width - 33)
449 const baseHead = baseline === 'unrouted' ? 'unrouted' : shortModel(baseline)
450
451 return (
452 <Box flexDirection="column" width={width}>
453 <Text bold>
454 {usd(t.cost)} spent · {usd(t.base)} {baselineName}
455 </Text>
456 <Text color={t.saved > 0 ? 'green' : undefined}>
457 {t.base > 0 ? `${Math.round(t.saved * 100)}% saved` : 'No model calls yet.'} · {l.routed} routed ·
458 classifier {usd(l.classifierCost)}
459 </Text>
460 {mode !== 'on' && <Text color="yellow">{modeLabel(mode)} (/router on to go back)</Text>}
461 <Text> </Text>
462 <Text dimColor wrap="truncate">
463 {'agent'.padEnd(labelWidth)} T model {'cost'.padStart(7)} {baseHead.padStart(8)}
464 </Text>
465 {rows.slice(0, room).map(r => (
466 <Text wrap="truncate" dimColor={r.via === 'unrouted'}>
467 {(r.label.length > labelWidth - 1 ? r.label.slice(0, labelWidth - 2) + '…' : r.label).padEnd(labelWidth)}
468 {tierTag(r)} {shortModel(r.model).padEnd(7)} {usd(r.cost).padStart(7)}{' '}
469 {usd(baselineCost(r, baseline)).padStart(8)}
470 </Text>
471 ))}
472 {rows.length > room && <Text dimColor>…{rows.length - room} more</Text>}
473 <Text> </Text>
474 <Text dimColor wrap="wrap">
475 T: S simple→Haiku 5.5 at low effort, M standard→Sonnet at medium effort at most, H hard→Opus,
476 L long→Fable (background only). ! risky, sent to Opus at least; ↓ one tier down for rate limits.
477 Dim rows are not routed. API list prices; plans billed by
478 subscription see rate-limit use, not dollars.
479 </Text>
480 </Box>
481 )
482 })
483}
484hooks/lib/agents.ts 36 lines1/** The four roles the router adds. Their model is decided per task at spawn time. */
2export const ROLES = [
3 {
4 name: 'architect',
5 description:
6 'Plans before code is written: reads the codebase, weighs designs, and returns a step-by-step plan with the files to touch. Use for new features, refactors and anything with design choices. Does not edit files.',
7 tools: ['Read', 'Grep', 'Glob', 'WebFetch', 'WebSearch'],
8 prompt: `You are the architect. Read the relevant code, then return a plan another agent can follow without re-deriving it.
9Output: the goal in one line; the files to change and why; ordered steps; risks and how to test the result.
10Do not edit files. Keep the plan as short as the task allows.`,
11 },
12 {
13 name: 'builder',
14 description:
15 'Writes and edits code to carry out a defined task or an architect plan. Use for implementing features, writing tests, and refactors with a clear target.',
16 prompt: `You are the builder. Make the change you are given, following the existing style of the codebase.
17Keep the diff focused. When done, list the files you changed and anything you could not finish.`,
18 },
19 {
20 name: 'runner',
21 description:
22 'Runs commands and reports results: tests, builds, linters, type checks, searches. Use when the job is to execute and summarize, not to change code.',
23 tools: ['Bash', 'Read', 'Grep', 'Glob'],
24 prompt: `You are the runner. Run what you are asked to run and report the result.
25Report: the command, pass or fail, and for each failure the test or file, the error, and the line. Quote errors exactly; do not paraphrase.
26Do not change code.`,
27 },
28 {
29 name: 'fixer',
30 description:
31 'Fixes what failed: takes a failing test, build error or bug report, finds the root cause, and changes the code until it passes. Use after a runner reports failures.',
32 prompt: `You are the fixer. Find the root cause of the failure you are given before changing anything.
33Fix the cause, not the symptom. Re-run the failing check to confirm. Report the cause in one or two lines and the files you changed.`,
34 },
35] as const
36hooks/lib/pricing.ts 268 lines1import type { Tier, Tokens } from '../../types'
2
3/** USD per million tokens, from platform.claude.com/docs/en/about-claude/pricing (Oct 2026). */
4export type Price = { input: number; output: number; cacheWrite: number; cacheRead: number }
5
6export const PRICES: Record<'haiku4' | 'haiku5' | 'sonnet' | 'opus' | 'fable', Price> = {
7 haiku4: { input: 1, output: 5, cacheWrite: 1.25, cacheRead: 0.1 },
8 // Prompts up to 100,000 tokens; HAIKU5_LONG above that.
9 haiku5: { input: 0.1, output: 0.5, cacheWrite: 0.125, cacheRead: 0.01 },
10 sonnet: { input: 2, output: 10, cacheWrite: 2.5, cacheRead: 0.1 },
11 opus: { input: 4, output: 20, cacheWrite: 5, cacheRead: 0.2 },
12 fable: { input: 10, output: 50, cacheWrite: 12.5, cacheRead: 0.25 },
13}
14
15/** Haiku 5.5 for a prompt (input, cache reads and cache writes together) over 100,000 tokens. */
16export const HAIKU5_LONG: Price = { input: 0.5, output: 2.5, cacheWrite: 0.625, cacheRead: 0.05 }
17export const HAIKU5_LONG_FROM = 100_000
18
19export type Family = keyof typeof PRICES
20
21/** Full model ids the router spawns on. */
22export const MODEL_FOR_TIER: Record<Tier, string> = {
23 simple: 'claude-haiku-5-5',
24 standard: 'claude-sonnet-5-5',
25 hard: 'claude-opus-5-5',
26 long: 'claude-fable-5-1',
27}
28
29/**
30 * The classifier stays on Haiku 4.5: it answers without thinking, so an 8-token cap holds.
31 * Haiku 5.5 thinks by default and would often come back empty at that cap.
32 */
33export const CLASSIFIER_MODEL = 'claude-haiku-4-5-20251001'
34
35export const ORDER: Tier[] = ['simple', 'standard', 'hard', 'long']
36
37export type Effort = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
38const EFFORTS: Effort[] = ['low', 'medium', 'high', 'xhigh', 'max']
39
40/**
41 * The most reasoning effort a routed agent's requests may ask for, by tier. Easy work
42 * thinks less; hard and long work keep whatever the session would have used.
43 */
44export const EFFORT_CAP: Record<Tier, Effort | undefined> = {
45 simple: 'low',
46 standard: 'medium',
47 hard: undefined,
48 long: undefined,
49}
50
51/**
52 * The lower of a request's effort and the cap. An integer budget or a level this
53 * table does not know is left alone, as is a request without effort.
54 */
55export function capEffort<E>(effort: E, cap: Effort | undefined): E | Effort {
56 if (!cap || typeof effort !== 'string') return effort
57 const i = EFFORTS.indexOf(effort as Effort)
58 return i < 0 || i <= EFFORTS.indexOf(cap) ? effort : cap
59}
60
61/** `/router` modes: `on` classifies, `off` leaves every spawn alone, a model name sends every routed spawn there. */
62export type Mode = 'on' | 'off' | 'haiku' | 'sonnet' | 'opus'
63export const MODES: Mode[] = ['on', 'off', 'haiku', 'sonnet', 'opus']
64
65export function parseMode(text: unknown): Mode | undefined {
66 const m = String(text ?? '').trim().toLowerCase()
67 return (MODES as string[]).includes(m) ? (m as Mode) : undefined
68}
69
70export const FORCED_TIER: Record<Exclude<Mode, 'on' | 'off'>, Tier> = {
71 haiku: 'simple',
72 sonnet: 'standard',
73 opus: 'hard',
74}
75
76/**
77 * Which price row a model id or alias falls in. Mythos bills as Fable.
78 * A bare `haiku` alias is the current Haiku (5.5). Unknown ids price as Opus.
79 */
80export function familyOf(model: string | undefined): Family {
81 const m = (model ?? '').toLowerCase()
82 if (m.includes('haiku')) return /haiku-[34]/.test(m) ? 'haiku4' : 'haiku5'
83 if (m.includes('sonnet')) return 'sonnet'
84 if (m.includes('fable') || m.includes('mythos')) return 'fable'
85 return 'opus'
86}
87
88const FAMILY_TIER: Record<Family, Tier> = {
89 haiku4: 'simple',
90 haiku5: 'simple',
91 sonnet: 'standard',
92 opus: 'hard',
93 fable: 'long',
94}
95
96/** The tier a model belongs to, for a model Claude named itself. */
97export function tierOfModel(model: string): Tier {
98 return FAMILY_TIER[familyOf(model)]
99}
100
101/**
102 * Cost of one request's tokens on a family. Haiku 5.5 switches price above a 100K-token prompt,
103 * so call it per request; on summed tokens it prices everything at the lower rate.
104 * Cache writes are priced as 5-minute writes; 1-hour writes cost more, so a session that uses
105 * them is under-counted here.
106 */
107export function costOf(t: Tokens, family: Family): number {
108 const prompt = t.input + t.cacheRead + t.cacheWrite
109 const p = family === 'haiku5' && prompt > HAIKU5_LONG_FROM ? HAIKU5_LONG : PRICES[family]
110 return (t.input * p.input + t.output * p.output + t.cacheWrite * p.cacheWrite + t.cacheRead * p.cacheRead) / 1e6
111}
112
113export function tokensFromUsage(u: {
114 input_tokens: number
115 output_tokens: number
116 cache_read_input_tokens: number
117 cache_creation_input_tokens: number
118}): Tokens {
119 return {
120 input: u.input_tokens ?? 0,
121 output: u.output_tokens ?? 0,
122 cacheRead: u.cache_read_input_tokens ?? 0,
123 cacheWrite: u.cache_creation_input_tokens ?? 0,
124 }
125}
126
127export function addTokens(a: Tokens, b: Tokens): Tokens {
128 return {
129 input: a.input + b.input,
130 output: a.output + b.output,
131 cacheRead: a.cacheRead + b.cacheRead,
132 cacheWrite: a.cacheWrite + b.cacheWrite,
133 }
134}
135
136export const ZERO: Tokens = { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }
137
138/**
139 * Role guard rails, applied after Haiku's answer.
140 * runner never goes above standard (running tests is not Opus work).
141 * architect and fixer never go below standard (a bad plan or a bad fix costs more than the tokens).
142 */
143export const ROLE_LIMITS: Record<string, { min?: Tier; max?: Tier; fallback: Tier }> = {
144 architect: { min: 'standard', fallback: 'hard' },
145 builder: { fallback: 'standard' },
146 runner: { max: 'standard', fallback: 'simple' },
147 fixer: { min: 'standard', fallback: 'standard' },
148}
149
150export function roleOf(agentType: string): string | undefined {
151 const m = /^model-router:(\w+)$/.exec(agentType)
152 return m ? m[1] : undefined
153}
154
155/**
156 * The role's limits, then `long` only for background work: a foreground task the
157 * parent waits on goes to Opus, never to Fable at 2.5 times the price.
158 */
159export function clampTier(tier: Tier, role: string | undefined, background = true): Tier {
160 let i = ORDER.indexOf(tier)
161 const lim = role ? ROLE_LIMITS[role] : undefined
162 if (lim?.min) i = Math.max(i, ORDER.indexOf(lim.min))
163 if (lim?.max) i = Math.min(i, ORDER.indexOf(lim.max))
164 if (!background) i = Math.min(i, ORDER.indexOf('hard'))
165 return ORDER[i]!
166}
167
168/**
169 * The tier after the risk floor and rate-limit pressure.
170 *
171 * Under pressure (a rate-limit window past the threshold) the tier drops one step and
172 * `long` is off the table; the role's limits still hold. A risky task then goes to at
173 * least `hard`, past a role's cap too: getting a destructive act wrong costs more than
174 * the tokens, so the risk floor wins over pressure.
175 */
176export function routeTier(
177 tier: Tier,
178 role: string | undefined,
179 background: boolean,
180 o: { risky?: boolean; underPressure?: boolean } = {},
181): Tier {
182 const i = ORDER.indexOf(tier)
183 let t = clampTier(o.underPressure ? ORDER[Math.max(0, i - 1)]! : tier, role, background && !o.underPressure)
184 if (o.risky && ORDER.indexOf(t) < ORDER.indexOf('hard')) t = 'hard'
185 return t
186}
187
188/**
189 * Destructive acts, for when no classifier answered. Narrow on purpose: it names the act
190 * (deploying to production, dropping a table, a forced push), not the subject, so code
191 * that merely deals with payments or databases does not match.
192 */
193const RISKY_ACT =
194 /\b(deploy|ship|release|roll\s?out)\w*\b[^.\n]{0,40}\b(to|on|in)\s+(prod|production|live)\b|\bdrop\s+(table|database|schema)\b|\btruncate\s+table\b|\brm\s+-rf\s+\/|\bgit\s+push\s+(-f|--force)\b|\bforce[- ]push\b|\b(delete|wipe|purge)\b[^.\n]{0,30}\b(prod|production)\s+(data|database|db|bucket|users?)\b/i
195
196export function looksRisky(text: string): boolean {
197 return RISKY_ACT.test(text)
198}
199
200/** Reads the classifier's risk flag: the word `risky` after the tier. */
201export function parseRisky(text: string): boolean {
202 return /\brisky\b/i.test(text)
203}
204
205/** The fullest rate-limit window, in percent; 0 off a subscription or before the first reading. */
206export function fullestWindow(limits: readonly { percentUsed: number }[] | undefined): number {
207 return (limits ?? []).reduce((m, l) => Math.max(m, l.percentUsed), 0)
208}
209
210/** The `limitPressure` setting as a percentage, or undefined when off. */
211export function pressureThreshold(setting: unknown): number | undefined {
212 const n = Number(setting ?? 80)
213 return Number.isFinite(n) && n > 0 && n <= 100 ? n : undefined
214}
215
216export function fallbackTier(role: string | undefined): Tier {
217 return (role && ROLE_LIMITS[role]?.fallback) || 'standard'
218}
219
220/** Reads the classifier's reply: the first tier word it names. */
221export function parseTier(text: string): Tier | undefined {
222 const m = /\b(simple|standard|hard|long)\b/i.exec(text)
223 return m ? (m[1]!.toLowerCase() as Tier) : undefined
224}
225
226export const CLASSIFIER_SYSTEM = `You route coding subagent tasks to a model. Read the task and reply with exactly one word.
227
228simple - mechanical, little judgment: rename, find/search/grep, list files, run an existing test or build command and report, format, small single-file edits with an obvious answer.
229standard - normal feature work: implement a function or component, write tests, a refactor within a few files, a bug with a clear repro.
230hard - needs deep reasoning: architecture or design decisions, an intermittent or cross-cutting bug, concurrency, security, performance work, a migration with subtle risk.
231long - a multi-hour autonomous run: a large multi-step build or migration across many files, explicitly long-running or background work that must keep going unattended.
232
233Pick the cheapest tier that will succeed. If unsure between two, pick the higher.
234
235Then add the word risky if carrying out the task itself could do costly or hard-to-reverse harm: deploying to production, running a migration or a destructive command against production or shared data, deleting data, rotating or exposing credentials, moving money, force-pushing over shared history. Judge the act, not the subject: writing, refactoring or testing code that deals with payments, databases or credentials is not risky.
236
237Reply with the tier word alone, or the tier word and risky: for example "standard" or "hard risky".`
238
239/** What the engine's built-in classifier reads when Haiku 4.5 gives no tier: no rubric, so keep it short. */
240export function builtinClassifierText(input: { agentType: string; description: string; prompt: string }): string {
241 const body = input.prompt.length > 2000 ? input.prompt.slice(0, 2000) + ' [...]' : input.prompt
242 return `How hard is this coding task for an AI agent? Agent: ${input.agentType}. ${input.description}. ${body}`
243}
244
245export function classifierPrompt(input: {
246 agentType: string
247 description: string
248 prompt: string
249 background: boolean
250}): string {
251 const body = input.prompt.length > 6000 ? input.prompt.slice(0, 6000) + '\n[...truncated]' : input.prompt
252 return [
253 `Agent type: ${input.agentType}`,
254 `Runs in background: ${input.background ? 'yes' : 'no'}`,
255 `Short description: ${input.description}`,
256 'Task:',
257 '<task>',
258 body,
259 '</task>',
260 ].join('\n')
261}
262
263export function usd(n: number): string {
264 if (n === 0) return '$0'
265 if (n < 0.01) return '<$0.01'
266 return '$' + n.toFixed(2)
267}
268types/index.d.ts 52 lines1export type Tier = 'simple' | 'standard' | 'hard' | 'long'
2
3export type Effort = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
4
5export type Tokens = {
6 input: number
7 output: number
8 cacheRead: number
9 cacheWrite: number
10}
11
12/** One subagent (or the main loop, id "main") as the bill pane shows it. */
13export type AgentRow = {
14 id: string
15 label: string
16 agentType: string
17 tier?: Tier
18 /**
19 * How the model was decided: the Haiku classifier, the engine's built-in classifier
20 * when Haiku gave no tier, the role fallback, a model Claude named, a `/router` mode
21 * that sends everything to one model, or not routed (main, forks, workflow agents,
22 * teammates, built-ins when routing is narrowed, everything under `/router off`).
23 */
24 via?: 'haiku' | 'builtin' | 'fallback' | 'explicit' | 'forced' | 'unrouted'
25 model?: string
26 /** True when the task was judged risky and sent to at least the hard tier. */
27 risky?: boolean
28 /** The fullest rate-limit window, in percent, when it pushed this agent a tier down. */
29 pressure?: number
30 /** The most effort this agent's requests may ask for; absent when the router leaves effort alone. */
31 effort?: Effort
32 /** What this agent would have run on without the router: the parent's model, or its own when unrouted. */
33 baselineModel?: string
34 tokens: Tokens
35 /** USD at the price of the model that actually answered each step. */
36 cost: number
37 startedAt: number
38}
39
40export type RouterLedger = {
41 rows: Record<string, AgentRow>
42 /** What the Haiku classifier calls themselves cost. */
43 classifierCost: number
44 routed: number
45}
46
47declare module 'claude-code' {
48 interface PluginState {
49 'model-router': { ledger: RouterLedger }
50 }
51}
52