Routes each coding subtask to the cheapest Claude model and effort that can do it reliably, logs real token cost per subagent automatically, and learns routing…

A Claude Code plugin that routes each coding subtask to the cheapest Claude model and effort that can do it reliably. It measures what every subagent really costs and learns routing rules from verified runs.
1. From the marketplace (recommended, gets updates). At the Claude Code prompt:
/plugin install model-router --marketplace DevstateBelgium/model-router
Answer y to add the marketplace, then pick a scope. Or in two steps:
/plugin marketplace add DevstateBelgium/model-router
/plugin install model-router@model-router
2. Drop-in, zero commands. Clone this repo into ~/.claude/skills/model-router/ (Windows: %USERPROFILE%\.claude\skills\model-router\) and start a new session:
git clone https://github.com/DevstateBelgium/model-router ~/.claude/skills/model-router
Claude Code loads any folder there that has .claude-plugin/plugin.json as a plugin (model-router@skills-dir), including its agents and hooks.
3. Whole team, per repository. Commit this to the project's .claude/settings.json. Teammates get it after they trust the folder:
{
"extraKnownMarketplaces": { "model-router": { "source": { "source": "github", "repo": "DevstateBelgium/model-router" } } },
"enabledPlugins": { "model-router@model-router": true }
}
4. Ask Claude to install it, in any session. Say "install the plugin from https://github.com/DevstateBelgium/model-router". Claude follows INSTALL.md, which runs node scripts/install.mjs --live.
The first time the plugin loads, it starts a short configuration run. In an interactive session the mod queues it automatically; in every other session the SessionStart hook tells Claude to run it on your first message. The run checks Node.js, Claude Code and gh, then asks one question: local-only, connect an existing knowledge repo, or create a new private one. Run it again any time with /model-router:setup.
A cloud session starts from a fresh clone of your repository, and Claude Code loads plugins only at session start. There are two ways to get model-router there:
git clone --depth 1 https://github.com/DevstateBelgium/model-router ~/.claude/skills/model-router || true ``--live mode also writes plain copies of the skill and the five tiers. Claude Code watches ~/.claude/skills/ and ~/.claude/agents/ and picks up new files within seconds, so routing works in the same session. Only folders that existed when the session started are watched. Cloud sessions start with a skills folder but no agents folder, so there the tiers are written as forked skills (context: fork) that run as subagents with the same pinned model and effort. When even the skills folder is missing, the installer says so and Claude reads the skill file directly. Hooks, telemetry, the mod and /model-router:* commands need a plugin reload. Claude can't trigger that itself, and over a remote connection it's refused, so those parts start in the next session. The live copies delete themselves once the plugin loads.Requirements: a recent Claude Code (tested on 2.1.292). Node.js 18+ on PATH for the hooks and scripts. Without Node.js the skill and agents still work, but automatic telemetry and push/pull are off. gh is needed only for push/pull.
Start a task. Ask for the work as you normally would. The skill loads by itself when Claude splits multi-step work across subagents. To force it, start the request with /model-router (or /model-router:run when the mod is not loaded):
/model-router move the date helpers into utils/ and add tests
Ask "which model should do X?" to get a routing suggestion without running anything. Orchestration works best when the main session runs Claude Opus 5.5 at xhigh effort, because the orchestrator does the planning and verification.
What happens during a run. Claude splits the request into tasks, each with the files it may touch and an acceptance check. Each task goes to the cheapest tier that fits:
| Task looks like | Tier |
|---|---|
| Reading, searching, renames, boilerplate, running tests | scout |
| Clear spec, 1 to 3 files, a test to pass | builder |
| Multi-file change, refactor, bug with an unknown cause | engineer |
| Engineer task that failed once | senior |
| Trade-off, security-sensitive logic, two failed attempts | architect |
Independent tasks run in parallel. Claude checks every result itself (reads the diff, runs the check). A failed task is retried once with a sharper brief, then moved one tier up. Small edits that take less time to do than to brief are done directly, without a subagent.
See what it cost. Type /router for this session's cost per tier and model (needs the mod, Claude Code 2.1.287+). /model-router:stats 7 shows subagent cost over the last 7 days from telemetry, with or without the mod.
How it learns. After each verified task, Claude adds one judgment row to pending-runs.md: tier, model, effort, whether the check passed, and whether the tier was too strong or too weak. When a run teaches something new, the reply ends with a one-line Model-router: learned … note. Once five rows are pending, or one pattern repeats three times, Claude proposes routing rules in boundaries.md, marked (proposed). A rule needs at least three consistent rows, and it takes effect only after you push.
Share what it learned. With a knowledge repo configured (/model-router:setup):
/model-router:push sends pending rows, telemetry and proposed rules to the repo. It lists the proposed rules first, so you see what goes live./model-router:pull brings in what other machines and cloud sessions learned. The SessionStart summary tells you when the local copy is stale.| Component | What it does | |
|---|---|---|
/model-router:run skill | Orchestration rules: split work, pick a tier, verify, escalate, log. Loads when Claude delegates work. | |
| 5 tier agents | model-router:scout (Claude Haiku 5.5, medium), builder (Claude Sonnet 5.5, high), engineer (Claude Opus 5.5, medium), senior (Claude Opus 5.5, xhigh), architect (Claude Fable 5.1, high). Each pins its model and effort. | |
| SessionStart hook | Seeds the data dir on first run, then adds a short summary to context: active boundaries, pending rows, a stale-mirror warning, a model-check reminder. | |
| PostToolUse + SubagentStop hooks | Log every subagent: resolved model, tokens summed over the whole subagent transcript, duration, estimated USD. No model effort needed. | |
Mod (hooks/router.tsx) | Runs inside Claude Code (needs v2.1.287+, otherwise ignored). Drops a model override on model-router:* spawns so the pinned model and effort always run. Meters every model request, orchestrator included, and shows router $… · subagents $… in the status line. Keeps per-session totals across sessions. | |
/router | Mod command. Opens a pane with this session's cost per tier and model, plus the last 7 days. Answers instantly with no model turn, and also works in claude -p. | |
/model-router <task> | Mod command. Short form of /model-router:run: forwards the task to the skill. Without the mod, use /model-router:run. | |
/model-router:stats [days] | Real cost per subagent run, from telemetry (works without the mod). | |
| `/model-router:setup [owner/repo \ | local]` | First-run configuration, also runnable any time: prerequisite check, then an optional private knowledge repo (created and seeded after confirmation). |
scripts/install.mjs | Cross-platform installer. --live makes the skill and agents work in the running session, cloud sessions included. | |
/model-router:pull, /model-router:push | Sync shared knowledge across machines and cloud sessions. Push detects conflicts instead of overwriting. |
State lives in ~/.claude/plugins/data/model-router-<origin>/ and survives plugin updates: mirror/ (copy of the knowledge repo), pending-runs.md (judgment rows), telemetry.jsonl (hook-written cost data), prices.json, state.json, and optional config.json.
The knowledge repo is optional. Set it in /config (model-router > Knowledge repo) or with /model-router:setup. Without one, everything runs locally.
bash tests/sync.test.sh runs push/pull end to end against a fake GitHub API (tests/fake-gh.mjs), so no network access is needed.bash tests/install.test.sh checks the installer, the live copies, first-run state and cleanup against a fake home directory.claude plugin test . runs the mod's tests (tests/router.test.ts).claude plugin validate . --strict checks the manifests and lists what the mod hooks and calls.claude --plugin-dir .) so Claude Code writes .claude-plugin/types/, then run tsc -p .. Both generated paths are git-ignored.hooks/router.tsx 191 lines1import { atom, read, update } from 'claude-code'
2import type { Register } from 'claude-code'
3
4import type { RouterRow, RouterSession } from '../types'
5
6type Price = { input: number; output: number; cache_write_multiplier: number; cache_read_multiplier: number }
7type Usage = { model: string; input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number }
8
9const PANE = 'model-router'
10const HISTORY_KEY = 'history'
11const HISTORY_MAX = 200
12const SETUP_KEY = 'setupStartedAt'
13export const FIRST_RUN_PROMPT =
14 'model-router was just installed. Run its first-run configuration now by invoking the model-router:setup skill, then tell me in one line what was set up.'
15const totals = atom({ plugin: 'model-router', key: 'totals' } as const, {})
16const agents = atom({ plugin: 'model-router', key: 'agents' } as const, {})
17const overrides = atom({ plugin: 'model-router', key: 'overrides' } as const, 0)
18
19let prices: Record<string, Price> = {}
20
21export function priceFor(table: Record<string, Price>, model: string): Price | undefined {
22 const key = Object.keys(table).find(k => model.startsWith(k))
23 return key ? table[key] : undefined
24}
25
26export function costOf(p: Price | undefined, u: Usage): number {
27 if (!p) return 0
28 return (
29 (u.input_tokens * p.input +
30 u.cache_creation_input_tokens * p.input * p.cache_write_multiplier +
31 u.cache_read_input_tokens * p.input * p.cache_read_multiplier +
32 u.output_tokens * p.output) /
33 1_000_000
34 )
35}
36
37export function addUsage(rows: Record<string, RouterRow>, who: string, u: Usage, p: Price | undefined) {
38 const key = `${who}|${u.model}`
39 const row: RouterRow = rows[key] ?? { who, model: u.model, requests: 0, input: 0, output: 0, cacheRead: 0, cacheWrite: 0, usd: 0, priced: true }
40 return {
41 ...rows,
42 [key]: {
43 ...row,
44 requests: row.requests + 1,
45 input: row.input + u.input_tokens,
46 output: row.output + u.output_tokens,
47 cacheRead: row.cacheRead + u.cache_read_input_tokens,
48 cacheWrite: row.cacheWrite + u.cache_creation_input_tokens,
49 usd: row.usd + costOf(p, u),
50 priced: row.priced && p !== undefined,
51 },
52 }
53}
54
55const tierName = (type: string) => type.replace(/^model-router:/, '')
56const money = (n: number) => `$${n < 10 ? n.toFixed(3) : n.toFixed(2)}`
57const kTok = (n: number) => (n >= 1000 ? `${(n / 1000).toFixed(1)}k` : String(n))
58
59export function summaryLine(rows: RouterRow[]): string {
60 const total = rows.reduce((s, r) => s + r.usd, 0)
61 const sub = rows.filter(r => r.who !== 'orchestrator').reduce((s, r) => s + r.usd, 0)
62 const spawned = new Set(rows.filter(r => r.who !== 'orchestrator').map(r => r.who)).size
63 return `router ${money(total)} · subagents ${money(sub)}${spawned ? ` (${spawned} tier${spawned > 1 ? 's' : ''})` : ''}`
64}
65
66export function table(rows: RouterRow[]): string {
67 if (!rows.length) return 'No model requests recorded in this session yet.'
68 const lines = ['| who | model | requests | input tok | output tok | est. $ |', '|---|---|---|---|---|---|']
69 for (const r of [...rows].sort((a, b) => b.usd - a.usd)) {
70 const input = r.input + r.cacheRead + r.cacheWrite
71 lines.push(`| ${r.who} | ${r.model} | ${r.requests} | ${kTok(input)} | ${kTok(r.output)} | ${r.priced ? money(r.usd) : 'n/a'} |`)
72 }
73 lines.push('', `${summaryLine(rows)}. List prices; cache reads and writes included.`)
74 return lines.join('\n')
75}
76
77export const register: Register = on => {
78 on('session.start', async ($, e, next) => {
79 try {
80 prices = JSON.parse(await $.fs.read(`${$.plugin.root}/seed/prices.json`)).models ?? {}
81 } catch {
82 prices = {}
83 }
84 try {
85 await $.command.register({ name: 'router', description: 'model-router: live cost per tier for this session and the last 7 days' })
86 await $.command.register({ name: 'model-router', description: 'model-router: route a task across the tiers (same as /model-router:run)', argumentHint: '<task>' })
87 } catch {}
88 // First load in an interactive session (install, enable, or first start): run the configuration once.
89 if (e.isInteractive && !(await $.store.get(SETUP_KEY))) {
90 await $.store.set(SETUP_KEY, new Date(await $.clock.now()).toISOString())
91 // session.start is awaited before the first prompt, so the prompt is queued from a timer.
92 $.clock.after(500, () => {
93 void $.prompt.submit({ text: FIRST_RUN_PROMPT })
94 })
95 }
96 return next(e)
97 })
98
99 // A model override would silently replace the tier's pinned model and evidence.
100 on('agent.spawn', async ($, e, next) => {
101 const pinned = e.subagentType.startsWith('model-router:') && e.model !== undefined
102 if (pinned) {
103 await update($, overrides, n => (n ?? 0) + 1)
104 $.ui.toast(`model-router: ignored model "${e.model}" for ${e.subagentType}; the tier's pinned model runs`)
105 }
106 const res = await next(pinned ? { ...e, model: undefined } : e)
107 if (res.agentId) {
108 const id = res.agentId
109 await update($, agents, m => ({ ...(m ?? {}), [id]: tierName(e.subagentType) }))
110 }
111 return res
112 })
113
114 on('turn.step', async function* ($, e, next) {
115 const res = yield* next(e)
116 if (res.usage) {
117 let who = 'orchestrator'
118 if (e.agentId) {
119 const known = (await read($, agents))[e.agentId]
120 who = known ?? tierName((await $.agent.list()).find(a => a.id === e.agentId)?.type ?? 'subagent')
121 }
122 const usage = res.usage
123 const rows = await update($, totals, t => addUsage(t ?? {}, who, usage, priceFor(prices, usage.model)))
124 $.ui.status(summaryLine(Object.values(rows)))
125 }
126 return res
127 })
128
129 on('session.end', async ($, e, next) => {
130 const rows = Object.values(await read($, totals))
131 if (rows.length) {
132 const history = ((await $.store.get(HISTORY_KEY)) as RouterSession[] | undefined) ?? []
133 await $.store.set(HISTORY_KEY, [...history, { endedAt: new Date(await $.clock.now()).toISOString(), rows }].slice(-HISTORY_MAX))
134 }
135 return next(e)
136 })
137
138 on('command.run', { command: 'router' }, async $ => {
139 const rows = Object.values(await read($, totals))
140 await $.ui.open({ id: PANE, title: 'model-router' })
141 return { text: table(rows) }
142 })
143
144 // Plugin skills are always namespaced, so the short name forwards to the skill.
145 // $.command.run rejects inside the hook this command waits on, so it is queued from a timer.
146 on('command.run', { command: 'model-router' }, async ($, e) => {
147 $.clock.after(0, () => {
148 void $.command.run({ command: 'model-router:run', args: e.args }).catch(() => {})
149 })
150 return {}
151 })
152
153 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
154 const { Box, Text } = $.ui.resolve(e)
155 const rows = Object.values(await read($, totals)).sort((a, b) => b.usd - a.usd)
156 const skipped = await read($, overrides)
157 const history = ((await $.store.get(HISTORY_KEY)) as RouterSession[] | undefined) ?? []
158 const weekAgo = (await $.clock.now()) - 7 * 86_400_000
159 const week = new Map<string, number>()
160 for (const s of history.filter(s => Date.parse(s.endedAt) >= weekAgo))
161 for (const r of s.rows) week.set(r.who, (week.get(r.who) ?? 0) + r.usd)
162 const weekRows = [...week].sort((a, b) => b[1] - a[1])
163 const width = Math.max(30, e.props.bodyColumns ?? 60)
164 const cell = (s: string, n: number) => (s.length > n ? s.slice(0, n - 1) + '…' : s.padEnd(n))
165 const modelCol = Math.max(10, width - 12 - 5 - 8 - 10 - 4)
166
167 return (
168 <Box flexDirection="column">
169 <Text bold>This session</Text>
170 {rows.length === 0 && <Text dimColor>No model requests yet.</Text>}
171 {rows.length > 0 && <Text dimColor>{cell('who', 12)} {cell('model', modelCol)} {cell('req', 5)} {cell('out', 8)} est. $</Text>}
172 {rows.map(r => (
173 <Text>
174 {cell(r.who, 12)} {cell(r.model, modelCol)} {cell(String(r.requests), 5)} {cell(kTok(r.output), 8)} {r.priced ? money(r.usd) : 'n/a'}
175 </Text>
176 ))}
177 {rows.length > 0 && <Text color="cyan">{summaryLine(rows)}</Text>}
178 {skipped > 0 && <Text color="yellow">Model overrides ignored on pinned tiers: {skipped}</Text>}
179 <Text> </Text>
180 <Text bold>Last 7 days (ended sessions)</Text>
181 {weekRows.length === 0 && <Text dimColor>Nothing recorded yet.</Text>}
182 {weekRows.map(([who, usd]) => (
183 <Text>
184 {cell(who, 12)} {money(usd)}
185 </Text>
186 ))}
187 </Box>
188 )
189 })
190}
191types/index.d.ts 27 lines1export type RouterRow = {
2 who: string;
3 model: string;
4 requests: number;
5 input: number;
6 output: number;
7 cacheRead: number;
8 cacheWrite: number;
9 usd: number;
10 priced: boolean;
11};
12
13export type RouterSession = {
14 endedAt: string;
15 rows: RouterRow[];
16};
17
18declare module 'claude-code' {
19 interface PluginState {
20 'model-router': {
21 totals: Record<string, RouterRow>;
22 agents: Record<string, string>;
23 overrides: number;
24 };
25 }
26}
27