Routes each turn and subagent to a model tier and reasoning effort using a classifier behind any server that speaks POST /v1/classify (model-router-api on…

model-router is a Claude Code mod that decides, request by request, which model answers and how much reasoning effort it spends, in the main conversation and in subagents alike. A classifier reads the situation (the prompt, recent conversation, subagent details) and returns a tier, an effort level and a risk score. The mod turns that into a model and effort, and leaves the request unchanged whenever anything fails.
The classifier sits behind any server that speaks POST /v1/classify. This repo ships one: model-router-api, a Cloudflare Worker running TypeSafe's Jev on Workers AI (worker/). With no API configured, the mod uses Claude Code's built-in classifier for subagents only: it can move a subagent's model up, never down, and the main loop is left alone unless routeMainModel is on.
The idea comes from jev-model-router.
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1The plugin lives in plugin/, and the repo is also a marketplace. From a terminal:
claude plugin marketplace add sftinc/claude-code-model-router
claude plugin install model-router@sftinc
Or inside Claude Code:
/plugin marketplace add sftinc/claude-code-model-router
/plugin install model-router@sftinc
If the repo is private, adding the marketplace needs git access to it (an SSH key or gh auth login).
Start Claude Code with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1. To run it from a clone instead:
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir /path/to/claude-code-model-router/plugin
Options live under pluginConfigs in ~/.claude/settings.json. The key is model-router@sftinc when the plugin is installed from the marketplace, model-router when it's loaded with --plugin-dir, and model-router@skills-dir when it's loaded from .claude/skills/.
{
"pluginConfigs": {
"model-router": {
"options": {
"apiUrl": "https://model-router-api.<subdomain>.workers.dev",
"apiSecret": "<ROUTER_SECRET>"
}
}
}
}
| Option | Default | Notes |
|---|---|---|
provider | auto | auto, api, builtin |
apiUrl / apiSecret | empty | Both needed for the API |
fastModel / balancedModel / deepModel | haiku / sonnet / opus | Alias or full id |
minUpgradeConfidence / minDowngradeConfidence | 0 / 0 | Minimum confidence for a move to a costlier / cheaper setting. The API applies its own thresholds, so 0 follows it; the built-in classifier gives no confidence, so it still never moves down |
riskyThreshold | 0.7 | Above it: deep tier, at least high effort |
contextMessages / contextChars | 30 / 60000 | Recent context sent with main-loop prompts |
routeSubagentModel | true | |
respectAgentModels | true | Keeps a model the Agent tool named, unless risk forces deep. The subagent is still classified, and its line shows both picks, such as model (agent: sonnet; router: haiku @ 97%) (* when the router would keep the agent's). A model set in an agent's definition isn't visible to the router and isn't protected |
routeMainEffort | true | |
routeMainModel | false | Switching models invalidates the prompt cache |
timeoutMs | 2000 | Wait budget for the classifier |
warmUp | true | Sends one throwaway classification when an interactive session starts, so the first prompt doesn't hit a cold classifier |
logDecisions | true | Shows one short line per routed turn or subagent in the transcript, such as main loop · model (sonnet), effort (low → xhigh @ 90%) or subagent Explore · model (opus → haiku @ 97%), plus the warm-up; the details, including whole api replies, go to the debug log (claude --debug). A line reporting a failure always shows |
A request routed to a model without effort support (Haiku) never carries an effort, whatever the switches say.
cd worker
npm install
npx wrangler secret put ROUTER_SECRET # generate with: openssl rand -base64 32
npm run deploy
ROUTER_API=https://model-router-api.<subdomain>.workers.dev ROUTER_SECRET=... scripts/live-check.sh
Before the first deploy, create an AI Gateway named model-router in the Cloudflare dashboard, with Workers AI set to Unified billing. COLLECT_LOG in wrangler.jsonc controls whether gateway logs store prompts and context.
Two type-declaration files are generated, not committed. Until they exist, your editor reports missing types.
/plugin-types plugin. It writes Claude Code's declarations to plugin/.claude/types/, which plugin/tsconfig.json reads. Run it again after updating Claude Code.npm run typecheck in worker/. It runs wrangler types, which writes worker/worker-configuration.d.ts, then typechecks.CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude plugin test plugin # the plugin
cd worker && npm test # the Worker
plugin/scripts/e2e.sh runs real headless sessions with the plugin loaded, in plan mode so nothing gets carried out. It checks the router's decisions in each session's debug log: a trivial subagent goes to haiku, a model named on the Agent tool is kept, and a risky prompt forces the deep tier. It needs the api configured, and each scenario is a real session that costs about $0.10–0.30. Pass scenario names to run only some of them.
With an API configured, each prompt and up to contextChars of recent conversation go to it. With COLLECT_LOG=true, the AI Gateway's logs keep them.
hooks/model-router.ts 472 lines1/**
2 * model-router — the hooks.
3 *
4 * prompt.submit classifies the main-loop prompt and leaves the answer for the
5 * next turn; turn.step applies it to the turn's first request and holds it for
6 * the rest of the turn; agent.spawn classifies and routes a subagent on the
7 * spot. Every failure leaves the request as it was.
8 */
9import type { AgentSpawnResult, Args, EngineInterface, Register } from 'claude-code'
10
11import { buildRecent, builtinDecision, endpoint, readVerdict, requestHeaders, selectProvider } from './client.ts'
12import type { RecentMessage, Situation } from './client.ts'
13import { DEFAULT_TIERS } from './models.ts'
14import {
15 TIER_ORDER,
16 changeBasis,
17 describeDecision,
18 describeMove,
19 describeSetup,
20 modelAlias,
21 pendingDecisions,
22 requestModelId,
23 route,
24 supportsEffort,
25} from './policy.ts'
26import type { Decision, Effort, PolicyConfig, Provider, Routing } from './policy.ts'
27
28/** The engine and the event payloads, as Claude Code declares them. */
29type Engine = EngineInterface
30type PromptEvent = Args<'prompt.submit'>
31type StepEvent = Args<'turn.step'>
32type SpawnEvent = Args<'agent.spawn'>
33
34type Change = { model?: string; effort?: Effort }
35
36// ---------------------------------------------------------------------------
37// Options: one table of defaults. A supplied value is used only when it has
38// the default's type, and a string must also be non-empty.
39
40type Options = {
41 provider: string
42 apiUrl: string
43 apiSecret: string
44 fastModel: string
45 balancedModel: string
46 deepModel: string
47 minUpgradeConfidence: number
48 minDowngradeConfidence: number
49 riskyThreshold: number
50 contextMessages: number
51 contextChars: number
52 routeSubagentModel: boolean
53 respectAgentModels: boolean
54 routeMainEffort: boolean
55 routeMainModel: boolean
56 timeoutMs: number
57 warmUp: boolean
58 logDecisions: boolean
59}
60
61const OPTION_DEFAULTS: Readonly<Options> = {
62 provider: 'auto',
63 apiUrl: '',
64 apiSecret: '',
65 fastModel: DEFAULT_TIERS.fast,
66 balancedModel: DEFAULT_TIERS.balanced,
67 deepModel: DEFAULT_TIERS.deep,
68 // The API applies its own thresholds before it answers, so by default a move follows it.
69 minUpgradeConfidence: 0,
70 minDowngradeConfidence: 0,
71 riskyThreshold: 0.7,
72 contextMessages: 30,
73 contextChars: 60000,
74 routeSubagentModel: true,
75 respectAgentModels: true,
76 routeMainEffort: true,
77 routeMainModel: false,
78 timeoutMs: 2000,
79 warmUp: true,
80 logDecisions: true,
81}
82
83function readOptions(given: Record<string, unknown>): Options {
84 const merged: Record<string, unknown> = { ...OPTION_DEFAULTS }
85 for (const [key, fallback] of Object.entries(OPTION_DEFAULTS)) {
86 const value = given[key]
87 const usable = typeof value === typeof fallback && value !== ''
88 if (usable) merged[key] = value
89 }
90 return merged as Options
91}
92
93// ---------------------------------------------------------------------------
94// Per-registration state
95
96type Runtime = {
97 opts: Options
98 policy: PolicyConfig
99 backend: Provider | null
100 url: string
101 mainCanChange: boolean
102 once: { setup: boolean; apiWarning: boolean; historyWarning: boolean }
103 pending: ReturnType<typeof pendingDecisions<Classified>>
104 turn: { id: string; change: Change | null } | null
105}
106
107function createRuntime(opts: Options): Runtime {
108 const backend = selectProvider(opts.provider, opts.apiUrl, opts.apiSecret)
109 return {
110 opts,
111 policy: {
112 tiers: { fast: opts.fastModel, balanced: opts.balancedModel, deep: opts.deepModel },
113 minUpgradeConfidence: opts.minUpgradeConfidence,
114 minDowngradeConfidence: opts.minDowngradeConfidence,
115 riskyThreshold: opts.riskyThreshold,
116 },
117 backend,
118 url: backend === 'api' ? endpoint(opts.apiUrl) : '',
119 // The built-in classifier gives no effort, so main effort needs the api.
120 mainCanChange: opts.routeMainModel || (opts.routeMainEffort && backend === 'api'),
121 once: { setup: false, apiWarning: false, historyWarning: false },
122 pending: pendingDecisions<Classified>(),
123 turn: null,
124 }
125}
126
127// ---------------------------------------------------------------------------
128// Logging. Every line starts with what it is about ("main loop · …"); the
129// engine already labels it with the plugin's name. The transcript gets one short line per routed turn or subagent,
130// and the warm-up's result (`tell`), so the router is visibly running without
131// cluttering the chat; the details behind them, such as whole api replies, go
132// to the debug log alone (`note`). logDecisions gates both. A line carrying a
133// failure (`warn`) is always written, so a router that couldn't do its job
134// says so.
135
136function tell($: Engine, rt: Runtime, text: string): void {
137 if (rt.opts.logDecisions) $.ui.log(text)
138}
139
140function note($: Engine, rt: Runtime, text: string): void {
141 if (rt.opts.logDecisions) $.ui.log(text, { to: 'debug' })
142}
143
144function warn($: Engine, text: string): void {
145 $.ui.log(text)
146}
147
148function setupOnce($: Engine, rt: Runtime): void {
149 if (rt.once.setup) return
150 rt.once.setup = true
151 const switches = {
152 subagentModel: rt.opts.routeSubagentModel,
153 mainEffort: rt.opts.routeMainEffort && rt.backend === 'api',
154 mainModel: rt.opts.routeMainModel,
155 }
156 note($, rt, 'setup · ' + describeSetup(rt.backend, rt.url, switches, rt.opts.provider === 'builtin'))
157}
158
159function apiWarningOnce($: Engine, rt: Runtime): void {
160 if (rt.once.apiWarning) return
161 rt.once.apiWarning = true
162 if (rt.opts.provider === 'api' && rt.backend === null) {
163 warn($, `setup · provider "api" needs both apiUrl and apiSecret; falling back to the engine's own classifier`)
164 }
165}
166
167// ---------------------------------------------------------------------------
168// Classification. Never throws; any trouble means no decision, and says why.
169
170/** A classification's outcome: the decision, or the short reason there is none. */
171type Classified = { decision: Decision | null; failure: string | null }
172
173const failed = (failure: string): Classified => ({ decision: null, failure })
174
175const errorText = (error: unknown): string => (error instanceof Error ? error.message : String(error))
176
177type Race<T> = { settled: true; value: T } | { settled: false }
178
179/** Whichever comes first: the answer or the timer. An answer after the timer is never seen. */
180function beforeTimer<T>(answer: Promise<T>, timer: Promise<void>): Promise<Race<T>> {
181 return Promise.race([
182 answer.then((value): Race<T> => ({ settled: true, value })),
183 timer.then((): Race<T> => ({ settled: false })),
184 ])
185}
186
187async function askApi($: Engine, rt: Runtime, situation: Situation, who: string): Promise<Classified> {
188 const { timeoutMs, apiSecret } = rt.opts
189 const reply = $.http.fetch(rt.url, {
190 method: 'POST',
191 headers: requestHeaders(apiSecret),
192 body: JSON.stringify(situation),
193 })
194 const outcome = await beforeTimer(reply, $.clock.sleep(timeoutMs))
195 if (!outcome.settled) return failed(`no reply from the api within ${timeoutMs}ms`)
196 const { ok, status, text } = outcome.value
197 note($, rt, `${who} · api reply: HTTP ${status} ${text}`)
198 if (!ok) return failed(`api returned HTTP ${status}`)
199 const decision = readVerdict(text)
200 return decision ? { decision, failure: null } : failed('api reply was a malformed verdict')
201}
202
203async function askBuiltin($: Engine, rt: Runtime, prompt: string, who: string): Promise<Classified> {
204 const { timeoutMs } = rt.opts
205 const label = $.model.classify(prompt, [...TIER_ORDER])
206 const outcome = await beforeTimer(label, $.clock.sleep(timeoutMs))
207 if (!outcome.settled) return failed(`no reply from the built-in classifier within ${timeoutMs}ms`)
208 note($, rt, `${who} · builtin label: ${String(outcome.value)}`)
209 return { decision: builtinDecision(outcome.value), failure: null }
210}
211
212async function classify($: Engine, rt: Runtime, situation: Situation, upOnly: boolean, who: string): Promise<Classified> {
213 try {
214 const got =
215 rt.backend === 'api' ? await askApi($, rt, situation, who) : await askBuiltin($, rt, situation.prompt, who)
216 return got.decision ? { decision: { ...got.decision, upOnly }, failure: null } : got
217 } catch (error) {
218 note($, rt, `${who} · classifier threw: ${error instanceof Error ? (error.stack ?? error.message) : String(error)}`)
219 return failed(`classifier threw: ${errorText(error)}`)
220 }
221}
222
223/** The throwaway situation a warm-up sends. */
224const WARM_UP: Situation = { source: 'main', prompt: 'warm-up' }
225
226/**
227 * Sends one throwaway classification in the background, so a classifier that
228 * has gone cold is warm again by the first prompt. The timer detaches it from
229 * the session.start dispatch; only how it went is reported.
230 */
231function warmUp($: Engine, rt: Runtime): void {
232 const { apiSecret } = rt.opts
233 $.clock.after(0, async () => {
234 const started = await $.clock.now()
235 try {
236 const reply = await $.http.fetch(rt.url, {
237 method: 'POST',
238 headers: requestHeaders(apiSecret),
239 body: JSON.stringify(WARM_UP),
240 })
241 const elapsed = (await $.clock.now()) - started
242 note($, rt, `warm-up · api reply: HTTP ${reply.status} ${reply.text}`)
243 if (reply.ok) tell($, rt, `warm-up · done in ${elapsed}ms`)
244 else warn($, `warm-up · failed (api returned HTTP ${reply.status})`)
245 } catch (error) {
246 warn($, `warm-up · failed (${errorText(error)})`)
247 }
248 })
249}
250
251/** Classifies, writes the verdict to the debug log, and returns the outcome. */
252async function classifyAndReport(
253 $: Engine,
254 rt: Runtime,
255 situation: Situation,
256 upOnly: boolean,
257 who: string,
258): Promise<Classified> {
259 const started = await $.clock.now()
260 const got = await classify($, rt, situation, upOnly, who)
261 const elapsed = (await $.clock.now()) - started
262 const via = rt.backend === 'api' ? 'api' : 'builtin'
263 note($, rt, `${who} · verdict via ${via}: ${describeDecision(got.decision, elapsed)}`)
264 return got
265}
266
267// ---------------------------------------------------------------------------
268// prompt.submit
269
270async function onPrompt($: Engine, rt: Runtime, e: PromptEvent): Promise<void> {
271 setupOnce($, rt)
272 if (!rt.mainCanChange) return
273 apiWarningOnce($, rt)
274
275 const prompt = typeof e.text === 'string' ? e.text : ''
276 if (prompt.trim() === '') {
277 rt.pending.put(null)
278 return
279 }
280
281 let recent: RecentMessage[] | undefined
282 let upOnly = false
283 if (rt.opts.contextMessages > 0) {
284 try {
285 const messages = await $.session.messages()
286 recent = buildRecent(messages, prompt, rt.opts.contextMessages, rt.opts.contextChars)
287 } catch {
288 upOnly = true
289 if (!rt.once.historyWarning) {
290 rt.once.historyWarning = true
291 warn($, 'setup · session history unavailable; prompts go to the classifier alone, raise-only')
292 }
293 }
294 }
295
296 const situation: Situation = recent ? { source: 'main', prompt, recent } : { source: 'main', prompt }
297 rt.pending.put(await classifyAndReport($, rt, situation, upOnly, 'main loop'))
298}
299
300// ---------------------------------------------------------------------------
301// turn.step
302
303type Shaped = { change: Change | null; stripped: boolean; sendsTo: string }
304
305/**
306 * What a routing becomes for this request under the switches. The effort is
307 * cleared outright (key present, value undefined) whenever the request carries
308 * one but its destination model takes none, regardless of the switches.
309 */
310function shapeChange(routing: Routing, e: StepEvent, opts: Options): Shaped {
311 const model = opts.routeMainModel && routing.model !== null ? requestModelId(routing.model) : undefined
312 const sendsTo = model ?? e.model
313 const stripped = e.effort !== undefined && !supportsEffort(sendsTo)
314
315 const modelPart = model === undefined ? null : { model }
316 let effortPart: Pick<Change, 'effort'> | null = null
317 if (stripped) effortPart = { effort: undefined }
318 else if (opts.routeMainEffort && routing.effort !== null) effortPart = { effort: routing.effort }
319
320 const change = modelPart || effortPart ? { ...modelPart, ...effortPart } : null
321 return { change, stripped, sendsTo }
322}
323
324/** The line's body: model and effort, each with an arrow when it changes. */
325function describeTurn(e: StepEvent, shaped: Shaped, decision: Decision | null, policy: PolicyConfig): string {
326 const model = shaped.change?.model
327 const parts = [
328 describeMove(
329 'model',
330 modelAlias(e.model),
331 model === undefined ? null : modelAlias(model),
332 changeBasis(decision, decision?.confidence ?? null, policy),
333 ),
334 ]
335 const effort = shaped.change?.effort
336 if (shaped.stripped) parts.push('effort (dropped)')
337 else if (effort !== undefined) {
338 const basis = changeBasis(decision, decision?.effortConfidence ?? null, policy)
339 parts.push(describeMove('effort', String(e.effort ?? 'unset'), effort, basis))
340 } else if (e.effort !== undefined) parts.push(describeMove('effort', String(e.effort)))
341 return parts.join(', ')
342}
343
344async function routeTurn($: Engine, rt: Runtime, e: StepEvent): Promise<Change | null> {
345 const { decision, failure } = rt.pending.take() ?? { decision: null, failure: null }
346 const routing = route(decision, { model: e.model, effort: e.effort }, rt.policy)
347 const shaped = shapeChange(routing, e, rt.opts)
348
349 if (shaped.change !== null || rt.mainCanChange) {
350 const line = `main loop · ${describeTurn(e, shaped, decision, rt.policy)}`
351 if (failure) warn($, `${line}, ${failure}`)
352 else tell($, rt, line)
353 }
354 if (rt.mainCanChange) {
355 const unapplied =
356 routing.model !== null && !rt.opts.routeMainModel
357 ? ` (${routing.model} not applied: main-model switching is disabled)`
358 : ''
359 note($, rt, `main loop · ${routing.reason}${unapplied}`)
360 }
361 return shaped.change
362}
363
364/**
365 * Leaves the effort this turn is sent with in ~/.claude/router/<session>.effort
366 * for a status line command to show, since the status line's own effort.level
367 * stays the session's setting. Empty when the request carries none.
368 */
369async function publishEffort($: Engine, e: StepEvent, change: Change | null): Promise<void> {
370 const effort = change && 'effort' in change ? change.effort : e.effort
371 const home = await $.env.get('HOME')
372 if (!home) return
373 await $.fs.write(`${home}/.claude/router/${await $.session.id()}.effort`, String(effort ?? ''))
374}
375
376// ---------------------------------------------------------------------------
377// agent.spawn
378
379/**
380 * Where the router sends a subagent: the model to set (null: leave it), the
381 * line's model part with any failure inside it (null: show the model the
382 * engine started it on), and the classifier's failure.
383 */
384type SpawnRouting = { model: string | null; shown: string | null; failure: string | null }
385
386async function routeSpawn($: Engine, rt: Runtime, e: SpawnEvent): Promise<SpawnRouting> {
387 const who = `subagent ${e.subagentType}`
388 const situation: Situation = { source: 'subagent', prompt: e.prompt }
389 if (typeof e.description === 'string') situation.description = e.description
390 if (typeof e.subagentType === 'string') situation.agentType = e.subagentType
391 const { decision, failure } = await classifyAndReport($, rt, situation, false, who)
392
393 const pinned = rt.opts.respectAgentModels && e.model !== undefined
394 const current = e.model ?? e.parentModel
395 const routing = route(decision, { model: current, pinned }, rt.policy)
396 note($, rt, `${who} · ${routing.reason}`)
397 const basis = changeBasis(decision, decision?.confidence ?? null, rt.policy)
398 if (routing.model !== null) {
399 return { model: routing.model, shown: describeMove('model', modelAlias(current), modelAlias(routing.model), basis), failure }
400 }
401 if (!pinned) return { model: null, shown: null, failure }
402
403 // Kept for the agent, but still show where the router would have sent it; * means the same model.
404 const agent = `agent: ${modelAlias(current)}`
405 if (!decision) return { model: null, shown: `model (${agent}; ${failure})`, failure }
406 const unpinned = route(decision, { model: current }, rt.policy).model
407 const router = unpinned === null ? '*' : modelAlias(unpinned)
408 const at = basis === null ? '' : ` @ ${basis}`
409 return { model: null, shown: `model (${agent}; router: ${router}${at})`, failure }
410}
411
412/**
413 * Reports what the subagent runs on. An unchanged one shows the model the
414 * engine resolved, since its definition may name one the router can't see.
415 */
416function reportSpawn($: Engine, rt: Runtime, e: SpawnEvent, sent: SpawnRouting, started: AgentSpawnResult): void {
417 const ranOn = started.deny === undefined ? started.model : undefined
418 const parts: string[] = []
419 if (sent.shown !== null) parts.push(sent.shown)
420 else {
421 if (ranOn !== undefined) parts.push(describeMove('model', modelAlias(ranOn)))
422 if (sent.failure !== null) parts.push(sent.failure)
423 }
424 if (parts.length === 0) return
425 const line = `subagent ${e.subagentType} · ${parts.join(', ')}`
426 if (sent.failure !== null) warn($, line)
427 else tell($, rt, line)
428}
429
430// ---------------------------------------------------------------------------
431
432export const register: Register = (on, options) => {
433 const rt = createRuntime(readOptions((options ?? {}) as Record<string, unknown>))
434
435 on('session.start', async ($, e, next) => {
436 const started = await next(e)
437 // Only with a person at the prompt: a -p run's first prompt arrives at once,
438 // so a warm-up would only race it.
439 if (e.isInteractive && rt.backend === 'api' && rt.opts.warmUp) warmUp($, rt)
440 return started
441 })
442
443 on('prompt.submit', async ($, e, next) => {
444 await onPrompt($, rt, e)
445 return next(e)
446 })
447
448 on('turn.step', async function* ($, e, next) {
449 if (e.agentId) return yield* next(e)
450
451 const sameTurn = e.index > 0 && rt.turn !== null && rt.turn.id === e.turnId
452 const change = sameTurn && rt.turn ? rt.turn.change : await routeTurn($, rt, e)
453 if (!sameTurn) {
454 rt.turn = { id: e.turnId, change }
455 // Only for the status line, so a failure here never holds up or changes the request.
456 publishEffort($, e, change).catch(() => undefined)
457 }
458 return yield* next(change ? { ...e, ...change } : e)
459 })
460
461 on('agent.spawn', async ($, e, next) => {
462 setupOnce($, rt)
463 if (!rt.opts.routeSubagentModel || e.fork) return next(e)
464 apiWarningOnce($, rt)
465 if (typeof e.prompt !== 'string' || e.prompt.trim() === '') return next(e)
466 const sent = await routeSpawn($, rt, e)
467 const started = await next(sent.model === null ? e : { ...e, model: sent.model })
468 reportSpawn($, rt, e, sent, started)
469 return started
470 })
471}
472hooks/client.ts 145 lines1/**
2 * model-router — the classifier's contract, on the mod's side.
3 *
4 * Pure functions only; the engine is never touched here. It picks the backend,
5 * builds the situation's `recent` from session messages, and reads a verdict
6 * into a Decision. The wire format is defined in worker/src/types.ts;
7 * nothing here knows which model answers.
8 */
9import { TIER_ORDER } from './policy.ts'
10import type { Decision, Provider, Tier } from './policy.ts'
11
12export type RecentMessage = { role: 'user' | 'assistant'; text: string }
13
14/** POST /v1/classify request body. */
15export type Situation = {
16 source: 'main' | 'subagent'
17 prompt: string
18 recent?: RecentMessage[]
19 description?: string
20 agentType?: string
21}
22
23/** The fields of a `$.session.messages()` entry that buildRecent reads. */
24export type SessionEntry = {
25 role: 'user' | 'assistant'
26 text: string
27 toolUses: readonly { tool_use_id: string; tool: string; input?: unknown }[]
28 toolResults?: readonly { tool_use_id: string; text: string; isError: boolean }[]
29}
30
31type ToolResult = NonNullable<SessionEntry['toolResults']>[number]
32
33/** `api` when both the URL and the secret are set and the built-in classifier wasn't chosen. */
34export function selectProvider(forced: string, url: string, secret: string): Provider | null {
35 if (forced === 'builtin') return null
36 return url && secret ? 'api' : null
37}
38
39/** The classify endpoint under a base URL, whatever its trailing slashes or a pasted `/v1/classify`. */
40export function endpoint(baseUrl: string): string {
41 const base = baseUrl.replace(/\/+$/, '').replace(/\/v1\/classify$/, '')
42 return `${base}/v1/classify`
43}
44
45export function requestHeaders(secret: string): Record<string, string> {
46 return { 'content-type': 'application/json', authorization: `Bearer ${secret}` }
47}
48
49/** A tool result as one short line: its first non-blank line, marked when it failed. */
50function outcomeOf(result: ToolResult | undefined): string {
51 if (!result) return 'no result'
52 const first = result.text.split('\n').find((line) => line.trim() !== '')?.trim() ?? ''
53 const line = first.length > 120 ? `${first.slice(0, 117)}...` : first
54 if (result.isError) return `error: ${line}`
55 return line || 'ok'
56}
57
58function toRecent(entry: SessionEntry, results: ReadonlyMap<string, ToolResult>): RecentMessage {
59 const lines = entry.text.trim() ? [entry.text.trim()] : []
60 for (const use of entry.toolUses) lines.push(`[${use.tool}] ${outcomeOf(results.get(use.tool_use_id))}`)
61 return { role: entry.role, text: lines.join('\n') }
62}
63
64/**
65 * The last `n` messages before the prompt, tool calls reduced to their name
66 * and a one-line outcome, within `chars`: the oldest go first, and a lone
67 * message still over keeps its end. Undefined when `n` is 0 (prompt only).
68 */
69export function buildRecent(
70 messages: readonly SessionEntry[],
71 prompt: string,
72 n: number,
73 chars: number,
74): RecentMessage[] | undefined {
75 if (n <= 0) return undefined
76
77 // Results can sit on a later message than the call they answer.
78 const results = new Map<string, ToolResult>()
79 for (const message of messages) for (const result of message.toolResults ?? []) results.set(result.tool_use_id, result)
80
81 let list = [...messages]
82 const last = list[list.length - 1]
83 if (last && last.role === 'user' && last.text.trim() === prompt.trim()) list = list.slice(0, -1)
84
85 const recent = list
86 .map((entry) => toRecent(entry, results))
87 .filter((message) => message.text !== '')
88 .slice(-n)
89
90 let total = recent.reduce((sum, message) => sum + message.text.length, 0)
91 while (recent.length > 1 && total > chars) total -= recent.shift()?.text.length ?? 0
92 const only = recent[0]
93 if (recent.length === 1 && only && total > chars) recent[0] = { ...only, text: only.text.slice(-chars) }
94 return recent
95}
96
97const isRecord = (value: unknown): value is Record<string, unknown> =>
98 typeof value === 'object' && value !== null && !Array.isArray(value)
99
100const isUnit = (value: unknown): value is number => typeof value === 'number' && value >= 0 && value <= 1
101
102const isConfidence = (value: unknown): value is number | null => value === null || isUnit(value)
103
104const isTier = (value: unknown): value is Tier => TIER_ORDER.includes(value as Tier)
105
106/** A verdict from /v1/classify as a Decision, or null for anything incomplete or out of range. */
107export function readVerdict(text: string): Decision | null {
108 let verdict: unknown
109 try {
110 verdict = JSON.parse(text)
111 } catch {
112 return null
113 }
114 if (!isRecord(verdict) || typeof verdict.classifier !== 'string' || verdict.classifier === '') return null
115 const { tier, effort, risky } = verdict
116 if (!isRecord(tier) || !isTier(tier.value) || !isConfidence(tier.confidence)) return null
117 if (
118 !isRecord(effort) ||
119 typeof effort.level !== 'number' ||
120 !Number.isInteger(effort.level) ||
121 effort.level < 0 ||
122 effort.level > 3 ||
123 !isConfidence(effort.confidence)
124 ) {
125 return null
126 }
127 if (!isRecord(risky) || !isUnit(risky.p)) return null
128
129 return {
130 tier: tier.value,
131 confidence: tier.confidence,
132 effort: effort.level,
133 effortConfidence: effort.confidence,
134 risky: risky.p,
135 classifier: verdict.classifier,
136 upOnly: false,
137 }
138}
139
140/** The built-in classifier's label as a Decision: a tier, no confidence, no effort or risk. */
141export function builtinDecision(label: string | undefined): Decision | null {
142 if (!isTier(label)) return null
143 return { tier: label, confidence: null, risky: null, effort: null, effortConfidence: null, classifier: null, upOnly: false }
144}
145hooks/models.ts 30 lines1/**
2 * model-router — the models it knows.
3 *
4 * The one place a model is described: its short alias, the full id sent for
5 * that alias, the tier it usually belongs to, and whether it takes a
6 * reasoning effort. Which model each tier sends is a setting (/config); this
7 * list is only what those settings can name by alias.
8 */
9import type { Tier, Tiers } from './policy.ts'
10
11export type Model = {
12 /** Short name used in settings; also the word looked for in a full model id. */
13 alias: string
14 /** The full id sent for the alias; absent when the family has no id yet. */
15 id?: string
16 tier: Tier
17 effort: boolean
18}
19
20export const MODELS: readonly Model[] = [
21 { alias: 'haiku', id: 'claude-haiku-4-5-20251001', tier: 'fast', effort: false },
22 { alias: 'sonnet', id: 'claude-sonnet-5', tier: 'balanced', effort: true },
23 { alias: 'opus', id: 'claude-opus-5-5', tier: 'deep', effort: true },
24 { alias: 'fable', id: 'claude-fable-5-1', tier: 'deep', effort: true },
25 { alias: 'mythos', tier: 'deep', effort: true },
26]
27
28/** What each tier sends when its setting is left alone; plugin.json repeats these. */
29export const DEFAULT_TIERS: Readonly<Tiers> = { fast: 'haiku', balanced: 'sonnet', deep: 'opus' }
30hooks/policy.ts 278 lines1/**
2 * model-router — the decision rules.
3 *
4 * Nothing in this file talks to the engine. Given a classifier's Decision and
5 * what a request currently carries, it works out the model and effort to send,
6 * and it formats the short lines the hooks write to the log and transcript.
7 */
8import { MODELS } from './models.ts'
9import type { Model } from './models.ts'
10
11export type Provider = 'api'
12
13export type Tier = 'fast' | 'balanced' | 'deep'
14
15/** Tiers from cheapest to most capable; a tier's index is its position. */
16export const TIER_ORDER: readonly Tier[] = ['fast', 'balanced', 'deep']
17
18/** The model alias or id configured for each tier. */
19export type Tiers = { fast: string; balanced: string; deep: string }
20
21/** The reasoning-effort ladder, lowest first. `max` sits above it but is never chosen here. */
22export const EFFORT_ORDER = ['low', 'medium', 'high', 'xhigh'] as const
23
24export type Effort = (typeof EFFORT_ORDER)[number]
25
26export type Decision = {
27 tier: Tier
28 confidence: number | null
29 risky: number | null
30 /** Integer effort level 0..3, or null when the classifier gave none. */
31 effort: number | null
32 effortConfidence: number | null
33 /** The model that answered; null for the engine's built-in classifier. */
34 classifier: string | null
35 /** When true, this decision may only raise the model or effort, never lower them. */
36 upOnly: boolean
37}
38
39export type PolicyConfig = {
40 tiers: Tiers
41 minUpgradeConfidence: number
42 minDowngradeConfidence: number
43 riskyThreshold: number
44}
45
46/** What to do with a request; null in either field means "leave it as it is". */
47export type Routing = { model: string | null; effort: Effort | null; reason: string }
48
49export type Current = { model: string; effort?: string | number; pinned?: boolean }
50
51// ---------------------------------------------------------------------------
52// Ladders and names
53
54const MAX_EFFORT_RANK = 4
55
56/** An effort level as a ladder name; fractions are dropped and the result is clamped to the ladder. */
57export function effortName(level: number): Effort {
58 const whole = Number.isNaN(level) ? 0 : Math.trunc(level)
59 const index = Math.min(EFFORT_ORDER.length - 1, Math.max(0, whole))
60 return EFFORT_ORDER[index] as Effort
61}
62
63/** Where an effort sits on the ladder: 0..3 for the named rungs, 4 for `max`, null for anything else. */
64export function effortRank(effort: string | number | undefined): number | null {
65 if (typeof effort !== 'string') return null
66 if (effort === 'max') return MAX_EFFORT_RANK
67 const index = (EFFORT_ORDER as readonly string[]).indexOf(effort)
68 return index === -1 ? null : index
69}
70
71/** The listed model whose alias appears in a model id or alias, if any. */
72function knownModel(model: string): Model | undefined {
73 const id = model.toLowerCase()
74 return MODELS.find((known) => id.includes(known.alias))
75}
76
77/**
78 * A model's tier position. Configured tier values win (checked cheapest
79 * first); failing that, the listed model's usual tier; otherwise unknown.
80 */
81export function rankOf(model: string, tiers: Tiers): number | null {
82 const id = model.toLowerCase()
83 for (const [position, tier] of TIER_ORDER.entries()) {
84 const configured = tiers[tier].trim().toLowerCase()
85 if (configured !== '' && id.includes(configured)) return position
86 }
87 const known = knownModel(model)
88 return known ? TIER_ORDER.indexOf(known.tier) : null
89}
90
91/** The full model id for a short alias; anything that isn't an alias comes back untouched. */
92export function requestModelId(model: string): string {
93 const alias = model.trim().toLowerCase()
94 return MODELS.find((known) => known.alias === alias)?.id ?? model
95}
96
97/** The name a line shows for a model: its alias when it's listed, otherwise as given. */
98export function modelAlias(model: string): string {
99 return knownModel(model)?.alias ?? model
100}
101
102/** Whether a model takes a reasoning effort; one that isn't listed is assumed to. */
103export function supportsEffort(model: string): boolean {
104 return knownModel(model)?.effort ?? true
105}
106
107// ---------------------------------------------------------------------------
108// Routing
109
110/**
111 * Whether a move from `from` to `to` clears the confidence bar. Only a move
112 * down from a known position is a downgrade; anything else is an upgrade.
113 * With no confidence, upgrades go through and downgrades don't.
114 */
115function clearsBar(from: number | null, to: number, confidence: number | null, config: PolicyConfig): boolean {
116 if (from === to) return false
117 const downgrade = from !== null && to < from
118 if (confidence === null) return !downgrade
119 return confidence >= (downgrade ? config.minDowngradeConfidence : config.minUpgradeConfidence)
120}
121
122const num = (value: number | null): string => (value === null ? '?' : value.toFixed(2))
123
124/** Whether a decision's risk forces the deep tier, whatever else it says. */
125function riskForced(decision: Decision, config: PolicyConfig): boolean {
126 return decision.risky !== null && decision.risky > config.riskyThreshold
127}
128
129/** Route one request: which model and effort the decision asks for, given what it carries now. */
130export function route(decision: Decision | null, current: Current, config: PolicyConfig): Routing {
131 if (!decision) return { model: null, effort: null, reason: 'no decision, request left as is' }
132
133 const forced = riskForced(decision, config)
134 const tier: Tier = forced ? 'deep' : decision.tier
135 const level = forced ? Math.max(decision.effort ?? 0, 2) : decision.effort
136
137 // Model.
138 const wantedModel = config.tiers[tier]
139 let model: string | null = null
140 let heldByPin = false
141 if (wantedModel !== '' && wantedModel !== current.model) {
142 const from = rankOf(current.model, config.tiers)
143 const to = TIER_ORDER.indexOf(tier)
144 const wouldLower = from === null || to < from
145 if (forced) {
146 if (from === null || from < to) model = wantedModel
147 } else if (current.pinned) {
148 heldByPin = true
149 } else if (decision.upOnly && wouldLower) {
150 // an up-only decision can't move to a lower or unknown place
151 } else if (clearsBar(from, to, decision.confidence, config)) {
152 model = wantedModel
153 }
154 }
155
156 // Effort: only for requests that carry a named effort.
157 let effort: Effort | null = null
158 if (level !== null && typeof current.effort === 'string') {
159 const from = effortRank(current.effort)
160 let to = EFFORT_ORDER.indexOf(effortName(level))
161 if (forced && from !== null) to = Math.max(to, from)
162 const lowering = from === null || to < from
163 const blockedByUpOnly = decision.upOnly && !forced && lowering
164 if (to !== from && !blockedByUpOnly && (forced || clearsBar(from, to, decision.effortConfidence, config))) {
165 effort = EFFORT_ORDER[to] ?? null
166 }
167 }
168
169 if (model !== null || effort !== null) {
170 const why = forced
171 ? `risk ${num(decision.risky)} is over the threshold, so deep is forced`
172 : `chosen by tier ${tier}, confidence ${num(decision.confidence)}`
173 return { model, effort, reason: why }
174 }
175
176 const flags = [
177 decision.confidence === null ? 'confidence unreported' : `confidence ${num(decision.confidence)}`,
178 decision.upOnly ? 'raise-only' : null,
179 forced ? 'risk-forced' : null,
180 ]
181 .filter((flag): flag is string => flag !== null)
182 .join('; ')
183
184 let reason: string
185 if (heldByPin) {
186 reason = `agent chose ${current.model}, left in place; router preferred ${wantedModel} (${tier}; ${flags})`
187 } else {
188 const target = wantedModel === '' ? 'no model set' : wantedModel
189 const effortWanted = level === null ? '' : `, effort ${effortName(level)}`
190 const effortNow = typeof current.effort === 'string' ? ` with effort ${current.effort}` : ''
191 reason = `no switch: router leaned ${tier} (${target}${effortWanted}; ${flags}), request stays on ${current.model}${effortNow}`
192 }
193 return { model: null, effort: null, reason }
194}
195
196// ---------------------------------------------------------------------------
197// Hand-off between prompt.submit and turn.step
198
199/**
200 * The prompts submitted since the last turn began, oldest first. A turn only
201 * gets a decision when exactly one prompt preceded it; with several, there's
202 * no telling which one it answers. A null entry is a prompt whose
203 * classification failed, and it still counts as a prompt.
204 */
205export function pendingDecisions<T = Decision>(): { put(d: T | null): void; take(): T | null } {
206 let since: (T | null)[] = []
207 return {
208 put(d) {
209 // Past two entries the answer is "nothing" either way; don't grow further.
210 if (since.length < 2) since.push(d)
211 },
212 take() {
213 const only = since.length === 1 ? (since[0] ?? null) : null
214 since = []
215 return only
216 },
217 }
218}
219
220// ---------------------------------------------------------------------------
221// Log text
222
223export type Switches = { subagentModel: boolean; mainEffort: boolean; mainModel: boolean }
224
225const SWITCH_NAMES: readonly (readonly [keyof Switches, string])[] = [
226 ['subagentModel', 'subagent model'],
227 ['mainEffort', 'main effort'],
228 ['mainModel', 'main model'],
229]
230
231/** Where classifications come from and which parts of a request the router may touch. */
232export function describeSetup(
233 provider: Provider | null,
234 url: string,
235 switches: Switches,
236 pickedBuiltin = false,
237): string {
238 let source = `api ${url}`
239 if (provider !== 'api') source = `engine built-in (${pickedBuiltin ? 'picked in options' : 'no API configured'})`
240 const routes = SWITCH_NAMES.filter(([key]) => switches[key]).map(([, name]) => name)
241 return `classifier = ${source}; routes: ${routes.length === 0 ? 'none (every switch is off)' : routes.join(' + ')}`
242}
243
244/** A classifier answer on one line, e.g. `balanced@0.82, effort 2=high@0.61, risk 0.04 (jev, 41ms)`. */
245export function describeDecision(decision: Decision | null, ms: number | null): string {
246 const time = ms === null ? null : `${Math.round(ms)}ms`
247 if (!decision) return time === null ? 'nothing came back' : `nothing came back after ${time}`
248
249 let line = `${decision.tier}@${num(decision.confidence)}`
250 if (decision.effort !== null) {
251 line += `, effort ${decision.effort}=${effortName(decision.effort)}@${num(decision.effortConfidence)}`
252 }
253 if (decision.risky !== null) line += `, risk ${num(decision.risky)}`
254 if (decision.upOnly) line += ', raise-only'
255 const trailer = [decision.classifier, time].filter((bit): bit is string => bit !== null && bit !== '')
256 return trailer.length === 0 ? line : `${line} (${trailer.join(', ')})`
257}
258
259const percent = (value: number): string => `${Math.round(value * 100)}%`
260
261/**
262 * What a change is credited to: the risk when it forced the deep tier,
263 * otherwise the confidence given; null when neither is known.
264 */
265export function changeBasis(decision: Decision | null, confidence: number | null, config: PolicyConfig): string | null {
266 if (decision && riskForced(decision, config)) return `risk ${percent(decision.risky ?? 0)}`
267 return confidence === null ? null : percent(confidence)
268}
269
270/**
271 * One part of a routed line: `model (sonnet)` when it stays,
272 * `model (sonnet → haiku @ 90%)` when it changes.
273 */
274export function describeMove(name: string, from: string, to: string | null = null, basis: string | null = null): string {
275 if (to === null) return `${name} (${from})`
276 return `${name} (${from} → ${to}${basis === null ? '' : ` @ ${basis}`})`
277}
278