SLOPSHOPPER

model-router

Routes each turn and subagent to a model tier and reasoning effort using a classifier behind any server that speaks POST /v1/classify (model-router-api on…

newpromptmodelnetworktimeragents
v?no licenseupdated 2026-10-05sftinc/claude-code-model-router/plugin
A shopper browsing a rack in a slop shop
README

model-router

model-router is a Claude Code mod that decides, request by request, which model answers and how much reasoning effort it spends, in the main conversation and in subagents alike. A classifier reads the situation (the prompt, recent conversation, subagent details) and returns a tier, an effort level and a risk score. The mod turns that into a model and effort, and leaves the request unchanged whenever anything fails.

The classifier sits behind any server that speaks POST /v1/classify. This repo ships one: model-router-api, a Cloudflare Worker running TypeSafe's Jev on Workers AI (worker/). With no API configured, the mod uses Claude Code's built-in classifier for subagents only: it can move a subagent's model up, never down, and the main loop is left alone unless routeMainModel is on.

The idea comes from jev-model-router.

Requirements

  • Claude Code 2.1.281 or later, with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1
  • For the Worker: a Cloudflare account, Node 24, and wrangler

Install

The plugin lives in plugin/, and the repo is also a marketplace. From a terminal:

claude plugin marketplace add sftinc/claude-code-model-router
claude plugin install model-router@sftinc

Or inside Claude Code:

/plugin marketplace add sftinc/claude-code-model-router
/plugin install model-router@sftinc

If the repo is private, adding the marketplace needs git access to it (an SSH key or gh auth login).

Start Claude Code with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1. To run it from a clone instead:

CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir /path/to/claude-code-model-router/plugin

Configure

Options live under pluginConfigs in ~/.claude/settings.json. The key is model-router@sftinc when the plugin is installed from the marketplace, model-router when it's loaded with --plugin-dir, and model-router@skills-dir when it's loaded from .claude/skills/.

{
  "pluginConfigs": {
    "model-router": {
      "options": {
        "apiUrl": "https://model-router-api.<subdomain>.workers.dev",
        "apiSecret": "<ROUTER_SECRET>"
      }
    }
  }
}
OptionDefaultNotes
providerautoauto, api, builtin
apiUrl / apiSecretemptyBoth needed for the API
fastModel / balancedModel / deepModelhaiku / sonnet / opusAlias or full id
minUpgradeConfidence / minDowngradeConfidence0 / 0Minimum confidence for a move to a costlier / cheaper setting. The API applies its own thresholds, so 0 follows it; the built-in classifier gives no confidence, so it still never moves down
riskyThreshold0.7Above it: deep tier, at least high effort
contextMessages / contextChars30 / 60000Recent context sent with main-loop prompts
routeSubagentModeltrue
respectAgentModelstrueKeeps a model the Agent tool named, unless risk forces deep. The subagent is still classified, and its line shows both picks, such as model (agent: sonnet; router: haiku @ 97%) (* when the router would keep the agent's). A model set in an agent's definition isn't visible to the router and isn't protected
routeMainEfforttrue
routeMainModelfalseSwitching models invalidates the prompt cache
timeoutMs2000Wait budget for the classifier
warmUptrueSends one throwaway classification when an interactive session starts, so the first prompt doesn't hit a cold classifier
logDecisionstrueShows one short line per routed turn or subagent in the transcript, such as main loop · model (sonnet), effort (low → xhigh @ 90%) or subagent Explore · model (opus → haiku @ 97%), plus the warm-up; the details, including whole api replies, go to the debug log (claude --debug). A line reporting a failure always shows

A request routed to a model without effort support (Haiku) never carries an effort, whatever the switches say.

The Worker

cd worker
npm install
npx wrangler secret put ROUTER_SECRET   # generate with: openssl rand -base64 32
npm run deploy
ROUTER_API=https://model-router-api.<subdomain>.workers.dev ROUTER_SECRET=... scripts/live-check.sh

Before the first deploy, create an AI Gateway named model-router in the Cloudflare dashboard, with Workers AI set to Unified billing. COLLECT_LOG in wrangler.jsonc controls whether gateway logs store prompts and context.

Development

Two type-declaration files are generated, not committed. Until they exist, your editor reports missing types.

  • Plugin: in Claude Code, run /plugin-types plugin. It writes Claude Code's declarations to plugin/.claude/types/, which plugin/tsconfig.json reads. Run it again after updating Claude Code.
  • Worker: run npm run typecheck in worker/. It runs wrangler types, which writes worker/worker-configuration.d.ts, then typechecks.

Tests

CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude plugin test plugin   # the plugin
cd worker && npm test                                              # the Worker

plugin/scripts/e2e.sh runs real headless sessions with the plugin loaded, in plan mode so nothing gets carried out. It checks the router's decisions in each session's debug log: a trivial subagent goes to haiku, a model named on the Agent tool is kept, and a risky prompt forces the deep tier. It needs the api configured, and each scenario is a real session that costs about $0.10–0.30. Pass scenario names to run only some of them.

Privacy

With an API configured, each prompt and up to contextChars of recent conversation go to it. With COLLECT_LOG=true, the AI Gateway's logs keep them.

Source 4 files
hooks/model-router.ts 472 lines
1/**
2 * model-router — the hooks.
3 *
4 * prompt.submit classifies the main-loop prompt and leaves the answer for the
5 * next turn; turn.step applies it to the turn's first request and holds it for
6 * the rest of the turn; agent.spawn classifies and routes a subagent on the
7 * spot. Every failure leaves the request as it was.
8 */
9import type { AgentSpawnResult, Args, EngineInterface, Register } from 'claude-code'
10
11import { buildRecent, builtinDecision, endpoint, readVerdict, requestHeaders, selectProvider } from './client.ts'
12import type { RecentMessage, Situation } from './client.ts'
13import { DEFAULT_TIERS } from './models.ts'
14import {
15  TIER_ORDER,
16  changeBasis,
17  describeDecision,
18  describeMove,
19  describeSetup,
20  modelAlias,
21  pendingDecisions,
22  requestModelId,
23  route,
24  supportsEffort,
25} from './policy.ts'
26import type { Decision, Effort, PolicyConfig, Provider, Routing } from './policy.ts'
27
28/** The engine and the event payloads, as Claude Code declares them. */
29type Engine = EngineInterface
30type PromptEvent = Args<'prompt.submit'>
31type StepEvent = Args<'turn.step'>
32type SpawnEvent = Args<'agent.spawn'>
33
34type Change = { model?: string; effort?: Effort }
35
36// ---------------------------------------------------------------------------
37// Options: one table of defaults. A supplied value is used only when it has
38// the default's type, and a string must also be non-empty.
39
40type Options = {
41  provider: string
42  apiUrl: string
43  apiSecret: string
44  fastModel: string
45  balancedModel: string
46  deepModel: string
47  minUpgradeConfidence: number
48  minDowngradeConfidence: number
49  riskyThreshold: number
50  contextMessages: number
51  contextChars: number
52  routeSubagentModel: boolean
53  respectAgentModels: boolean
54  routeMainEffort: boolean
55  routeMainModel: boolean
56  timeoutMs: number
57  warmUp: boolean
58  logDecisions: boolean
59}
60
61const OPTION_DEFAULTS: Readonly<Options> = {
62  provider: 'auto',
63  apiUrl: '',
64  apiSecret: '',
65  fastModel: DEFAULT_TIERS.fast,
66  balancedModel: DEFAULT_TIERS.balanced,
67  deepModel: DEFAULT_TIERS.deep,
68  // The API applies its own thresholds before it answers, so by default a move follows it.
69  minUpgradeConfidence: 0,
70  minDowngradeConfidence: 0,
71  riskyThreshold: 0.7,
72  contextMessages: 30,
73  contextChars: 60000,
74  routeSubagentModel: true,
75  respectAgentModels: true,
76  routeMainEffort: true,
77  routeMainModel: false,
78  timeoutMs: 2000,
79  warmUp: true,
80  logDecisions: true,
81}
82
83function readOptions(given: Record<string, unknown>): Options {
84  const merged: Record<string, unknown> = { ...OPTION_DEFAULTS }
85  for (const [key, fallback] of Object.entries(OPTION_DEFAULTS)) {
86    const value = given[key]
87    const usable = typeof value === typeof fallback && value !== ''
88    if (usable) merged[key] = value
89  }
90  return merged as Options
91}
92
93// ---------------------------------------------------------------------------
94// Per-registration state
95
96type Runtime = {
97  opts: Options
98  policy: PolicyConfig
99  backend: Provider | null
100  url: string
101  mainCanChange: boolean
102  once: { setup: boolean; apiWarning: boolean; historyWarning: boolean }
103  pending: ReturnType<typeof pendingDecisions<Classified>>
104  turn: { id: string; change: Change | null } | null
105}
106
107function createRuntime(opts: Options): Runtime {
108  const backend = selectProvider(opts.provider, opts.apiUrl, opts.apiSecret)
109  return {
110    opts,
111    policy: {
112      tiers: { fast: opts.fastModel, balanced: opts.balancedModel, deep: opts.deepModel },
113      minUpgradeConfidence: opts.minUpgradeConfidence,
114      minDowngradeConfidence: opts.minDowngradeConfidence,
115      riskyThreshold: opts.riskyThreshold,
116    },
117    backend,
118    url: backend === 'api' ? endpoint(opts.apiUrl) : '',
119    // The built-in classifier gives no effort, so main effort needs the api.
120    mainCanChange: opts.routeMainModel || (opts.routeMainEffort && backend === 'api'),
121    once: { setup: false, apiWarning: false, historyWarning: false },
122    pending: pendingDecisions<Classified>(),
123    turn: null,
124  }
125}
126
127// ---------------------------------------------------------------------------
128// Logging. Every line starts with what it is about ("main loop · …"); the
129// engine already labels it with the plugin's name. The transcript gets one short line per routed turn or subagent,
130// and the warm-up's result (`tell`), so the router is visibly running without
131// cluttering the chat; the details behind them, such as whole api replies, go
132// to the debug log alone (`note`). logDecisions gates both. A line carrying a
133// failure (`warn`) is always written, so a router that couldn't do its job
134// says so.
135
136function tell($: Engine, rt: Runtime, text: string): void {
137  if (rt.opts.logDecisions) $.ui.log(text)
138}
139
140function note($: Engine, rt: Runtime, text: string): void {
141  if (rt.opts.logDecisions) $.ui.log(text, { to: 'debug' })
142}
143
144function warn($: Engine, text: string): void {
145  $.ui.log(text)
146}
147
148function setupOnce($: Engine, rt: Runtime): void {
149  if (rt.once.setup) return
150  rt.once.setup = true
151  const switches = {
152    subagentModel: rt.opts.routeSubagentModel,
153    mainEffort: rt.opts.routeMainEffort && rt.backend === 'api',
154    mainModel: rt.opts.routeMainModel,
155  }
156  note($, rt, 'setup · ' + describeSetup(rt.backend, rt.url, switches, rt.opts.provider === 'builtin'))
157}
158
159function apiWarningOnce($: Engine, rt: Runtime): void {
160  if (rt.once.apiWarning) return
161  rt.once.apiWarning = true
162  if (rt.opts.provider === 'api' && rt.backend === null) {
163    warn($, `setup · provider "api" needs both apiUrl and apiSecret; falling back to the engine's own classifier`)
164  }
165}
166
167// ---------------------------------------------------------------------------
168// Classification. Never throws; any trouble means no decision, and says why.
169
170/** A classification's outcome: the decision, or the short reason there is none. */
171type Classified = { decision: Decision | null; failure: string | null }
172
173const failed = (failure: string): Classified => ({ decision: null, failure })
174
175const errorText = (error: unknown): string => (error instanceof Error ? error.message : String(error))
176
177type Race<T> = { settled: true; value: T } | { settled: false }
178
179/** Whichever comes first: the answer or the timer. An answer after the timer is never seen. */
180function beforeTimer<T>(answer: Promise<T>, timer: Promise<void>): Promise<Race<T>> {
181  return Promise.race([
182    answer.then((value): Race<T> => ({ settled: true, value })),
183    timer.then((): Race<T> => ({ settled: false })),
184  ])
185}
186
187async function askApi($: Engine, rt: Runtime, situation: Situation, who: string): Promise<Classified> {
188  const { timeoutMs, apiSecret } = rt.opts
189  const reply = $.http.fetch(rt.url, {
190    method: 'POST',
191    headers: requestHeaders(apiSecret),
192    body: JSON.stringify(situation),
193  })
194  const outcome = await beforeTimer(reply, $.clock.sleep(timeoutMs))
195  if (!outcome.settled) return failed(`no reply from the api within ${timeoutMs}ms`)
196  const { ok, status, text } = outcome.value
197  note($, rt, `${who} · api reply: HTTP ${status} ${text}`)
198  if (!ok) return failed(`api returned HTTP ${status}`)
199  const decision = readVerdict(text)
200  return decision ? { decision, failure: null } : failed('api reply was a malformed verdict')
201}
202
203async function askBuiltin($: Engine, rt: Runtime, prompt: string, who: string): Promise<Classified> {
204  const { timeoutMs } = rt.opts
205  const label = $.model.classify(prompt, [...TIER_ORDER])
206  const outcome = await beforeTimer(label, $.clock.sleep(timeoutMs))
207  if (!outcome.settled) return failed(`no reply from the built-in classifier within ${timeoutMs}ms`)
208  note($, rt, `${who} · builtin label: ${String(outcome.value)}`)
209  return { decision: builtinDecision(outcome.value), failure: null }
210}
211
212async function classify($: Engine, rt: Runtime, situation: Situation, upOnly: boolean, who: string): Promise<Classified> {
213  try {
214    const got =
215      rt.backend === 'api' ? await askApi($, rt, situation, who) : await askBuiltin($, rt, situation.prompt, who)
216    return got.decision ? { decision: { ...got.decision, upOnly }, failure: null } : got
217  } catch (error) {
218    note($, rt, `${who} · classifier threw: ${error instanceof Error ? (error.stack ?? error.message) : String(error)}`)
219    return failed(`classifier threw: ${errorText(error)}`)
220  }
221}
222
223/** The throwaway situation a warm-up sends. */
224const WARM_UP: Situation = { source: 'main', prompt: 'warm-up' }
225
226/**
227 * Sends one throwaway classification in the background, so a classifier that
228 * has gone cold is warm again by the first prompt. The timer detaches it from
229 * the session.start dispatch; only how it went is reported.
230 */
231function warmUp($: Engine, rt: Runtime): void {
232  const { apiSecret } = rt.opts
233  $.clock.after(0, async () => {
234    const started = await $.clock.now()
235    try {
236      const reply = await $.http.fetch(rt.url, {
237        method: 'POST',
238        headers: requestHeaders(apiSecret),
239        body: JSON.stringify(WARM_UP),
240      })
241      const elapsed = (await $.clock.now()) - started
242      note($, rt, `warm-up · api reply: HTTP ${reply.status} ${reply.text}`)
243      if (reply.ok) tell($, rt, `warm-up · done in ${elapsed}ms`)
244      else warn($, `warm-up · failed (api returned HTTP ${reply.status})`)
245    } catch (error) {
246      warn($, `warm-up · failed (${errorText(error)})`)
247    }
248  })
249}
250
251/** Classifies, writes the verdict to the debug log, and returns the outcome. */
252async function classifyAndReport(
253  $: Engine,
254  rt: Runtime,
255  situation: Situation,
256  upOnly: boolean,
257  who: string,
258): Promise<Classified> {
259  const started = await $.clock.now()
260  const got = await classify($, rt, situation, upOnly, who)
261  const elapsed = (await $.clock.now()) - started
262  const via = rt.backend === 'api' ? 'api' : 'builtin'
263  note($, rt, `${who} · verdict via ${via}: ${describeDecision(got.decision, elapsed)}`)
264  return got
265}
266
267// ---------------------------------------------------------------------------
268// prompt.submit
269
270async function onPrompt($: Engine, rt: Runtime, e: PromptEvent): Promise<void> {
271  setupOnce($, rt)
272  if (!rt.mainCanChange) return
273  apiWarningOnce($, rt)
274
275  const prompt = typeof e.text === 'string' ? e.text : ''
276  if (prompt.trim() === '') {
277    rt.pending.put(null)
278    return
279  }
280
281  let recent: RecentMessage[] | undefined
282  let upOnly = false
283  if (rt.opts.contextMessages > 0) {
284    try {
285      const messages = await $.session.messages()
286      recent = buildRecent(messages, prompt, rt.opts.contextMessages, rt.opts.contextChars)
287    } catch {
288      upOnly = true
289      if (!rt.once.historyWarning) {
290        rt.once.historyWarning = true
291        warn($, 'setup · session history unavailable; prompts go to the classifier alone, raise-only')
292      }
293    }
294  }
295
296  const situation: Situation = recent ? { source: 'main', prompt, recent } : { source: 'main', prompt }
297  rt.pending.put(await classifyAndReport($, rt, situation, upOnly, 'main loop'))
298}
299
300// ---------------------------------------------------------------------------
301// turn.step
302
303type Shaped = { change: Change | null; stripped: boolean; sendsTo: string }
304
305/**
306 * What a routing becomes for this request under the switches. The effort is
307 * cleared outright (key present, value undefined) whenever the request carries
308 * one but its destination model takes none, regardless of the switches.
309 */
310function shapeChange(routing: Routing, e: StepEvent, opts: Options): Shaped {
311  const model = opts.routeMainModel && routing.model !== null ? requestModelId(routing.model) : undefined
312  const sendsTo = model ?? e.model
313  const stripped = e.effort !== undefined && !supportsEffort(sendsTo)
314
315  const modelPart = model === undefined ? null : { model }
316  let effortPart: Pick<Change, 'effort'> | null = null
317  if (stripped) effortPart = { effort: undefined }
318  else if (opts.routeMainEffort && routing.effort !== null) effortPart = { effort: routing.effort }
319
320  const change = modelPart || effortPart ? { ...modelPart, ...effortPart } : null
321  return { change, stripped, sendsTo }
322}
323
324/** The line's body: model and effort, each with an arrow when it changes. */
325function describeTurn(e: StepEvent, shaped: Shaped, decision: Decision | null, policy: PolicyConfig): string {
326  const model = shaped.change?.model
327  const parts = [
328    describeMove(
329      'model',
330      modelAlias(e.model),
331      model === undefined ? null : modelAlias(model),
332      changeBasis(decision, decision?.confidence ?? null, policy),
333    ),
334  ]
335  const effort = shaped.change?.effort
336  if (shaped.stripped) parts.push('effort (dropped)')
337  else if (effort !== undefined) {
338    const basis = changeBasis(decision, decision?.effortConfidence ?? null, policy)
339    parts.push(describeMove('effort', String(e.effort ?? 'unset'), effort, basis))
340  } else if (e.effort !== undefined) parts.push(describeMove('effort', String(e.effort)))
341  return parts.join(', ')
342}
343
344async function routeTurn($: Engine, rt: Runtime, e: StepEvent): Promise<Change | null> {
345  const { decision, failure } = rt.pending.take() ?? { decision: null, failure: null }
346  const routing = route(decision, { model: e.model, effort: e.effort }, rt.policy)
347  const shaped = shapeChange(routing, e, rt.opts)
348
349  if (shaped.change !== null || rt.mainCanChange) {
350    const line = `main loop · ${describeTurn(e, shaped, decision, rt.policy)}`
351    if (failure) warn($, `${line}, ${failure}`)
352    else tell($, rt, line)
353  }
354  if (rt.mainCanChange) {
355    const unapplied =
356      routing.model !== null && !rt.opts.routeMainModel
357        ? ` (${routing.model} not applied: main-model switching is disabled)`
358        : ''
359    note($, rt, `main loop · ${routing.reason}${unapplied}`)
360  }
361  return shaped.change
362}
363
364/**
365 * Leaves the effort this turn is sent with in ~/.claude/router/<session>.effort
366 * for a status line command to show, since the status line's own effort.level
367 * stays the session's setting. Empty when the request carries none.
368 */
369async function publishEffort($: Engine, e: StepEvent, change: Change | null): Promise<void> {
370  const effort = change && 'effort' in change ? change.effort : e.effort
371  const home = await $.env.get('HOME')
372  if (!home) return
373  await $.fs.write(`${home}/.claude/router/${await $.session.id()}.effort`, String(effort ?? ''))
374}
375
376// ---------------------------------------------------------------------------
377// agent.spawn
378
379/**
380 * Where the router sends a subagent: the model to set (null: leave it), the
381 * line's model part with any failure inside it (null: show the model the
382 * engine started it on), and the classifier's failure.
383 */
384type SpawnRouting = { model: string | null; shown: string | null; failure: string | null }
385
386async function routeSpawn($: Engine, rt: Runtime, e: SpawnEvent): Promise<SpawnRouting> {
387  const who = `subagent ${e.subagentType}`
388  const situation: Situation = { source: 'subagent', prompt: e.prompt }
389  if (typeof e.description === 'string') situation.description = e.description
390  if (typeof e.subagentType === 'string') situation.agentType = e.subagentType
391  const { decision, failure } = await classifyAndReport($, rt, situation, false, who)
392
393  const pinned = rt.opts.respectAgentModels && e.model !== undefined
394  const current = e.model ?? e.parentModel
395  const routing = route(decision, { model: current, pinned }, rt.policy)
396  note($, rt, `${who} · ${routing.reason}`)
397  const basis = changeBasis(decision, decision?.confidence ?? null, rt.policy)
398  if (routing.model !== null) {
399    return { model: routing.model, shown: describeMove('model', modelAlias(current), modelAlias(routing.model), basis), failure }
400  }
401  if (!pinned) return { model: null, shown: null, failure }
402
403  // Kept for the agent, but still show where the router would have sent it; * means the same model.
404  const agent = `agent: ${modelAlias(current)}`
405  if (!decision) return { model: null, shown: `model (${agent}; ${failure})`, failure }
406  const unpinned = route(decision, { model: current }, rt.policy).model
407  const router = unpinned === null ? '*' : modelAlias(unpinned)
408  const at = basis === null ? '' : ` @ ${basis}`
409  return { model: null, shown: `model (${agent}; router: ${router}${at})`, failure }
410}
411
412/**
413 * Reports what the subagent runs on. An unchanged one shows the model the
414 * engine resolved, since its definition may name one the router can't see.
415 */
416function reportSpawn($: Engine, rt: Runtime, e: SpawnEvent, sent: SpawnRouting, started: AgentSpawnResult): void {
417  const ranOn = started.deny === undefined ? started.model : undefined
418  const parts: string[] = []
419  if (sent.shown !== null) parts.push(sent.shown)
420  else {
421    if (ranOn !== undefined) parts.push(describeMove('model', modelAlias(ranOn)))
422    if (sent.failure !== null) parts.push(sent.failure)
423  }
424  if (parts.length === 0) return
425  const line = `subagent ${e.subagentType} · ${parts.join(', ')}`
426  if (sent.failure !== null) warn($, line)
427  else tell($, rt, line)
428}
429
430// ---------------------------------------------------------------------------
431
432export const register: Register = (on, options) => {
433  const rt = createRuntime(readOptions((options ?? {}) as Record<string, unknown>))
434
435  on('session.start', async ($, e, next) => {
436    const started = await next(e)
437    // Only with a person at the prompt: a -p run's first prompt arrives at once,
438    // so a warm-up would only race it.
439    if (e.isInteractive && rt.backend === 'api' && rt.opts.warmUp) warmUp($, rt)
440    return started
441  })
442
443  on('prompt.submit', async ($, e, next) => {
444    await onPrompt($, rt, e)
445    return next(e)
446  })
447
448  on('turn.step', async function* ($, e, next) {
449    if (e.agentId) return yield* next(e)
450
451    const sameTurn = e.index > 0 && rt.turn !== null && rt.turn.id === e.turnId
452    const change = sameTurn && rt.turn ? rt.turn.change : await routeTurn($, rt, e)
453    if (!sameTurn) {
454      rt.turn = { id: e.turnId, change }
455      // Only for the status line, so a failure here never holds up or changes the request.
456      publishEffort($, e, change).catch(() => undefined)
457    }
458    return yield* next(change ? { ...e, ...change } : e)
459  })
460
461  on('agent.spawn', async ($, e, next) => {
462    setupOnce($, rt)
463    if (!rt.opts.routeSubagentModel || e.fork) return next(e)
464    apiWarningOnce($, rt)
465    if (typeof e.prompt !== 'string' || e.prompt.trim() === '') return next(e)
466    const sent = await routeSpawn($, rt, e)
467    const started = await next(sent.model === null ? e : { ...e, model: sent.model })
468    reportSpawn($, rt, e, sent, started)
469    return started
470  })
471}
472
hooks/client.ts 145 lines
1/**
2 * model-router — the classifier's contract, on the mod's side.
3 *
4 * Pure functions only; the engine is never touched here. It picks the backend,
5 * builds the situation's `recent` from session messages, and reads a verdict
6 * into a Decision. The wire format is defined in worker/src/types.ts;
7 * nothing here knows which model answers.
8 */
9import { TIER_ORDER } from './policy.ts'
10import type { Decision, Provider, Tier } from './policy.ts'
11
12export type RecentMessage = { role: 'user' | 'assistant'; text: string }
13
14/** POST /v1/classify request body. */
15export type Situation = {
16  source: 'main' | 'subagent'
17  prompt: string
18  recent?: RecentMessage[]
19  description?: string
20  agentType?: string
21}
22
23/** The fields of a `$.session.messages()` entry that buildRecent reads. */
24export type SessionEntry = {
25  role: 'user' | 'assistant'
26  text: string
27  toolUses: readonly { tool_use_id: string; tool: string; input?: unknown }[]
28  toolResults?: readonly { tool_use_id: string; text: string; isError: boolean }[]
29}
30
31type ToolResult = NonNullable<SessionEntry['toolResults']>[number]
32
33/** `api` when both the URL and the secret are set and the built-in classifier wasn't chosen. */
34export function selectProvider(forced: string, url: string, secret: string): Provider | null {
35  if (forced === 'builtin') return null
36  return url && secret ? 'api' : null
37}
38
39/** The classify endpoint under a base URL, whatever its trailing slashes or a pasted `/v1/classify`. */
40export function endpoint(baseUrl: string): string {
41  const base = baseUrl.replace(/\/+$/, '').replace(/\/v1\/classify$/, '')
42  return `${base}/v1/classify`
43}
44
45export function requestHeaders(secret: string): Record<string, string> {
46  return { 'content-type': 'application/json', authorization: `Bearer ${secret}` }
47}
48
49/** A tool result as one short line: its first non-blank line, marked when it failed. */
50function outcomeOf(result: ToolResult | undefined): string {
51  if (!result) return 'no result'
52  const first = result.text.split('\n').find((line) => line.trim() !== '')?.trim() ?? ''
53  const line = first.length > 120 ? `${first.slice(0, 117)}...` : first
54  if (result.isError) return `error: ${line}`
55  return line || 'ok'
56}
57
58function toRecent(entry: SessionEntry, results: ReadonlyMap<string, ToolResult>): RecentMessage {
59  const lines = entry.text.trim() ? [entry.text.trim()] : []
60  for (const use of entry.toolUses) lines.push(`[${use.tool}] ${outcomeOf(results.get(use.tool_use_id))}`)
61  return { role: entry.role, text: lines.join('\n') }
62}
63
64/**
65 * The last `n` messages before the prompt, tool calls reduced to their name
66 * and a one-line outcome, within `chars`: the oldest go first, and a lone
67 * message still over keeps its end. Undefined when `n` is 0 (prompt only).
68 */
69export function buildRecent(
70  messages: readonly SessionEntry[],
71  prompt: string,
72  n: number,
73  chars: number,
74): RecentMessage[] | undefined {
75  if (n <= 0) return undefined
76
77  // Results can sit on a later message than the call they answer.
78  const results = new Map<string, ToolResult>()
79  for (const message of messages) for (const result of message.toolResults ?? []) results.set(result.tool_use_id, result)
80
81  let list = [...messages]
82  const last = list[list.length - 1]
83  if (last && last.role === 'user' && last.text.trim() === prompt.trim()) list = list.slice(0, -1)
84
85  const recent = list
86    .map((entry) => toRecent(entry, results))
87    .filter((message) => message.text !== '')
88    .slice(-n)
89
90  let total = recent.reduce((sum, message) => sum + message.text.length, 0)
91  while (recent.length > 1 && total > chars) total -= recent.shift()?.text.length ?? 0
92  const only = recent[0]
93  if (recent.length === 1 && only && total > chars) recent[0] = { ...only, text: only.text.slice(-chars) }
94  return recent
95}
96
97const isRecord = (value: unknown): value is Record<string, unknown> =>
98  typeof value === 'object' && value !== null && !Array.isArray(value)
99
100const isUnit = (value: unknown): value is number => typeof value === 'number' && value >= 0 && value <= 1
101
102const isConfidence = (value: unknown): value is number | null => value === null || isUnit(value)
103
104const isTier = (value: unknown): value is Tier => TIER_ORDER.includes(value as Tier)
105
106/** A verdict from /v1/classify as a Decision, or null for anything incomplete or out of range. */
107export function readVerdict(text: string): Decision | null {
108  let verdict: unknown
109  try {
110    verdict = JSON.parse(text)
111  } catch {
112    return null
113  }
114  if (!isRecord(verdict) || typeof verdict.classifier !== 'string' || verdict.classifier === '') return null
115  const { tier, effort, risky } = verdict
116  if (!isRecord(tier) || !isTier(tier.value) || !isConfidence(tier.confidence)) return null
117  if (
118    !isRecord(effort) ||
119    typeof effort.level !== 'number' ||
120    !Number.isInteger(effort.level) ||
121    effort.level < 0 ||
122    effort.level > 3 ||
123    !isConfidence(effort.confidence)
124  ) {
125    return null
126  }
127  if (!isRecord(risky) || !isUnit(risky.p)) return null
128
129  return {
130    tier: tier.value,
131    confidence: tier.confidence,
132    effort: effort.level,
133    effortConfidence: effort.confidence,
134    risky: risky.p,
135    classifier: verdict.classifier,
136    upOnly: false,
137  }
138}
139
140/** The built-in classifier's label as a Decision: a tier, no confidence, no effort or risk. */
141export function builtinDecision(label: string | undefined): Decision | null {
142  if (!isTier(label)) return null
143  return { tier: label, confidence: null, risky: null, effort: null, effortConfidence: null, classifier: null, upOnly: false }
144}
145
hooks/models.ts 30 lines
1/**
2 * model-router — the models it knows.
3 *
4 * The one place a model is described: its short alias, the full id sent for
5 * that alias, the tier it usually belongs to, and whether it takes a
6 * reasoning effort. Which model each tier sends is a setting (/config); this
7 * list is only what those settings can name by alias.
8 */
9import type { Tier, Tiers } from './policy.ts'
10
11export type Model = {
12  /** Short name used in settings; also the word looked for in a full model id. */
13  alias: string
14  /** The full id sent for the alias; absent when the family has no id yet. */
15  id?: string
16  tier: Tier
17  effort: boolean
18}
19
20export const MODELS: readonly Model[] = [
21  { alias: 'haiku', id: 'claude-haiku-4-5-20251001', tier: 'fast', effort: false },
22  { alias: 'sonnet', id: 'claude-sonnet-5', tier: 'balanced', effort: true },
23  { alias: 'opus', id: 'claude-opus-5-5', tier: 'deep', effort: true },
24  { alias: 'fable', id: 'claude-fable-5-1', tier: 'deep', effort: true },
25  { alias: 'mythos', tier: 'deep', effort: true },
26]
27
28/** What each tier sends when its setting is left alone; plugin.json repeats these. */
29export const DEFAULT_TIERS: Readonly<Tiers> = { fast: 'haiku', balanced: 'sonnet', deep: 'opus' }
30
hooks/policy.ts 278 lines
1/**
2 * model-router — the decision rules.
3 *
4 * Nothing in this file talks to the engine. Given a classifier's Decision and
5 * what a request currently carries, it works out the model and effort to send,
6 * and it formats the short lines the hooks write to the log and transcript.
7 */
8import { MODELS } from './models.ts'
9import type { Model } from './models.ts'
10
11export type Provider = 'api'
12
13export type Tier = 'fast' | 'balanced' | 'deep'
14
15/** Tiers from cheapest to most capable; a tier's index is its position. */
16export const TIER_ORDER: readonly Tier[] = ['fast', 'balanced', 'deep']
17
18/** The model alias or id configured for each tier. */
19export type Tiers = { fast: string; balanced: string; deep: string }
20
21/** The reasoning-effort ladder, lowest first. `max` sits above it but is never chosen here. */
22export const EFFORT_ORDER = ['low', 'medium', 'high', 'xhigh'] as const
23
24export type Effort = (typeof EFFORT_ORDER)[number]
25
26export type Decision = {
27  tier: Tier
28  confidence: number | null
29  risky: number | null
30  /** Integer effort level 0..3, or null when the classifier gave none. */
31  effort: number | null
32  effortConfidence: number | null
33  /** The model that answered; null for the engine's built-in classifier. */
34  classifier: string | null
35  /** When true, this decision may only raise the model or effort, never lower them. */
36  upOnly: boolean
37}
38
39export type PolicyConfig = {
40  tiers: Tiers
41  minUpgradeConfidence: number
42  minDowngradeConfidence: number
43  riskyThreshold: number
44}
45
46/** What to do with a request; null in either field means "leave it as it is". */
47export type Routing = { model: string | null; effort: Effort | null; reason: string }
48
49export type Current = { model: string; effort?: string | number; pinned?: boolean }
50
51// ---------------------------------------------------------------------------
52// Ladders and names
53
54const MAX_EFFORT_RANK = 4
55
56/** An effort level as a ladder name; fractions are dropped and the result is clamped to the ladder. */
57export function effortName(level: number): Effort {
58  const whole = Number.isNaN(level) ? 0 : Math.trunc(level)
59  const index = Math.min(EFFORT_ORDER.length - 1, Math.max(0, whole))
60  return EFFORT_ORDER[index] as Effort
61}
62
63/** Where an effort sits on the ladder: 0..3 for the named rungs, 4 for `max`, null for anything else. */
64export function effortRank(effort: string | number | undefined): number | null {
65  if (typeof effort !== 'string') return null
66  if (effort === 'max') return MAX_EFFORT_RANK
67  const index = (EFFORT_ORDER as readonly string[]).indexOf(effort)
68  return index === -1 ? null : index
69}
70
71/** The listed model whose alias appears in a model id or alias, if any. */
72function knownModel(model: string): Model | undefined {
73  const id = model.toLowerCase()
74  return MODELS.find((known) => id.includes(known.alias))
75}
76
77/**
78 * A model's tier position. Configured tier values win (checked cheapest
79 * first); failing that, the listed model's usual tier; otherwise unknown.
80 */
81export function rankOf(model: string, tiers: Tiers): number | null {
82  const id = model.toLowerCase()
83  for (const [position, tier] of TIER_ORDER.entries()) {
84    const configured = tiers[tier].trim().toLowerCase()
85    if (configured !== '' && id.includes(configured)) return position
86  }
87  const known = knownModel(model)
88  return known ? TIER_ORDER.indexOf(known.tier) : null
89}
90
91/** The full model id for a short alias; anything that isn't an alias comes back untouched. */
92export function requestModelId(model: string): string {
93  const alias = model.trim().toLowerCase()
94  return MODELS.find((known) => known.alias === alias)?.id ?? model
95}
96
97/** The name a line shows for a model: its alias when it's listed, otherwise as given. */
98export function modelAlias(model: string): string {
99  return knownModel(model)?.alias ?? model
100}
101
102/** Whether a model takes a reasoning effort; one that isn't listed is assumed to. */
103export function supportsEffort(model: string): boolean {
104  return knownModel(model)?.effort ?? true
105}
106
107// ---------------------------------------------------------------------------
108// Routing
109
110/**
111 * Whether a move from `from` to `to` clears the confidence bar. Only a move
112 * down from a known position is a downgrade; anything else is an upgrade.
113 * With no confidence, upgrades go through and downgrades don't.
114 */
115function clearsBar(from: number | null, to: number, confidence: number | null, config: PolicyConfig): boolean {
116  if (from === to) return false
117  const downgrade = from !== null && to < from
118  if (confidence === null) return !downgrade
119  return confidence >= (downgrade ? config.minDowngradeConfidence : config.minUpgradeConfidence)
120}
121
122const num = (value: number | null): string => (value === null ? '?' : value.toFixed(2))
123
124/** Whether a decision's risk forces the deep tier, whatever else it says. */
125function riskForced(decision: Decision, config: PolicyConfig): boolean {
126  return decision.risky !== null && decision.risky > config.riskyThreshold
127}
128
129/** Route one request: which model and effort the decision asks for, given what it carries now. */
130export function route(decision: Decision | null, current: Current, config: PolicyConfig): Routing {
131  if (!decision) return { model: null, effort: null, reason: 'no decision, request left as is' }
132
133  const forced = riskForced(decision, config)
134  const tier: Tier = forced ? 'deep' : decision.tier
135  const level = forced ? Math.max(decision.effort ?? 0, 2) : decision.effort
136
137  // Model.
138  const wantedModel = config.tiers[tier]
139  let model: string | null = null
140  let heldByPin = false
141  if (wantedModel !== '' && wantedModel !== current.model) {
142    const from = rankOf(current.model, config.tiers)
143    const to = TIER_ORDER.indexOf(tier)
144    const wouldLower = from === null || to < from
145    if (forced) {
146      if (from === null || from < to) model = wantedModel
147    } else if (current.pinned) {
148      heldByPin = true
149    } else if (decision.upOnly && wouldLower) {
150      // an up-only decision can't move to a lower or unknown place
151    } else if (clearsBar(from, to, decision.confidence, config)) {
152      model = wantedModel
153    }
154  }
155
156  // Effort: only for requests that carry a named effort.
157  let effort: Effort | null = null
158  if (level !== null && typeof current.effort === 'string') {
159    const from = effortRank(current.effort)
160    let to = EFFORT_ORDER.indexOf(effortName(level))
161    if (forced && from !== null) to = Math.max(to, from)
162    const lowering = from === null || to < from
163    const blockedByUpOnly = decision.upOnly && !forced && lowering
164    if (to !== from && !blockedByUpOnly && (forced || clearsBar(from, to, decision.effortConfidence, config))) {
165      effort = EFFORT_ORDER[to] ?? null
166    }
167  }
168
169  if (model !== null || effort !== null) {
170    const why = forced
171      ? `risk ${num(decision.risky)} is over the threshold, so deep is forced`
172      : `chosen by tier ${tier}, confidence ${num(decision.confidence)}`
173    return { model, effort, reason: why }
174  }
175
176  const flags = [
177    decision.confidence === null ? 'confidence unreported' : `confidence ${num(decision.confidence)}`,
178    decision.upOnly ? 'raise-only' : null,
179    forced ? 'risk-forced' : null,
180  ]
181    .filter((flag): flag is string => flag !== null)
182    .join('; ')
183
184  let reason: string
185  if (heldByPin) {
186    reason = `agent chose ${current.model}, left in place; router preferred ${wantedModel} (${tier}; ${flags})`
187  } else {
188    const target = wantedModel === '' ? 'no model set' : wantedModel
189    const effortWanted = level === null ? '' : `, effort ${effortName(level)}`
190    const effortNow = typeof current.effort === 'string' ? ` with effort ${current.effort}` : ''
191    reason = `no switch: router leaned ${tier} (${target}${effortWanted}; ${flags}), request stays on ${current.model}${effortNow}`
192  }
193  return { model: null, effort: null, reason }
194}
195
196// ---------------------------------------------------------------------------
197// Hand-off between prompt.submit and turn.step
198
199/**
200 * The prompts submitted since the last turn began, oldest first. A turn only
201 * gets a decision when exactly one prompt preceded it; with several, there's
202 * no telling which one it answers. A null entry is a prompt whose
203 * classification failed, and it still counts as a prompt.
204 */
205export function pendingDecisions<T = Decision>(): { put(d: T | null): void; take(): T | null } {
206  let since: (T | null)[] = []
207  return {
208    put(d) {
209      // Past two entries the answer is "nothing" either way; don't grow further.
210      if (since.length < 2) since.push(d)
211    },
212    take() {
213      const only = since.length === 1 ? (since[0] ?? null) : null
214      since = []
215      return only
216    },
217  }
218}
219
220// ---------------------------------------------------------------------------
221// Log text
222
223export type Switches = { subagentModel: boolean; mainEffort: boolean; mainModel: boolean }
224
225const SWITCH_NAMES: readonly (readonly [keyof Switches, string])[] = [
226  ['subagentModel', 'subagent model'],
227  ['mainEffort', 'main effort'],
228  ['mainModel', 'main model'],
229]
230
231/** Where classifications come from and which parts of a request the router may touch. */
232export function describeSetup(
233  provider: Provider | null,
234  url: string,
235  switches: Switches,
236  pickedBuiltin = false,
237): string {
238  let source = `api ${url}`
239  if (provider !== 'api') source = `engine built-in (${pickedBuiltin ? 'picked in options' : 'no API configured'})`
240  const routes = SWITCH_NAMES.filter(([key]) => switches[key]).map(([, name]) => name)
241  return `classifier = ${source}; routes: ${routes.length === 0 ? 'none (every switch is off)' : routes.join(' + ')}`
242}
243
244/** A classifier answer on one line, e.g. `balanced@0.82, effort 2=high@0.61, risk 0.04 (jev, 41ms)`. */
245export function describeDecision(decision: Decision | null, ms: number | null): string {
246  const time = ms === null ? null : `${Math.round(ms)}ms`
247  if (!decision) return time === null ? 'nothing came back' : `nothing came back after ${time}`
248
249  let line = `${decision.tier}@${num(decision.confidence)}`
250  if (decision.effort !== null) {
251    line += `, effort ${decision.effort}=${effortName(decision.effort)}@${num(decision.effortConfidence)}`
252  }
253  if (decision.risky !== null) line += `, risk ${num(decision.risky)}`
254  if (decision.upOnly) line += ', raise-only'
255  const trailer = [decision.classifier, time].filter((bit): bit is string => bit !== null && bit !== '')
256  return trailer.length === 0 ? line : `${line} (${trailer.join(', ')})`
257}
258
259const percent = (value: number): string => `${Math.round(value * 100)}%`
260
261/**
262 * What a change is credited to: the risk when it forced the deep tier,
263 * otherwise the confidence given; null when neither is known.
264 */
265export function changeBasis(decision: Decision | null, confidence: number | null, config: PolicyConfig): string | null {
266  if (decision && riskForced(decision, config)) return `risk ${percent(decision.risky ?? 0)}`
267  return confidence === null ? null : percent(confidence)
268}
269
270/**
271 * One part of a routed line: `model (sonnet)` when it stays,
272 * `model (sonnet → haiku @ 90%)` when it changes.
273 */
274export function describeMove(name: string, from: string, to: string | null = null, basis: string | null = null): string {
275  if (to === null) return `${name} (${from})`
276  return `${name} (${from} → ${to}${basis === null ? '' : ` @ ${basis}`})`
277}
278