SLOPSHOPPER

model-router

Picks the model and the reasoning effort for each task with TypeSafe's Jev, a System One decision model reached either through TypeSafe's own API or the Vercel…

newtoaststatuspromptmodelnetwork
v0.4.2MITupdated 2026-10-09GambetaClub/model-router
A shopper browsing a rack in a slop shop
README

Model Router

A Claude Code mod that picks the model and reasoning effort for every prompt. It asks Jev, TypeSafe's small decision model, how hard the task is, then sends the turn to the model that fits.

Forked from jev-model-router in claude-code-templates.

What it does

TierModelFor
fastHaiku 5.5Small, mechanical work: read a file, rename a symbol
balancedSonnet 5.5Everyday changes across a few files
deepOpus 5.5Design, unknown bugs, security, migrations
superDeepFable 5.1Big refactors and features that span many files
  • It sets the effort for each prompt: low, medium, high or xhigh.
  • It picks the main model on the first turn and keeps it until /clear or a compaction. Switching models mid-chat would throw away the prompt cache.
  • Each subagent gets its own model. Subagents start fresh, so a different model costs nothing there.
  • When a prompt has nothing to do with the current chat, the mod holds it back and suggests a new session, with a ready-to-paste claude --model … '<prompt>' command. Send the prompt again to run it in the current chat.
  • A task that would change production, move money or destroy data gets at least deep and high effort.

Install

git clone https://github.com/GambetaClub/model-router .claude/skills/model-router

You need Claude Code 2.1.287 or newer. On an older version, start it with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude. Accept the trust prompt the first time you open the project.

Setup

Add this to ~/.claude/settings.json (user settings, not project settings):

{
  "pluginConfigs": {
    "model-router@skills-dir": {
      "options": { "typesafeApiKey": "…", "timeoutMs": 2000 }
    }
  }
}

With no key, the mod falls back to Claude Code's built-in classifier (Haiku). It works, but the classifier only sees the tier names, gives no confidence score, never moves you to a smaller model, and can't do the new-task check.

What you'll see

[Jev Model Router] jev: tier fast (0.99) · effort 0.0 → low (1.00) · risky 0.07 · 396ms
[Jev Model Router] main loop → claude-haiku-5-5, effort low: fast (confidence 0.99)

The status line under the prompt reads like Haiku 5.5 · low effort · fast 99%.

/model and the model's own answer will keep naming your session model. The mod changes each request, not the setting. To see which model really answered, check the session log:

jq -r 'select(.type=="assistant") | .message.model' ~/.claude/projects/<project>/<session>.jsonl | uniq -c

Options

OptionDefaultWhat it does
typesafeApiKey / gatewayApiKeynoneJev through TypeSafe (preferred, gives confidence) or the Vercel AI Gateway
fastModel, balancedModel, deepModel, superDeepModelhaiku, sonnet, opus, fableAlias or full model id per tier
routeMainModeltruePick the main chat's model from its first prompt
routeMainEfforttruePick the main chat's effort
routeSubagentModeltruePick each subagent's model
minUpgradeConfidence / minDowngradeConfidence0.3 / 0.6How sure Jev must be to move up or down
timeoutMs800Wait limit per classification; 2000 avoids timeouts on the first call
logDecisionstrueShow the log lines above

Privacy

With a key set, your prompt goes to TypeSafe or Vercel, along with up to five earlier prompts (500 characters each) for the new-task check. For subagents it sends their prompt, description and type. With no key, nothing goes anywhere except Anthropic.

Tests

claude plugin test .
node eval/eval.ts

eval/eval.ts checks the first-prompt pick and the no-switch rule: real Jev calls through the hook, plus an audit of your past sessions. It needs Node 22.6 or newer.

MIT licensed.

Source 2 files
hooks/model-router.ts 404 lines
1/**
2 * model-router — Claude Mod
3 *
4 * Picks the model each task runs on with TypeSafe's Jev, a System One
5 * decision model: unstructured state in, a typed choice with a probability
6 * distribution out.
7 *
8 * Jev is reached one of two ways, whichever key is configured: TypeSafe's
9 * own API (`typesafeApiKey`), which reports a calibrated confidence per
10 * answer, or the Vercel AI Gateway (`gatewayApiKey`), which does not. With
11 * neither, the engine's own `$.model.classify` stands in, so the mod is
12 * useful without any account.
13 *
14 * Three things it can set, each on its own switch:
15 *   agent.spawn  — the model of each subagent (on by default)
16 *   turn.step    — the reasoning effort of the main loop (on by default)
17 *   turn.step    — the model of the main loop (off by default: switching
18 *                  models mid-session invalidates the prompt cache, which can
19 *                  cost more than the cheaper tier saves)
20 *
21 * Every one of them moves in both directions: a task the decision model reads
22 * as mechanical is routed down, one it reads as hard is routed up. The two
23 * mistakes do not cost the same, so they do not clear the same confidence bar
24 * (see `minUpgradeConfidence` / `minDowngradeConfidence` in policy.ts).
25 *
26 * The Agent tool has no effort parameter, so a subagent's effort is not ours
27 * to set; only its model is.
28 *
29 * The prompt is classified at `prompt.submit`, which runs before the turn
30 * starts, and the decision is applied at the turn's first request.
31 *
32 * Every failure path is fail-open: a classification that errors or runs past
33 * the latency budget leaves the request exactly as the engine built it.
34 *
35 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
36 * or "gatewayApiKey"). Never hardcode it in this file.
37 *
38 * Needs Claude Code >= 2.1.287. Typed
39 * against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
40 *
41 * Privacy: with a key set, the prompt text is sent to whichever backend the
42 * key belongs to.
43 */
44import type { Register } from 'claude-code'
45import {
46  DEFAULT_BASE_URL,
47  DEFAULT_MODEL,
48  describeDecision,
49  describeSetup,
50  describeStatus,
51  endpoint,
52  pendingDecisions,
53  readDecision,
54  selectProvider,
55  requestBody,
56  requestHeaders,
57  requestModelId,
58  route,
59  TIER_ORDER,
60  bareCommand,
61  NEW_TASK_THRESHOLD,
62  newSessionAdvice,
63} from './policy.ts'
64import type { Decision, Effort, PolicyConfig, Provider, Tier } from './policy.ts'
65
66export const register: Register = (on, options) => {
67  const text = (key: string, fallback: string) =>
68    typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
69  const number = (key: string, fallback: number) =>
70    typeof options[key] === 'number' ? (options[key] as number) : fallback
71  const flag = (key: string, fallback: boolean) =>
72    typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
73
74  // TypeSafe's own API is preferred when both keys are set: it is the only
75  // one that reports a calibrated confidence, which the policy's threshold
76  // reads. `provider` forces one, including "builtin" to use neither.
77  const typesafeKey = text('typesafeApiKey', '')
78  const gatewayKey = text('gatewayApiKey', '')
79  const forced = text('provider', 'auto')
80  const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey)
81
82  // Each backend keeps its own URL and model, so an override written for one
83  // can never be sent to the other when `auto` picks differently than expected.
84  const apiKey = active === 'typesafe' ? typesafeKey : active === 'gateway' ? gatewayKey : ''
85  const modelId = !active
86    ? ''
87    : active === 'typesafe'
88      ? text('typesafeModel', DEFAULT_MODEL.typesafe)
89      : text('gatewayModel', DEFAULT_MODEL.gateway)
90  const url = !active
91    ? ''
92    : active === 'typesafe'
93      ? endpoint('typesafe', text('typesafeBaseUrl', DEFAULT_BASE_URL.typesafe))
94      : endpoint('gateway', text('gatewayBaseUrl', DEFAULT_BASE_URL.gateway))
95
96  // A backend named in the options but missing its key degrades to the
97  // built-in classifier, which is silent; say so once, when a hook first runs.
98  let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
99
100  const timeoutMs = number('timeoutMs', 800)
101  const routeSubagentModel = flag('routeSubagentModel', true)
102  const routeMainEffort = flag('routeMainEffort', true)
103  const routeMainModel = flag('routeMainModel', true)
104  const routeMainLoop = routeMainEffort || routeMainModel
105  const logDecisions = flag('logDecisions', true)
106
107  const policy: PolicyConfig = {
108    tiers: {
109      fast: text('fastModel', 'haiku'),
110      balanced: text('balancedModel', 'sonnet'),
111      deep: text('deepModel', 'opus'),
112      superDeep: text('superDeepModel', 'fable'),
113    },
114    minUpgradeConfidence: number('minUpgradeConfidence', 0.3),
115    minDowngradeConfidence: number('minDowngradeConfidence', 0.6),
116  }
117
118  // The classification waiting for the turn that reads its prompt, and what
119  // the current turn settled on. Both are single slots: main-loop turns run
120  // one at a time, so nothing accumulates over a long session. `pending`
121  // reports no decision when two prompts are waiting at once, rather than
122  // routing a turn on a decision made for a different prompt.
123  const pending = pendingDecisions()
124  // Said once, the first time a hook runs. A router that loaded and one that
125  // never loaded are otherwise told apart only by the absence of later lines,
126  // and absence is not evidence: the policy leaves most turns alone anyway.
127  let announced = false
128  let appliedTurnId: string | undefined
129  let applied: { model?: string; effort?: Effort } | null = null
130  // The main model is picked on a conversation's first turn and kept: a switch later re-reads it all uncached.
131  let modelSettled = false
132  let sessionModel: string | undefined
133  // This session's prompts, so Jev can tell a new task from the current one.
134  let earlier: string[] = []
135  let held: string | undefined
136  let stopNextTurn: string[] | undefined
137
138  on('session.end', ($, e, next) => {
139    modelSettled = false
140    sessionModel = undefined
141    earlier = []
142    held = undefined
143    stopNextTurn = undefined
144    return next(e)
145  })
146  on('session.compact', ($, e, next) => {
147    modelSettled = false
148    sessionModel = undefined
149    return next(e)
150  })
151
152  on('prompt.submit', async ($, e, next) => {
153    // Before the routing guards: a module whose switches are all off has still
154    // loaded, and that is exactly when its silence is most misleading.
155    if (!announced) {
156      announced = true
157      if (logDecisions) {
158        $.ui.log(
159          `[Jev Model Router] ${describeSetup(
160            active,
161            url,
162            {
163              subagentModel: routeSubagentModel,
164              mainEffort: routeMainEffort,
165              mainModel: routeMainModel,
166            },
167            forced === 'builtin',
168          )}`,
169        )
170      }
171    }
172    if (!routeMainLoop) return next(e)
173
174    if (!unusableReported) {
175      unusableReported = true
176      $.ui.log(`[Jev Model Router] provider "${forced}" has no key set; using the built-in classifier`)
177    }
178
179    // A slash command alone gives the decision model only the command's name.
180    // Its turn keeps the session's model and effort; the null put keeps a
181    // previous prompt's decision from reaching it.
182    if (bareCommand(e.text)) {
183      if (logDecisions) $.ui.log('[Jev Model Router] a command with nothing after it; leaving the turn alone')
184      pending.put(null)
185      return next(e)
186    }
187
188    const startedAt = await $.clock.now()
189    let decision: Decision | null = null
190    if (active) {
191      try {
192        const response = await Promise.race([
193          $.http.fetch(url, {
194            method: 'POST',
195            headers: requestHeaders(active, apiKey, modelId),
196            body: requestBody(
197              active,
198              earlier.length > 0
199                ? { prompt: e.text, earlierPrompts: earlier.slice(-5).map(text => text.slice(0, 500)) }
200                : { prompt: e.text },
201              modelId,
202            ),
203          }),
204          $.clock.sleep(timeoutMs),
205        ])
206        if (response && response.ok) decision = readDecision(response.text)
207        else if (response) $.ui.log(`[Jev Model Router] ${active} responded ${response.status}`)
208        else $.ui.log(`[Jev Model Router] classification passed ${timeoutMs}ms; leaving the turn alone`)
209      } catch (error) {
210        $.ui.log(`[Jev Model Router] classification failed: ${String(error)}`)
211      }
212    } else {
213      // No backend: the engine's own small-model classifier answers the same
214      // question, without the confidence the policy's threshold reads.
215      try {
216        const label = await $.model.classify(e.text, TIER_ORDER)
217        if (label) {
218          decision = {
219            tier: label as Tier,
220            confidence: null,
221            risky: null,
222            effort: null,
223            effortConfidence: null,
224          }
225        }
226      } catch (error) {
227        $.ui.log(`[Jev Model Router] built-in classifier failed: ${String(error)}`)
228      }
229    }
230
231    // What the decision model actually answered, whatever the policy then
232    // does with it. This is the line that proves the classification ran.
233    if (logDecisions) {
234      const ms = (await $.clock.now()) - startedAt
235      $.ui.log(`[Jev Model Router] jev: ${describeDecision(decision, ms)}`)
236    }
237
238    const resent = held === e.text
239    held = undefined
240    if (
241      !resent &&
242      e.origin.kind === 'composer' &&
243      decision?.newTask != null &&
244      decision.newTask >= NEW_TASK_THRESHOLD
245    ) {
246      held = e.text
247      // Keep the prompt in the chat; turn.start shows this under it and cancels the turn.
248      stopNextTurn = newSessionAdvice(e.text, decision.tier, policy.tiers[decision.tier])
249      return next(e)
250    }
251
252    earlier.push(e.text)
253    pending.put(decision)
254    return next(e)
255  })
256
257  on('turn.start', async ($, e, next) => {
258    const result = await next(e)
259    if (stopNextTurn) {
260      const advice = stopNextTurn
261      stopNextTurn = undefined
262      for (const line of advice) $.ui.log(`[Jev Model Router] ${line}`)
263      $.ui.toast('Jev: new task. /clear and resend it, or send it again to run it here.', { timeoutMs: 15000 })
264      $.ui.status(advice[0])
265      await $.turn.abort({ turnId: e.turnId })
266    }
267    return result
268  })
269
270  on('turn.step', async function* ($, e, next) {
271    if (!routeMainLoop || e.agentId) return yield* next(e)
272
273    // Every request after the first reuses what the turn settled on, so
274    // neither the model nor the effort changes under its own tool loop.
275    if (e.index > 0 && e.turnId === appliedTurnId) {
276      return yield* next(applied ? { ...e, ...applied } : e)
277    }
278
279    const decision = pending.take()
280    const routing = route(decision, { model: e.model, effort: e.effort }, policy)
281    const change: { model?: string; effort?: Effort } = {}
282    // The main loop's `model` is sent to the API as written, so an alias
283    // becomes its id here; a subagent's (agent.spawn) may stay an alias.
284    if (routeMainModel && routing.model && !modelSettled) sessionModel = requestModelId(routing.model)
285    if (routeMainModel && sessionModel && sessionModel !== e.model) change.model = sessionModel
286    const keptMidway = routeMainModel && routing.model && modelSettled
287    modelSettled = true
288    if (routeMainEffort && routing.effort) change.effort = routing.effort
289
290    appliedTurnId = e.turnId
291    applied = Object.keys(change).length > 0 ? change : null
292    // A row in the transcript scrolls away; this line stays on screen.
293    if (logDecisions && decision) $.ui.status(describeStatus(decision, applied, { model: e.model, effort: e.effort }))
294
295    if (!applied) {
296      // A turn left alone is the common case, and it used to be silent, which
297      // made a working mod look like one that never loaded. Say what happened.
298      if (logDecisions) {
299        const suppressed =
300          routing.model && !routeMainModel
301            ? ' (main-loop model routing off)'
302            : keptMidway
303              ? ' (model kept mid-conversation)'
304              : ''
305        $.ui.log(`[Jev Model Router] main loop: ${routing.reason}${suppressed}`)
306      }
307      return yield* next(e)
308    }
309    if (logDecisions) {
310      const what = [change.model, change.effort && `effort ${change.effort}`]
311        .filter(Boolean)
312        .join(', ')
313      $.ui.log(`[Jev Model Router] main loop → ${what}: ${routing.reason}${keptMidway ? ' (model kept mid-conversation)' : ''}`)
314    }
315    return yield* next({ ...e, ...change })
316  })
317
318  on('agent.spawn', async ($, e, next) => {
319    // Before the routing guards: a module whose switches are all off has still
320    // loaded, and that is exactly when its silence is most misleading.
321    if (!announced) {
322      announced = true
323      if (logDecisions) {
324        $.ui.log(
325          `[Jev Model Router] ${describeSetup(
326            active,
327            url,
328            {
329              subagentModel: routeSubagentModel,
330              mainEffort: routeMainEffort,
331              mainModel: routeMainModel,
332            },
333            forced === 'builtin',
334          )}`,
335        )
336      }
337    }
338
339    // A fork inherits its parent's model; `model` is ignored for it.
340    if (!routeSubagentModel || e.fork) return next(e)
341
342    if (!unusableReported) {
343      unusableReported = true
344      $.ui.log(`[Jev Model Router] provider "${forced}" has no key set; using the built-in classifier`)
345    }
346
347    const startedAt = await $.clock.now()
348    let decision: Decision | null = null
349    if (active) {
350      try {
351        const response = await Promise.race([
352          $.http.fetch(url, {
353            method: 'POST',
354            headers: requestHeaders(active, apiKey, modelId),
355            body: requestBody(
356              active,
357              { prompt: e.prompt, description: e.description, agentType: e.subagentType },
358              modelId,
359            ),
360          }),
361          $.clock.sleep(timeoutMs),
362        ])
363        if (response && response.ok) decision = readDecision(response.text)
364        else if (response) $.ui.log(`[Jev Model Router] ${active} responded ${response.status}`)
365        else $.ui.log(`[Jev Model Router] classification passed ${timeoutMs}ms; leaving the subagent alone`)
366      } catch (error) {
367        $.ui.log(`[Jev Model Router] classification failed: ${String(error)}`)
368      }
369    } else {
370      try {
371        const label = await $.model.classify(e.prompt, TIER_ORDER)
372        if (label) {
373          decision = {
374            tier: label as Tier,
375            confidence: null,
376            risky: null,
377            effort: null,
378            effortConfidence: null,
379          }
380        }
381      } catch (error) {
382        $.ui.log(`[Jev Model Router] built-in classifier failed: ${String(error)}`)
383      }
384    }
385
386    if (logDecisions) {
387      const ms = (await $.clock.now()) - startedAt
388      $.ui.log(`[Jev Model Router] jev (${e.subagentType}): ${describeDecision(decision, ms)}`)
389    }
390
391    // The subagent's own model wins when the caller named one; otherwise it
392    // would inherit the parent's, so that is what a change is measured from.
393    // The Agent tool takes no effort, so only the model is ours to set here.
394    const current = e.model ?? e.parentModel
395    const { model, reason } = route(decision, { model: current }, policy)
396    if (!model) {
397      if (logDecisions) $.ui.log(`[Jev Model Router] ${e.subagentType}: ${reason}`)
398      return next(e)
399    }
400    if (logDecisions) $.ui.log(`[Jev Model Router] ${e.subagentType} → ${model}: ${reason}`)
401    return next({ ...e, model })
402  })
403}
404
hooks/policy.ts 572 lines
1/**
2 * model-router — pure decision logic.
3 *
4 * No `$` and no I/O here: this module only builds the request the decision
5 * API takes, reads its answer, and turns that answer into a model id. The
6 * hooks module does every call on `$` at its own call site.
7 *
8 * Two backends speak to the same model with different wire shapes:
9 *
10 *   typesafe  POST https://api.typesafe.ai/v1/systemone
11 *             `{ model, state, questions }`; a yes/no question is a `noul`
12 *             and every answer carries its own `confidence`.
13 *   gateway   POST https://ai-gateway.vercel.sh/v4/ai/evaluation-model
14 *             `{ state, questions }` with the model in a header; a yes/no
15 *             question is a `boolean`, and there is no `confidence` field —
16 *             it has to be derived from an optional distribution.
17 *
18 * The Gateway shape is not documented publicly; it was read from
19 * @ai-sdk/gateway and @ai-sdk/provider.
20 */
21
22export type Provider = 'typesafe' | 'gateway'
23
24export type Tier = 'fast' | 'balanced' | 'deep' | 'superDeep'
25
26export interface Tiers {
27  fast: string
28  balanced: string
29  deep: string
30  superDeep: string
31}
32
33export interface Decision {
34  tier: Tier
35  /** Confidence in the tier, or null when the backend reported none. */
36  confidence: number | null
37  /** P(true) that carrying the task out would itself be costly or final. */
38  risky: number | null
39  /** 0..3 along the effort rubric, or null when absent. */
40  effort: number | null
41  /** Confidence in the effort, or null when the backend reported none. */
42  effortConfidence: number | null
43  /** P(true) that the prompt starts a task unrelated to the session's earlier prompts. */
44  newTask?: number | null
45}
46
47/** The reasoning levels a turn can ask for, cheapest first. */
48export const EFFORT_ORDER = ['low', 'medium', 'high', 'xhigh'] as const
49
50export type Effort = (typeof EFFORT_ORDER)[number]
51
52export const TIER_ORDER: readonly Tier[] = ['fast', 'balanced', 'deep', 'superDeep']
53
54/**
55 * How each tier is described to the decision model. Deliberately about the
56 * shape of the work, not about model names: the model never sees an id.
57 */
58const TIER_CRITERIA: Record<Tier, string> = {
59  fast: 'Mechanical and local: read or summarise a file, run one command, rename a symbol, answer something already in context.',
60  balanced:
61    'Ordinary engineering: implement a well-specified change across a few files, write tests, fix a clearly described bug, review a small diff.',
62  deep: 'Hard or high-stakes: architecture and design, debugging a failure whose cause is unknown, security, data migrations, concurrency, anything touching production or money.',
63  superDeep:
64    'Large work to orchestrate: a big refactor or a feature spanning many files or modules, planned and split across several steps or subagents.',
65}
66
67const EFFORT_RUBRIC = ['almost none', 'some', 'a lot', 'as much as possible'] as const
68
69export const DEFAULT_BASE_URL: Record<Provider, string> = {
70  typesafe: 'https://api.typesafe.ai',
71  gateway: 'https://ai-gateway.vercel.sh/v4/ai',
72}
73
74/**
75 * The Gateway's own protocol version, sent as `ai-gateway-protocol-version`.
76 * Tracks the `AI_GATEWAY_PROTOCOL_VERSION` of `@ai-sdk/gateway` (4.0.87).
77 */
78const AI_GATEWAY_PROTOCOL_VERSION = '0.0.1'
79
80export const DEFAULT_MODEL: Record<Provider, string> = {
81  typesafe: 'jev-latest',
82  gateway: 'typesafe-ai/jev',
83}
84
85/**
86 * Which backend a configuration asks for, or null for the built-in
87 * classifier. `auto` prefers TypeSafe, since it is the only one that reports
88 * a calibrated confidence; a forced backend whose key is missing resolves to
89 * null rather than falling through to the other one's key.
90 */
91export function selectProvider(
92  forced: string,
93  typesafeKey: string,
94  gatewayKey: string,
95): Provider | null {
96  if (forced === 'builtin') return null
97  if (forced === 'typesafe') return typesafeKey ? 'typesafe' : null
98  if (forced === 'gateway') return gatewayKey ? 'gateway' : null
99  if (typesafeKey) return 'typesafe'
100  if (gatewayKey) return 'gateway'
101  return null
102}
103
104/** The full endpoint a backend posts to. */
105export function endpoint(provider: Provider, baseUrl: string): string {
106  const root = baseUrl.replace(/\/+$/, '')
107  return provider === 'typesafe' ? `${root}/v1/systemone` : `${root}/evaluation-model`
108}
109
110/** The `questions` map, in the shape the backend's schema names. */
111export function questions(provider: Provider, withNewTask = false): Record<string, unknown> {
112  const yesNo = provider === 'typesafe' ? 'noul' : 'boolean'
113  return {
114    ...(withNewTask && {
115      newTask: {
116        type: yesNo,
117        instructions:
118          'The prompt starts a new task unrelated to the earlier prompts of this session, so nothing said in that conversation would help with it.',
119      },
120    }),
121    tier: {
122      type: 'choice',
123      instructions: 'Which is the cheapest tier that can complete this coding task well?',
124      criteria: TIER_CRITERIA,
125    },
126    effort: {
127      type: 'score',
128      instructions: 'How much step-by-step reasoning does this task need?',
129      criteria: EFFORT_RUBRIC,
130    },
131    risky: {
132      // The same question under two names: `noul` on TypeSafe's own API,
133      // `boolean` in the AI SDK's evaluation schema.
134      type: provider === 'typesafe' ? 'noul' : 'boolean',
135      // Asked about the act, not the subject. The first wording ("the task
136      // touches production, money, credentials") scored 0.96 on "add a
137      // refund endpoint that calls Stripe" — ordinary code that happens to be
138      // about money — and would have escalated it past a 0.98-confidence
139      // answer of the balanced tier.
140      instructions:
141        'Carrying out this task would itself change production, move real money, or alter data that cannot be restored. Writing or testing code that deals with such things, without running it against the real system, does not count.',
142    },
143  }
144}
145
146/** The request body. The Gateway carries the model in a header instead. */
147export function requestBody(
148  provider: Provider,
149  state: Record<string, unknown>,
150  model: string,
151): string {
152  const body =
153    provider === 'typesafe'
154      ? { model, state, questions: questions(provider, 'earlierPrompts' in state) }
155      : { state, questions: questions(provider, 'earlierPrompts' in state) }
156  return JSON.stringify(body)
157}
158
159/** The request headers. */
160export function requestHeaders(
161  provider: Provider,
162  apiKey: string,
163  model: string,
164): Record<string, string> {
165  const common = { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` }
166  if (provider === 'typesafe') return common
167  return {
168    ...common,
169    'ai-gateway-auth-method': 'api-key',
170    'ai-model-id': model,
171    // The Gateway rejects any request that does not name the protocol it
172    // speaks: 400 "Unsupported gateway protocol version". Every other header
173    // here is accepted without it, so the omission fails the whole backend.
174    'ai-gateway-protocol-version': AI_GATEWAY_PROTOCOL_VERSION,
175    'ai-evaluation-model-specification-version': '4',
176  }
177}
178
179function isTier(value: unknown): value is Tier {
180  return TIER_ORDER.includes(value as Tier)
181}
182
183/**
184 * Reads a response from either backend.
185 *
186 * TypeSafe's own API reports a `confidence` per answer and a `noul` number
187 * for a yes/no question. The Gateway reports neither: confidence has to come
188 * from the highest probability of a distribution that is itself optional, and
189 * a yes/no answer arrives as `probability`. Both are handled, and a missing
190 * confidence reads as null rather than as a number the policy would trust.
191 */
192export function readDecision(responseText: string): Decision | null {
193  let parsed: unknown
194  try {
195    parsed = JSON.parse(responseText)
196  } catch {
197    return null
198  }
199  const answers = (parsed as { answers?: Record<string, Record<string, unknown>> }).answers
200  if (!answers) return null
201
202  const tierAnswer = answers.tier
203  if (!tierAnswer || !isTier(tierAnswer.choice)) return null
204
205  const effortAnswer = answers.effort
206  const yesOf = (answer: Record<string, unknown> | undefined) =>
207    typeof answer?.noul === 'number'
208      ? answer.noul
209      : typeof answer?.probability === 'number'
210        ? answer.probability
211        : null
212  const risky = yesOf(answers.risky)
213
214  return {
215    tier: tierAnswer.choice,
216    confidence: confidenceOf(tierAnswer),
217    effort: typeof effortAnswer?.score === 'number' ? effortAnswer.score : null,
218    effortConfidence: effortAnswer ? confidenceOf(effortAnswer) : null,
219    risky,
220    newTask: yesOf(answers.newTask),
221  }
222}
223
224/** How sure Jev must be that a prompt is a new task before it is held back. */
225export const NEW_TASK_THRESHOLD = 0.8
226
227/** What to show instead of running a prompt that belongs in a new session. */
228export function newSessionAdvice(text: string, tier: Tier, model: string): string[] {
229  const quoted = `'${text.replace(/'/g, `'\\''`)}'`
230  return [
231    `New task, unrelated to this chat. Jev picks ${tier} (${model}).`,
232    'Start fresh: /clear, then send it again.',
233    `Or keep this chat and open a new tab: claude --model ${requestModelId(model)} ${quoted}`,
234    'Jev wrong? Send the same prompt again to run it here.',
235  ]
236}
237
238/**
239 * How sure an answer is. TypeSafe reports it; the Gateway does not, so there
240 * it is the highest probability of a distribution that is itself optional.
241 */
242function confidenceOf(answer: Record<string, unknown>): number | null {
243  if (typeof answer.confidence === 'number') return answer.confidence
244  const probabilities = answer.probabilities as Record<string, number> | undefined
245  const values = probabilities ? Object.values(probabilities) : []
246  return values.length > 0 ? Math.max(...values) : null
247}
248
249/** The rubric score (0..3) as a reasoning level. */
250export function effortLevel(score: number): Effort {
251  const index = Math.min(EFFORT_ORDER.length - 1, Math.max(0, Math.round(score)))
252  return EFFORT_ORDER[index] as Effort
253}
254
255/**
256 * Where a reasoning level sits on the ladder, or null when its place cannot
257 * be known. `max` is above every rung the rubric can produce, so it ranks
258 * above them without joining EFFORT_ORDER, which is also the set of values
259 * this router is allowed to ask for.
260 */
261export function effortRank(effort: string | number | undefined): number | null {
262  if (typeof effort !== 'string') return null
263  if (effort === 'max') return EFFORT_ORDER.length
264  const index = EFFORT_ORDER.indexOf(effort as Effort)
265  return index === -1 ? null : index
266}
267
268/**
269 * Where a model id sits on the tier ladder, by matching it against the
270 * configured tier names first and then the family words. Null when it matches
271 * none, in which case the change is treated as an upgrade rather than guessed
272 * at: an unrecognised id gets the gentler threshold, never the strict one.
273 */
274export function rankOf(model: string, tiers: Tiers): number | null {
275  const lowered = model.toLowerCase()
276  for (let index = 0; index < TIER_ORDER.length; index++) {
277    const tier = TIER_ORDER[index] as Tier
278    const configured = tiers[tier].toLowerCase()
279    if (configured && lowered.includes(configured)) return index
280  }
281  if (lowered.includes('haiku')) return 0
282  if (lowered.includes('sonnet')) return 1
283  if (lowered.includes('opus')) return 2
284  if (lowered.includes('fable')) return 3
285  return null
286}
287
288/**
289 * The full id a family alias names on the main loop.
290 *
291 * `agent.spawn` takes an alias (`haiku`) the way the Agent tool does, but
292 * `turn.step`'s `model` is the id the engine already resolved for the request
293 * and goes to the API as written: an alias there is refused ("There's an
294 * issue with the selected model (haiku)"). So the tiers stay aliases in the
295 * options, and only a main-loop rewrite resolves them, here.
296 */
297const ALIAS_IDS: Record<string, string> = {
298  haiku: 'claude-haiku-5-5',
299  sonnet: 'claude-sonnet-5-5',
300  opus: 'claude-opus-5-5',
301  fable: 'claude-fable-5-1',
302}
303
304/**
305 * What to write into `turn.step`'s `model`: a full id as given, or the id
306 * behind a family alias. Anything else is returned unchanged for the engine
307 * to judge.
308 */
309export function requestModelId(model: string): string {
310  return ALIAS_IDS[model.trim().toLowerCase()] ?? model
311}
312
313export interface PolicyConfig {
314  tiers: Tiers
315  /**
316   * How sure the decision must be to spend more (a bigger model, more
317   * reasoning). Being wrong here costs money, so the bar is low.
318   */
319  minUpgradeConfidence: number
320  /**
321   * How sure it must be to spend less. Being wrong here means a task handled
322   * by too small a model or too little thought, so the bar is high.
323   */
324  minDowngradeConfidence: number
325}
326
327export interface Routing {
328  /** The model to run on, or null to leave the request as it is. */
329  model: string | null
330  /** The reasoning level to ask for, or null to leave it as it is. */
331  effort: Effort | null
332  /** Why, for the log line. */
333  reason: string
334}
335
336const NOTHING: Routing = { model: null, effort: null, reason: 'no decision' }
337
338/**
339 * Whether a change of rank passes its threshold. Both directions are allowed;
340 * they just do not have to clear the same bar, because the two mistakes do not
341 * cost the same. A move whose direction cannot be told (an unrecognised
342 * current value) is treated as an upgrade.
343 */
344function allowed(
345  wanted: number,
346  current: number | null,
347  confidence: number | null,
348  config: PolicyConfig,
349): boolean {
350  if (current !== null && wanted === current) return false
351  const isDowngrade = current !== null && wanted < current
352  const bar = isDowngrade ? config.minDowngradeConfidence : config.minUpgradeConfidence
353  // A backend that reports no confidence (the Gateway without a distribution,
354  // or the built-in classifier) clears the upgrade bar but never the
355  // downgrade one: spending less on an unmeasured hunch is the bad trade.
356  if (confidence === null) return !isDowngrade
357  return confidence >= bar
358}
359
360/**
361 * Turns a decision into a model and a reasoning level, either of which may be
362 * null to leave the request as it is. Both can move in either direction.
363 */
364export function route(
365  decision: Decision | null,
366  current: { model: string; effort?: string | number },
367  config: PolicyConfig,
368): Routing {
369  if (!decision) return NOTHING
370
371  let tier = decision.tier
372  let effortScore = decision.effort
373  let forced = false
374
375  // Carrying out something final is never worth the saving: take the deep
376  // tier and real reasoning, whatever the cheaper answer said, and skip the
377  // thresholds — this is the one case that is not a confidence question.
378  if (decision.risky !== null && decision.risky > 0.7) {
379    if (TIER_ORDER.indexOf(tier) < TIER_ORDER.indexOf('deep')) tier = 'deep'
380    effortScore = Math.max(effortScore ?? 0, 2)
381    forced = true
382  }
383
384  const wantedTier = TIER_ORDER.indexOf(tier)
385  const currentTier = rankOf(current.model, config.tiers)
386  const wantedModel = config.tiers[tier]
387
388  const model =
389    wantedModel &&
390    wantedModel !== current.model &&
391    // Risk raises the model floor too; forcing never pulls superDeep down to deep.
392    (forced
393      ? currentTier === null || wantedTier > currentTier
394      : allowed(wantedTier, currentTier, decision.confidence, config))
395      ? wantedModel
396      : null
397
398  let effort: Effort | null = null
399  if (effortScore !== null) {
400    const currentRank = effortRank(current.effort)
401    let wantedRank = EFFORT_ORDER.indexOf(effortLevel(effortScore))
402
403    // Risk raises the floor; it must never lower one. Forcing only skips the
404    // thresholds, so without this clamp a task already at `xhigh` or `max`
405    // and rated mechanically simple would be pulled down to `high` with no
406    // confidence check at all — the opposite of what the rule is for.
407    if (forced && currentRank !== null) wantedRank = Math.max(wantedRank, currentRank)
408
409    // A numeric effort is the caller's own scale, not this ladder; leave it.
410    const comparable = typeof current.effort !== 'number'
411    const wanted = EFFORT_ORDER[Math.min(EFFORT_ORDER.length - 1, wantedRank)] as Effort
412    if (
413      comparable &&
414      wantedRank !== currentRank &&
415      (forced || allowed(wantedRank, currentRank, decision.effortConfidence, config))
416    ) {
417      effort = wanted
418    }
419  }
420
421  const said = decision.confidence === null ? 'confidence n/d' : `confidence ${decision.confidence.toFixed(2)}`
422
423  if (!model && !effort) {
424    // Naming what it wanted and what it kept is the whole point of this line.
425    // Without it, a mod that classified and decided to leave the request alone
426    // is indistinguishable from one that never loaded.
427    const wantedEffort = effortScore === null ? null : effortLevel(effortScore)
428    const kept = `${current.model}${current.effort === undefined ? '' : `/${current.effort}`}`
429    const modelDiffers = Boolean(wantedModel) && wantedModel !== current.model
430    const effortDiffers = wantedEffort !== null && wantedEffort !== current.effort
431    if (!modelDiffers && !effortDiffers) {
432      return { model: null, effort: null, reason: `stayed on ${kept}: already the right fit` }
433    }
434    const target = modelDiffers ? wantedModel : `effort ${wantedEffort}`
435    const sure = modelDiffers ? decision.confidence : decision.effortConfidence
436    const why =
437      sure === null
438        ? `no confidence score, so no switch to ${target}`
439        : `too unsure to switch to ${target} (${Math.round(sure * 100)}% sure)`
440    return { model: null, effort: null, reason: `stayed on ${kept}: ${why}` }
441  }
442
443  return { model, effort, reason: forced ? `${tier}, forced by risk` : `${tier} (${said})` }
444}
445
446/**
447 * Whether a prompt is a slash command and nothing else (`/simplify`). The
448 * decision model sees only the name, never the skill or command it runs:
449 * measured on TypeSafe (three calls each), `/simplify`, `/run` and `/github`
450 * came back fast with effort 0.1 to 0.5, low effort for a multi-step skill,
451 * while `/code-review` and `/security-review` came back balanced and deep.
452 * With text after the name there is a task to read, and it is classified.
453 * A one-segment path alone (`/etc`) matches too; nobody sends one as a task.
454 */
455export function bareCommand(text: string): boolean {
456  return /^\/[^\s/]+$/.test(text.trim())
457}
458
459/**
460 * Holds a prompt's classification until the turn that reads that prompt
461 * starts.
462 *
463 * Nothing ties a decision to the turn it belongs to. Prompts can be queued
464 * while the model is busy, a peer session's message can be delivered inside a
465 * running turn, and `prompt.submit` carries no turn id at all while the
466 * session is idle. So when more than one prompt is waiting, `take` reports
467 * none: running a turn on another prompt's decision is a worse outcome than
468 * not routing it, and not routing is what every other failure path here does.
469 */
470export function pendingDecisions(): {
471  put(decision: Decision | null): void
472  take(): Decision | null
473} {
474  let held: Decision | null = null
475  let waiting = 0
476
477  return {
478    put(decision) {
479      waiting += 1
480      // Past the first, which prompt a turn will read is unknowable, so the
481      // slot is emptied instead of holding a decision that may not fit.
482      held = waiting === 1 ? decision : null
483    },
484    take() {
485      const decision = waiting === 1 ? held : null
486      held = null
487      waiting = 0
488      return decision
489    },
490  }
491}
492
493/** A number for the log, or `n/d` when the backend reported none. */
494function reported(value: number | null): string {
495  return value === null ? 'n/d' : value.toFixed(2)
496}
497
498/**
499 * The one-time line that says the router is alive, which backend answers it,
500 * and which of the three switches are on.
501 *
502 * Without this, a router that loaded and a router that never loaded are told
503 * apart only by the absence of later lines, which is not evidence of anything.
504 */
505export function describeSetup(
506  provider: Provider | null,
507  url: string,
508  switches: { subagentModel: boolean; mainEffort: boolean; mainModel: boolean },
509  // `provider: "builtin"` is a choice, not a missing key. Reporting it as a
510  // credential problem sends someone hunting for a key they meant to omit.
511  builtinByChoice = false,
512): string {
513  const backend = provider
514    ? `${provider} (${url})`
515    : builtinByChoice
516      ? 'the built-in classifier, by choice'
517      : 'the built-in classifier, no key set'
518  const on = [
519    switches.subagentModel && 'subagent model',
520    switches.mainEffort && 'main effort',
521    switches.mainModel && 'main model',
522  ].filter(Boolean)
523  return `ready on ${backend}; routing ${on.length > 0 ? on.join(', ') : 'nothing, every switch is off'}`
524}
525
526/**
527 * What the decision model answered, before any policy touches it: the raw
528 * tier, effort and risk with their confidences, and how long it took.
529 *
530 * This is the line that shows the classification happened at all, separately
531 * from whether the policy then decided to act on it.
532 */
533export function describeDecision(decision: Decision | null, ms: number | null): string {
534  const took = ms === null ? '' : ` · ${Math.round(ms)}ms`
535  if (!decision) return `no answer${took}`
536
537  const parts = [`tier ${decision.tier} (${reported(decision.confidence)})`]
538  if (decision.effort !== null) {
539    parts.push(
540      `effort ${decision.effort.toFixed(1)} → ${effortLevel(decision.effort)} (${reported(decision.effortConfidence)})`,
541    )
542  }
543  if (decision.risky !== null) parts.push(`risky ${reported(decision.risky)}`)
544  if (decision.newTask != null) parts.push(`new task ${reported(decision.newTask)}`)
545  return parts.join(' · ') + took
546}
547
548/**
549 * The persistent status line: the last thing the router did, short enough to
550 * sit on screen beside the engine's own notices.
551 */
552export function describeStatus(
553  decision: Decision | null,
554  change: { model?: string; effort?: Effort } | null,
555  current: { model: string; effort?: string | number },
556): string {
557  const model = modelName(change?.model ?? current.model)
558  const effort = change?.effort ?? current.effort
559  const running = effort === undefined ? model : `${model} · ${effort} effort`
560  if (!decision) return `${running} · no answer`
561  const confidence = decision.confidence === null ? '' : ` ${Math.round(decision.confidence * 100)}%`
562  return `${running} · ${decision.tier}${confidence}`
563}
564
565/** `claude-haiku-5-5` as "Haiku 5.5"; an alias or unknown id is only capitalised. */
566export function modelName(id: string): string {
567  const match = /^claude-([a-z]+)-(\d+)(?:-(\d{1,2})(?!\d))?/.exec(id)
568  const capital = (word: string) => word.charAt(0).toUpperCase() + word.slice(1)
569  if (!match) return capital(id)
570  return `${capital(match[1] as string)} ${match[2]}${match[3] ? `.${match[3]}` : ''}`
571}
572