Picks the model and the reasoning effort for each task with TypeSafe's Jev, a System One decision model reached either through TypeSafe's own API or the Vercel…

A Claude Code mod that picks the model and reasoning effort for every prompt. It asks Jev, TypeSafe's small decision model, how hard the task is, then sends the turn to the model that fits.
Forked from jev-model-router in claude-code-templates.
| Tier | Model | For |
|---|---|---|
fast | Haiku 5.5 | Small, mechanical work: read a file, rename a symbol |
balanced | Sonnet 5.5 | Everyday changes across a few files |
deep | Opus 5.5 | Design, unknown bugs, security, migrations |
superDeep | Fable 5.1 | Big refactors and features that span many files |
low, medium, high or xhigh./clear or a compaction. Switching models mid-chat would throw away the prompt cache.claude --model … '<prompt>' command. Send the prompt again to run it in the current chat.deep and high effort.git clone https://github.com/GambetaClub/model-router .claude/skills/model-router
You need Claude Code 2.1.287 or newer. On an older version, start it with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude. Accept the trust prompt the first time you open the project.
Add this to ~/.claude/settings.json (user settings, not project settings):
{
"pluginConfigs": {
"model-router@skills-dir": {
"options": { "typesafeApiKey": "…", "timeoutMs": 2000 }
}
}
}
With no key, the mod falls back to Claude Code's built-in classifier (Haiku). It works, but the classifier only sees the tier names, gives no confidence score, never moves you to a smaller model, and can't do the new-task check.
[Jev Model Router] jev: tier fast (0.99) · effort 0.0 → low (1.00) · risky 0.07 · 396ms
[Jev Model Router] main loop → claude-haiku-5-5, effort low: fast (confidence 0.99)
The status line under the prompt reads like Haiku 5.5 · low effort · fast 99%.
/model and the model's own answer will keep naming your session model. The mod changes each request, not the setting. To see which model really answered, check the session log:
jq -r 'select(.type=="assistant") | .message.model' ~/.claude/projects/<project>/<session>.jsonl | uniq -c
| Option | Default | What it does |
|---|---|---|
typesafeApiKey / gatewayApiKey | none | Jev through TypeSafe (preferred, gives confidence) or the Vercel AI Gateway |
fastModel, balancedModel, deepModel, superDeepModel | haiku, sonnet, opus, fable | Alias or full model id per tier |
routeMainModel | true | Pick the main chat's model from its first prompt |
routeMainEffort | true | Pick the main chat's effort |
routeSubagentModel | true | Pick each subagent's model |
minUpgradeConfidence / minDowngradeConfidence | 0.3 / 0.6 | How sure Jev must be to move up or down |
timeoutMs | 800 | Wait limit per classification; 2000 avoids timeouts on the first call |
logDecisions | true | Show the log lines above |
With a key set, your prompt goes to TypeSafe or Vercel, along with up to five earlier prompts (500 characters each) for the new-task check. For subagents it sends their prompt, description and type. With no key, nothing goes anywhere except Anthropic.
claude plugin test .
node eval/eval.ts
eval/eval.ts checks the first-prompt pick and the no-switch rule: real Jev calls through the hook, plus an audit of your past sessions. It needs Node 22.6 or newer.
MIT licensed.
hooks/model-router.ts 404 lines1/**
2 * model-router — Claude Mod
3 *
4 * Picks the model each task runs on with TypeSafe's Jev, a System One
5 * decision model: unstructured state in, a typed choice with a probability
6 * distribution out.
7 *
8 * Jev is reached one of two ways, whichever key is configured: TypeSafe's
9 * own API (`typesafeApiKey`), which reports a calibrated confidence per
10 * answer, or the Vercel AI Gateway (`gatewayApiKey`), which does not. With
11 * neither, the engine's own `$.model.classify` stands in, so the mod is
12 * useful without any account.
13 *
14 * Three things it can set, each on its own switch:
15 * agent.spawn — the model of each subagent (on by default)
16 * turn.step — the reasoning effort of the main loop (on by default)
17 * turn.step — the model of the main loop (off by default: switching
18 * models mid-session invalidates the prompt cache, which can
19 * cost more than the cheaper tier saves)
20 *
21 * Every one of them moves in both directions: a task the decision model reads
22 * as mechanical is routed down, one it reads as hard is routed up. The two
23 * mistakes do not cost the same, so they do not clear the same confidence bar
24 * (see `minUpgradeConfidence` / `minDowngradeConfidence` in policy.ts).
25 *
26 * The Agent tool has no effort parameter, so a subagent's effort is not ours
27 * to set; only its model is.
28 *
29 * The prompt is classified at `prompt.submit`, which runs before the turn
30 * starts, and the decision is applied at the turn's first request.
31 *
32 * Every failure path is fail-open: a classification that errors or runs past
33 * the latency budget leaves the request exactly as the engine built it.
34 *
35 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
36 * or "gatewayApiKey"). Never hardcode it in this file.
37 *
38 * Needs Claude Code >= 2.1.287. Typed
39 * against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
40 *
41 * Privacy: with a key set, the prompt text is sent to whichever backend the
42 * key belongs to.
43 */
44import type { Register } from 'claude-code'
45import {
46 DEFAULT_BASE_URL,
47 DEFAULT_MODEL,
48 describeDecision,
49 describeSetup,
50 describeStatus,
51 endpoint,
52 pendingDecisions,
53 readDecision,
54 selectProvider,
55 requestBody,
56 requestHeaders,
57 requestModelId,
58 route,
59 TIER_ORDER,
60 bareCommand,
61 NEW_TASK_THRESHOLD,
62 newSessionAdvice,
63} from './policy.ts'
64import type { Decision, Effort, PolicyConfig, Provider, Tier } from './policy.ts'
65
66export const register: Register = (on, options) => {
67 const text = (key: string, fallback: string) =>
68 typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
69 const number = (key: string, fallback: number) =>
70 typeof options[key] === 'number' ? (options[key] as number) : fallback
71 const flag = (key: string, fallback: boolean) =>
72 typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
73
74 // TypeSafe's own API is preferred when both keys are set: it is the only
75 // one that reports a calibrated confidence, which the policy's threshold
76 // reads. `provider` forces one, including "builtin" to use neither.
77 const typesafeKey = text('typesafeApiKey', '')
78 const gatewayKey = text('gatewayApiKey', '')
79 const forced = text('provider', 'auto')
80 const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey)
81
82 // Each backend keeps its own URL and model, so an override written for one
83 // can never be sent to the other when `auto` picks differently than expected.
84 const apiKey = active === 'typesafe' ? typesafeKey : active === 'gateway' ? gatewayKey : ''
85 const modelId = !active
86 ? ''
87 : active === 'typesafe'
88 ? text('typesafeModel', DEFAULT_MODEL.typesafe)
89 : text('gatewayModel', DEFAULT_MODEL.gateway)
90 const url = !active
91 ? ''
92 : active === 'typesafe'
93 ? endpoint('typesafe', text('typesafeBaseUrl', DEFAULT_BASE_URL.typesafe))
94 : endpoint('gateway', text('gatewayBaseUrl', DEFAULT_BASE_URL.gateway))
95
96 // A backend named in the options but missing its key degrades to the
97 // built-in classifier, which is silent; say so once, when a hook first runs.
98 let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
99
100 const timeoutMs = number('timeoutMs', 800)
101 const routeSubagentModel = flag('routeSubagentModel', true)
102 const routeMainEffort = flag('routeMainEffort', true)
103 const routeMainModel = flag('routeMainModel', true)
104 const routeMainLoop = routeMainEffort || routeMainModel
105 const logDecisions = flag('logDecisions', true)
106
107 const policy: PolicyConfig = {
108 tiers: {
109 fast: text('fastModel', 'haiku'),
110 balanced: text('balancedModel', 'sonnet'),
111 deep: text('deepModel', 'opus'),
112 superDeep: text('superDeepModel', 'fable'),
113 },
114 minUpgradeConfidence: number('minUpgradeConfidence', 0.3),
115 minDowngradeConfidence: number('minDowngradeConfidence', 0.6),
116 }
117
118 // The classification waiting for the turn that reads its prompt, and what
119 // the current turn settled on. Both are single slots: main-loop turns run
120 // one at a time, so nothing accumulates over a long session. `pending`
121 // reports no decision when two prompts are waiting at once, rather than
122 // routing a turn on a decision made for a different prompt.
123 const pending = pendingDecisions()
124 // Said once, the first time a hook runs. A router that loaded and one that
125 // never loaded are otherwise told apart only by the absence of later lines,
126 // and absence is not evidence: the policy leaves most turns alone anyway.
127 let announced = false
128 let appliedTurnId: string | undefined
129 let applied: { model?: string; effort?: Effort } | null = null
130 // The main model is picked on a conversation's first turn and kept: a switch later re-reads it all uncached.
131 let modelSettled = false
132 let sessionModel: string | undefined
133 // This session's prompts, so Jev can tell a new task from the current one.
134 let earlier: string[] = []
135 let held: string | undefined
136 let stopNextTurn: string[] | undefined
137
138 on('session.end', ($, e, next) => {
139 modelSettled = false
140 sessionModel = undefined
141 earlier = []
142 held = undefined
143 stopNextTurn = undefined
144 return next(e)
145 })
146 on('session.compact', ($, e, next) => {
147 modelSettled = false
148 sessionModel = undefined
149 return next(e)
150 })
151
152 on('prompt.submit', async ($, e, next) => {
153 // Before the routing guards: a module whose switches are all off has still
154 // loaded, and that is exactly when its silence is most misleading.
155 if (!announced) {
156 announced = true
157 if (logDecisions) {
158 $.ui.log(
159 `[Jev Model Router] ${describeSetup(
160 active,
161 url,
162 {
163 subagentModel: routeSubagentModel,
164 mainEffort: routeMainEffort,
165 mainModel: routeMainModel,
166 },
167 forced === 'builtin',
168 )}`,
169 )
170 }
171 }
172 if (!routeMainLoop) return next(e)
173
174 if (!unusableReported) {
175 unusableReported = true
176 $.ui.log(`[Jev Model Router] provider "${forced}" has no key set; using the built-in classifier`)
177 }
178
179 // A slash command alone gives the decision model only the command's name.
180 // Its turn keeps the session's model and effort; the null put keeps a
181 // previous prompt's decision from reaching it.
182 if (bareCommand(e.text)) {
183 if (logDecisions) $.ui.log('[Jev Model Router] a command with nothing after it; leaving the turn alone')
184 pending.put(null)
185 return next(e)
186 }
187
188 const startedAt = await $.clock.now()
189 let decision: Decision | null = null
190 if (active) {
191 try {
192 const response = await Promise.race([
193 $.http.fetch(url, {
194 method: 'POST',
195 headers: requestHeaders(active, apiKey, modelId),
196 body: requestBody(
197 active,
198 earlier.length > 0
199 ? { prompt: e.text, earlierPrompts: earlier.slice(-5).map(text => text.slice(0, 500)) }
200 : { prompt: e.text },
201 modelId,
202 ),
203 }),
204 $.clock.sleep(timeoutMs),
205 ])
206 if (response && response.ok) decision = readDecision(response.text)
207 else if (response) $.ui.log(`[Jev Model Router] ${active} responded ${response.status}`)
208 else $.ui.log(`[Jev Model Router] classification passed ${timeoutMs}ms; leaving the turn alone`)
209 } catch (error) {
210 $.ui.log(`[Jev Model Router] classification failed: ${String(error)}`)
211 }
212 } else {
213 // No backend: the engine's own small-model classifier answers the same
214 // question, without the confidence the policy's threshold reads.
215 try {
216 const label = await $.model.classify(e.text, TIER_ORDER)
217 if (label) {
218 decision = {
219 tier: label as Tier,
220 confidence: null,
221 risky: null,
222 effort: null,
223 effortConfidence: null,
224 }
225 }
226 } catch (error) {
227 $.ui.log(`[Jev Model Router] built-in classifier failed: ${String(error)}`)
228 }
229 }
230
231 // What the decision model actually answered, whatever the policy then
232 // does with it. This is the line that proves the classification ran.
233 if (logDecisions) {
234 const ms = (await $.clock.now()) - startedAt
235 $.ui.log(`[Jev Model Router] jev: ${describeDecision(decision, ms)}`)
236 }
237
238 const resent = held === e.text
239 held = undefined
240 if (
241 !resent &&
242 e.origin.kind === 'composer' &&
243 decision?.newTask != null &&
244 decision.newTask >= NEW_TASK_THRESHOLD
245 ) {
246 held = e.text
247 // Keep the prompt in the chat; turn.start shows this under it and cancels the turn.
248 stopNextTurn = newSessionAdvice(e.text, decision.tier, policy.tiers[decision.tier])
249 return next(e)
250 }
251
252 earlier.push(e.text)
253 pending.put(decision)
254 return next(e)
255 })
256
257 on('turn.start', async ($, e, next) => {
258 const result = await next(e)
259 if (stopNextTurn) {
260 const advice = stopNextTurn
261 stopNextTurn = undefined
262 for (const line of advice) $.ui.log(`[Jev Model Router] ${line}`)
263 $.ui.toast('Jev: new task. /clear and resend it, or send it again to run it here.', { timeoutMs: 15000 })
264 $.ui.status(advice[0])
265 await $.turn.abort({ turnId: e.turnId })
266 }
267 return result
268 })
269
270 on('turn.step', async function* ($, e, next) {
271 if (!routeMainLoop || e.agentId) return yield* next(e)
272
273 // Every request after the first reuses what the turn settled on, so
274 // neither the model nor the effort changes under its own tool loop.
275 if (e.index > 0 && e.turnId === appliedTurnId) {
276 return yield* next(applied ? { ...e, ...applied } : e)
277 }
278
279 const decision = pending.take()
280 const routing = route(decision, { model: e.model, effort: e.effort }, policy)
281 const change: { model?: string; effort?: Effort } = {}
282 // The main loop's `model` is sent to the API as written, so an alias
283 // becomes its id here; a subagent's (agent.spawn) may stay an alias.
284 if (routeMainModel && routing.model && !modelSettled) sessionModel = requestModelId(routing.model)
285 if (routeMainModel && sessionModel && sessionModel !== e.model) change.model = sessionModel
286 const keptMidway = routeMainModel && routing.model && modelSettled
287 modelSettled = true
288 if (routeMainEffort && routing.effort) change.effort = routing.effort
289
290 appliedTurnId = e.turnId
291 applied = Object.keys(change).length > 0 ? change : null
292 // A row in the transcript scrolls away; this line stays on screen.
293 if (logDecisions && decision) $.ui.status(describeStatus(decision, applied, { model: e.model, effort: e.effort }))
294
295 if (!applied) {
296 // A turn left alone is the common case, and it used to be silent, which
297 // made a working mod look like one that never loaded. Say what happened.
298 if (logDecisions) {
299 const suppressed =
300 routing.model && !routeMainModel
301 ? ' (main-loop model routing off)'
302 : keptMidway
303 ? ' (model kept mid-conversation)'
304 : ''
305 $.ui.log(`[Jev Model Router] main loop: ${routing.reason}${suppressed}`)
306 }
307 return yield* next(e)
308 }
309 if (logDecisions) {
310 const what = [change.model, change.effort && `effort ${change.effort}`]
311 .filter(Boolean)
312 .join(', ')
313 $.ui.log(`[Jev Model Router] main loop → ${what}: ${routing.reason}${keptMidway ? ' (model kept mid-conversation)' : ''}`)
314 }
315 return yield* next({ ...e, ...change })
316 })
317
318 on('agent.spawn', async ($, e, next) => {
319 // Before the routing guards: a module whose switches are all off has still
320 // loaded, and that is exactly when its silence is most misleading.
321 if (!announced) {
322 announced = true
323 if (logDecisions) {
324 $.ui.log(
325 `[Jev Model Router] ${describeSetup(
326 active,
327 url,
328 {
329 subagentModel: routeSubagentModel,
330 mainEffort: routeMainEffort,
331 mainModel: routeMainModel,
332 },
333 forced === 'builtin',
334 )}`,
335 )
336 }
337 }
338
339 // A fork inherits its parent's model; `model` is ignored for it.
340 if (!routeSubagentModel || e.fork) return next(e)
341
342 if (!unusableReported) {
343 unusableReported = true
344 $.ui.log(`[Jev Model Router] provider "${forced}" has no key set; using the built-in classifier`)
345 }
346
347 const startedAt = await $.clock.now()
348 let decision: Decision | null = null
349 if (active) {
350 try {
351 const response = await Promise.race([
352 $.http.fetch(url, {
353 method: 'POST',
354 headers: requestHeaders(active, apiKey, modelId),
355 body: requestBody(
356 active,
357 { prompt: e.prompt, description: e.description, agentType: e.subagentType },
358 modelId,
359 ),
360 }),
361 $.clock.sleep(timeoutMs),
362 ])
363 if (response && response.ok) decision = readDecision(response.text)
364 else if (response) $.ui.log(`[Jev Model Router] ${active} responded ${response.status}`)
365 else $.ui.log(`[Jev Model Router] classification passed ${timeoutMs}ms; leaving the subagent alone`)
366 } catch (error) {
367 $.ui.log(`[Jev Model Router] classification failed: ${String(error)}`)
368 }
369 } else {
370 try {
371 const label = await $.model.classify(e.prompt, TIER_ORDER)
372 if (label) {
373 decision = {
374 tier: label as Tier,
375 confidence: null,
376 risky: null,
377 effort: null,
378 effortConfidence: null,
379 }
380 }
381 } catch (error) {
382 $.ui.log(`[Jev Model Router] built-in classifier failed: ${String(error)}`)
383 }
384 }
385
386 if (logDecisions) {
387 const ms = (await $.clock.now()) - startedAt
388 $.ui.log(`[Jev Model Router] jev (${e.subagentType}): ${describeDecision(decision, ms)}`)
389 }
390
391 // The subagent's own model wins when the caller named one; otherwise it
392 // would inherit the parent's, so that is what a change is measured from.
393 // The Agent tool takes no effort, so only the model is ours to set here.
394 const current = e.model ?? e.parentModel
395 const { model, reason } = route(decision, { model: current }, policy)
396 if (!model) {
397 if (logDecisions) $.ui.log(`[Jev Model Router] ${e.subagentType}: ${reason}`)
398 return next(e)
399 }
400 if (logDecisions) $.ui.log(`[Jev Model Router] ${e.subagentType} → ${model}: ${reason}`)
401 return next({ ...e, model })
402 })
403}
404hooks/policy.ts 572 lines1/**
2 * model-router — pure decision logic.
3 *
4 * No `$` and no I/O here: this module only builds the request the decision
5 * API takes, reads its answer, and turns that answer into a model id. The
6 * hooks module does every call on `$` at its own call site.
7 *
8 * Two backends speak to the same model with different wire shapes:
9 *
10 * typesafe POST https://api.typesafe.ai/v1/systemone
11 * `{ model, state, questions }`; a yes/no question is a `noul`
12 * and every answer carries its own `confidence`.
13 * gateway POST https://ai-gateway.vercel.sh/v4/ai/evaluation-model
14 * `{ state, questions }` with the model in a header; a yes/no
15 * question is a `boolean`, and there is no `confidence` field —
16 * it has to be derived from an optional distribution.
17 *
18 * The Gateway shape is not documented publicly; it was read from
19 * @ai-sdk/gateway and @ai-sdk/provider.
20 */
21
22export type Provider = 'typesafe' | 'gateway'
23
24export type Tier = 'fast' | 'balanced' | 'deep' | 'superDeep'
25
26export interface Tiers {
27 fast: string
28 balanced: string
29 deep: string
30 superDeep: string
31}
32
33export interface Decision {
34 tier: Tier
35 /** Confidence in the tier, or null when the backend reported none. */
36 confidence: number | null
37 /** P(true) that carrying the task out would itself be costly or final. */
38 risky: number | null
39 /** 0..3 along the effort rubric, or null when absent. */
40 effort: number | null
41 /** Confidence in the effort, or null when the backend reported none. */
42 effortConfidence: number | null
43 /** P(true) that the prompt starts a task unrelated to the session's earlier prompts. */
44 newTask?: number | null
45}
46
47/** The reasoning levels a turn can ask for, cheapest first. */
48export const EFFORT_ORDER = ['low', 'medium', 'high', 'xhigh'] as const
49
50export type Effort = (typeof EFFORT_ORDER)[number]
51
52export const TIER_ORDER: readonly Tier[] = ['fast', 'balanced', 'deep', 'superDeep']
53
54/**
55 * How each tier is described to the decision model. Deliberately about the
56 * shape of the work, not about model names: the model never sees an id.
57 */
58const TIER_CRITERIA: Record<Tier, string> = {
59 fast: 'Mechanical and local: read or summarise a file, run one command, rename a symbol, answer something already in context.',
60 balanced:
61 'Ordinary engineering: implement a well-specified change across a few files, write tests, fix a clearly described bug, review a small diff.',
62 deep: 'Hard or high-stakes: architecture and design, debugging a failure whose cause is unknown, security, data migrations, concurrency, anything touching production or money.',
63 superDeep:
64 'Large work to orchestrate: a big refactor or a feature spanning many files or modules, planned and split across several steps or subagents.',
65}
66
67const EFFORT_RUBRIC = ['almost none', 'some', 'a lot', 'as much as possible'] as const
68
69export const DEFAULT_BASE_URL: Record<Provider, string> = {
70 typesafe: 'https://api.typesafe.ai',
71 gateway: 'https://ai-gateway.vercel.sh/v4/ai',
72}
73
74/**
75 * The Gateway's own protocol version, sent as `ai-gateway-protocol-version`.
76 * Tracks the `AI_GATEWAY_PROTOCOL_VERSION` of `@ai-sdk/gateway` (4.0.87).
77 */
78const AI_GATEWAY_PROTOCOL_VERSION = '0.0.1'
79
80export const DEFAULT_MODEL: Record<Provider, string> = {
81 typesafe: 'jev-latest',
82 gateway: 'typesafe-ai/jev',
83}
84
85/**
86 * Which backend a configuration asks for, or null for the built-in
87 * classifier. `auto` prefers TypeSafe, since it is the only one that reports
88 * a calibrated confidence; a forced backend whose key is missing resolves to
89 * null rather than falling through to the other one's key.
90 */
91export function selectProvider(
92 forced: string,
93 typesafeKey: string,
94 gatewayKey: string,
95): Provider | null {
96 if (forced === 'builtin') return null
97 if (forced === 'typesafe') return typesafeKey ? 'typesafe' : null
98 if (forced === 'gateway') return gatewayKey ? 'gateway' : null
99 if (typesafeKey) return 'typesafe'
100 if (gatewayKey) return 'gateway'
101 return null
102}
103
104/** The full endpoint a backend posts to. */
105export function endpoint(provider: Provider, baseUrl: string): string {
106 const root = baseUrl.replace(/\/+$/, '')
107 return provider === 'typesafe' ? `${root}/v1/systemone` : `${root}/evaluation-model`
108}
109
110/** The `questions` map, in the shape the backend's schema names. */
111export function questions(provider: Provider, withNewTask = false): Record<string, unknown> {
112 const yesNo = provider === 'typesafe' ? 'noul' : 'boolean'
113 return {
114 ...(withNewTask && {
115 newTask: {
116 type: yesNo,
117 instructions:
118 'The prompt starts a new task unrelated to the earlier prompts of this session, so nothing said in that conversation would help with it.',
119 },
120 }),
121 tier: {
122 type: 'choice',
123 instructions: 'Which is the cheapest tier that can complete this coding task well?',
124 criteria: TIER_CRITERIA,
125 },
126 effort: {
127 type: 'score',
128 instructions: 'How much step-by-step reasoning does this task need?',
129 criteria: EFFORT_RUBRIC,
130 },
131 risky: {
132 // The same question under two names: `noul` on TypeSafe's own API,
133 // `boolean` in the AI SDK's evaluation schema.
134 type: provider === 'typesafe' ? 'noul' : 'boolean',
135 // Asked about the act, not the subject. The first wording ("the task
136 // touches production, money, credentials") scored 0.96 on "add a
137 // refund endpoint that calls Stripe" — ordinary code that happens to be
138 // about money — and would have escalated it past a 0.98-confidence
139 // answer of the balanced tier.
140 instructions:
141 'Carrying out this task would itself change production, move real money, or alter data that cannot be restored. Writing or testing code that deals with such things, without running it against the real system, does not count.',
142 },
143 }
144}
145
146/** The request body. The Gateway carries the model in a header instead. */
147export function requestBody(
148 provider: Provider,
149 state: Record<string, unknown>,
150 model: string,
151): string {
152 const body =
153 provider === 'typesafe'
154 ? { model, state, questions: questions(provider, 'earlierPrompts' in state) }
155 : { state, questions: questions(provider, 'earlierPrompts' in state) }
156 return JSON.stringify(body)
157}
158
159/** The request headers. */
160export function requestHeaders(
161 provider: Provider,
162 apiKey: string,
163 model: string,
164): Record<string, string> {
165 const common = { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` }
166 if (provider === 'typesafe') return common
167 return {
168 ...common,
169 'ai-gateway-auth-method': 'api-key',
170 'ai-model-id': model,
171 // The Gateway rejects any request that does not name the protocol it
172 // speaks: 400 "Unsupported gateway protocol version". Every other header
173 // here is accepted without it, so the omission fails the whole backend.
174 'ai-gateway-protocol-version': AI_GATEWAY_PROTOCOL_VERSION,
175 'ai-evaluation-model-specification-version': '4',
176 }
177}
178
179function isTier(value: unknown): value is Tier {
180 return TIER_ORDER.includes(value as Tier)
181}
182
183/**
184 * Reads a response from either backend.
185 *
186 * TypeSafe's own API reports a `confidence` per answer and a `noul` number
187 * for a yes/no question. The Gateway reports neither: confidence has to come
188 * from the highest probability of a distribution that is itself optional, and
189 * a yes/no answer arrives as `probability`. Both are handled, and a missing
190 * confidence reads as null rather than as a number the policy would trust.
191 */
192export function readDecision(responseText: string): Decision | null {
193 let parsed: unknown
194 try {
195 parsed = JSON.parse(responseText)
196 } catch {
197 return null
198 }
199 const answers = (parsed as { answers?: Record<string, Record<string, unknown>> }).answers
200 if (!answers) return null
201
202 const tierAnswer = answers.tier
203 if (!tierAnswer || !isTier(tierAnswer.choice)) return null
204
205 const effortAnswer = answers.effort
206 const yesOf = (answer: Record<string, unknown> | undefined) =>
207 typeof answer?.noul === 'number'
208 ? answer.noul
209 : typeof answer?.probability === 'number'
210 ? answer.probability
211 : null
212 const risky = yesOf(answers.risky)
213
214 return {
215 tier: tierAnswer.choice,
216 confidence: confidenceOf(tierAnswer),
217 effort: typeof effortAnswer?.score === 'number' ? effortAnswer.score : null,
218 effortConfidence: effortAnswer ? confidenceOf(effortAnswer) : null,
219 risky,
220 newTask: yesOf(answers.newTask),
221 }
222}
223
224/** How sure Jev must be that a prompt is a new task before it is held back. */
225export const NEW_TASK_THRESHOLD = 0.8
226
227/** What to show instead of running a prompt that belongs in a new session. */
228export function newSessionAdvice(text: string, tier: Tier, model: string): string[] {
229 const quoted = `'${text.replace(/'/g, `'\\''`)}'`
230 return [
231 `New task, unrelated to this chat. Jev picks ${tier} (${model}).`,
232 'Start fresh: /clear, then send it again.',
233 `Or keep this chat and open a new tab: claude --model ${requestModelId(model)} ${quoted}`,
234 'Jev wrong? Send the same prompt again to run it here.',
235 ]
236}
237
238/**
239 * How sure an answer is. TypeSafe reports it; the Gateway does not, so there
240 * it is the highest probability of a distribution that is itself optional.
241 */
242function confidenceOf(answer: Record<string, unknown>): number | null {
243 if (typeof answer.confidence === 'number') return answer.confidence
244 const probabilities = answer.probabilities as Record<string, number> | undefined
245 const values = probabilities ? Object.values(probabilities) : []
246 return values.length > 0 ? Math.max(...values) : null
247}
248
249/** The rubric score (0..3) as a reasoning level. */
250export function effortLevel(score: number): Effort {
251 const index = Math.min(EFFORT_ORDER.length - 1, Math.max(0, Math.round(score)))
252 return EFFORT_ORDER[index] as Effort
253}
254
255/**
256 * Where a reasoning level sits on the ladder, or null when its place cannot
257 * be known. `max` is above every rung the rubric can produce, so it ranks
258 * above them without joining EFFORT_ORDER, which is also the set of values
259 * this router is allowed to ask for.
260 */
261export function effortRank(effort: string | number | undefined): number | null {
262 if (typeof effort !== 'string') return null
263 if (effort === 'max') return EFFORT_ORDER.length
264 const index = EFFORT_ORDER.indexOf(effort as Effort)
265 return index === -1 ? null : index
266}
267
268/**
269 * Where a model id sits on the tier ladder, by matching it against the
270 * configured tier names first and then the family words. Null when it matches
271 * none, in which case the change is treated as an upgrade rather than guessed
272 * at: an unrecognised id gets the gentler threshold, never the strict one.
273 */
274export function rankOf(model: string, tiers: Tiers): number | null {
275 const lowered = model.toLowerCase()
276 for (let index = 0; index < TIER_ORDER.length; index++) {
277 const tier = TIER_ORDER[index] as Tier
278 const configured = tiers[tier].toLowerCase()
279 if (configured && lowered.includes(configured)) return index
280 }
281 if (lowered.includes('haiku')) return 0
282 if (lowered.includes('sonnet')) return 1
283 if (lowered.includes('opus')) return 2
284 if (lowered.includes('fable')) return 3
285 return null
286}
287
288/**
289 * The full id a family alias names on the main loop.
290 *
291 * `agent.spawn` takes an alias (`haiku`) the way the Agent tool does, but
292 * `turn.step`'s `model` is the id the engine already resolved for the request
293 * and goes to the API as written: an alias there is refused ("There's an
294 * issue with the selected model (haiku)"). So the tiers stay aliases in the
295 * options, and only a main-loop rewrite resolves them, here.
296 */
297const ALIAS_IDS: Record<string, string> = {
298 haiku: 'claude-haiku-5-5',
299 sonnet: 'claude-sonnet-5-5',
300 opus: 'claude-opus-5-5',
301 fable: 'claude-fable-5-1',
302}
303
304/**
305 * What to write into `turn.step`'s `model`: a full id as given, or the id
306 * behind a family alias. Anything else is returned unchanged for the engine
307 * to judge.
308 */
309export function requestModelId(model: string): string {
310 return ALIAS_IDS[model.trim().toLowerCase()] ?? model
311}
312
313export interface PolicyConfig {
314 tiers: Tiers
315 /**
316 * How sure the decision must be to spend more (a bigger model, more
317 * reasoning). Being wrong here costs money, so the bar is low.
318 */
319 minUpgradeConfidence: number
320 /**
321 * How sure it must be to spend less. Being wrong here means a task handled
322 * by too small a model or too little thought, so the bar is high.
323 */
324 minDowngradeConfidence: number
325}
326
327export interface Routing {
328 /** The model to run on, or null to leave the request as it is. */
329 model: string | null
330 /** The reasoning level to ask for, or null to leave it as it is. */
331 effort: Effort | null
332 /** Why, for the log line. */
333 reason: string
334}
335
336const NOTHING: Routing = { model: null, effort: null, reason: 'no decision' }
337
338/**
339 * Whether a change of rank passes its threshold. Both directions are allowed;
340 * they just do not have to clear the same bar, because the two mistakes do not
341 * cost the same. A move whose direction cannot be told (an unrecognised
342 * current value) is treated as an upgrade.
343 */
344function allowed(
345 wanted: number,
346 current: number | null,
347 confidence: number | null,
348 config: PolicyConfig,
349): boolean {
350 if (current !== null && wanted === current) return false
351 const isDowngrade = current !== null && wanted < current
352 const bar = isDowngrade ? config.minDowngradeConfidence : config.minUpgradeConfidence
353 // A backend that reports no confidence (the Gateway without a distribution,
354 // or the built-in classifier) clears the upgrade bar but never the
355 // downgrade one: spending less on an unmeasured hunch is the bad trade.
356 if (confidence === null) return !isDowngrade
357 return confidence >= bar
358}
359
360/**
361 * Turns a decision into a model and a reasoning level, either of which may be
362 * null to leave the request as it is. Both can move in either direction.
363 */
364export function route(
365 decision: Decision | null,
366 current: { model: string; effort?: string | number },
367 config: PolicyConfig,
368): Routing {
369 if (!decision) return NOTHING
370
371 let tier = decision.tier
372 let effortScore = decision.effort
373 let forced = false
374
375 // Carrying out something final is never worth the saving: take the deep
376 // tier and real reasoning, whatever the cheaper answer said, and skip the
377 // thresholds — this is the one case that is not a confidence question.
378 if (decision.risky !== null && decision.risky > 0.7) {
379 if (TIER_ORDER.indexOf(tier) < TIER_ORDER.indexOf('deep')) tier = 'deep'
380 effortScore = Math.max(effortScore ?? 0, 2)
381 forced = true
382 }
383
384 const wantedTier = TIER_ORDER.indexOf(tier)
385 const currentTier = rankOf(current.model, config.tiers)
386 const wantedModel = config.tiers[tier]
387
388 const model =
389 wantedModel &&
390 wantedModel !== current.model &&
391 // Risk raises the model floor too; forcing never pulls superDeep down to deep.
392 (forced
393 ? currentTier === null || wantedTier > currentTier
394 : allowed(wantedTier, currentTier, decision.confidence, config))
395 ? wantedModel
396 : null
397
398 let effort: Effort | null = null
399 if (effortScore !== null) {
400 const currentRank = effortRank(current.effort)
401 let wantedRank = EFFORT_ORDER.indexOf(effortLevel(effortScore))
402
403 // Risk raises the floor; it must never lower one. Forcing only skips the
404 // thresholds, so without this clamp a task already at `xhigh` or `max`
405 // and rated mechanically simple would be pulled down to `high` with no
406 // confidence check at all — the opposite of what the rule is for.
407 if (forced && currentRank !== null) wantedRank = Math.max(wantedRank, currentRank)
408
409 // A numeric effort is the caller's own scale, not this ladder; leave it.
410 const comparable = typeof current.effort !== 'number'
411 const wanted = EFFORT_ORDER[Math.min(EFFORT_ORDER.length - 1, wantedRank)] as Effort
412 if (
413 comparable &&
414 wantedRank !== currentRank &&
415 (forced || allowed(wantedRank, currentRank, decision.effortConfidence, config))
416 ) {
417 effort = wanted
418 }
419 }
420
421 const said = decision.confidence === null ? 'confidence n/d' : `confidence ${decision.confidence.toFixed(2)}`
422
423 if (!model && !effort) {
424 // Naming what it wanted and what it kept is the whole point of this line.
425 // Without it, a mod that classified and decided to leave the request alone
426 // is indistinguishable from one that never loaded.
427 const wantedEffort = effortScore === null ? null : effortLevel(effortScore)
428 const kept = `${current.model}${current.effort === undefined ? '' : `/${current.effort}`}`
429 const modelDiffers = Boolean(wantedModel) && wantedModel !== current.model
430 const effortDiffers = wantedEffort !== null && wantedEffort !== current.effort
431 if (!modelDiffers && !effortDiffers) {
432 return { model: null, effort: null, reason: `stayed on ${kept}: already the right fit` }
433 }
434 const target = modelDiffers ? wantedModel : `effort ${wantedEffort}`
435 const sure = modelDiffers ? decision.confidence : decision.effortConfidence
436 const why =
437 sure === null
438 ? `no confidence score, so no switch to ${target}`
439 : `too unsure to switch to ${target} (${Math.round(sure * 100)}% sure)`
440 return { model: null, effort: null, reason: `stayed on ${kept}: ${why}` }
441 }
442
443 return { model, effort, reason: forced ? `${tier}, forced by risk` : `${tier} (${said})` }
444}
445
446/**
447 * Whether a prompt is a slash command and nothing else (`/simplify`). The
448 * decision model sees only the name, never the skill or command it runs:
449 * measured on TypeSafe (three calls each), `/simplify`, `/run` and `/github`
450 * came back fast with effort 0.1 to 0.5, low effort for a multi-step skill,
451 * while `/code-review` and `/security-review` came back balanced and deep.
452 * With text after the name there is a task to read, and it is classified.
453 * A one-segment path alone (`/etc`) matches too; nobody sends one as a task.
454 */
455export function bareCommand(text: string): boolean {
456 return /^\/[^\s/]+$/.test(text.trim())
457}
458
459/**
460 * Holds a prompt's classification until the turn that reads that prompt
461 * starts.
462 *
463 * Nothing ties a decision to the turn it belongs to. Prompts can be queued
464 * while the model is busy, a peer session's message can be delivered inside a
465 * running turn, and `prompt.submit` carries no turn id at all while the
466 * session is idle. So when more than one prompt is waiting, `take` reports
467 * none: running a turn on another prompt's decision is a worse outcome than
468 * not routing it, and not routing is what every other failure path here does.
469 */
470export function pendingDecisions(): {
471 put(decision: Decision | null): void
472 take(): Decision | null
473} {
474 let held: Decision | null = null
475 let waiting = 0
476
477 return {
478 put(decision) {
479 waiting += 1
480 // Past the first, which prompt a turn will read is unknowable, so the
481 // slot is emptied instead of holding a decision that may not fit.
482 held = waiting === 1 ? decision : null
483 },
484 take() {
485 const decision = waiting === 1 ? held : null
486 held = null
487 waiting = 0
488 return decision
489 },
490 }
491}
492
493/** A number for the log, or `n/d` when the backend reported none. */
494function reported(value: number | null): string {
495 return value === null ? 'n/d' : value.toFixed(2)
496}
497
498/**
499 * The one-time line that says the router is alive, which backend answers it,
500 * and which of the three switches are on.
501 *
502 * Without this, a router that loaded and a router that never loaded are told
503 * apart only by the absence of later lines, which is not evidence of anything.
504 */
505export function describeSetup(
506 provider: Provider | null,
507 url: string,
508 switches: { subagentModel: boolean; mainEffort: boolean; mainModel: boolean },
509 // `provider: "builtin"` is a choice, not a missing key. Reporting it as a
510 // credential problem sends someone hunting for a key they meant to omit.
511 builtinByChoice = false,
512): string {
513 const backend = provider
514 ? `${provider} (${url})`
515 : builtinByChoice
516 ? 'the built-in classifier, by choice'
517 : 'the built-in classifier, no key set'
518 const on = [
519 switches.subagentModel && 'subagent model',
520 switches.mainEffort && 'main effort',
521 switches.mainModel && 'main model',
522 ].filter(Boolean)
523 return `ready on ${backend}; routing ${on.length > 0 ? on.join(', ') : 'nothing, every switch is off'}`
524}
525
526/**
527 * What the decision model answered, before any policy touches it: the raw
528 * tier, effort and risk with their confidences, and how long it took.
529 *
530 * This is the line that shows the classification happened at all, separately
531 * from whether the policy then decided to act on it.
532 */
533export function describeDecision(decision: Decision | null, ms: number | null): string {
534 const took = ms === null ? '' : ` · ${Math.round(ms)}ms`
535 if (!decision) return `no answer${took}`
536
537 const parts = [`tier ${decision.tier} (${reported(decision.confidence)})`]
538 if (decision.effort !== null) {
539 parts.push(
540 `effort ${decision.effort.toFixed(1)} → ${effortLevel(decision.effort)} (${reported(decision.effortConfidence)})`,
541 )
542 }
543 if (decision.risky !== null) parts.push(`risky ${reported(decision.risky)}`)
544 if (decision.newTask != null) parts.push(`new task ${reported(decision.newTask)}`)
545 return parts.join(' · ') + took
546}
547
548/**
549 * The persistent status line: the last thing the router did, short enough to
550 * sit on screen beside the engine's own notices.
551 */
552export function describeStatus(
553 decision: Decision | null,
554 change: { model?: string; effort?: Effort } | null,
555 current: { model: string; effort?: string | number },
556): string {
557 const model = modelName(change?.model ?? current.model)
558 const effort = change?.effort ?? current.effort
559 const running = effort === undefined ? model : `${model} · ${effort} effort`
560 if (!decision) return `${running} · no answer`
561 const confidence = decision.confidence === null ? '' : ` ${Math.round(decision.confidence * 100)}%`
562 return `${running} · ${decision.tier}${confidence}`
563}
564
565/** `claude-haiku-5-5` as "Haiku 5.5"; an alias or unknown id is only capitalised. */
566export function modelName(id: string): string {
567 const match = /^claude-([a-z]+)-(\d+)(?:-(\d{1,2})(?!\d))?/.exec(id)
568 const capital = (word: string) => word.charAt(0).toUpperCase() + word.slice(1)
569 if (!match) return capital(id)
570 return `${capital(match[1] as string)} ${match[2]}${match[3] ? `.${match[3]}` : ''}`
571}
572