Jev (TypeSafe) picks the model and effort per task type; ported from jev-model-router with main-loop model switching on

Picks the model and the reasoning effort each task runs with, using Jev, TypeSafe's System One decision model: unstructured state in, a typed choice with a probability distribution out, no free-form text.
Two backends, chosen by whichever key is set:
| Backend | Endpoint | Model | Confidence |
|---|---|---|---|
typesafe | POST api.typesafe.ai/v1/systemone | jev-latest | reported per answer |
gateway | POST ai-gateway.vercel.sh/v4/ai/evaluation-model | typesafe-ai/jev | derived from an optional distribution |
TypeSafe's own API wins when both keys are set: it is the only one that reports a calibrated confidence, which is what the confidence bars below read. Set provider to force one, or to builtin to use neither. Each backend keeps its own URL and model option, so an override written for one is never sent to the other. A provider forced onto a backend whose key is missing degrades to the built-in classifier and says so once in the log.
Three switches, and they are not equally safe:
| Switch | What it sets | Default |
|---|---|---|
routeSubagentModel | the model of each subagent, at agent.spawn | on |
routeMainEffort | the reasoning effort of the main conversation, at turn.step | on |
routeMainModel | the model of the main conversation, at turn.step | off |
A subagent starts with its own context, so routing its model costs nothing beyond the classification. Changing the main loop's model mid-session is the expensive one: it invalidates the prompt cache, and on a long context re-caching can cost more than the cheaper tier saves. Turn it on once you have measured your own sessions, not before.
The Agent tool has no effort parameter, so a subagent's effort is not this mod's to set.
Both directions, both dimensions. A task read as mechanical is routed down; one read as hard is routed up — model and effort alike.
The prompt is classified at prompt.submit, which runs before the turn starts, and the decision is applied to the turn's first model request and reused by the rest of that turn.
With no key configured the mod still works: it falls back to the engine's own $.model.classify, which answers the same question with the small fast model. That path reports no confidence, so the threshold does not apply to it.
One request, three questions evaluated in parallel:
tier — a choice between three descriptions of the work (mechanical and local / ordinary engineering / hard or high-stakes). The decision model never sees a model name.effort — a score on a four-level rubric, for how much step-by-step reasoning the task needs.risky — whether the task touches production, money, credentials, or state that cannot be undone. A noul on TypeSafe's API, a boolean on the Gateway: the same question under two names.TypeSafe's API reports a confidence per answer. The Gateway's answer shape carries no confidence field, so on that backend confidence is read as the highest probability in the distribution — and that distribution is itself optional in the schema, in which case confidence is absent and the threshold does not fire.
The two mistakes do not cost the same, so they do not clear the same bar:
minUpgradeConfidence, 0.3 by default. Being wrong costs money.minDowngradeConfidence, 0.6 by default. Being wrong means a task handled by too small a model or too little thought.risky above 0.7 takes the deep tier and real reasoning, past both bars. That one is not a confidence question.Every other failure — a non-2xx response, a timeout, a malformed body, a thrown error — leaves the request exactly as the engine built it. The router never blocks a turn.
A slash command with nothing after it (/simplify) is not classified: the decision model would see only the command's name, never what the command does. Its turn keeps the session's model and effort. With text after the name, the prompt is classified like any other.
With logDecisions on (the default), the router reports every step of its own work, because nothing else in Claude Code shows it: the model and effort it rewrites are parameters of each request, not the session's settings, so the status line, the header and the effort box never move whatever it decides.
[model-router] ready on typesafe (https://api.typesafe.ai/v1/systemone); routing subagent model, main effort
[model-router] jev: tier fast (0.87) · effort 0.4 → low (0.71) · risky 0.02 · 249ms
[model-router] main loop → effort low: fast (confidence 0.87)
[model-router] jev: tier fast (0.41) · effort 0.4 → low (0.38) · risky 0.01 · 210ms
[model-router] main loop: kept opus/medium, wanted haiku/low (confidence 0.41)
jev: line is what the decision model replied, before any policy is applied — the tier, the effort score and the risk, each with its confidence, and how long the call took.minDowngradeConfidence.It also keeps a one-line status on screen, replaced as it goes:
jev · fast 0.87 → haiku/low
jev · fast 0.41 · unchanged
confidence n/d means the backend reported no confidence, which the built-in classifier never does and the Gateway does whenever its probability distribution is absent; the ready on line says which one answered.
No lines at all has three causes, and only the last is the module failing to load. Check them in this order:
claude -p (or the SDK). A headless run has no transcript and no status row: every line still goes to the debug log, ~/.claude/debug/<session-id>.txt (a .txt, not a .log; latest is a symlink to the newest), and an SDK host receives each one as ui_log. The router is working; look there..claude/skills/ (where --mod writes it) only once the project is trusted: it is repository content, so an untrusted folder's .claude/ is not read at all, and claude -p never asks. Open claude interactively in the folder and accept the trust prompt, or name the plugin explicitly with --plugin-dir (see Install). claude --debug settles it: a loaded module prints hooks module model-router loaded (worker, …); events: prompt.submit,turn.step,agent.spawn (@inline when loaded with --plugin-dir); Found N plugins without it means the plugin is not in the session.installed plugins' hooks modules not loaded: rollout flag (tengu_plugin_hooks_modules) is off. Update Claude Code.A ready on the built-in classifier, no key set line when you did set a key means the key sits under the wrong pluginConfigs entry: the key must match the plugin's id, which depends on how it was loaded (see Options).
With a key set, the prompt text leaves the machine and goes to whichever backend the key belongs to. The main-loop path sends the prompt; the subagent path sends the subagent's prompt, its description and its agent type. Nothing else. With no key set, nothing leaves the machine.
typesafeApiKey: string TypeSafe API key (preferred: it reports a confidence)
gatewayApiKey: string Vercel AI Gateway key
provider: string "auto" | "typesafe" | "gateway" | "builtin"
typesafeBaseUrl: string empty uses https://api.typesafe.ai
typesafeModel: string empty uses jev-latest
gatewayBaseUrl: string empty uses https://ai-gateway.vercel.sh/v4/ai
gatewayModel: string empty uses typesafe-ai/jev
fastModel: string fast tier, alias or full id (default "haiku")
balancedModel: string balanced tier, alias or full id (default "sonnet")
deepModel: string deep tier, alias or full id (default "opus")
minUpgradeConfidence: number bar to spend more (default 0.3)
minDowngradeConfidence: number bar to spend less (default 0.6)
routeSubagentModel: boolean model of each subagent (default true)
routeMainEffort: boolean effort of the main loop (default true)
routeMainModel: boolean model of the main loop (default false)
timeoutMs: number latency budget per classification (default 800)
logDecisions: boolean log each decision (default true)
The three tiers take an alias (haiku, sonnet, opus) or a full model id. A subagent is spawned with the name as given, the way the Agent tool takes it; the main loop's request needs an id, so there an alias is resolved to the family's current id (haiku → claude-haiku-4-5-20251001, sonnet → claude-sonnet-5, opus → claude-opus-5). Set a full id to pin a specific version. A decision for the tier the session already runs is not a change, so a session on claude-opus-5[1m] keeps its 1M-context id.
Declared in .claude-plugin/plugin.json (userConfig). Set them in /config, in user settings (~/.claude/settings.json, not project settings), with --settings <file> or in managed settings:
{ "pluginConfigs": { "model-router": { "options": { "typesafeApiKey": "" } } } }
The entry's key is the plugin's id, and the id follows how the plugin was loaded: "model-router" when auto-loaded from .claude/skills/ (the --mod install), "model-router" with --plugin-dir. Under the wrong key every option stays at its default, and the ready on line reports no key set.
/plugin install model-router --marketplace muhx/model-router
Answer y to add the marketplace, then pick a scope (user scope loads it in every session). Then set your key in /config, or in ~/.claude/settings.json:
{ "pluginConfigs": { "model-router": { "options": { "typesafeApiKey": "" } } } }
Get a key from console.typesafe.ai. With no key the mod still runs, falling back to the engine's own classifier.
For development from a folder:
claude --plugin-dir /path/to/model-router
Loaded that way the plugin id is model-router@inline, and options go under the "model-router" key either way.
bun test cli-tool/components/mods/productivity/model-router/tests
Requirements. Mods are on by default in Claude Code 2.1.287+. Typed against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
A mod runs without node_modules, so neither @typesafe-ai/sdk nor the AI SDK is available here: both backends are spoken to over HTTP through $.http.fetch. The TypeSafe wire shape was read from @typesafe-ai/sdk v0.6.0; the Gateway's, which is experimental in the AI SDK (experimental_evaluate, 7.0.105+) and not documented publicly, from @ai-sdk/gateway v4.0.86 and @ai-sdk/provider v4.0.17. Either may change.
The policy layer has unit tests (42) and the Jev request shape has been verified against the live API. The mod has not yet been exercised across a long real session, so treat the confidence thresholds as starting points and watch the [model-router] log lines before trusting it on a large context.
A derivative of jev-model-router by claude-code-templates, used under the MIT License. The policy engine, the Jev wire shapes and the confidence model are theirs; this fork renames the mod and turns main-loop model switching on by default. Decisions come from Jev, TypeSafe's System One model.
hooks/model-router.ts 338 lines1/**
2 * model-router — Claude Mod
3 *
4 * Picks the model each task runs on with TypeSafe's Jev, a System One
5 * decision model: unstructured state in, a typed choice with a probability
6 * distribution out.
7 *
8 * Jev is reached one of two ways, whichever key is configured: TypeSafe's
9 * own API (`typesafeApiKey`), which reports a calibrated confidence per
10 * answer, or the Vercel AI Gateway (`gatewayApiKey`), which does not. With
11 * neither, the engine's own `$.model.classify` stands in, so the mod is
12 * useful without any account.
13 *
14 * Three things it can set, each on its own switch:
15 * agent.spawn — the model of each subagent (on by default)
16 * turn.step — the reasoning effort of the main loop (on by default)
17 * turn.step — the model of the main loop (off by default: switching
18 * models mid-session invalidates the prompt cache, which can
19 * cost more than the cheaper tier saves)
20 *
21 * Every one of them moves in both directions: a task the decision model reads
22 * as mechanical is routed down, one it reads as hard is routed up. The two
23 * mistakes do not cost the same, so they do not clear the same confidence bar
24 * (see `minUpgradeConfidence` / `minDowngradeConfidence` in policy.ts).
25 *
26 * The Agent tool has no effort parameter, so a subagent's effort is not ours
27 * to set; only its model is.
28 *
29 * The prompt is classified at `prompt.submit`, which runs before the turn
30 * starts, and the decision is applied at the turn's first request.
31 *
32 * Every failure path is fail-open: a classification that errors or runs past
33 * the latency budget leaves the request exactly as the engine built it.
34 *
35 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
36 * or "gatewayApiKey"). Never hardcode it in this file.
37 *
38 * Needs Claude Code >= 2.1.287. Typed
39 * against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
40 *
41 * Privacy: with a key set, the prompt text is sent to whichever backend the
42 * key belongs to.
43 */
44import type { Register } from 'claude-code'
45import {
46 DEFAULT_BASE_URL,
47 DEFAULT_MODEL,
48 describeDecision,
49 describeSetup,
50 describeStatus,
51 endpoint,
52 pendingDecisions,
53 readDecision,
54 selectProvider,
55 requestBody,
56 requestHeaders,
57 requestModelId,
58 route,
59 TIER_ORDER,
60 bareCommand,
61} from './policy.ts'
62import type { Decision, Effort, PolicyConfig, Provider, Tier } from './policy.ts'
63
64export const register: Register = (on, options) => {
65 const text = (key: string, fallback: string) =>
66 typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
67 const number = (key: string, fallback: number) =>
68 typeof options[key] === 'number' ? (options[key] as number) : fallback
69 const flag = (key: string, fallback: boolean) =>
70 typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
71
72 // TypeSafe's own API is preferred when both keys are set: it is the only
73 // one that reports a calibrated confidence, which the policy's threshold
74 // reads. `provider` forces one, including "builtin" to use neither.
75 const typesafeKey = text('typesafeApiKey', '')
76 const gatewayKey = text('gatewayApiKey', '')
77 const forced = text('provider', 'auto')
78 const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey)
79
80 // Each backend keeps its own URL and model, so an override written for one
81 // can never be sent to the other when `auto` picks differently than expected.
82 const apiKey = active === 'typesafe' ? typesafeKey : active === 'gateway' ? gatewayKey : ''
83 const modelId = !active
84 ? ''
85 : active === 'typesafe'
86 ? text('typesafeModel', DEFAULT_MODEL.typesafe)
87 : text('gatewayModel', DEFAULT_MODEL.gateway)
88 const url = !active
89 ? ''
90 : active === 'typesafe'
91 ? endpoint('typesafe', text('typesafeBaseUrl', DEFAULT_BASE_URL.typesafe))
92 : endpoint('gateway', text('gatewayBaseUrl', DEFAULT_BASE_URL.gateway))
93
94 // A backend named in the options but missing its key degrades to the
95 // built-in classifier, which is silent; say so once, when a hook first runs.
96 let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
97
98 const timeoutMs = number('timeoutMs', 800)
99 const routeSubagentModel = flag('routeSubagentModel', true)
100 const routeMainEffort = flag('routeMainEffort', true)
101 const routeMainModel = flag('routeMainModel', true)
102 const routeMainLoop = routeMainEffort || routeMainModel
103 const logDecisions = flag('logDecisions', true)
104
105 const policy: PolicyConfig = {
106 tiers: {
107 fast: text('fastModel', 'haiku'),
108 balanced: text('balancedModel', 'sonnet'),
109 deep: text('deepModel', 'opus'),
110 },
111 minUpgradeConfidence: number('minUpgradeConfidence', 0.3),
112 minDowngradeConfidence: number('minDowngradeConfidence', 0.6),
113 }
114
115 // The classification waiting for the turn that reads its prompt, and what
116 // the current turn settled on. Both are single slots: main-loop turns run
117 // one at a time, so nothing accumulates over a long session. `pending`
118 // reports no decision when two prompts are waiting at once, rather than
119 // routing a turn on a decision made for a different prompt.
120 const pending = pendingDecisions()
121 // Said once, the first time a hook runs. A router that loaded and one that
122 // never loaded are otherwise told apart only by the absence of later lines,
123 // and absence is not evidence: the policy leaves most turns alone anyway.
124 let announced = false
125 let appliedTurnId: string | undefined
126 let applied: { model?: string; effort?: Effort } | null = null
127
128 on('prompt.submit', async ($, e, next) => {
129 // Before the routing guards: a module whose switches are all off has still
130 // loaded, and that is exactly when its silence is most misleading.
131 if (!announced) {
132 announced = true
133 if (logDecisions) {
134 $.ui.log(
135 `[model-router] ${describeSetup(
136 active,
137 url,
138 {
139 subagentModel: routeSubagentModel,
140 mainEffort: routeMainEffort,
141 mainModel: routeMainModel,
142 },
143 forced === 'builtin',
144 )}`,
145 )
146 }
147 }
148 if (!routeMainLoop) return next(e)
149
150 if (!unusableReported) {
151 unusableReported = true
152 $.ui.log(`[model-router] provider "${forced}" has no key set; using the built-in classifier`)
153 }
154
155 // A slash command alone gives the decision model only the command's name.
156 // Its turn keeps the session's model and effort; the null put keeps a
157 // previous prompt's decision from reaching it.
158 if (bareCommand(e.text)) {
159 if (logDecisions) $.ui.log('[model-router] a command with nothing after it; leaving the turn alone')
160 pending.put(null)
161 return next(e)
162 }
163
164 const startedAt = await $.clock.now()
165 let decision: Decision | null = null
166 if (active) {
167 try {
168 const response = await Promise.race([
169 $.http.fetch(url, {
170 method: 'POST',
171 headers: requestHeaders(active, apiKey, modelId),
172 body: requestBody(active, { prompt: e.text }, modelId),
173 }),
174 $.clock.sleep(timeoutMs),
175 ])
176 if (response && response.ok) decision = readDecision(response.text)
177 else if (response) $.ui.log(`[model-router] ${active} responded ${response.status}`)
178 else $.ui.log(`[model-router] classification passed ${timeoutMs}ms; leaving the turn alone`)
179 } catch (error) {
180 $.ui.log(`[model-router] classification failed: ${String(error)}`)
181 }
182 } else {
183 // No backend: the engine's own small-model classifier answers the same
184 // question, without the confidence the policy's threshold reads.
185 try {
186 const label = await $.model.classify(e.text, TIER_ORDER)
187 if (label) {
188 decision = {
189 tier: label as Tier,
190 confidence: null,
191 risky: null,
192 effort: null,
193 effortConfidence: null,
194 }
195 }
196 } catch (error) {
197 $.ui.log(`[model-router] built-in classifier failed: ${String(error)}`)
198 }
199 }
200
201 // What the decision model actually answered, whatever the policy then
202 // does with it. This is the line that proves the classification ran.
203 if (logDecisions) {
204 const ms = (await $.clock.now()) - startedAt
205 $.ui.log(`[model-router] jev: ${describeDecision(decision, ms)}`)
206 }
207
208 pending.put(decision)
209 return next(e)
210 })
211
212 on('turn.step', async function* ($, e, next) {
213 if (!routeMainLoop || e.agentId) return yield* next(e)
214
215 // Every request after the first reuses what the turn settled on, so
216 // neither the model nor the effort changes under its own tool loop.
217 if (e.index > 0 && e.turnId === appliedTurnId) {
218 return yield* next(applied ? { ...e, ...applied } : e)
219 }
220
221 const decision = pending.take()
222 const routing = route(decision, { model: e.model, effort: e.effort }, policy)
223 const change: { model?: string; effort?: Effort } = {}
224 // The main loop's `model` is sent to the API as written, so an alias
225 // becomes its id here; a subagent's (agent.spawn) may stay an alias.
226 if (routeMainModel && routing.model) change.model = requestModelId(routing.model)
227 if (routeMainEffort && routing.effort) change.effort = routing.effort
228
229 appliedTurnId = e.turnId
230 applied = Object.keys(change).length > 0 ? change : null
231 // A row in the transcript scrolls away; this line stays on screen.
232 if (logDecisions) $.ui.status(describeStatus(decision, applied))
233
234 if (!applied) {
235 // A turn left alone is the common case, and it used to be silent, which
236 // made a working mod look like one that never loaded. Say what happened.
237 if (logDecisions) {
238 const suppressed = routing.model && !routeMainModel ? ' (main-loop model routing off)' : ''
239 $.ui.log(`[model-router] main loop: ${routing.reason}${suppressed}`)
240 }
241 return yield* next(e)
242 }
243 if (logDecisions) {
244 const what = [change.model, change.effort && `effort ${change.effort}`]
245 .filter(Boolean)
246 .join(', ')
247 $.ui.log(`[model-router] main loop → ${what}: ${routing.reason}`)
248 }
249 return yield* next({ ...e, ...change })
250 })
251
252 on('agent.spawn', async ($, e, next) => {
253 // Before the routing guards: a module whose switches are all off has still
254 // loaded, and that is exactly when its silence is most misleading.
255 if (!announced) {
256 announced = true
257 if (logDecisions) {
258 $.ui.log(
259 `[model-router] ${describeSetup(
260 active,
261 url,
262 {
263 subagentModel: routeSubagentModel,
264 mainEffort: routeMainEffort,
265 mainModel: routeMainModel,
266 },
267 forced === 'builtin',
268 )}`,
269 )
270 }
271 }
272
273 // A fork inherits its parent's model; `model` is ignored for it.
274 if (!routeSubagentModel || e.fork) return next(e)
275
276 if (!unusableReported) {
277 unusableReported = true
278 $.ui.log(`[model-router] provider "${forced}" has no key set; using the built-in classifier`)
279 }
280
281 const startedAt = await $.clock.now()
282 let decision: Decision | null = null
283 if (active) {
284 try {
285 const response = await Promise.race([
286 $.http.fetch(url, {
287 method: 'POST',
288 headers: requestHeaders(active, apiKey, modelId),
289 body: requestBody(
290 active,
291 { prompt: e.prompt, description: e.description, agentType: e.subagentType },
292 modelId,
293 ),
294 }),
295 $.clock.sleep(timeoutMs),
296 ])
297 if (response && response.ok) decision = readDecision(response.text)
298 else if (response) $.ui.log(`[model-router] ${active} responded ${response.status}`)
299 else $.ui.log(`[model-router] classification passed ${timeoutMs}ms; leaving the subagent alone`)
300 } catch (error) {
301 $.ui.log(`[model-router] classification failed: ${String(error)}`)
302 }
303 } else {
304 try {
305 const label = await $.model.classify(e.prompt, TIER_ORDER)
306 if (label) {
307 decision = {
308 tier: label as Tier,
309 confidence: null,
310 risky: null,
311 effort: null,
312 effortConfidence: null,
313 }
314 }
315 } catch (error) {
316 $.ui.log(`[model-router] built-in classifier failed: ${String(error)}`)
317 }
318 }
319
320 if (logDecisions) {
321 const ms = (await $.clock.now()) - startedAt
322 $.ui.log(`[model-router] jev (${e.subagentType}): ${describeDecision(decision, ms)}`)
323 }
324
325 // The subagent's own model wins when the caller named one; otherwise it
326 // would inherit the parent's, so that is what a change is measured from.
327 // The Agent tool takes no effort, so only the model is ours to set here.
328 const current = e.model ?? e.parentModel
329 const { model, reason } = route(decision, { model: current }, policy)
330 if (!model) {
331 if (logDecisions) $.ui.log(`[model-router] ${e.subagentType}: ${reason}`)
332 return next(e)
333 }
334 if (logDecisions) $.ui.log(`[model-router] ${e.subagentType} → ${model}: ${reason}`)
335 return next({ ...e, model })
336 })
337}
338hooks/policy.ts 518 lines1/**
2 * jev-model-router — pure decision logic.
3 *
4 * No `$` and no I/O here: this module only builds the request the decision
5 * API takes, reads its answer, and turns that answer into a model id. The
6 * hooks module does every call on `$` at its own call site.
7 *
8 * Two backends speak to the same model with different wire shapes:
9 *
10 * typesafe POST https://api.typesafe.ai/v1/systemone
11 * `{ model, state, questions }`; a yes/no question is a `noul`
12 * and every answer carries its own `confidence`.
13 * gateway POST https://ai-gateway.vercel.sh/v4/ai/evaluation-model
14 * `{ state, questions }` with the model in a header; a yes/no
15 * question is a `boolean`, and there is no `confidence` field —
16 * it has to be derived from an optional distribution.
17 *
18 * The Gateway shape is not documented publicly; it was read from
19 * @ai-sdk/gateway and @ai-sdk/provider.
20 */
21
22export type Provider = 'typesafe' | 'gateway'
23
24export type Tier = 'fast' | 'balanced' | 'deep'
25
26export interface Tiers {
27 fast: string
28 balanced: string
29 deep: string
30}
31
32export interface Decision {
33 tier: Tier
34 /** Confidence in the tier, or null when the backend reported none. */
35 confidence: number | null
36 /** P(true) that carrying the task out would itself be costly or final. */
37 risky: number | null
38 /** 0..3 along the effort rubric, or null when absent. */
39 effort: number | null
40 /** Confidence in the effort, or null when the backend reported none. */
41 effortConfidence: number | null
42}
43
44/** The reasoning levels a turn can ask for, cheapest first. */
45export const EFFORT_ORDER = ['low', 'medium', 'high', 'xhigh'] as const
46
47export type Effort = (typeof EFFORT_ORDER)[number]
48
49export const TIER_ORDER: readonly Tier[] = ['fast', 'balanced', 'deep']
50
51/**
52 * How each tier is described to the decision model. Deliberately about the
53 * shape of the work, not about model names: the model never sees an id.
54 */
55const TIER_CRITERIA: Record<Tier, string> = {
56 fast: 'Mechanical and local: read or summarise a file, run one command, rename a symbol, answer something already in context.',
57 balanced:
58 'Ordinary engineering: implement a well-specified change across a few files, write tests, fix a clearly described bug, review a small diff.',
59 deep: 'Hard or high-stakes: architecture and design, debugging a failure whose cause is unknown, security, data migrations, concurrency, anything touching production or money.',
60}
61
62const EFFORT_RUBRIC = ['almost none', 'some', 'a lot', 'as much as possible'] as const
63
64export const DEFAULT_BASE_URL: Record<Provider, string> = {
65 typesafe: 'https://api.typesafe.ai',
66 gateway: 'https://ai-gateway.vercel.sh/v4/ai',
67}
68
69/**
70 * The Gateway's own protocol version, sent as `ai-gateway-protocol-version`.
71 * Tracks the `AI_GATEWAY_PROTOCOL_VERSION` of `@ai-sdk/gateway` (4.0.87).
72 */
73const AI_GATEWAY_PROTOCOL_VERSION = '0.0.1'
74
75export const DEFAULT_MODEL: Record<Provider, string> = {
76 typesafe: 'jev-latest',
77 gateway: 'typesafe-ai/jev',
78}
79
80/**
81 * Which backend a configuration asks for, or null for the built-in
82 * classifier. `auto` prefers TypeSafe, since it is the only one that reports
83 * a calibrated confidence; a forced backend whose key is missing resolves to
84 * null rather than falling through to the other one's key.
85 */
86export function selectProvider(
87 forced: string,
88 typesafeKey: string,
89 gatewayKey: string,
90): Provider | null {
91 if (forced === 'builtin') return null
92 if (forced === 'typesafe') return typesafeKey ? 'typesafe' : null
93 if (forced === 'gateway') return gatewayKey ? 'gateway' : null
94 if (typesafeKey) return 'typesafe'
95 if (gatewayKey) return 'gateway'
96 return null
97}
98
99/** The full endpoint a backend posts to. */
100export function endpoint(provider: Provider, baseUrl: string): string {
101 const root = baseUrl.replace(/\/+$/, '')
102 return provider === 'typesafe' ? `${root}/v1/systemone` : `${root}/evaluation-model`
103}
104
105/** The `questions` map, in the shape the backend's schema names. */
106export function questions(provider: Provider): Record<string, unknown> {
107 return {
108 tier: {
109 type: 'choice',
110 instructions: 'Which is the cheapest tier that can complete this coding task well?',
111 criteria: TIER_CRITERIA,
112 },
113 effort: {
114 type: 'score',
115 instructions: 'How much step-by-step reasoning does this task need?',
116 criteria: EFFORT_RUBRIC,
117 },
118 risky: {
119 // The same question under two names: `noul` on TypeSafe's own API,
120 // `boolean` in the AI SDK's evaluation schema.
121 type: provider === 'typesafe' ? 'noul' : 'boolean',
122 // Asked about the act, not the subject. The first wording ("the task
123 // touches production, money, credentials") scored 0.96 on "add a
124 // refund endpoint that calls Stripe" — ordinary code that happens to be
125 // about money — and would have escalated it past a 0.98-confidence
126 // answer of the balanced tier.
127 instructions:
128 'Carrying out this task would itself change production, move real money, or alter data that cannot be restored. Writing or testing code that deals with such things, without running it against the real system, does not count.',
129 },
130 }
131}
132
133/** The request body. The Gateway carries the model in a header instead. */
134export function requestBody(
135 provider: Provider,
136 state: Record<string, unknown>,
137 model: string,
138): string {
139 const body =
140 provider === 'typesafe'
141 ? { model, state, questions: questions(provider) }
142 : { state, questions: questions(provider) }
143 return JSON.stringify(body)
144}
145
146/** The request headers. */
147export function requestHeaders(
148 provider: Provider,
149 apiKey: string,
150 model: string,
151): Record<string, string> {
152 const common = { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` }
153 if (provider === 'typesafe') return common
154 return {
155 ...common,
156 'ai-gateway-auth-method': 'api-key',
157 'ai-model-id': model,
158 // The Gateway rejects any request that does not name the protocol it
159 // speaks: 400 "Unsupported gateway protocol version". Every other header
160 // here is accepted without it, so the omission fails the whole backend.
161 'ai-gateway-protocol-version': AI_GATEWAY_PROTOCOL_VERSION,
162 'ai-evaluation-model-specification-version': '4',
163 }
164}
165
166function isTier(value: unknown): value is Tier {
167 return value === 'fast' || value === 'balanced' || value === 'deep'
168}
169
170/**
171 * Reads a response from either backend.
172 *
173 * TypeSafe's own API reports a `confidence` per answer and a `noul` number
174 * for a yes/no question. The Gateway reports neither: confidence has to come
175 * from the highest probability of a distribution that is itself optional, and
176 * a yes/no answer arrives as `probability`. Both are handled, and a missing
177 * confidence reads as null rather than as a number the policy would trust.
178 */
179export function readDecision(responseText: string): Decision | null {
180 let parsed: unknown
181 try {
182 parsed = JSON.parse(responseText)
183 } catch {
184 return null
185 }
186 const answers = (parsed as { answers?: Record<string, Record<string, unknown>> }).answers
187 if (!answers) return null
188
189 const tierAnswer = answers.tier
190 if (!tierAnswer || !isTier(tierAnswer.choice)) return null
191
192 const effortAnswer = answers.effort
193 const riskyAnswer = answers.risky
194 const risky =
195 typeof riskyAnswer?.noul === 'number'
196 ? riskyAnswer.noul
197 : typeof riskyAnswer?.probability === 'number'
198 ? riskyAnswer.probability
199 : null
200
201 return {
202 tier: tierAnswer.choice,
203 confidence: confidenceOf(tierAnswer),
204 effort: typeof effortAnswer?.score === 'number' ? effortAnswer.score : null,
205 effortConfidence: effortAnswer ? confidenceOf(effortAnswer) : null,
206 risky,
207 }
208}
209
210/**
211 * How sure an answer is. TypeSafe reports it; the Gateway does not, so there
212 * it is the highest probability of a distribution that is itself optional.
213 */
214function confidenceOf(answer: Record<string, unknown>): number | null {
215 if (typeof answer.confidence === 'number') return answer.confidence
216 const probabilities = answer.probabilities as Record<string, number> | undefined
217 const values = probabilities ? Object.values(probabilities) : []
218 return values.length > 0 ? Math.max(...values) : null
219}
220
221/** The rubric score (0..3) as a reasoning level. */
222export function effortLevel(score: number): Effort {
223 const index = Math.min(EFFORT_ORDER.length - 1, Math.max(0, Math.round(score)))
224 return EFFORT_ORDER[index] as Effort
225}
226
227/**
228 * Where a reasoning level sits on the ladder, or null when its place cannot
229 * be known. `max` is above every rung the rubric can produce, so it ranks
230 * above them without joining EFFORT_ORDER, which is also the set of values
231 * this router is allowed to ask for.
232 */
233export function effortRank(effort: string | number | undefined): number | null {
234 if (typeof effort !== 'string') return null
235 if (effort === 'max') return EFFORT_ORDER.length
236 const index = EFFORT_ORDER.indexOf(effort as Effort)
237 return index === -1 ? null : index
238}
239
240/**
241 * Where a model id sits on the tier ladder, by matching it against the
242 * configured tier names first and then the family words. Null when it matches
243 * none, in which case the change is treated as an upgrade rather than guessed
244 * at: an unrecognised id gets the gentler threshold, never the strict one.
245 */
246export function rankOf(model: string, tiers: Tiers): number | null {
247 const lowered = model.toLowerCase()
248 for (let index = 0; index < TIER_ORDER.length; index++) {
249 const tier = TIER_ORDER[index] as Tier
250 const configured = tiers[tier].toLowerCase()
251 if (configured && lowered.includes(configured)) return index
252 }
253 if (lowered.includes('haiku')) return 0
254 if (lowered.includes('sonnet')) return 1
255 if (lowered.includes('opus')) return 2
256 return null
257}
258
259/**
260 * The full id a family alias names on the main loop.
261 *
262 * `agent.spawn` takes an alias (`haiku`) the way the Agent tool does, but
263 * `turn.step`'s `model` is the id the engine already resolved for the request
264 * and goes to the API as written: an alias there is refused ("There's an
265 * issue with the selected model (haiku)"). So the tiers stay aliases in the
266 * options, and only a main-loop rewrite resolves them, here.
267 */
268const ALIAS_IDS: Record<string, string> = {
269 haiku: 'claude-haiku-4-5-20251001',
270 sonnet: 'claude-sonnet-5',
271 opus: 'claude-opus-5',
272}
273
274/**
275 * What to write into `turn.step`'s `model`: a full id as given, or the id
276 * behind a family alias. Anything else is returned unchanged for the engine
277 * to judge.
278 */
279export function requestModelId(model: string): string {
280 return ALIAS_IDS[model.trim().toLowerCase()] ?? model
281}
282
283export interface PolicyConfig {
284 tiers: Tiers
285 /**
286 * How sure the decision must be to spend more (a bigger model, more
287 * reasoning). Being wrong here costs money, so the bar is low.
288 */
289 minUpgradeConfidence: number
290 /**
291 * How sure it must be to spend less. Being wrong here means a task handled
292 * by too small a model or too little thought, so the bar is high.
293 */
294 minDowngradeConfidence: number
295}
296
297export interface Routing {
298 /** The model to run on, or null to leave the request as it is. */
299 model: string | null
300 /** The reasoning level to ask for, or null to leave it as it is. */
301 effort: Effort | null
302 /** Why, for the log line. */
303 reason: string
304}
305
306const NOTHING: Routing = { model: null, effort: null, reason: 'no decision' }
307
308/**
309 * Whether a change of rank passes its threshold. Both directions are allowed;
310 * they just do not have to clear the same bar, because the two mistakes do not
311 * cost the same. A move whose direction cannot be told (an unrecognised
312 * current value) is treated as an upgrade.
313 */
314function allowed(
315 wanted: number,
316 current: number | null,
317 confidence: number | null,
318 config: PolicyConfig,
319): boolean {
320 if (current !== null && wanted === current) return false
321 const isDowngrade = current !== null && wanted < current
322 const bar = isDowngrade ? config.minDowngradeConfidence : config.minUpgradeConfidence
323 // A backend that reports no confidence (the Gateway without a distribution,
324 // or the built-in classifier) clears the upgrade bar but never the
325 // downgrade one: spending less on an unmeasured hunch is the bad trade.
326 if (confidence === null) return !isDowngrade
327 return confidence >= bar
328}
329
330/**
331 * Turns a decision into a model and a reasoning level, either of which may be
332 * null to leave the request as it is. Both can move in either direction.
333 */
334export function route(
335 decision: Decision | null,
336 current: { model: string; effort?: string | number },
337 config: PolicyConfig,
338): Routing {
339 if (!decision) return NOTHING
340
341 let tier = decision.tier
342 let effortScore = decision.effort
343 let forced = false
344
345 // Carrying out something final is never worth the saving: take the deep
346 // tier and real reasoning, whatever the cheaper answer said, and skip the
347 // thresholds — this is the one case that is not a confidence question.
348 if (decision.risky !== null && decision.risky > 0.7) {
349 tier = 'deep'
350 effortScore = Math.max(effortScore ?? 0, 2)
351 forced = true
352 }
353
354 const wantedTier = TIER_ORDER.indexOf(tier)
355 const currentTier = rankOf(current.model, config.tiers)
356 const wantedModel = config.tiers[tier]
357
358 const model =
359 wantedModel &&
360 wantedModel !== current.model &&
361 (forced || allowed(wantedTier, currentTier, decision.confidence, config))
362 ? wantedModel
363 : null
364
365 let effort: Effort | null = null
366 if (effortScore !== null) {
367 const currentRank = effortRank(current.effort)
368 let wantedRank = EFFORT_ORDER.indexOf(effortLevel(effortScore))
369
370 // Risk raises the floor; it must never lower one. Forcing only skips the
371 // thresholds, so without this clamp a task already at `xhigh` or `max`
372 // and rated mechanically simple would be pulled down to `high` with no
373 // confidence check at all — the opposite of what the rule is for.
374 if (forced && currentRank !== null) wantedRank = Math.max(wantedRank, currentRank)
375
376 // A numeric effort is the caller's own scale, not this ladder; leave it.
377 const comparable = typeof current.effort !== 'number'
378 const wanted = EFFORT_ORDER[Math.min(EFFORT_ORDER.length - 1, wantedRank)] as Effort
379 if (
380 comparable &&
381 wantedRank !== currentRank &&
382 (forced || allowed(wantedRank, currentRank, decision.effortConfidence, config))
383 ) {
384 effort = wanted
385 }
386 }
387
388 const said = decision.confidence === null ? 'confidence n/d' : `confidence ${decision.confidence.toFixed(2)}`
389
390 if (!model && !effort) {
391 // Naming what it wanted and what it kept is the whole point of this line.
392 // Without it, a mod that classified and decided to leave the request alone
393 // is indistinguishable from one that never loaded.
394 const wantedEffort = effortScore === null ? null : effortLevel(effortScore)
395 const kept = `${current.model}${current.effort === undefined ? '' : `/${current.effort}`}`
396 const wanted = `${wantedModel}${wantedEffort ? `/${wantedEffort}` : ''}`
397 return { model: null, effort: null, reason: `kept ${kept}, wanted ${wanted} (${said})` }
398 }
399
400 return { model, effort, reason: forced ? `${tier}, forced by risk` : `${tier} (${said})` }
401}
402
403/**
404 * Whether a prompt is a slash command and nothing else (`/simplify`). The
405 * decision model sees only the name, never the skill or command it runs:
406 * measured on TypeSafe (three calls each), `/simplify`, `/run` and `/github`
407 * came back fast with effort 0.1 to 0.5, low effort for a multi-step skill,
408 * while `/code-review` and `/security-review` came back balanced and deep.
409 * With text after the name there is a task to read, and it is classified.
410 * A one-segment path alone (`/etc`) matches too; nobody sends one as a task.
411 */
412export function bareCommand(text: string): boolean {
413 return /^\/[^\s/]+$/.test(text.trim())
414}
415
416/**
417 * Holds a prompt's classification until the turn that reads that prompt
418 * starts.
419 *
420 * Nothing ties a decision to the turn it belongs to. Prompts can be queued
421 * while the model is busy, a peer session's message can be delivered inside a
422 * running turn, and `prompt.submit` carries no turn id at all while the
423 * session is idle. So when more than one prompt is waiting, `take` reports
424 * none: running a turn on another prompt's decision is a worse outcome than
425 * not routing it, and not routing is what every other failure path here does.
426 */
427export function pendingDecisions(): {
428 put(decision: Decision | null): void
429 take(): Decision | null
430} {
431 let held: Decision | null = null
432 let waiting = 0
433
434 return {
435 put(decision) {
436 waiting += 1
437 // Past the first, which prompt a turn will read is unknowable, so the
438 // slot is emptied instead of holding a decision that may not fit.
439 held = waiting === 1 ? decision : null
440 },
441 take() {
442 const decision = waiting === 1 ? held : null
443 held = null
444 waiting = 0
445 return decision
446 },
447 }
448}
449
450/** A number for the log, or `n/d` when the backend reported none. */
451function reported(value: number | null): string {
452 return value === null ? 'n/d' : value.toFixed(2)
453}
454
455/**
456 * The one-time line that says the router is alive, which backend answers it,
457 * and which of the three switches are on.
458 *
459 * Without this, a router that loaded and a router that never loaded are told
460 * apart only by the absence of later lines, which is not evidence of anything.
461 */
462export function describeSetup(
463 provider: Provider | null,
464 url: string,
465 switches: { subagentModel: boolean; mainEffort: boolean; mainModel: boolean },
466 // `provider: "builtin"` is a choice, not a missing key. Reporting it as a
467 // credential problem sends someone hunting for a key they meant to omit.
468 builtinByChoice = false,
469): string {
470 const backend = provider
471 ? `${provider} (${url})`
472 : builtinByChoice
473 ? 'the built-in classifier, by choice'
474 : 'the built-in classifier, no key set'
475 const on = [
476 switches.subagentModel && 'subagent model',
477 switches.mainEffort && 'main effort',
478 switches.mainModel && 'main model',
479 ].filter(Boolean)
480 return `ready on ${backend}; routing ${on.length > 0 ? on.join(', ') : 'nothing, every switch is off'}`
481}
482
483/**
484 * What the decision model answered, before any policy touches it: the raw
485 * tier, effort and risk with their confidences, and how long it took.
486 *
487 * This is the line that shows the classification happened at all, separately
488 * from whether the policy then decided to act on it.
489 */
490export function describeDecision(decision: Decision | null, ms: number | null): string {
491 const took = ms === null ? '' : ` · ${Math.round(ms)}ms`
492 if (!decision) return `no answer${took}`
493
494 const parts = [`tier ${decision.tier} (${reported(decision.confidence)})`]
495 if (decision.effort !== null) {
496 parts.push(
497 `effort ${decision.effort.toFixed(1)} → ${effortLevel(decision.effort)} (${reported(decision.effortConfidence)})`,
498 )
499 }
500 if (decision.risky !== null) parts.push(`risky ${reported(decision.risky)}`)
501 return parts.join(' · ') + took
502}
503
504/**
505 * The persistent status line: the last thing the router did, short enough to
506 * sit on screen beside the engine's own notices.
507 */
508export function describeStatus(
509 decision: Decision | null,
510 change: { model?: string; effort?: Effort } | null,
511): string {
512 if (!decision) return 'jev · no answer'
513 const asked = `${decision.tier} ${reported(decision.confidence)}`
514 if (!change) return `jev · ${asked} · unchanged`
515 const to = [change.model, change.effort].filter(Boolean).join('/')
516 return `jev · ${asked} → ${to}`
517}
518