Picks the model and the reasoning effort for each task with TypeSafe's Jev, a System One decision model reached either through TypeSafe's own API or the Vercel…

Picks the model and the reasoning effort each task runs with, using Jev, TypeSafe's System One decision model: unstructured state in, a typed choice with a probability distribution out, no free-form text.
Two backends, chosen by whichever key is set:
| Backend | Endpoint | Model | Confidence |
|---|---|---|---|
typesafe | POST api.typesafe.ai/v1/systemone | jev-latest | reported per answer |
gateway | POST ai-gateway.vercel.sh/v4/ai/evaluation-model | typesafe-ai/jev | derived from an optional distribution |
TypeSafe's own API wins when both keys are set: it is the only one that reports a calibrated confidence, which is what the confidence bars below read. Set provider to force one, or to builtin to use neither. Each backend keeps its own URL and model option, so an override written for one is never sent to the other. A provider forced onto a backend whose key is missing degrades to the built-in classifier and says so once in the log.
Three switches, and they are not equally safe:
| Switch | What it sets | Default |
|---|---|---|
routeSubagentModel | the model of each subagent, at agent.spawn | on |
routeMainEffort | the reasoning effort of the main conversation, at turn.step | on |
routeMainModel | the model of the main conversation, at turn.step | off |
A subagent starts with its own context, so routing its model costs nothing beyond the classification. Changing the main loop's model mid-session is the expensive one: it invalidates the prompt cache, and on a long context re-caching can cost more than the cheaper tier saves. Turn it on once you have measured your own sessions, not before.
The Agent tool has no effort parameter, so a subagent's effort is not this mod's to set.
Both directions, both dimensions. A task read as mechanical is routed down; one read as hard is routed up — model and effort alike.
The prompt is classified at prompt.submit, which runs before the turn starts, and the decision is applied to the turn's first model request and reused by the rest of that turn.
With no key configured the mod still works: it falls back to the engine's own $.model.classify, which answers the same question with the small fast model. That path reports no confidence, so the threshold does not apply to it.
One request, three questions evaluated in parallel:
tier — a choice between three descriptions of the work (mechanical and local / ordinary engineering / hard or high-stakes). The decision model never sees a model name.effort — a score on a four-level rubric, for how much step-by-step reasoning the task needs.risky — whether the task touches production, money, credentials, or state that cannot be undone. A noul on TypeSafe's API, a boolean on the Gateway: the same question under two names.TypeSafe's API reports a confidence per answer. The Gateway's answer shape carries no confidence field, so on that backend confidence is read as the highest probability in the distribution — and that distribution is itself optional in the schema, in which case confidence is absent and the threshold does not fire.
The two mistakes do not cost the same, so they do not clear the same bar:
minUpgradeConfidence, 0.3 by default. Being wrong costs money.minDowngradeConfidence, 0.6 by default. Being wrong means a task handled by too small a model or too little thought.risky above 0.7 takes the deep tier and real reasoning, past both bars. That one is not a confidence question.Every other failure — a non-2xx response, a timeout, a malformed body, a thrown error — leaves the request exactly as the engine built it. The router never blocks a turn.
With logDecisions on (the default), the router reports every step of its own work, because nothing else in Claude Code shows it: the model and effort it rewrites are parameters of each request, not the session's settings, so the status line, the header and the effort box never move whatever it decides.
[jev-model-router] ready on typesafe (https://api.typesafe.ai/v1/systemone); routing subagent model, main effort
[jev-model-router] jev: tier fast (0.87) · effort 0.4 → low (0.71) · risky 0.02 · 249ms
[jev-model-router] main loop → effort low: fast (confidence 0.87)
[jev-model-router] jev: tier fast (0.41) · effort 0.4 → low (0.38) · risky 0.01 · 210ms
[jev-model-router] main loop: kept opus/medium, wanted haiku/low (confidence 0.41)
jev: line is what the decision model replied, before any policy is applied — the tier, the effort score and the risk, each with its confidence, and how long the call took.minDowngradeConfidence.It also keeps a one-line status on screen, replaced as it goes:
jev · fast 0.87 → haiku/low
jev · fast 0.41 · unchanged
confidence n/d means the backend reported no confidence, which the built-in classifier never does and the Gateway does whenever its probability distribution is absent; the ready on line says which one answered.
No lines at all has three causes, and only the last is the module failing to load. Check them in this order:
claude -p (or the SDK). A headless run has no transcript and no status row: every line still goes to the debug log, ~/.claude/debug/<session-id>.txt (a .txt, not a .log; latest is a symlink to the newest), and an SDK host receives each one as ui_log. The router is working; look there..claude/skills/ (where --mod writes it) only once the project is trusted: it is repository content, so an untrusted folder's .claude/ is not read at all, and claude -p never asks. Open claude interactively in the folder and accept the trust prompt, or name the plugin explicitly with --plugin-dir (see Install). claude --debug settles it: a loaded module prints hooks module jev-model-router@skills-dir loaded (worker, …); events: prompt.submit,turn.step,agent.spawn (@inline when loaded with --plugin-dir); Found N plugins without it means the plugin is not in the session.CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 the debug log says installed plugins' hooks modules not loaded: rollout flag (tengu_plugin_hooks_modules) is off. Set the flag; Claude Code must be 2.1.259+.A ready on the built-in classifier, no key set line when you did set a key means the key sits under the wrong pluginConfigs entry: the key must match the plugin's id, which depends on how it was loaded (see Options).
With a key set, the prompt text leaves the machine and goes to whichever backend the key belongs to. The main-loop path sends the prompt; the subagent path sends the subagent's prompt, its description and its agent type. Nothing else. With no key set, nothing leaves the machine.
typesafeApiKey: string TypeSafe API key (preferred: it reports a confidence)
gatewayApiKey: string Vercel AI Gateway key
provider: string "auto" | "typesafe" | "gateway" | "builtin"
typesafeBaseUrl: string empty uses https://api.typesafe.ai
typesafeModel: string empty uses jev-latest
gatewayBaseUrl: string empty uses https://ai-gateway.vercel.sh/v4/ai
gatewayModel: string empty uses typesafe-ai/jev
fastModel: string fast tier, alias or full id (default "haiku")
balancedModel: string balanced tier, alias or full id (default "sonnet")
deepModel: string deep tier, alias or full id (default "opus")
minUpgradeConfidence: number bar to spend more (default 0.3)
minDowngradeConfidence: number bar to spend less (default 0.6)
routeSubagentModel: boolean model of each subagent (default true)
routeMainEffort: boolean effort of the main loop (default true)
routeMainModel: boolean model of the main loop (default false)
timeoutMs: number latency budget per classification (default 800)
logDecisions: boolean log each decision (default true)
The three tiers take an alias (haiku, sonnet, opus) or a full model id. A subagent is spawned with the name as given, the way the Agent tool takes it; the main loop's request needs an id, so there an alias is resolved to the family's current id (haiku → claude-haiku-4-5-20251001, sonnet → claude-sonnet-5, opus → claude-opus-5). Set a full id to pin a specific version. A decision for the tier the session already runs is not a change, so a session on claude-opus-5[1m] keeps its 1M-context id.
Declared in .claude-plugin/plugin.json (userConfig). Set them in /config, in user settings (~/.claude/settings.json, not project settings), with --settings <file> or in managed settings:
{ "pluginConfigs": { "jev-model-router": { "options": { "typesafeApiKey": "" } } } }
The entry's key is the plugin's id, and the id follows how the plugin was loaded: "jev-model-router@skills-dir" when auto-loaded from .claude/skills/ (the --mod install), "jev-model-router" with --plugin-dir. Under the wrong key every option stays at its default, and the ready on line reports no key set.
npx claude-code-templates@latest --mod productivity/jev-model-router
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude
--mod writes the plugin to .claude/skills/jev-model-router/ in the project, and Claude Code auto-loads it as jev-model-router@skills-dir in a trusted project: a folder's .claude/ is repository content and is not read until you accept the trust prompt on the first interactive claude there (-p never asks, so a headless run in a fresh folder never sees it). The options then go under the "jev-model-router@skills-dir" key in pluginConfigs (see Options).
For one session with hot reload, or in a folder you do not want to trust, name it on the command line instead — it loads as jev-model-router@inline and reads options from the "jev-model-router" key:
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir .claude/skills/jev-model-router
Either way, claude plugin validate .claude/skills/jev-model-router prints every event it hooks and every $ call it makes.
bun test cli-tool/components/mods/productivity/jev-model-router/tests
Early access. Mods need Claude Code 2.1.259+ with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1; the $ API may change between releases. Typed against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
A mod runs without node_modules, so neither @typesafe-ai/sdk nor the AI SDK is available here: both backends are spoken to over HTTP through $.http.fetch. The TypeSafe wire shape was read from @typesafe-ai/sdk v0.6.0; the Gateway's, which is experimental in the AI SDK (experimental_evaluate, 7.0.105+) and not documented publicly, from @ai-sdk/gateway v4.0.86 and @ai-sdk/provider v4.0.17. Either may change.
hooks/jev-model-router.ts 328 lines1/**
2 * jev-model-router — Claude Mod (EARLY ACCESS)
3 *
4 * Picks the model each task runs on with TypeSafe's Jev, a System One
5 * decision model: unstructured state in, a typed choice with a probability
6 * distribution out.
7 *
8 * Jev is reached one of two ways, whichever key is configured: TypeSafe's
9 * own API (`typesafeApiKey`), which reports a calibrated confidence per
10 * answer, or the Vercel AI Gateway (`gatewayApiKey`), which does not. With
11 * neither, the engine's own `$.model.classify` stands in, so the mod is
12 * useful without any account.
13 *
14 * Three things it can set, each on its own switch:
15 * agent.spawn — the model of each subagent (on by default)
16 * turn.step — the reasoning effort of the main loop (on by default)
17 * turn.step — the model of the main loop (off by default: switching
18 * models mid-session invalidates the prompt cache, which can
19 * cost more than the cheaper tier saves)
20 *
21 * Every one of them moves in both directions: a task the decision model reads
22 * as mechanical is routed down, one it reads as hard is routed up. The two
23 * mistakes do not cost the same, so they do not clear the same confidence bar
24 * (see `minUpgradeConfidence` / `minDowngradeConfidence` in policy.ts).
25 *
26 * The Agent tool has no effort parameter, so a subagent's effort is not ours
27 * to set; only its model is.
28 *
29 * The prompt is classified at `prompt.submit`, which runs before the turn
30 * starts, and the decision is applied at the turn's first request.
31 *
32 * Every failure path is fail-open: a classification that errors or runs past
33 * the latency budget leaves the request exactly as the engine built it.
34 *
35 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
36 * or "gatewayApiKey"). Never hardcode it in this file.
37 *
38 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (Claude Code >= 2.1.259). Typed
39 * against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
40 *
41 * Privacy: with a key set, the prompt text is sent to whichever backend the
42 * key belongs to.
43 */
44import type { Register } from 'claude-code'
45import {
46 DEFAULT_BASE_URL,
47 DEFAULT_MODEL,
48 describeDecision,
49 describeSetup,
50 describeStatus,
51 endpoint,
52 pendingDecisions,
53 readDecision,
54 selectProvider,
55 requestBody,
56 requestHeaders,
57 requestModelId,
58 route,
59 TIER_ORDER,
60} from './policy.ts'
61import type { Decision, Effort, PolicyConfig, Provider, Tier } from './policy.ts'
62
63export const register: Register = (on, options) => {
64 const text = (key: string, fallback: string) =>
65 typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
66 const number = (key: string, fallback: number) =>
67 typeof options[key] === 'number' ? (options[key] as number) : fallback
68 const flag = (key: string, fallback: boolean) =>
69 typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
70
71 // TypeSafe's own API is preferred when both keys are set: it is the only
72 // one that reports a calibrated confidence, which the policy's threshold
73 // reads. `provider` forces one, including "builtin" to use neither.
74 const typesafeKey = text('typesafeApiKey', '')
75 const gatewayKey = text('gatewayApiKey', '')
76 const forced = text('provider', 'auto')
77 const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey)
78
79 // Each backend keeps its own URL and model, so an override written for one
80 // can never be sent to the other when `auto` picks differently than expected.
81 const apiKey = active === 'typesafe' ? typesafeKey : active === 'gateway' ? gatewayKey : ''
82 const modelId = !active
83 ? ''
84 : active === 'typesafe'
85 ? text('typesafeModel', DEFAULT_MODEL.typesafe)
86 : text('gatewayModel', DEFAULT_MODEL.gateway)
87 const url = !active
88 ? ''
89 : active === 'typesafe'
90 ? endpoint('typesafe', text('typesafeBaseUrl', DEFAULT_BASE_URL.typesafe))
91 : endpoint('gateway', text('gatewayBaseUrl', DEFAULT_BASE_URL.gateway))
92
93 // A backend named in the options but missing its key degrades to the
94 // built-in classifier, which is silent; say so once, when a hook first runs.
95 let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
96
97 const timeoutMs = number('timeoutMs', 800)
98 const routeSubagentModel = flag('routeSubagentModel', true)
99 const routeMainEffort = flag('routeMainEffort', true)
100 const routeMainModel = flag('routeMainModel', false)
101 const routeMainLoop = routeMainEffort || routeMainModel
102 const logDecisions = flag('logDecisions', true)
103
104 const policy: PolicyConfig = {
105 tiers: {
106 fast: text('fastModel', 'haiku'),
107 balanced: text('balancedModel', 'sonnet'),
108 deep: text('deepModel', 'opus'),
109 },
110 minUpgradeConfidence: number('minUpgradeConfidence', 0.3),
111 minDowngradeConfidence: number('minDowngradeConfidence', 0.6),
112 }
113
114 // The classification waiting for the turn that reads its prompt, and what
115 // the current turn settled on. Both are single slots: main-loop turns run
116 // one at a time, so nothing accumulates over a long session. `pending`
117 // reports no decision when two prompts are waiting at once, rather than
118 // routing a turn on a decision made for a different prompt.
119 const pending = pendingDecisions()
120 // Said once, the first time a hook runs. A router that loaded and one that
121 // never loaded are otherwise told apart only by the absence of later lines,
122 // and absence is not evidence: the policy leaves most turns alone anyway.
123 let announced = false
124 let appliedTurnId: string | undefined
125 let applied: { model?: string; effort?: Effort } | null = null
126
127 on('prompt.submit', async ($, e, next) => {
128 // Before the routing guards: a module whose switches are all off has still
129 // loaded, and that is exactly when its silence is most misleading.
130 if (!announced) {
131 announced = true
132 if (logDecisions) {
133 $.ui.log(
134 `[jev-model-router] ${describeSetup(
135 active,
136 url,
137 {
138 subagentModel: routeSubagentModel,
139 mainEffort: routeMainEffort,
140 mainModel: routeMainModel,
141 },
142 forced === 'builtin',
143 )}`,
144 )
145 }
146 }
147 if (!routeMainLoop) return next(e)
148
149 if (!unusableReported) {
150 unusableReported = true
151 $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
152 }
153
154 const startedAt = await $.clock.now()
155 let decision: Decision | null = null
156 if (active) {
157 try {
158 const response = await Promise.race([
159 $.http.fetch(url, {
160 method: 'POST',
161 headers: requestHeaders(active, apiKey, modelId),
162 body: requestBody(active, { prompt: e.text }, modelId),
163 }),
164 $.clock.sleep(timeoutMs),
165 ])
166 if (response && response.ok) decision = readDecision(response.text)
167 else if (response) $.ui.log(`[jev-model-router] ${active} responded ${response.status}`)
168 else $.ui.log(`[jev-model-router] classification passed ${timeoutMs}ms; leaving the turn alone`)
169 } catch (error) {
170 $.ui.log(`[jev-model-router] classification failed: ${String(error)}`)
171 }
172 } else {
173 // No backend: the engine's own small-model classifier answers the same
174 // question, without the confidence the policy's threshold reads.
175 try {
176 const label = await $.model.classify(e.text, TIER_ORDER)
177 if (label) {
178 decision = {
179 tier: label as Tier,
180 confidence: null,
181 risky: null,
182 effort: null,
183 effortConfidence: null,
184 }
185 }
186 } catch (error) {
187 $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
188 }
189 }
190
191 // What the decision model actually answered, whatever the policy then
192 // does with it. This is the line that proves the classification ran.
193 if (logDecisions) {
194 const ms = (await $.clock.now()) - startedAt
195 $.ui.log(`[jev-model-router] jev: ${describeDecision(decision, ms)}`)
196 }
197
198 pending.put(decision)
199 return next(e)
200 })
201
202 on('turn.step', async function* ($, e, next) {
203 if (!routeMainLoop || e.agentId) return yield* next(e)
204
205 // Every request after the first reuses what the turn settled on, so
206 // neither the model nor the effort changes under its own tool loop.
207 if (e.index > 0 && e.turnId === appliedTurnId) {
208 return yield* next(applied ? { ...e, ...applied } : e)
209 }
210
211 const decision = pending.take()
212 const routing = route(decision, { model: e.model, effort: e.effort }, policy)
213 const change: { model?: string; effort?: Effort } = {}
214 // The main loop's `model` is sent to the API as written, so an alias
215 // becomes its id here; a subagent's (agent.spawn) may stay an alias.
216 if (routeMainModel && routing.model) change.model = requestModelId(routing.model)
217 if (routeMainEffort && routing.effort) change.effort = routing.effort
218
219 appliedTurnId = e.turnId
220 applied = Object.keys(change).length > 0 ? change : null
221 // A row in the transcript scrolls away; this line stays on screen.
222 if (logDecisions) $.ui.status(describeStatus(decision, applied))
223
224 if (!applied) {
225 // A turn left alone is the common case, and it used to be silent, which
226 // made a working mod look like one that never loaded. Say what happened.
227 if (logDecisions) {
228 const suppressed = routing.model && !routeMainModel ? ' (main-loop model routing off)' : ''
229 $.ui.log(`[jev-model-router] main loop: ${routing.reason}${suppressed}`)
230 }
231 return yield* next(e)
232 }
233 if (logDecisions) {
234 const what = [change.model, change.effort && `effort ${change.effort}`]
235 .filter(Boolean)
236 .join(', ')
237 $.ui.log(`[jev-model-router] main loop → ${what}: ${routing.reason}`)
238 }
239 return yield* next({ ...e, ...change })
240 })
241
242 on('agent.spawn', async ($, e, next) => {
243 // Before the routing guards: a module whose switches are all off has still
244 // loaded, and that is exactly when its silence is most misleading.
245 if (!announced) {
246 announced = true
247 if (logDecisions) {
248 $.ui.log(
249 `[jev-model-router] ${describeSetup(
250 active,
251 url,
252 {
253 subagentModel: routeSubagentModel,
254 mainEffort: routeMainEffort,
255 mainModel: routeMainModel,
256 },
257 forced === 'builtin',
258 )}`,
259 )
260 }
261 }
262
263 // A fork inherits its parent's model; `model` is ignored for it.
264 if (!routeSubagentModel || e.fork) return next(e)
265
266 if (!unusableReported) {
267 unusableReported = true
268 $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
269 }
270
271 const startedAt = await $.clock.now()
272 let decision: Decision | null = null
273 if (active) {
274 try {
275 const response = await Promise.race([
276 $.http.fetch(url, {
277 method: 'POST',
278 headers: requestHeaders(active, apiKey, modelId),
279 body: requestBody(
280 active,
281 { prompt: e.prompt, description: e.description, agentType: e.subagentType },
282 modelId,
283 ),
284 }),
285 $.clock.sleep(timeoutMs),
286 ])
287 if (response && response.ok) decision = readDecision(response.text)
288 else if (response) $.ui.log(`[jev-model-router] ${active} responded ${response.status}`)
289 else $.ui.log(`[jev-model-router] classification passed ${timeoutMs}ms; leaving the subagent alone`)
290 } catch (error) {
291 $.ui.log(`[jev-model-router] classification failed: ${String(error)}`)
292 }
293 } else {
294 try {
295 const label = await $.model.classify(e.prompt, TIER_ORDER)
296 if (label) {
297 decision = {
298 tier: label as Tier,
299 confidence: null,
300 risky: null,
301 effort: null,
302 effortConfidence: null,
303 }
304 }
305 } catch (error) {
306 $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
307 }
308 }
309
310 if (logDecisions) {
311 const ms = (await $.clock.now()) - startedAt
312 $.ui.log(`[jev-model-router] jev (${e.subagentType}): ${describeDecision(decision, ms)}`)
313 }
314
315 // The subagent's own model wins when the caller named one; otherwise it
316 // would inherit the parent's, so that is what a change is measured from.
317 // The Agent tool takes no effort, so only the model is ours to set here.
318 const current = e.model ?? e.parentModel
319 const { model, reason } = route(decision, { model: current }, policy)
320 if (!model) {
321 if (logDecisions) $.ui.log(`[jev-model-router] ${e.subagentType}: ${reason}`)
322 return next(e)
323 }
324 if (logDecisions) $.ui.log(`[jev-model-router] ${e.subagentType} → ${model}: ${reason}`)
325 return next({ ...e, model })
326 })
327}
328hooks/policy.ts 505 lines1/**
2 * jev-model-router — pure decision logic.
3 *
4 * No `$` and no I/O here: this module only builds the request the decision
5 * API takes, reads its answer, and turns that answer into a model id. The
6 * hooks module does every call on `$` at its own call site.
7 *
8 * Two backends speak to the same model with different wire shapes:
9 *
10 * typesafe POST https://api.typesafe.ai/v1/systemone
11 * `{ model, state, questions }`; a yes/no question is a `noul`
12 * and every answer carries its own `confidence`.
13 * gateway POST https://ai-gateway.vercel.sh/v4/ai/evaluation-model
14 * `{ state, questions }` with the model in a header; a yes/no
15 * question is a `boolean`, and there is no `confidence` field —
16 * it has to be derived from an optional distribution.
17 *
18 * The Gateway shape is not documented publicly; it was read from
19 * @ai-sdk/gateway and @ai-sdk/provider.
20 */
21
22export type Provider = 'typesafe' | 'gateway'
23
24export type Tier = 'fast' | 'balanced' | 'deep'
25
26export interface Tiers {
27 fast: string
28 balanced: string
29 deep: string
30}
31
32export interface Decision {
33 tier: Tier
34 /** Confidence in the tier, or null when the backend reported none. */
35 confidence: number | null
36 /** P(true) that carrying the task out would itself be costly or final. */
37 risky: number | null
38 /** 0..3 along the effort rubric, or null when absent. */
39 effort: number | null
40 /** Confidence in the effort, or null when the backend reported none. */
41 effortConfidence: number | null
42}
43
44/** The reasoning levels a turn can ask for, cheapest first. */
45export const EFFORT_ORDER = ['low', 'medium', 'high', 'xhigh'] as const
46
47export type Effort = (typeof EFFORT_ORDER)[number]
48
49export const TIER_ORDER: readonly Tier[] = ['fast', 'balanced', 'deep']
50
51/**
52 * How each tier is described to the decision model. Deliberately about the
53 * shape of the work, not about model names: the model never sees an id.
54 */
55const TIER_CRITERIA: Record<Tier, string> = {
56 fast: 'Mechanical and local: read or summarise a file, run one command, rename a symbol, answer something already in context.',
57 balanced:
58 'Ordinary engineering: implement a well-specified change across a few files, write tests, fix a clearly described bug, review a small diff.',
59 deep: 'Hard or high-stakes: architecture and design, debugging a failure whose cause is unknown, security, data migrations, concurrency, anything touching production or money.',
60}
61
62const EFFORT_RUBRIC = ['almost none', 'some', 'a lot', 'as much as possible'] as const
63
64export const DEFAULT_BASE_URL: Record<Provider, string> = {
65 typesafe: 'https://api.typesafe.ai',
66 gateway: 'https://ai-gateway.vercel.sh/v4/ai',
67}
68
69/**
70 * The Gateway's own protocol version, sent as `ai-gateway-protocol-version`.
71 * Tracks the `AI_GATEWAY_PROTOCOL_VERSION` of `@ai-sdk/gateway` (4.0.87).
72 */
73const AI_GATEWAY_PROTOCOL_VERSION = '0.0.1'
74
75export const DEFAULT_MODEL: Record<Provider, string> = {
76 typesafe: 'jev-latest',
77 gateway: 'typesafe-ai/jev',
78}
79
80/**
81 * Which backend a configuration asks for, or null for the built-in
82 * classifier. `auto` prefers TypeSafe, since it is the only one that reports
83 * a calibrated confidence; a forced backend whose key is missing resolves to
84 * null rather than falling through to the other one's key.
85 */
86export function selectProvider(
87 forced: string,
88 typesafeKey: string,
89 gatewayKey: string,
90): Provider | null {
91 if (forced === 'builtin') return null
92 if (forced === 'typesafe') return typesafeKey ? 'typesafe' : null
93 if (forced === 'gateway') return gatewayKey ? 'gateway' : null
94 if (typesafeKey) return 'typesafe'
95 if (gatewayKey) return 'gateway'
96 return null
97}
98
99/** The full endpoint a backend posts to. */
100export function endpoint(provider: Provider, baseUrl: string): string {
101 const root = baseUrl.replace(/\/+$/, '')
102 return provider === 'typesafe' ? `${root}/v1/systemone` : `${root}/evaluation-model`
103}
104
105/** The `questions` map, in the shape the backend's schema names. */
106export function questions(provider: Provider): Record<string, unknown> {
107 return {
108 tier: {
109 type: 'choice',
110 instructions: 'Which is the cheapest tier that can complete this coding task well?',
111 criteria: TIER_CRITERIA,
112 },
113 effort: {
114 type: 'score',
115 instructions: 'How much step-by-step reasoning does this task need?',
116 criteria: EFFORT_RUBRIC,
117 },
118 risky: {
119 // The same question under two names: `noul` on TypeSafe's own API,
120 // `boolean` in the AI SDK's evaluation schema.
121 type: provider === 'typesafe' ? 'noul' : 'boolean',
122 // Asked about the act, not the subject. The first wording ("the task
123 // touches production, money, credentials") scored 0.96 on "add a
124 // refund endpoint that calls Stripe" — ordinary code that happens to be
125 // about money — and would have escalated it past a 0.98-confidence
126 // answer of the balanced tier.
127 instructions:
128 'Carrying out this task would itself change production, move real money, or alter data that cannot be restored. Writing or testing code that deals with such things, without running it against the real system, does not count.',
129 },
130 }
131}
132
133/** The request body. The Gateway carries the model in a header instead. */
134export function requestBody(
135 provider: Provider,
136 state: Record<string, unknown>,
137 model: string,
138): string {
139 const body =
140 provider === 'typesafe'
141 ? { model, state, questions: questions(provider) }
142 : { state, questions: questions(provider) }
143 return JSON.stringify(body)
144}
145
146/** The request headers. */
147export function requestHeaders(
148 provider: Provider,
149 apiKey: string,
150 model: string,
151): Record<string, string> {
152 const common = { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` }
153 if (provider === 'typesafe') return common
154 return {
155 ...common,
156 'ai-gateway-auth-method': 'api-key',
157 'ai-model-id': model,
158 // The Gateway rejects any request that does not name the protocol it
159 // speaks: 400 "Unsupported gateway protocol version". Every other header
160 // here is accepted without it, so the omission fails the whole backend.
161 'ai-gateway-protocol-version': AI_GATEWAY_PROTOCOL_VERSION,
162 'ai-evaluation-model-specification-version': '4',
163 }
164}
165
166function isTier(value: unknown): value is Tier {
167 return value === 'fast' || value === 'balanced' || value === 'deep'
168}
169
170/**
171 * Reads a response from either backend.
172 *
173 * TypeSafe's own API reports a `confidence` per answer and a `noul` number
174 * for a yes/no question. The Gateway reports neither: confidence has to come
175 * from the highest probability of a distribution that is itself optional, and
176 * a yes/no answer arrives as `probability`. Both are handled, and a missing
177 * confidence reads as null rather than as a number the policy would trust.
178 */
179export function readDecision(responseText: string): Decision | null {
180 let parsed: unknown
181 try {
182 parsed = JSON.parse(responseText)
183 } catch {
184 return null
185 }
186 const answers = (parsed as { answers?: Record<string, Record<string, unknown>> }).answers
187 if (!answers) return null
188
189 const tierAnswer = answers.tier
190 if (!tierAnswer || !isTier(tierAnswer.choice)) return null
191
192 const effortAnswer = answers.effort
193 const riskyAnswer = answers.risky
194 const risky =
195 typeof riskyAnswer?.noul === 'number'
196 ? riskyAnswer.noul
197 : typeof riskyAnswer?.probability === 'number'
198 ? riskyAnswer.probability
199 : null
200
201 return {
202 tier: tierAnswer.choice,
203 confidence: confidenceOf(tierAnswer),
204 effort: typeof effortAnswer?.score === 'number' ? effortAnswer.score : null,
205 effortConfidence: effortAnswer ? confidenceOf(effortAnswer) : null,
206 risky,
207 }
208}
209
210/**
211 * How sure an answer is. TypeSafe reports it; the Gateway does not, so there
212 * it is the highest probability of a distribution that is itself optional.
213 */
214function confidenceOf(answer: Record<string, unknown>): number | null {
215 if (typeof answer.confidence === 'number') return answer.confidence
216 const probabilities = answer.probabilities as Record<string, number> | undefined
217 const values = probabilities ? Object.values(probabilities) : []
218 return values.length > 0 ? Math.max(...values) : null
219}
220
221/** The rubric score (0..3) as a reasoning level. */
222export function effortLevel(score: number): Effort {
223 const index = Math.min(EFFORT_ORDER.length - 1, Math.max(0, Math.round(score)))
224 return EFFORT_ORDER[index] as Effort
225}
226
227/**
228 * Where a reasoning level sits on the ladder, or null when its place cannot
229 * be known. `max` is above every rung the rubric can produce, so it ranks
230 * above them without joining EFFORT_ORDER, which is also the set of values
231 * this router is allowed to ask for.
232 */
233export function effortRank(effort: string | number | undefined): number | null {
234 if (typeof effort !== 'string') return null
235 if (effort === 'max') return EFFORT_ORDER.length
236 const index = EFFORT_ORDER.indexOf(effort as Effort)
237 return index === -1 ? null : index
238}
239
240/**
241 * Where a model id sits on the tier ladder, by matching it against the
242 * configured tier names first and then the family words. Null when it matches
243 * none, in which case the change is treated as an upgrade rather than guessed
244 * at: an unrecognised id gets the gentler threshold, never the strict one.
245 */
246export function rankOf(model: string, tiers: Tiers): number | null {
247 const lowered = model.toLowerCase()
248 for (let index = 0; index < TIER_ORDER.length; index++) {
249 const tier = TIER_ORDER[index] as Tier
250 const configured = tiers[tier].toLowerCase()
251 if (configured && lowered.includes(configured)) return index
252 }
253 if (lowered.includes('haiku')) return 0
254 if (lowered.includes('sonnet')) return 1
255 if (lowered.includes('opus')) return 2
256 return null
257}
258
259/**
260 * The full id a family alias names on the main loop.
261 *
262 * `agent.spawn` takes an alias (`haiku`) the way the Agent tool does, but
263 * `turn.step`'s `model` is the id the engine already resolved for the request
264 * and goes to the API as written: an alias there is refused ("There's an
265 * issue with the selected model (haiku)"). So the tiers stay aliases in the
266 * options, and only a main-loop rewrite resolves them, here.
267 */
268const ALIAS_IDS: Record<string, string> = {
269 haiku: 'claude-haiku-4-5-20251001',
270 sonnet: 'claude-sonnet-5',
271 opus: 'claude-opus-5',
272}
273
274/**
275 * What to write into `turn.step`'s `model`: a full id as given, or the id
276 * behind a family alias. Anything else is returned unchanged for the engine
277 * to judge.
278 */
279export function requestModelId(model: string): string {
280 return ALIAS_IDS[model.trim().toLowerCase()] ?? model
281}
282
283export interface PolicyConfig {
284 tiers: Tiers
285 /**
286 * How sure the decision must be to spend more (a bigger model, more
287 * reasoning). Being wrong here costs money, so the bar is low.
288 */
289 minUpgradeConfidence: number
290 /**
291 * How sure it must be to spend less. Being wrong here means a task handled
292 * by too small a model or too little thought, so the bar is high.
293 */
294 minDowngradeConfidence: number
295}
296
297export interface Routing {
298 /** The model to run on, or null to leave the request as it is. */
299 model: string | null
300 /** The reasoning level to ask for, or null to leave it as it is. */
301 effort: Effort | null
302 /** Why, for the log line. */
303 reason: string
304}
305
306const NOTHING: Routing = { model: null, effort: null, reason: 'no decision' }
307
308/**
309 * Whether a change of rank passes its threshold. Both directions are allowed;
310 * they just do not have to clear the same bar, because the two mistakes do not
311 * cost the same. A move whose direction cannot be told (an unrecognised
312 * current value) is treated as an upgrade.
313 */
314function allowed(
315 wanted: number,
316 current: number | null,
317 confidence: number | null,
318 config: PolicyConfig,
319): boolean {
320 if (current !== null && wanted === current) return false
321 const isDowngrade = current !== null && wanted < current
322 const bar = isDowngrade ? config.minDowngradeConfidence : config.minUpgradeConfidence
323 // A backend that reports no confidence (the Gateway without a distribution,
324 // or the built-in classifier) clears the upgrade bar but never the
325 // downgrade one: spending less on an unmeasured hunch is the bad trade.
326 if (confidence === null) return !isDowngrade
327 return confidence >= bar
328}
329
330/**
331 * Turns a decision into a model and a reasoning level, either of which may be
332 * null to leave the request as it is. Both can move in either direction.
333 */
334export function route(
335 decision: Decision | null,
336 current: { model: string; effort?: string | number },
337 config: PolicyConfig,
338): Routing {
339 if (!decision) return NOTHING
340
341 let tier = decision.tier
342 let effortScore = decision.effort
343 let forced = false
344
345 // Carrying out something final is never worth the saving: take the deep
346 // tier and real reasoning, whatever the cheaper answer said, and skip the
347 // thresholds — this is the one case that is not a confidence question.
348 if (decision.risky !== null && decision.risky > 0.7) {
349 tier = 'deep'
350 effortScore = Math.max(effortScore ?? 0, 2)
351 forced = true
352 }
353
354 const wantedTier = TIER_ORDER.indexOf(tier)
355 const currentTier = rankOf(current.model, config.tiers)
356 const wantedModel = config.tiers[tier]
357
358 const model =
359 wantedModel &&
360 wantedModel !== current.model &&
361 (forced || allowed(wantedTier, currentTier, decision.confidence, config))
362 ? wantedModel
363 : null
364
365 let effort: Effort | null = null
366 if (effortScore !== null) {
367 const currentRank = effortRank(current.effort)
368 let wantedRank = EFFORT_ORDER.indexOf(effortLevel(effortScore))
369
370 // Risk raises the floor; it must never lower one. Forcing only skips the
371 // thresholds, so without this clamp a task already at `xhigh` or `max`
372 // and rated mechanically simple would be pulled down to `high` with no
373 // confidence check at all — the opposite of what the rule is for.
374 if (forced && currentRank !== null) wantedRank = Math.max(wantedRank, currentRank)
375
376 // A numeric effort is the caller's own scale, not this ladder; leave it.
377 const comparable = typeof current.effort !== 'number'
378 const wanted = EFFORT_ORDER[Math.min(EFFORT_ORDER.length - 1, wantedRank)] as Effort
379 if (
380 comparable &&
381 wantedRank !== currentRank &&
382 (forced || allowed(wantedRank, currentRank, decision.effortConfidence, config))
383 ) {
384 effort = wanted
385 }
386 }
387
388 const said = decision.confidence === null ? 'confidence n/d' : `confidence ${decision.confidence.toFixed(2)}`
389
390 if (!model && !effort) {
391 // Naming what it wanted and what it kept is the whole point of this line.
392 // Without it, a mod that classified and decided to leave the request alone
393 // is indistinguishable from one that never loaded.
394 const wantedEffort = effortScore === null ? null : effortLevel(effortScore)
395 const kept = `${current.model}${current.effort === undefined ? '' : `/${current.effort}`}`
396 const wanted = `${wantedModel}${wantedEffort ? `/${wantedEffort}` : ''}`
397 return { model: null, effort: null, reason: `kept ${kept}, wanted ${wanted} (${said})` }
398 }
399
400 return { model, effort, reason: forced ? `${tier}, forced by risk` : `${tier} (${said})` }
401}
402
403/**
404 * Holds a prompt's classification until the turn that reads that prompt
405 * starts.
406 *
407 * Nothing ties a decision to the turn it belongs to. Prompts can be queued
408 * while the model is busy, a peer session's message can be delivered inside a
409 * running turn, and `prompt.submit` carries no turn id at all while the
410 * session is idle. So when more than one prompt is waiting, `take` reports
411 * none: running a turn on another prompt's decision is a worse outcome than
412 * not routing it, and not routing is what every other failure path here does.
413 */
414export function pendingDecisions(): {
415 put(decision: Decision | null): void
416 take(): Decision | null
417} {
418 let held: Decision | null = null
419 let waiting = 0
420
421 return {
422 put(decision) {
423 waiting += 1
424 // Past the first, which prompt a turn will read is unknowable, so the
425 // slot is emptied instead of holding a decision that may not fit.
426 held = waiting === 1 ? decision : null
427 },
428 take() {
429 const decision = waiting === 1 ? held : null
430 held = null
431 waiting = 0
432 return decision
433 },
434 }
435}
436
437/** A number for the log, or `n/d` when the backend reported none. */
438function reported(value: number | null): string {
439 return value === null ? 'n/d' : value.toFixed(2)
440}
441
442/**
443 * The one-time line that says the router is alive, which backend answers it,
444 * and which of the three switches are on.
445 *
446 * Without this, a router that loaded and a router that never loaded are told
447 * apart only by the absence of later lines, which is not evidence of anything.
448 */
449export function describeSetup(
450 provider: Provider | null,
451 url: string,
452 switches: { subagentModel: boolean; mainEffort: boolean; mainModel: boolean },
453 // `provider: "builtin"` is a choice, not a missing key. Reporting it as a
454 // credential problem sends someone hunting for a key they meant to omit.
455 builtinByChoice = false,
456): string {
457 const backend = provider
458 ? `${provider} (${url})`
459 : builtinByChoice
460 ? 'the built-in classifier, by choice'
461 : 'the built-in classifier, no key set'
462 const on = [
463 switches.subagentModel && 'subagent model',
464 switches.mainEffort && 'main effort',
465 switches.mainModel && 'main model',
466 ].filter(Boolean)
467 return `ready on ${backend}; routing ${on.length > 0 ? on.join(', ') : 'nothing, every switch is off'}`
468}
469
470/**
471 * What the decision model answered, before any policy touches it: the raw
472 * tier, effort and risk with their confidences, and how long it took.
473 *
474 * This is the line that shows the classification happened at all, separately
475 * from whether the policy then decided to act on it.
476 */
477export function describeDecision(decision: Decision | null, ms: number | null): string {
478 const took = ms === null ? '' : ` · ${Math.round(ms)}ms`
479 if (!decision) return `no answer${took}`
480
481 const parts = [`tier ${decision.tier} (${reported(decision.confidence)})`]
482 if (decision.effort !== null) {
483 parts.push(
484 `effort ${decision.effort.toFixed(1)} → ${effortLevel(decision.effort)} (${reported(decision.effortConfidence)})`,
485 )
486 }
487 if (decision.risky !== null) parts.push(`risky ${reported(decision.risky)}`)
488 return parts.join(' · ') + took
489}
490
491/**
492 * The persistent status line: the last thing the router did, short enough to
493 * sit on screen beside the engine's own notices.
494 */
495export function describeStatus(
496 decision: Decision | null,
497 change: { model?: string; effort?: Effort } | null,
498): string {
499 if (!decision) return 'jev · no answer'
500 const asked = `${decision.tier} ${reported(decision.confidence)}`
501 if (!change) return `jev · ${asked} · unchanged`
502 const to = [change.model, change.effort].filter(Boolean).join('/')
503 return `jev · ${asked} → ${to}`
504}
505