SLOPSHOPPER

jev-model-router

Picks the model and the reasoning effort for each task with TypeSafe's Jev, a System One decision model reached either through TypeSafe's own API or the Vercel…

newstatuspromptmodelnetworkagents
★ 1v0.4.2MITupdated 2026-10-01Kiy-K/oxide/.claude/skills/jev-model-router
A shopper browsing a rack in a slop shop
README

jev-model-router

Picks the model and the reasoning effort each task runs with, using Jev, TypeSafe's System One decision model: unstructured state in, a typed choice with a probability distribution out, no free-form text.

Two backends, chosen by whichever key is set:

BackendEndpointModelConfidence
typesafePOST api.typesafe.ai/v1/systemonejev-latestreported per answer
gatewayPOST ai-gateway.vercel.sh/v4/ai/evaluation-modeltypesafe-ai/jevderived from an optional distribution

TypeSafe's own API wins when both keys are set: it is the only one that reports a calibrated confidence, which is what the confidence bars below read. Set provider to force one, or to builtin to use neither. Each backend keeps its own URL and model option, so an override written for one is never sent to the other. A provider forced onto a backend whose key is missing degrades to the built-in classifier and says so once in the log.

Three switches, and they are not equally safe:

SwitchWhat it setsDefault
routeSubagentModelthe model of each subagent, at agent.spawnon
routeMainEffortthe reasoning effort of the main conversation, at turn.stepon
routeMainModelthe model of the main conversation, at turn.stepoff

A subagent starts with its own context, so routing its model costs nothing beyond the classification. Changing the main loop's model mid-session is the expensive one: it invalidates the prompt cache, and on a long context re-caching can cost more than the cheaper tier saves. Turn it on once you have measured your own sessions, not before.

The Agent tool has no effort parameter, so a subagent's effort is not this mod's to set.

Both directions, both dimensions. A task read as mechanical is routed down; one read as hard is routed up — model and effort alike.

The prompt is classified at prompt.submit, which runs before the turn starts, and the decision is applied to the turn's first model request and reused by the rest of that turn.

With no key configured the mod still works: it falls back to the engine's own $.model.classify, which answers the same question with the small fast model. That path reports no confidence, so the threshold does not apply to it.

What it asks

One request, three questions evaluated in parallel:

  • tier — a choice between three descriptions of the work (mechanical and local / ordinary engineering / hard or high-stakes). The decision model never sees a model name.
  • effort — a score on a four-level rubric, for how much step-by-step reasoning the task needs.
  • risky — whether the task touches production, money, credentials, or state that cannot be undone. A noul on TypeSafe's API, a boolean on the Gateway: the same question under two names.

How it decides

TypeSafe's API reports a confidence per answer. The Gateway's answer shape carries no confidence field, so on that backend confidence is read as the highest probability in the distribution — and that distribution is itself optional in the schema, in which case confidence is absent and the threshold does not fire.

The two mistakes do not cost the same, so they do not clear the same bar:

  • Spending more (a bigger model, more reasoning) needs minUpgradeConfidence, 0.3 by default. Being wrong costs money.
  • Spending less needs minDowngradeConfidence, 0.6 by default. Being wrong means a task handled by too small a model or too little thought.
  • risky above 0.7 takes the deep tier and real reasoning, past both bars. That one is not a confidence question.
  • A backend that reports no confidence at all — the Gateway without a distribution, or the built-in classifier — may only move a request up. Spending less on an unmeasured hunch is the bad trade.
  • A model id matching no tier, or a numeric effort (the caller's own scale), has no knowable direction: the model gets the gentler upgrade bar, and a numeric effort is left alone.

Every other failure — a non-2xx response, a timeout, a malformed body, a thrown error — leaves the request exactly as the engine built it. The router never blocks a turn.

What you see in the transcript

With logDecisions on (the default), the router reports every step of its own work, because nothing else in Claude Code shows it: the model and effort it rewrites are parameters of each request, not the session's settings, so the status line, the header and the effort box never move whatever it decides.

[jev-model-router] ready on typesafe (https://api.typesafe.ai/v1/systemone); routing subagent model, main effort
[jev-model-router] jev: tier fast (0.87) · effort 0.4 → low (0.71) · risky 0.02 · 249ms
[jev-model-router] main loop → effort low: fast (confidence 0.87)
[jev-model-router] jev: tier fast (0.41) · effort 0.4 → low (0.38) · risky 0.01 · 210ms
[jev-model-router] main loop: kept opus/medium, wanted haiku/low (confidence 0.41)
  • The first line appears once per session, the first time a hook runs. It is the proof the module loaded and which backend answers it.
  • A jev: line is what the decision model replied, before any policy is applied — the tier, the effort score and the risk, each with its confidence, and how long the call took.
  • The line under it is what the policy then did. The last pair above is a working router declining to act: it wanted to spend less but did not clear minDowngradeConfidence.

It also keeps a one-line status on screen, replaced as it goes:

jev · fast 0.87 → haiku/low
jev · fast 0.41 · unchanged

confidence n/d means the backend reported no confidence, which the built-in classifier never does and the Gateway does whenever its probability distribution is absent; the ready on line says which one answered.

No lines at all has three causes, and only the last is the module failing to load. Check them in this order:

  1. You ran claude -p (or the SDK). A headless run has no transcript and no status row: every line still goes to the debug log, ~/.claude/debug/<session-id>.txt (a .txt, not a .log; latest is a symlink to the newest), and an SDK host receives each one as ui_log. The router is working; look there.
  2. The plugin was never loaded. Claude Code adopts a plugin from a project's .claude/skills/ (where --mod writes it) only once the project is trusted: it is repository content, so an untrusted folder's .claude/ is not read at all, and claude -p never asks. Open claude interactively in the folder and accept the trust prompt, or name the plugin explicitly with --plugin-dir (see Install). claude --debug settles it: a loaded module prints hooks module jev-model-router@skills-dir loaded (worker, …); events: prompt.submit,turn.step,agent.spawn (@inline when loaded with --plugin-dir); Found N plugins without it means the plugin is not in the session.
  3. Function hooks are off. Without CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 the debug log says installed plugins' hooks modules not loaded: rollout flag (tengu_plugin_hooks_modules) is off. Set the flag; Claude Code must be 2.1.259+.

A ready on the built-in classifier, no key set line when you did set a key means the key sits under the wrong pluginConfigs entry: the key must match the plugin's id, which depends on how it was loaded (see Options).

Privacy

With a key set, the prompt text leaves the machine and goes to whichever backend the key belongs to. The main-loop path sends the prompt; the subagent path sends the subagent's prompt, its description and its agent type. Nothing else. With no key set, nothing leaves the machine.

Options

  typesafeApiKey:         string  TypeSafe API key (preferred: it reports a confidence)
  gatewayApiKey:          string  Vercel AI Gateway key
  provider:               string  "auto" | "typesafe" | "gateway" | "builtin"
  typesafeBaseUrl:        string  empty uses https://api.typesafe.ai
  typesafeModel:          string  empty uses jev-latest
  gatewayBaseUrl:         string  empty uses https://ai-gateway.vercel.sh/v4/ai
  gatewayModel:           string  empty uses typesafe-ai/jev
  fastModel:              string  fast tier, alias or full id (default "haiku")
  balancedModel:          string  balanced tier, alias or full id (default "sonnet")
  deepModel:              string  deep tier, alias or full id (default "opus")
  minUpgradeConfidence:   number  bar to spend more (default 0.3)
  minDowngradeConfidence: number  bar to spend less (default 0.6)
  routeSubagentModel:     boolean model of each subagent (default true)
  routeMainEffort:        boolean effort of the main loop (default true)
  routeMainModel:         boolean model of the main loop (default false)
  timeoutMs:              number  latency budget per classification (default 800)
  logDecisions:           boolean log each decision (default true)

The three tiers take an alias (haiku, sonnet, opus) or a full model id. A subagent is spawned with the name as given, the way the Agent tool takes it; the main loop's request needs an id, so there an alias is resolved to the family's current id (haiku → claude-haiku-4-5-20251001, sonnet → claude-sonnet-5, opus → claude-opus-5). Set a full id to pin a specific version. A decision for the tier the session already runs is not a change, so a session on claude-opus-5[1m] keeps its 1M-context id.

Declared in .claude-plugin/plugin.json (userConfig). Set them in /config, in user settings (~/.claude/settings.json, not project settings), with --settings <file> or in managed settings:

{ "pluginConfigs": { "jev-model-router": { "options": { "typesafeApiKey": "" } } } }

The entry's key is the plugin's id, and the id follows how the plugin was loaded: "jev-model-router@skills-dir" when auto-loaded from .claude/skills/ (the --mod install), "jev-model-router" with --plugin-dir. Under the wrong key every option stays at its default, and the ready on line reports no key set.

Install

npx claude-code-templates@latest --mod productivity/jev-model-router
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude

--mod writes the plugin to .claude/skills/jev-model-router/ in the project, and Claude Code auto-loads it as jev-model-router@skills-dir in a trusted project: a folder's .claude/ is repository content and is not read until you accept the trust prompt on the first interactive claude there (-p never asks, so a headless run in a fresh folder never sees it). The options then go under the "jev-model-router@skills-dir" key in pluginConfigs (see Options).

For one session with hot reload, or in a folder you do not want to trust, name it on the command line instead — it loads as jev-model-router@inline and reads options from the "jev-model-router" key:

CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir .claude/skills/jev-model-router

Either way, claude plugin validate .claude/skills/jev-model-router prints every event it hooks and every $ call it makes.

Tests

bun test cli-tool/components/mods/productivity/jev-model-router/tests

Early access. Mods need Claude Code 2.1.259+ with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1; the $ API may change between releases. Typed against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods

A mod runs without node_modules, so neither @typesafe-ai/sdk nor the AI SDK is available here: both backends are spoken to over HTTP through $.http.fetch. The TypeSafe wire shape was read from @typesafe-ai/sdk v0.6.0; the Gateway's, which is experimental in the AI SDK (experimental_evaluate, 7.0.105+) and not documented publicly, from @ai-sdk/gateway v4.0.86 and @ai-sdk/provider v4.0.17. Either may change.

Source 2 files
hooks/jev-model-router.ts 328 lines
1/**
2 * jev-model-router — Claude Mod (EARLY ACCESS)
3 *
4 * Picks the model each task runs on with TypeSafe's Jev, a System One
5 * decision model: unstructured state in, a typed choice with a probability
6 * distribution out.
7 *
8 * Jev is reached one of two ways, whichever key is configured: TypeSafe's
9 * own API (`typesafeApiKey`), which reports a calibrated confidence per
10 * answer, or the Vercel AI Gateway (`gatewayApiKey`), which does not. With
11 * neither, the engine's own `$.model.classify` stands in, so the mod is
12 * useful without any account.
13 *
14 * Three things it can set, each on its own switch:
15 *   agent.spawn  — the model of each subagent (on by default)
16 *   turn.step    — the reasoning effort of the main loop (on by default)
17 *   turn.step    — the model of the main loop (off by default: switching
18 *                  models mid-session invalidates the prompt cache, which can
19 *                  cost more than the cheaper tier saves)
20 *
21 * Every one of them moves in both directions: a task the decision model reads
22 * as mechanical is routed down, one it reads as hard is routed up. The two
23 * mistakes do not cost the same, so they do not clear the same confidence bar
24 * (see `minUpgradeConfidence` / `minDowngradeConfidence` in policy.ts).
25 *
26 * The Agent tool has no effort parameter, so a subagent's effort is not ours
27 * to set; only its model is.
28 *
29 * The prompt is classified at `prompt.submit`, which runs before the turn
30 * starts, and the decision is applied at the turn's first request.
31 *
32 * Every failure path is fail-open: a classification that errors or runs past
33 * the latency budget leaves the request exactly as the engine built it.
34 *
35 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
36 * or "gatewayApiKey"). Never hardcode it in this file.
37 *
38 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (Claude Code >= 2.1.259). Typed
39 * against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
40 *
41 * Privacy: with a key set, the prompt text is sent to whichever backend the
42 * key belongs to.
43 */
44import type { Register } from 'claude-code'
45import {
46  DEFAULT_BASE_URL,
47  DEFAULT_MODEL,
48  describeDecision,
49  describeSetup,
50  describeStatus,
51  endpoint,
52  pendingDecisions,
53  readDecision,
54  selectProvider,
55  requestBody,
56  requestHeaders,
57  requestModelId,
58  route,
59  TIER_ORDER,
60} from './policy.ts'
61import type { Decision, Effort, PolicyConfig, Provider, Tier } from './policy.ts'
62
63export const register: Register = (on, options) => {
64  const text = (key: string, fallback: string) =>
65    typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
66  const number = (key: string, fallback: number) =>
67    typeof options[key] === 'number' ? (options[key] as number) : fallback
68  const flag = (key: string, fallback: boolean) =>
69    typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
70
71  // TypeSafe's own API is preferred when both keys are set: it is the only
72  // one that reports a calibrated confidence, which the policy's threshold
73  // reads. `provider` forces one, including "builtin" to use neither.
74  const typesafeKey = text('typesafeApiKey', '')
75  const gatewayKey = text('gatewayApiKey', '')
76  const forced = text('provider', 'auto')
77  const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey)
78
79  // Each backend keeps its own URL and model, so an override written for one
80  // can never be sent to the other when `auto` picks differently than expected.
81  const apiKey = active === 'typesafe' ? typesafeKey : active === 'gateway' ? gatewayKey : ''
82  const modelId = !active
83    ? ''
84    : active === 'typesafe'
85      ? text('typesafeModel', DEFAULT_MODEL.typesafe)
86      : text('gatewayModel', DEFAULT_MODEL.gateway)
87  const url = !active
88    ? ''
89    : active === 'typesafe'
90      ? endpoint('typesafe', text('typesafeBaseUrl', DEFAULT_BASE_URL.typesafe))
91      : endpoint('gateway', text('gatewayBaseUrl', DEFAULT_BASE_URL.gateway))
92
93  // A backend named in the options but missing its key degrades to the
94  // built-in classifier, which is silent; say so once, when a hook first runs.
95  let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
96
97  const timeoutMs = number('timeoutMs', 800)
98  const routeSubagentModel = flag('routeSubagentModel', true)
99  const routeMainEffort = flag('routeMainEffort', true)
100  const routeMainModel = flag('routeMainModel', false)
101  const routeMainLoop = routeMainEffort || routeMainModel
102  const logDecisions = flag('logDecisions', true)
103
104  const policy: PolicyConfig = {
105    tiers: {
106      fast: text('fastModel', 'haiku'),
107      balanced: text('balancedModel', 'sonnet'),
108      deep: text('deepModel', 'opus'),
109    },
110    minUpgradeConfidence: number('minUpgradeConfidence', 0.3),
111    minDowngradeConfidence: number('minDowngradeConfidence', 0.6),
112  }
113
114  // The classification waiting for the turn that reads its prompt, and what
115  // the current turn settled on. Both are single slots: main-loop turns run
116  // one at a time, so nothing accumulates over a long session. `pending`
117  // reports no decision when two prompts are waiting at once, rather than
118  // routing a turn on a decision made for a different prompt.
119  const pending = pendingDecisions()
120  // Said once, the first time a hook runs. A router that loaded and one that
121  // never loaded are otherwise told apart only by the absence of later lines,
122  // and absence is not evidence: the policy leaves most turns alone anyway.
123  let announced = false
124  let appliedTurnId: string | undefined
125  let applied: { model?: string; effort?: Effort } | null = null
126
127  on('prompt.submit', async ($, e, next) => {
128    // Before the routing guards: a module whose switches are all off has still
129    // loaded, and that is exactly when its silence is most misleading.
130    if (!announced) {
131      announced = true
132      if (logDecisions) {
133        $.ui.log(
134          `[jev-model-router] ${describeSetup(
135            active,
136            url,
137            {
138              subagentModel: routeSubagentModel,
139              mainEffort: routeMainEffort,
140              mainModel: routeMainModel,
141            },
142            forced === 'builtin',
143          )}`,
144        )
145      }
146    }
147    if (!routeMainLoop) return next(e)
148
149    if (!unusableReported) {
150      unusableReported = true
151      $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
152    }
153
154    const startedAt = await $.clock.now()
155    let decision: Decision | null = null
156    if (active) {
157      try {
158        const response = await Promise.race([
159          $.http.fetch(url, {
160            method: 'POST',
161            headers: requestHeaders(active, apiKey, modelId),
162            body: requestBody(active, { prompt: e.text }, modelId),
163          }),
164          $.clock.sleep(timeoutMs),
165        ])
166        if (response && response.ok) decision = readDecision(response.text)
167        else if (response) $.ui.log(`[jev-model-router] ${active} responded ${response.status}`)
168        else $.ui.log(`[jev-model-router] classification passed ${timeoutMs}ms; leaving the turn alone`)
169      } catch (error) {
170        $.ui.log(`[jev-model-router] classification failed: ${String(error)}`)
171      }
172    } else {
173      // No backend: the engine's own small-model classifier answers the same
174      // question, without the confidence the policy's threshold reads.
175      try {
176        const label = await $.model.classify(e.text, TIER_ORDER)
177        if (label) {
178          decision = {
179            tier: label as Tier,
180            confidence: null,
181            risky: null,
182            effort: null,
183            effortConfidence: null,
184          }
185        }
186      } catch (error) {
187        $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
188      }
189    }
190
191    // What the decision model actually answered, whatever the policy then
192    // does with it. This is the line that proves the classification ran.
193    if (logDecisions) {
194      const ms = (await $.clock.now()) - startedAt
195      $.ui.log(`[jev-model-router] jev: ${describeDecision(decision, ms)}`)
196    }
197
198    pending.put(decision)
199    return next(e)
200  })
201
202  on('turn.step', async function* ($, e, next) {
203    if (!routeMainLoop || e.agentId) return yield* next(e)
204
205    // Every request after the first reuses what the turn settled on, so
206    // neither the model nor the effort changes under its own tool loop.
207    if (e.index > 0 && e.turnId === appliedTurnId) {
208      return yield* next(applied ? { ...e, ...applied } : e)
209    }
210
211    const decision = pending.take()
212    const routing = route(decision, { model: e.model, effort: e.effort }, policy)
213    const change: { model?: string; effort?: Effort } = {}
214    // The main loop's `model` is sent to the API as written, so an alias
215    // becomes its id here; a subagent's (agent.spawn) may stay an alias.
216    if (routeMainModel && routing.model) change.model = requestModelId(routing.model)
217    if (routeMainEffort && routing.effort) change.effort = routing.effort
218
219    appliedTurnId = e.turnId
220    applied = Object.keys(change).length > 0 ? change : null
221    // A row in the transcript scrolls away; this line stays on screen.
222    if (logDecisions) $.ui.status(describeStatus(decision, applied))
223
224    if (!applied) {
225      // A turn left alone is the common case, and it used to be silent, which
226      // made a working mod look like one that never loaded. Say what happened.
227      if (logDecisions) {
228        const suppressed = routing.model && !routeMainModel ? ' (main-loop model routing off)' : ''
229        $.ui.log(`[jev-model-router] main loop: ${routing.reason}${suppressed}`)
230      }
231      return yield* next(e)
232    }
233    if (logDecisions) {
234      const what = [change.model, change.effort && `effort ${change.effort}`]
235        .filter(Boolean)
236        .join(', ')
237      $.ui.log(`[jev-model-router] main loop → ${what}: ${routing.reason}`)
238    }
239    return yield* next({ ...e, ...change })
240  })
241
242  on('agent.spawn', async ($, e, next) => {
243    // Before the routing guards: a module whose switches are all off has still
244    // loaded, and that is exactly when its silence is most misleading.
245    if (!announced) {
246      announced = true
247      if (logDecisions) {
248        $.ui.log(
249          `[jev-model-router] ${describeSetup(
250            active,
251            url,
252            {
253              subagentModel: routeSubagentModel,
254              mainEffort: routeMainEffort,
255              mainModel: routeMainModel,
256            },
257            forced === 'builtin',
258          )}`,
259        )
260      }
261    }
262
263    // A fork inherits its parent's model; `model` is ignored for it.
264    if (!routeSubagentModel || e.fork) return next(e)
265
266    if (!unusableReported) {
267      unusableReported = true
268      $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
269    }
270
271    const startedAt = await $.clock.now()
272    let decision: Decision | null = null
273    if (active) {
274      try {
275        const response = await Promise.race([
276          $.http.fetch(url, {
277            method: 'POST',
278            headers: requestHeaders(active, apiKey, modelId),
279            body: requestBody(
280              active,
281              { prompt: e.prompt, description: e.description, agentType: e.subagentType },
282              modelId,
283            ),
284          }),
285          $.clock.sleep(timeoutMs),
286        ])
287        if (response && response.ok) decision = readDecision(response.text)
288        else if (response) $.ui.log(`[jev-model-router] ${active} responded ${response.status}`)
289        else $.ui.log(`[jev-model-router] classification passed ${timeoutMs}ms; leaving the subagent alone`)
290      } catch (error) {
291        $.ui.log(`[jev-model-router] classification failed: ${String(error)}`)
292      }
293    } else {
294      try {
295        const label = await $.model.classify(e.prompt, TIER_ORDER)
296        if (label) {
297          decision = {
298            tier: label as Tier,
299            confidence: null,
300            risky: null,
301            effort: null,
302            effortConfidence: null,
303          }
304        }
305      } catch (error) {
306        $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
307      }
308    }
309
310    if (logDecisions) {
311      const ms = (await $.clock.now()) - startedAt
312      $.ui.log(`[jev-model-router] jev (${e.subagentType}): ${describeDecision(decision, ms)}`)
313    }
314
315    // The subagent's own model wins when the caller named one; otherwise it
316    // would inherit the parent's, so that is what a change is measured from.
317    // The Agent tool takes no effort, so only the model is ours to set here.
318    const current = e.model ?? e.parentModel
319    const { model, reason } = route(decision, { model: current }, policy)
320    if (!model) {
321      if (logDecisions) $.ui.log(`[jev-model-router] ${e.subagentType}: ${reason}`)
322      return next(e)
323    }
324    if (logDecisions) $.ui.log(`[jev-model-router] ${e.subagentType} → ${model}: ${reason}`)
325    return next({ ...e, model })
326  })
327}
328
hooks/policy.ts 505 lines
1/**
2 * jev-model-router — pure decision logic.
3 *
4 * No `$` and no I/O here: this module only builds the request the decision
5 * API takes, reads its answer, and turns that answer into a model id. The
6 * hooks module does every call on `$` at its own call site.
7 *
8 * Two backends speak to the same model with different wire shapes:
9 *
10 *   typesafe  POST https://api.typesafe.ai/v1/systemone
11 *             `{ model, state, questions }`; a yes/no question is a `noul`
12 *             and every answer carries its own `confidence`.
13 *   gateway   POST https://ai-gateway.vercel.sh/v4/ai/evaluation-model
14 *             `{ state, questions }` with the model in a header; a yes/no
15 *             question is a `boolean`, and there is no `confidence` field —
16 *             it has to be derived from an optional distribution.
17 *
18 * The Gateway shape is not documented publicly; it was read from
19 * @ai-sdk/gateway and @ai-sdk/provider.
20 */
21
22export type Provider = 'typesafe' | 'gateway'
23
24export type Tier = 'fast' | 'balanced' | 'deep'
25
26export interface Tiers {
27  fast: string
28  balanced: string
29  deep: string
30}
31
32export interface Decision {
33  tier: Tier
34  /** Confidence in the tier, or null when the backend reported none. */
35  confidence: number | null
36  /** P(true) that carrying the task out would itself be costly or final. */
37  risky: number | null
38  /** 0..3 along the effort rubric, or null when absent. */
39  effort: number | null
40  /** Confidence in the effort, or null when the backend reported none. */
41  effortConfidence: number | null
42}
43
44/** The reasoning levels a turn can ask for, cheapest first. */
45export const EFFORT_ORDER = ['low', 'medium', 'high', 'xhigh'] as const
46
47export type Effort = (typeof EFFORT_ORDER)[number]
48
49export const TIER_ORDER: readonly Tier[] = ['fast', 'balanced', 'deep']
50
51/**
52 * How each tier is described to the decision model. Deliberately about the
53 * shape of the work, not about model names: the model never sees an id.
54 */
55const TIER_CRITERIA: Record<Tier, string> = {
56  fast: 'Mechanical and local: read or summarise a file, run one command, rename a symbol, answer something already in context.',
57  balanced:
58    'Ordinary engineering: implement a well-specified change across a few files, write tests, fix a clearly described bug, review a small diff.',
59  deep: 'Hard or high-stakes: architecture and design, debugging a failure whose cause is unknown, security, data migrations, concurrency, anything touching production or money.',
60}
61
62const EFFORT_RUBRIC = ['almost none', 'some', 'a lot', 'as much as possible'] as const
63
64export const DEFAULT_BASE_URL: Record<Provider, string> = {
65  typesafe: 'https://api.typesafe.ai',
66  gateway: 'https://ai-gateway.vercel.sh/v4/ai',
67}
68
69/**
70 * The Gateway's own protocol version, sent as `ai-gateway-protocol-version`.
71 * Tracks the `AI_GATEWAY_PROTOCOL_VERSION` of `@ai-sdk/gateway` (4.0.87).
72 */
73const AI_GATEWAY_PROTOCOL_VERSION = '0.0.1'
74
75export const DEFAULT_MODEL: Record<Provider, string> = {
76  typesafe: 'jev-latest',
77  gateway: 'typesafe-ai/jev',
78}
79
80/**
81 * Which backend a configuration asks for, or null for the built-in
82 * classifier. `auto` prefers TypeSafe, since it is the only one that reports
83 * a calibrated confidence; a forced backend whose key is missing resolves to
84 * null rather than falling through to the other one's key.
85 */
86export function selectProvider(
87  forced: string,
88  typesafeKey: string,
89  gatewayKey: string,
90): Provider | null {
91  if (forced === 'builtin') return null
92  if (forced === 'typesafe') return typesafeKey ? 'typesafe' : null
93  if (forced === 'gateway') return gatewayKey ? 'gateway' : null
94  if (typesafeKey) return 'typesafe'
95  if (gatewayKey) return 'gateway'
96  return null
97}
98
99/** The full endpoint a backend posts to. */
100export function endpoint(provider: Provider, baseUrl: string): string {
101  const root = baseUrl.replace(/\/+$/, '')
102  return provider === 'typesafe' ? `${root}/v1/systemone` : `${root}/evaluation-model`
103}
104
105/** The `questions` map, in the shape the backend's schema names. */
106export function questions(provider: Provider): Record<string, unknown> {
107  return {
108    tier: {
109      type: 'choice',
110      instructions: 'Which is the cheapest tier that can complete this coding task well?',
111      criteria: TIER_CRITERIA,
112    },
113    effort: {
114      type: 'score',
115      instructions: 'How much step-by-step reasoning does this task need?',
116      criteria: EFFORT_RUBRIC,
117    },
118    risky: {
119      // The same question under two names: `noul` on TypeSafe's own API,
120      // `boolean` in the AI SDK's evaluation schema.
121      type: provider === 'typesafe' ? 'noul' : 'boolean',
122      // Asked about the act, not the subject. The first wording ("the task
123      // touches production, money, credentials") scored 0.96 on "add a
124      // refund endpoint that calls Stripe" — ordinary code that happens to be
125      // about money — and would have escalated it past a 0.98-confidence
126      // answer of the balanced tier.
127      instructions:
128        'Carrying out this task would itself change production, move real money, or alter data that cannot be restored. Writing or testing code that deals with such things, without running it against the real system, does not count.',
129    },
130  }
131}
132
133/** The request body. The Gateway carries the model in a header instead. */
134export function requestBody(
135  provider: Provider,
136  state: Record<string, unknown>,
137  model: string,
138): string {
139  const body =
140    provider === 'typesafe'
141      ? { model, state, questions: questions(provider) }
142      : { state, questions: questions(provider) }
143  return JSON.stringify(body)
144}
145
146/** The request headers. */
147export function requestHeaders(
148  provider: Provider,
149  apiKey: string,
150  model: string,
151): Record<string, string> {
152  const common = { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` }
153  if (provider === 'typesafe') return common
154  return {
155    ...common,
156    'ai-gateway-auth-method': 'api-key',
157    'ai-model-id': model,
158    // The Gateway rejects any request that does not name the protocol it
159    // speaks: 400 "Unsupported gateway protocol version". Every other header
160    // here is accepted without it, so the omission fails the whole backend.
161    'ai-gateway-protocol-version': AI_GATEWAY_PROTOCOL_VERSION,
162    'ai-evaluation-model-specification-version': '4',
163  }
164}
165
166function isTier(value: unknown): value is Tier {
167  return value === 'fast' || value === 'balanced' || value === 'deep'
168}
169
170/**
171 * Reads a response from either backend.
172 *
173 * TypeSafe's own API reports a `confidence` per answer and a `noul` number
174 * for a yes/no question. The Gateway reports neither: confidence has to come
175 * from the highest probability of a distribution that is itself optional, and
176 * a yes/no answer arrives as `probability`. Both are handled, and a missing
177 * confidence reads as null rather than as a number the policy would trust.
178 */
179export function readDecision(responseText: string): Decision | null {
180  let parsed: unknown
181  try {
182    parsed = JSON.parse(responseText)
183  } catch {
184    return null
185  }
186  const answers = (parsed as { answers?: Record<string, Record<string, unknown>> }).answers
187  if (!answers) return null
188
189  const tierAnswer = answers.tier
190  if (!tierAnswer || !isTier(tierAnswer.choice)) return null
191
192  const effortAnswer = answers.effort
193  const riskyAnswer = answers.risky
194  const risky =
195    typeof riskyAnswer?.noul === 'number'
196      ? riskyAnswer.noul
197      : typeof riskyAnswer?.probability === 'number'
198        ? riskyAnswer.probability
199        : null
200
201  return {
202    tier: tierAnswer.choice,
203    confidence: confidenceOf(tierAnswer),
204    effort: typeof effortAnswer?.score === 'number' ? effortAnswer.score : null,
205    effortConfidence: effortAnswer ? confidenceOf(effortAnswer) : null,
206    risky,
207  }
208}
209
210/**
211 * How sure an answer is. TypeSafe reports it; the Gateway does not, so there
212 * it is the highest probability of a distribution that is itself optional.
213 */
214function confidenceOf(answer: Record<string, unknown>): number | null {
215  if (typeof answer.confidence === 'number') return answer.confidence
216  const probabilities = answer.probabilities as Record<string, number> | undefined
217  const values = probabilities ? Object.values(probabilities) : []
218  return values.length > 0 ? Math.max(...values) : null
219}
220
221/** The rubric score (0..3) as a reasoning level. */
222export function effortLevel(score: number): Effort {
223  const index = Math.min(EFFORT_ORDER.length - 1, Math.max(0, Math.round(score)))
224  return EFFORT_ORDER[index] as Effort
225}
226
227/**
228 * Where a reasoning level sits on the ladder, or null when its place cannot
229 * be known. `max` is above every rung the rubric can produce, so it ranks
230 * above them without joining EFFORT_ORDER, which is also the set of values
231 * this router is allowed to ask for.
232 */
233export function effortRank(effort: string | number | undefined): number | null {
234  if (typeof effort !== 'string') return null
235  if (effort === 'max') return EFFORT_ORDER.length
236  const index = EFFORT_ORDER.indexOf(effort as Effort)
237  return index === -1 ? null : index
238}
239
240/**
241 * Where a model id sits on the tier ladder, by matching it against the
242 * configured tier names first and then the family words. Null when it matches
243 * none, in which case the change is treated as an upgrade rather than guessed
244 * at: an unrecognised id gets the gentler threshold, never the strict one.
245 */
246export function rankOf(model: string, tiers: Tiers): number | null {
247  const lowered = model.toLowerCase()
248  for (let index = 0; index < TIER_ORDER.length; index++) {
249    const tier = TIER_ORDER[index] as Tier
250    const configured = tiers[tier].toLowerCase()
251    if (configured && lowered.includes(configured)) return index
252  }
253  if (lowered.includes('haiku')) return 0
254  if (lowered.includes('sonnet')) return 1
255  if (lowered.includes('opus')) return 2
256  return null
257}
258
259/**
260 * The full id a family alias names on the main loop.
261 *
262 * `agent.spawn` takes an alias (`haiku`) the way the Agent tool does, but
263 * `turn.step`'s `model` is the id the engine already resolved for the request
264 * and goes to the API as written: an alias there is refused ("There's an
265 * issue with the selected model (haiku)"). So the tiers stay aliases in the
266 * options, and only a main-loop rewrite resolves them, here.
267 */
268const ALIAS_IDS: Record<string, string> = {
269  haiku: 'claude-haiku-4-5-20251001',
270  sonnet: 'claude-sonnet-5',
271  opus: 'claude-opus-5',
272}
273
274/**
275 * What to write into `turn.step`'s `model`: a full id as given, or the id
276 * behind a family alias. Anything else is returned unchanged for the engine
277 * to judge.
278 */
279export function requestModelId(model: string): string {
280  return ALIAS_IDS[model.trim().toLowerCase()] ?? model
281}
282
283export interface PolicyConfig {
284  tiers: Tiers
285  /**
286   * How sure the decision must be to spend more (a bigger model, more
287   * reasoning). Being wrong here costs money, so the bar is low.
288   */
289  minUpgradeConfidence: number
290  /**
291   * How sure it must be to spend less. Being wrong here means a task handled
292   * by too small a model or too little thought, so the bar is high.
293   */
294  minDowngradeConfidence: number
295}
296
297export interface Routing {
298  /** The model to run on, or null to leave the request as it is. */
299  model: string | null
300  /** The reasoning level to ask for, or null to leave it as it is. */
301  effort: Effort | null
302  /** Why, for the log line. */
303  reason: string
304}
305
306const NOTHING: Routing = { model: null, effort: null, reason: 'no decision' }
307
308/**
309 * Whether a change of rank passes its threshold. Both directions are allowed;
310 * they just do not have to clear the same bar, because the two mistakes do not
311 * cost the same. A move whose direction cannot be told (an unrecognised
312 * current value) is treated as an upgrade.
313 */
314function allowed(
315  wanted: number,
316  current: number | null,
317  confidence: number | null,
318  config: PolicyConfig,
319): boolean {
320  if (current !== null && wanted === current) return false
321  const isDowngrade = current !== null && wanted < current
322  const bar = isDowngrade ? config.minDowngradeConfidence : config.minUpgradeConfidence
323  // A backend that reports no confidence (the Gateway without a distribution,
324  // or the built-in classifier) clears the upgrade bar but never the
325  // downgrade one: spending less on an unmeasured hunch is the bad trade.
326  if (confidence === null) return !isDowngrade
327  return confidence >= bar
328}
329
330/**
331 * Turns a decision into a model and a reasoning level, either of which may be
332 * null to leave the request as it is. Both can move in either direction.
333 */
334export function route(
335  decision: Decision | null,
336  current: { model: string; effort?: string | number },
337  config: PolicyConfig,
338): Routing {
339  if (!decision) return NOTHING
340
341  let tier = decision.tier
342  let effortScore = decision.effort
343  let forced = false
344
345  // Carrying out something final is never worth the saving: take the deep
346  // tier and real reasoning, whatever the cheaper answer said, and skip the
347  // thresholds — this is the one case that is not a confidence question.
348  if (decision.risky !== null && decision.risky > 0.7) {
349    tier = 'deep'
350    effortScore = Math.max(effortScore ?? 0, 2)
351    forced = true
352  }
353
354  const wantedTier = TIER_ORDER.indexOf(tier)
355  const currentTier = rankOf(current.model, config.tiers)
356  const wantedModel = config.tiers[tier]
357
358  const model =
359    wantedModel &&
360    wantedModel !== current.model &&
361    (forced || allowed(wantedTier, currentTier, decision.confidence, config))
362      ? wantedModel
363      : null
364
365  let effort: Effort | null = null
366  if (effortScore !== null) {
367    const currentRank = effortRank(current.effort)
368    let wantedRank = EFFORT_ORDER.indexOf(effortLevel(effortScore))
369
370    // Risk raises the floor; it must never lower one. Forcing only skips the
371    // thresholds, so without this clamp a task already at `xhigh` or `max`
372    // and rated mechanically simple would be pulled down to `high` with no
373    // confidence check at all — the opposite of what the rule is for.
374    if (forced && currentRank !== null) wantedRank = Math.max(wantedRank, currentRank)
375
376    // A numeric effort is the caller's own scale, not this ladder; leave it.
377    const comparable = typeof current.effort !== 'number'
378    const wanted = EFFORT_ORDER[Math.min(EFFORT_ORDER.length - 1, wantedRank)] as Effort
379    if (
380      comparable &&
381      wantedRank !== currentRank &&
382      (forced || allowed(wantedRank, currentRank, decision.effortConfidence, config))
383    ) {
384      effort = wanted
385    }
386  }
387
388  const said = decision.confidence === null ? 'confidence n/d' : `confidence ${decision.confidence.toFixed(2)}`
389
390  if (!model && !effort) {
391    // Naming what it wanted and what it kept is the whole point of this line.
392    // Without it, a mod that classified and decided to leave the request alone
393    // is indistinguishable from one that never loaded.
394    const wantedEffort = effortScore === null ? null : effortLevel(effortScore)
395    const kept = `${current.model}${current.effort === undefined ? '' : `/${current.effort}`}`
396    const wanted = `${wantedModel}${wantedEffort ? `/${wantedEffort}` : ''}`
397    return { model: null, effort: null, reason: `kept ${kept}, wanted ${wanted} (${said})` }
398  }
399
400  return { model, effort, reason: forced ? `${tier}, forced by risk` : `${tier} (${said})` }
401}
402
403/**
404 * Holds a prompt's classification until the turn that reads that prompt
405 * starts.
406 *
407 * Nothing ties a decision to the turn it belongs to. Prompts can be queued
408 * while the model is busy, a peer session's message can be delivered inside a
409 * running turn, and `prompt.submit` carries no turn id at all while the
410 * session is idle. So when more than one prompt is waiting, `take` reports
411 * none: running a turn on another prompt's decision is a worse outcome than
412 * not routing it, and not routing is what every other failure path here does.
413 */
414export function pendingDecisions(): {
415  put(decision: Decision | null): void
416  take(): Decision | null
417} {
418  let held: Decision | null = null
419  let waiting = 0
420
421  return {
422    put(decision) {
423      waiting += 1
424      // Past the first, which prompt a turn will read is unknowable, so the
425      // slot is emptied instead of holding a decision that may not fit.
426      held = waiting === 1 ? decision : null
427    },
428    take() {
429      const decision = waiting === 1 ? held : null
430      held = null
431      waiting = 0
432      return decision
433    },
434  }
435}
436
437/** A number for the log, or `n/d` when the backend reported none. */
438function reported(value: number | null): string {
439  return value === null ? 'n/d' : value.toFixed(2)
440}
441
442/**
443 * The one-time line that says the router is alive, which backend answers it,
444 * and which of the three switches are on.
445 *
446 * Without this, a router that loaded and a router that never loaded are told
447 * apart only by the absence of later lines, which is not evidence of anything.
448 */
449export function describeSetup(
450  provider: Provider | null,
451  url: string,
452  switches: { subagentModel: boolean; mainEffort: boolean; mainModel: boolean },
453  // `provider: "builtin"` is a choice, not a missing key. Reporting it as a
454  // credential problem sends someone hunting for a key they meant to omit.
455  builtinByChoice = false,
456): string {
457  const backend = provider
458    ? `${provider} (${url})`
459    : builtinByChoice
460      ? 'the built-in classifier, by choice'
461      : 'the built-in classifier, no key set'
462  const on = [
463    switches.subagentModel && 'subagent model',
464    switches.mainEffort && 'main effort',
465    switches.mainModel && 'main model',
466  ].filter(Boolean)
467  return `ready on ${backend}; routing ${on.length > 0 ? on.join(', ') : 'nothing, every switch is off'}`
468}
469
470/**
471 * What the decision model answered, before any policy touches it: the raw
472 * tier, effort and risk with their confidences, and how long it took.
473 *
474 * This is the line that shows the classification happened at all, separately
475 * from whether the policy then decided to act on it.
476 */
477export function describeDecision(decision: Decision | null, ms: number | null): string {
478  const took = ms === null ? '' : ` · ${Math.round(ms)}ms`
479  if (!decision) return `no answer${took}`
480
481  const parts = [`tier ${decision.tier} (${reported(decision.confidence)})`]
482  if (decision.effort !== null) {
483    parts.push(
484      `effort ${decision.effort.toFixed(1)} → ${effortLevel(decision.effort)} (${reported(decision.effortConfidence)})`,
485    )
486  }
487  if (decision.risky !== null) parts.push(`risky ${reported(decision.risky)}`)
488  return parts.join(' · ') + took
489}
490
491/**
492 * The persistent status line: the last thing the router did, short enough to
493 * sit on screen beside the engine's own notices.
494 */
495export function describeStatus(
496  decision: Decision | null,
497  change: { model?: string; effort?: Effort } | null,
498): string {
499  if (!decision) return 'jev · no answer'
500  const asked = `${decision.tier} ${reported(decision.confidence)}`
501  if (!change) return `jev · ${asked} · unchanged`
502  const to = [change.model, change.effort].filter(Boolean).join('/')
503  return `jev · ${asked} → ${to}`
504}
505