Picks the model and reasoning effort for each task with TypeSafe's Jev. Routes subagent models at agent.spawn and subagent and main-loop effort at turn.step…

Per-task reasoning effort for Claude Code, decided by a typed decision model.
thinkdial is a Claude Code mod (a plugin of function hooks). Before each turn it asks Jev, TypeSafe's System One decision model, three typed questions about the task: how hard is it, how much reasoning does it need, and is it risky. It then sets the reasoning effort of the request:
codex:codex-rescue), as --effort.It does not change which model answers, unless you turn that on (see why not). Every failure path is fail-open: if the classification is slow, errors out or makes no sense, the request goes out exactly as Claude Code built it.
The plugin id is still
jev-model-router(that is what Claude Code and yourpluginConfigsknow it by). The project isthinkdial.
The router started by routing models as well: haiku for mechanical turns, opus for hard ones. One user's week of real sessions, 1,174 routed main-loop turns across 133 sessions in late September 2026, showed why that loses money in a long Claude Code session:
| What changed at the start of a turn | Turns | Prompt-cache miss |
|---|---|---|
| nothing | 433 | 1 % |
| effort only | 495 | 1 % |
| model | 16 | 81 % |
| model + effort | 20 | 70 % |
Model routing is still in the code, behind routeMainModel and routeSubagentModel, for anyone whose sessions look different. Measure your own before you turn it on: the decision log below has everything you need.
flowchart LR
subgraph Main["Main conversation"]
A[prompt.submit] -->|prompt text| J1{{Jev}}
J1 -->|tier · effort · risky| P1[policy]
P1 --> T1["turn.step<br/>effort for this turn"]
end
subgraph Sub["Claude subagent"]
S[agent.spawn] -->|prompt · description · type| J2{{Jev}}
J2 --> K[(decision by agentId)]
K --> T2["subagent's turn.step<br/>effort for every request"]
end
subgraph Codex["Codex delegation"]
C[agent.spawn<br/>codex:codex-rescue] --> J3{{Jev}}
J3 --> F["prompt += --model gpt-6-sol --effort X"]
end
prompt.submit, before the turn starts. A subagent's prompt is classified at agent.spawn.turn.step) and reused by every later request of that turn, or of that subagent. Effort never changes inside a tool loop.Requirements: Claude Code 2.1.259+, function hooks enabled, and a TypeSafe API key. Without a key the router still runs on Claude Code's built-in classifier, which reports no confidence and so can only raise effort, never lower it.
git clone https://github.com/DmitryBMsk/thinkdial.git ~/src/thinkdial
# load it for every project: a user-level skills folder is auto-loaded as jev-model-router@skills-dir
ln -s ~/src/thinkdial ~/.claude/skills/jev-model-router
Add this to ~/.claude/settings.json. These are the recommended effort-only settings, and the defaults in the code differ from them, as noted below:
{
"env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" },
"pluginConfigs": {
"jev-model-router@skills-dir": {
"options": {
"typesafeApiKey": "<your TypeSafe key>",
"provider": "typesafe",
"timeoutMs": 1500,
"routeMainEffort": true,
"routeSubagentEffort": true,
"routeMainModel": false, // default false
"routeSubagentModel": false, // default TRUE in code: set false for effort-only
"codexFastModel": "gpt-6-sol", // default gpt-6-luna: pin Codex to one model
"mainEffortFloor": "medium",
"mainBalancedEffortFloor": "high",
"mainEffortCeiling": "high"
}
}
}
}
Start claude. The first routed turn prints [jev-model-router] ready on typesafe …; routing subagent effort, main effort.
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir ~/src/thinkdial
Loaded this way the plugin's id is plain jev-model-router, so its options go under "pluginConfigs": { "jev-model-router": { … } }, for example in a file passed with --settings. For an isolated claude -p worker use --restricted, not --safe-mode: --safe-mode also switches off hooks loaded with --plugin-dir.
claude plugin validate ~/src/thinkdial prints every event the plugin hooks and every $ call it makes.
Jev answers three questions in a single request:
| Question | Shape | Meaning |
|---|---|---|
tier | choice of 3 | mechanical and local · ordinary engineering · hard or high-stakes. The decision model never sees a model name. |
effort | score 0–3 | how much step-by-step reasoning the task needs: low · medium · high · xhigh |
risky | probability | does it touch production, money, credentials or state that cannot be undone? |
The two kinds of mistake do not cost the same, so they do not clear the same bar:
minUpgradeConfidence (0.3). A wrong call costs money.minDowngradeConfidence (0.6). A wrong call is a task done with too little thought.risky above 0.7 forces at least high effort past both bars. It can raise effort but never lower it.On top of the decision sit three policy limits. They apply without a confidence check:
| Option | Effect |
|---|---|
mainEffortFloor | lowest effort a turn may get, so a short "check this" prompt is not sent at low |
mainBalancedEffortFloor | higher floor for a task read as ordinary engineering |
mainEffortCeiling | highest effort below the deep tier, so xhigh/max are kept for opus-class models |
A slash command with nothing after it (/simplify) is not classified, because Jev would only see the command's name. That turn keeps the session's effort.
The Agent tool has no effort parameter, so the router sets a subagent's effort from inside its loop:
agent.spawn it classifies the subagent's prompt, description and type.agentId.turn.step it applies the effort, then reuses it for every later request and turn of that agent. If the first request overtakes the spawn, it waits for the spawn once, up to timeoutMs.Subagents use the same floor and ceiling as the main loop. Some subagents are left alone:
| Case | Behaviour |
|---|---|
definition pins effort: in its .md file | kept: effort pinned by definition (medium) |
--agents / registered definition pins an effort on the session's model | kept, detected when its effort differs from the session's: effort set by agent definition (high, session medium) |
| model takes no effort (Haiku 4.5) | nothing sent: model takes no effort |
fork, Codex delegation, or a plugin's own $.agent.spawn | untouched |
Known limits. The hook API does not say whether a subagent's effort was pinned. The router infers it, so a definition pinned to the same level as the session looks unpinned and may be routed. A pin on a different model given through --agents is not detected either. To make a pin certain, put effort: in the agent's .md file.
A codex:codex-rescue spawn gets --model and --effort flags prepended to its prompt, taken from the same decision. The rescue agent passes them on to Codex. With codexFastModel set to the same model as the other rungs, only the effort varies:
| Decision | Codex flags |
|---|---|
| no decision / ordinary | gpt-6-sol · low |
| ordinary, high effort | gpt-6-sol · medium |
| hard | gpt-6-sol · high (xhigh when the effort reads xhigh) |
risky > 0.7 | gpt-6-sol · xhigh |
A prompt that already names --model or --effort is left alone: an explicit choice wins.
The effort the router sets is a parameter of each request, so Claude Code's status line and effort box never show it. The router reports its own work in three places.
Transcript lines (when logDecisions is on):
[jev-model-router] ready on typesafe (https://api.typesafe.ai/v1/systemone); routing subagent effort, main effort
[jev-model-router] jev: tier deep (0.98) · effort 2.1 → high (0.70) · risky 0.09 · 612ms
[jev-model-router] main loop → effort high: deep (confidence 0.98)
[jev-model-router] general-purpose → effort high: deep (confidence 0.98)
[jev-model-router] ra-impl-medium: effort pinned by definition (medium)
A status line, replaced as it goes: jev · deep 0.98 → high.
A decision log, one JSONL file per session in ~/.claude/jev-router/<session-id>.jsonl (no --debug needed, no prompt text):
{"ts":"2026-10-04T10:12:41.204Z","event":"subagent","agentId":"ad05ecf8…","agentType":"general-purpose",
"tier":"deep","confidence":0.98,"effort":2.02,"effortConfidence":0.68,"risky":0.07,
"from":{"model":"claude-sonnet-5-5","effort":"medium","source":"parent"},
"applied":{"effort":"high"},"reason":"deep (confidence 0.98)"}
from is what the request would have used untouched. source says where its model came from: call, definition or parent. ran appears when the first request used a different model. applied is what the router changed. For cost analysis, join these records to the per-request usage in your session transcripts by timestamp.
Set them under pluginConfigs.<id>.options in user settings (~/.claude/settings.json), with --settings <file>, in managed settings or through /config. The <id> depends on how the plugin was loaded: jev-model-router@skills-dir when it is auto-loaded from a skills folder, plain jev-model-router with --plugin-dir. Under the wrong key every option stays at its default, and the ready on line says no key set.
| Option | Default | |
|---|---|---|
typesafeApiKey | — | TypeSafe API key. Preferred, because it reports a calibrated confidence. |
gatewayApiKey | — | Vercel AI Gateway key. Confidence is read from the optional distribution. |
provider | auto | auto · typesafe · gateway · builtin |
typesafeBaseUrl / typesafeModel | https://api.typesafe.ai / jev-latest | |
gatewayBaseUrl / gatewayModel | https://ai-gateway.vercel.sh/v4/ai / typesafe-ai/jev | |
timeoutMs | 800 | time limit for each classification. Past it the request goes out unchanged. |
| Option | Default | |
|---|---|---|
routeMainEffort | true | effort of the main conversation |
routeSubagentEffort | true | effort of each Claude subagent |
routeCodexDelegation | true | --model / --effort for codex:codex-rescue spawns |
routeMainModel | false | model of the main conversation (breaks the prompt cache, see above) |
routeSubagentModel | true | model of each subagent. Set false for effort-only. |
| Option | Default | |
|---|---|---|
minUpgradeConfidence | 0.3 | confidence needed to spend more |
minDowngradeConfidence | 0.6 | confidence needed to spend less |
mainEffortFloor | — | lowest effort: low·medium·high·xhigh |
mainBalancedEffortFloor | — | floor for a task read as ordinary engineering |
mainEffortCeiling | — | highest effort below the deep tier |
mainModelFloor | balanced | lowest tier the main loop's model may go to (model routing only) |
mainMinModelDowngradeConfidence | 0.9 | confidence needed for a main-loop model downgrade |
mainModelMaxContextTokens | 80000 | above this context size, a main-loop model switch is held, except a return to the session's starting model |
fastModel / balancedModel / deepModel | haiku / sonnet / opus | the three tiers, as aliases or full ids |
| Option | Default | |
|---|---|---|
codexAgentTypes | codex:codex-rescue | comma-separated agent types treated as a Codex delegation |
codexFastModel / codexBalancedModel / codexDeepModel | gpt-6-luna / gpt-6-sol / gpt-6-sol | the Codex ladder. Make all three the same to route effort only. |
logDecisions | true | transcript lines and the status line |
decisionLog | true | the per-session JSONL decision log |
decisionLogDir | ~/.claude/jev-router | where it is written |
No [jev-model-router] lines at all. Check these in order:
claude -p and the SDK have no transcript, so every line goes to ~/.claude/debug/<session-id>.txt. The decision log is written either way.claude --debug should print hooks module jev-model-router@skills-dir loaded …; events: prompt.submit,turn.step,agent.spawn. A plugin in a project's .claude/skills/ is only read once that project is trusted; a user-level ~/.claude/skills/ link has no such step.rollout flag (tengu_plugin_hooks_modules) is off. Set CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1.ready on the built-in classifier, no key set when you did set a key: the key is under the wrong pluginConfigs id. See Options.
Effort never moves: check the reason in the decision log. kept …, wanted … means a floor or ceiling held it, or the confidence was below the bar. no decision means the classification timed out: raise timeoutMs.
timeoutMs.bun test # pure policy tests: hooks/policy.ts
hooks/
jev-model-router.ts wiring: prompt.submit · turn.step · agent.spawn
policy.ts every decision as a pure function (route, effort ceiling, pins, Codex flags)
hooks.json module manifest
tests/policy.test.ts bun:test
.claude-plugin/
plugin.json manifest and userConfig
All decision logic lives in policy.ts and is unit-tested. The hook wiring has no automated tests and is checked with live claude -p sessions against the decision log. .types/ and .claude-plugin/types/ are type declarations that Claude Code writes when it loads the plugin, so they are not in the repository, and tsconfig.json resolves only after a first load.
This is an early-access API: mods need Claude Code 2.1.259+, and the $ API may change between releases. The plugin is typed against Anthropic's declarations. It talks to both backends over $.http.fetch, because a mod runs without node_modules. The TypeSafe wire shape follows @typesafe-ai/sdk v0.6.0.
Based on the jev-model-router mod from davila7/claude-code-templates by Daniel (San) Ávila. This fork adds effort-only routing, subagent effort through turn.step, the Codex effort ladder, definition-pin detection and the measurements above.
MIT. See LICENSE.
hooks/jev-model-router.ts 678 lines1/**
2 * jev-model-router — Claude Mod (EARLY ACCESS)
3 *
4 * Picks the model each task runs on with TypeSafe's Jev, a System One
5 * decision model: unstructured state in, a typed choice with a probability
6 * distribution out.
7 *
8 * Jev is reached one of two ways, whichever key is configured: TypeSafe's
9 * own API (`typesafeApiKey`), which reports a calibrated confidence per
10 * answer, or the Vercel AI Gateway (`gatewayApiKey`), which does not. With
11 * neither, the engine's own `$.model.classify` stands in, so the mod is
12 * useful without any account.
13 *
14 * Four things it can set, each on its own switch:
15 * agent.spawn — the model of each subagent (on by default)
16 * turn.step — the reasoning effort of each Claude subagent (on by default)
17 * turn.step — the reasoning effort of the main loop (on by default)
18 * turn.step — the model of the main loop (off by default: switching
19 * models mid-session invalidates the prompt cache, which can
20 * cost more than the cheaper tier saves)
21 *
22 * Every one of them moves in both directions: a task the decision model reads
23 * as mechanical is routed down, one it reads as hard is routed up. The two
24 * mistakes do not cost the same, so they do not clear the same confidence bar
25 * (see `minUpgradeConfidence` / `minDowngradeConfidence` in policy.ts).
26 *
27 * The Agent tool has no effort parameter; a subagent's first turn.step sets it.
28 *
29 * The prompt is classified at `prompt.submit`, which runs before the turn
30 * starts, and the decision is applied at the turn's first request.
31 *
32 * Every failure path is fail-open: a classification that errors or runs past
33 * the latency budget leaves the request exactly as the engine built it.
34 *
35 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
36 * or "gatewayApiKey"). Never hardcode it in this file.
37 *
38 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (Claude Code >= 2.1.259). Typed
39 * against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
40 *
41 * Privacy: with a key set, the prompt text is sent to whichever backend the
42 * key belongs to.
43 */
44import type { EngineInterface, Register } from 'claude-code'
45import {
46 DEFAULT_BASE_URL,
47 DEFAULT_MODEL,
48 definitionDirs,
49 definitionMatches,
50 definitionModel,
51 definitionEffort,
52 pluginAgentDirs,
53 describeDecision,
54 describeSetup,
55 codexFlags,
56 describeStatus,
57 gateMainModel,
58 rankOf,
59 endpoint,
60 effortForModel,
61 reuseForStep,
62 pendingDecisions,
63 readDecision,
64 selectProvider,
65 requestBody,
66 requestHeaders,
67 requestModelId,
68 route,
69 subagentEffortRouting,
70 EFFORT_ORDER,
71 TIER_ORDER,
72 bareCommand,
73} from './policy.ts'
74import type { Decision, Effort, PolicyConfig, Provider, Tier } from './policy.ts'
75
76export const register: Register = (on, options) => {
77 const text = (key: string, fallback: string) =>
78 typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
79 const number = (key: string, fallback: number) =>
80 typeof options[key] === 'number' ? (options[key] as number) : fallback
81 const flag = (key: string, fallback: boolean) =>
82 typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
83
84 // TypeSafe's own API is preferred when both keys are set: it is the only
85 // one that reports a calibrated confidence, which the policy's threshold
86 // reads. `provider` forces one, including "builtin" to use neither.
87 const typesafeKey = text('typesafeApiKey', '')
88 const gatewayKey = text('gatewayApiKey', '')
89 const forced = text('provider', 'auto')
90 const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey)
91
92 // Each backend keeps its own URL and model, so an override written for one
93 // can never be sent to the other when `auto` picks differently than expected.
94 const apiKey = active === 'typesafe' ? typesafeKey : active === 'gateway' ? gatewayKey : ''
95 const modelId = !active
96 ? ''
97 : active === 'typesafe'
98 ? text('typesafeModel', DEFAULT_MODEL.typesafe)
99 : text('gatewayModel', DEFAULT_MODEL.gateway)
100 const url = !active
101 ? ''
102 : active === 'typesafe'
103 ? endpoint('typesafe', text('typesafeBaseUrl', DEFAULT_BASE_URL.typesafe))
104 : endpoint('gateway', text('gatewayBaseUrl', DEFAULT_BASE_URL.gateway))
105
106 // A backend named in the options but missing its key degrades to the
107 // built-in classifier, which is silent; say so once, when a hook first runs.
108 let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
109
110 const timeoutMs = number('timeoutMs', 800)
111 const routeSubagentModel = flag('routeSubagentModel', true)
112 const routeSubagentEffort = flag('routeSubagentEffort', true)
113 // A spawn of the Codex rescue agent gets `--model` / `--effort` for Codex
114 // itself from the same decision; the agent forwards them. Its own model is
115 // left alone: it only relays, so a bigger one would be money for nothing.
116 const routeCodexDelegation = flag('routeCodexDelegation', true)
117 const codexAgentTypes = text('codexAgentTypes', 'codex:codex-rescue').split(',').map((type) => type.trim())
118 const codexLadder: Record<Tier, string> = {
119 fast: text('codexFastModel', 'gpt-6-luna'),
120 balanced: text('codexBalancedModel', 'gpt-6-sol'),
121 deep: text('codexDeepModel', 'gpt-6-sol'),
122 }
123 const routeMainEffort = flag('routeMainEffort', true)
124 const routeMainModel = flag('routeMainModel', false)
125 // Above this many context tokens a main-loop model switch likely costs more
126 // than it saves (a heuristic, see gateMainModel); only a return to the
127 // starting model, a floor raise or a risk-forced switch passes.
128 const mainModelMaxContextTokens = number('mainModelMaxContextTokens', 80_000)
129 const routeMainLoop = routeMainEffort || routeMainModel
130 const switches = {
131 subagentModel: routeSubagentModel,
132 subagentEffort: routeSubagentEffort,
133 mainEffort: routeMainEffort,
134 mainModel: routeMainModel,
135 }
136 const logDecisions = flag('logDecisions', true)
137 // The debug log exists only under --debug, so an interactive session leaves
138 // no trace of what was routed. This one does: one JSONL file per session in
139 // `decisionLogDir` (default ~/.claude/jev-router), one line per decision.
140 // `$.fs` has no append, so writes are chained and the file is rewritten
141 // whole; a session is the only writer of its own file.
142 const decisionLog: DecisionLog = {
143 enabled: flag('decisionLog', true),
144 dir: text('decisionLogDir', ''),
145 queue: Promise.resolve(),
146 }
147
148 const policy: PolicyConfig = {
149 tiers: {
150 fast: text('fastModel', 'haiku'),
151 balanced: text('balancedModel', 'sonnet'),
152 deep: text('deepModel', 'opus'),
153 },
154 minUpgradeConfidence: number('minUpgradeConfidence', 0.3),
155 minDowngradeConfidence: number('minDowngradeConfidence', 0.6),
156 }
157 // The main loop answers the person directly and carries the whole session,
158 // so its model has a floor and a stricter bar to go down. Its effort has
159 // a floor and ceiling; subagent effort uses those same bounds.
160 const mainFloor = text('mainModelFloor', 'balanced')
161 const effortOption = (key: string): Effort | undefined => {
162 const value = text(key, '')
163 return (EFFORT_ORDER as readonly string[]).includes(value) ? (value as Effort) : undefined
164 }
165 const mainPolicy: PolicyConfig = {
166 ...policy,
167 routeModel: routeMainModel,
168 modelFloor: (TIER_ORDER as readonly string[]).includes(mainFloor) ? (mainFloor as Tier) : undefined,
169 minModelDowngradeConfidence: number('mainMinModelDowngradeConfidence', 0.9),
170 effortFloor: effortOption('mainEffortFloor'),
171 balancedEffortFloor: effortOption('mainBalancedEffortFloor'),
172 effortCeiling: effortOption('mainEffortCeiling'),
173 }
174 const subagentEffortPolicy: PolicyConfig = { ...mainPolicy, routeModel: false }
175 // A spawn's classification and origin wait for its first request, where
176 // the resolved model and runtime definition are finally available.
177 const pendingSpawns = new Map<string, SpawnDecision>()
178 // Null also records a miss, so unknown agents never wait a second time.
179 const subagentEffortById = new Map<string, Effort | null>()
180 let sessionEffort: string | number | undefined
181 let sessionModel: string | undefined
182 // A first step can overtake the agent.spawn continuation that gives us its id.
183 const startingSpawns = new Set<Promise<unknown>>()
184
185 // The classification waiting for the turn that reads its prompt, and what
186 // the current turn settled on. Both are single slots: main-loop turns run
187 // one at a time, so nothing accumulates over a long session. `pending`
188 // reports no decision when two prompts are waiting at once, rather than
189 // routing a turn on a decision made for a different prompt.
190 const pending = pendingDecisions()
191 // Said once, the first time a hook runs. A router that loaded and one that
192 // never loaded are otherwise told apart only by the absence of later lines,
193 // and absence is not evidence: the policy leaves most turns alone anyway.
194 let announced = false
195 let appliedTurnId: string | undefined
196 let applied: { model?: string; effort?: Effort } | null = null
197 // The main loop's model at its first routed turn, before any rewrite: the
198 // one a late switch may still return to. Kept in `$.store` by session id,
199 // since a resume or a reload starts this module over on whatever model the
200 // last turn was switched to.
201 let baseModel: string | undefined
202
203 on('prompt.submit', async ($, e, next) => {
204 // Before the routing guards: a module whose switches are all off has still
205 // loaded, and that is exactly when its silence is most misleading.
206 if (!announced) {
207 announced = true
208 if (logDecisions) {
209 $.ui.log(
210 `[jev-model-router] ${describeSetup(
211 active,
212 url,
213 switches,
214 forced === 'builtin',
215 )}`,
216 )
217 }
218 }
219 if (!routeMainLoop) return next(e)
220
221 if (!unusableReported) {
222 unusableReported = true
223 $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
224 }
225
226 // A slash command alone gives the decision model only the command's name.
227 // Its turn keeps the session's model and effort; the null put keeps a
228 // previous prompt's decision from reaching it.
229 if (bareCommand(e.text)) {
230 if (logDecisions) $.ui.log('[jev-model-router] a command with nothing after it; leaving the turn alone')
231 pending.put(null)
232 return next(e)
233 }
234
235 const startedAt = await $.clock.now()
236 let decision: Decision | null = null
237 if (active) {
238 try {
239 const response = await Promise.race([
240 $.http.fetch(url, {
241 method: 'POST',
242 headers: requestHeaders(active, apiKey, modelId),
243 body: requestBody(active, { prompt: e.text }, modelId),
244 }),
245 $.clock.sleep(timeoutMs),
246 ])
247 if (response && response.ok) decision = readDecision(response.text)
248 else if (response) $.ui.log(`[jev-model-router] ${active} responded ${response.status}`)
249 else $.ui.log(`[jev-model-router] classification passed ${timeoutMs}ms; leaving the turn alone`)
250 } catch (error) {
251 $.ui.log(`[jev-model-router] classification failed: ${String(error)}`)
252 }
253 } else {
254 // No backend: the engine's own small-model classifier answers the same
255 // question, without the confidence the policy's threshold reads.
256 try {
257 const label = await $.model.classify(e.text, TIER_ORDER)
258 if (label) {
259 decision = {
260 tier: label as Tier,
261 confidence: null,
262 risky: null,
263 effort: null,
264 effortConfidence: null,
265 }
266 }
267 } catch (error) {
268 $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
269 }
270 }
271
272 // What the decision model actually answered, whatever the policy then
273 // does with it. This is the line that proves the classification ran.
274 if (logDecisions) {
275 const ms = (await $.clock.now()) - startedAt
276 $.ui.log(`[jev-model-router] jev: ${describeDecision(decision, ms)}`)
277 }
278
279 pending.put(decision)
280 return next(e)
281 })
282
283 on('turn.step', async function* ($, e, next) {
284 if (e.agentId) {
285 if (!routeSubagentEffort) return yield* next(e)
286 let request = e
287 try {
288 if (!subagentEffortById.has(e.agentId)) {
289 if (!pendingSpawns.has(e.agentId) && startingSpawns.size) {
290 await Promise.race([Promise.allSettled([...startingSpawns]), $.clock.sleep(timeoutMs)])
291 }
292 const spawn = pendingSpawns.get(e.agentId)
293 // Consume once, including a missing spawn: later steps must be stable.
294 pendingSpawns.delete(e.agentId)
295 subagentEffortById.set(e.agentId, null)
296 if (spawn) {
297 const routing = subagentEffortRouting(spawn, { model: e.model, effort: e.effort }, subagentEffortPolicy, { model: sessionModel, effort: sessionEffort })
298 const effort = routing.effort
299 const applied = spawn.model || effort
300 ? { ...(spawn.model ? { model: spawn.model } : {}), ...(effort ? { effort } : {}) }
301 : null
302 const reason = spawn.modelReason
303 ? `model: ${spawn.modelReason}; effort: ${routing.reason}` : routing.reason
304 record($, decisionLog, {
305 event: 'subagent', agentId: e.agentId, agentType: spawn.agentType, ...spawn.decision,
306 from: { model: spawn.fromModel, effort: e.effort, source: spawn.source },
307 ...(e.model !== spawn.fromModel ? { ran: e.model } : {}),
308 applied, reason,
309 })
310 if (logDecisions) $.ui.log(
311 `[jev-model-router] ${spawn.agentType}${effort ? ` → effort ${effort}` : ''}: ${reason}`,
312 )
313 subagentEffortById.set(e.agentId, effort)
314 }
315 }
316 const effort = subagentEffortById.get(e.agentId)
317 if (effort && e.effort !== undefined) request = { ...e, effort: effortForModel(effort, e.model, subagentEffortPolicy)! }
318 } catch (error) {
319 subagentEffortById.set(e.agentId, null)
320 try {
321 $.ui.log(`[jev-model-router] subagent effort routing failed: ${String(error)}`)
322 } catch {
323 // A failed diagnostic must not block the agent's request.
324 }
325 }
326 return yield* next(request)
327 }
328 if (e.index === 0) {
329 sessionEffort = e.effort
330 sessionModel = e.model
331 }
332 if (!routeMainLoop) return yield* next(e)
333
334 // Every request after the first reuses what the turn settled on, so
335 // neither the model nor the effort changes under its own tool loop.
336 if (e.index > 0 && e.turnId === appliedTurnId) {
337 // A later request may have fallen back to a model without effort.
338 const reused = reuseForStep(applied, e, mainPolicy)
339 return yield* next(reused ? { ...e, ...reused } : e)
340 }
341
342 baseModel ??= await startingModel($, e.model)
343 const decision = pending.take()
344 const routing = route(decision, { model: e.model, effort: e.effort }, mainPolicy)
345 const change: { model?: string; effort?: Effort } = {}
346 let held: string | undefined
347 // The main loop's `model` is sent to the API as written, so an alias
348 // becomes its id here; a subagent's (agent.spawn) may stay an alias.
349 if (routeMainModel && routing.model) {
350 let contextTokens: number | undefined
351 try {
352 contextTokens = (await $.session.usage()).context.tokens
353 } catch {
354 // No reading: gate on nothing rather than block the turn.
355 }
356 // A loop already below its floor (switched there before the floor
357 // existed) is raised whatever the context: the floor is about quality.
358 const current = rankOf(e.model, policy.tiers)
359 const floor = mainPolicy.modelFloor ? TIER_ORDER.indexOf(mainPolicy.modelFloor) : 0
360 if (current !== null && current < floor) contextTokens = undefined
361 // Risk forcing the deep tier is not a cost question; the limit is.
362 if (routing.forced) contextTokens = undefined
363 const gated = gateMainModel(requestModelId(routing.model), contextTokens, mainModelMaxContextTokens, baseModel)
364 if (gated.model && gated.model !== e.model) change.model = gated.model
365 held = gated.reason
366 }
367 if (routeMainEffort && routing.effort) change.effort = routing.effort
368
369 appliedTurnId = e.turnId
370 applied = Object.keys(change).length > 0 ? change : null
371 // A row in the transcript scrolls away; this line stays on screen.
372 if (logDecisions) $.ui.status(describeStatus(decision, applied))
373 record($, decisionLog, { event: 'main', ...decision, from: { model: e.model, effort: e.effort }, applied, reason: held ? `${routing.reason}; ${held}` : routing.reason })
374
375 if (!applied) {
376 // A turn left alone is the common case, and it used to be silent, which
377 // made a working mod look like one that never loaded. Say what happened.
378 if (logDecisions) {
379 const suppressed = held ? ` (${held})` : routing.wantedModel && !routeMainModel ? ' (main-loop model routing off)' : ''
380 $.ui.log(`[jev-model-router] main loop: ${routing.reason}${suppressed}`)
381 }
382 return yield* next(e)
383 }
384 if (logDecisions) {
385 const what = [change.model, change.effort && `effort ${change.effort}`]
386 .filter(Boolean)
387 .join(', ')
388 $.ui.log(`[jev-model-router] main loop → ${what}: ${routing.reason}`)
389 }
390 return yield* next({ ...e, ...change })
391 })
392
393 on('agent.spawn', async ($, e, next) => {
394 // Before the routing guards: a module whose switches are all off has still
395 // loaded, and that is exactly when its silence is most misleading.
396 if (!announced) {
397 announced = true
398 if (logDecisions) {
399 $.ui.log(
400 `[jev-model-router] ${describeSetup(
401 active,
402 url,
403 switches,
404 forced === 'builtin',
405 )}`,
406 )
407 }
408 }
409
410 // A fork inherits its parent's model; `model` is ignored for it.
411 const codexType = codexAgentTypes.includes(e.subagentType)
412 const codexSpawn = routeCodexDelegation && codexType
413 const effortEnabled = routeSubagentEffort && !codexType
414 if ((!routeSubagentModel && !effortEnabled && !codexSpawn) || e.fork) return next(e)
415
416 if (!unusableReported) {
417 unusableReported = true
418 $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
419 }
420
421 const startedAt = await $.clock.now()
422 let decision: Decision | null = null
423 if (active) {
424 try {
425 const response = await Promise.race([
426 $.http.fetch(url, {
427 method: 'POST',
428 headers: requestHeaders(active, apiKey, modelId),
429 body: requestBody(
430 active,
431 { prompt: e.prompt, description: e.description, agentType: e.subagentType },
432 modelId,
433 ),
434 }),
435 $.clock.sleep(timeoutMs),
436 ])
437 if (response && response.ok) decision = readDecision(response.text)
438 else if (response) $.ui.log(`[jev-model-router] ${active} responded ${response.status}`)
439 else $.ui.log(`[jev-model-router] classification passed ${timeoutMs}ms; leaving the subagent alone`)
440 } catch (error) {
441 $.ui.log(`[jev-model-router] classification failed: ${String(error)}`)
442 }
443 } else {
444 try {
445 const label = await $.model.classify(e.prompt, TIER_ORDER)
446 if (label) {
447 decision = {
448 tier: label as Tier,
449 confidence: null,
450 risky: null,
451 effort: null,
452 effortConfidence: null,
453 }
454 }
455 } catch (error) {
456 $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
457 }
458 }
459
460 if (logDecisions) {
461 const ms = (await $.clock.now()) - startedAt
462 $.ui.log(`[jev-model-router] jev (${e.subagentType}): ${describeDecision(decision, ms)}`)
463 }
464
465 // The caller's model wins. Without one, the definition may pin a model;
466 // otherwise the agent inherits its parent's model. Keep that origin for
467 // the first-step record even if agent.spawn rewrites the request.
468 const definition = e.model && !effortEnabled ? null : await definedAgent($, e.subagentType)
469 const defined = definition?.model ?? null
470 const current = e.model ?? defined ?? e.parentModel
471 const source = e.model ? 'call' : defined ? 'definition' : 'parent'
472 const routed = codexSpawn || !routeSubagentModel ? null : route(decision, { model: current }, policy)
473 const model = routed?.model ?? null
474 const codex = codexSpawn ? codexFlags(decision, e.prompt, codexLadder) : null
475 const reason = codexSpawn
476 ? codex
477 ? `codex ${codex.model}/${codex.effort} (${decision?.tier ?? 'no decision, default'})`
478 : 'codex flags already in the prompt'
479 : (routed?.reason ?? 'no decision')
480 const applied = model || codex ? { ...(model ? { model } : {}), ...(codex ? { codex } : {}) } : null
481 if (!effortEnabled) {
482 record($, decisionLog, { event: 'subagent', agentType: e.subagentType, ...decision, from: { model: current, source }, applied, reason })
483 }
484 if (!applied) {
485 if (logDecisions && (!effortEnabled || routed)) $.ui.log(`[jev-model-router] ${e.subagentType}: ${reason}`)
486 } else if (logDecisions) {
487 $.ui.log(`[jev-model-router] ${e.subagentType} → ${[model, codex && `codex ${codex.model}/${codex.effort}`].filter(Boolean).join(', ')}: ${reason}`)
488 }
489 const changed = {
490 ...e,
491 ...(model ? { model } : {}),
492 ...(codex ? { prompt: `--model ${codex.model} --effort ${codex.effort} ${e.prompt}` } : {}),
493 }
494 if (!effortEnabled) return next(applied ? changed : e)
495 const startedPromise = next(changed).then((started) => {
496 // A timed-out first step has already cached a miss; do not leave a
497 // pending decision that no later step will consume.
498 if (started.agentId && !subagentEffortById.has(started.agentId)) pendingSpawns.set(started.agentId, {
499 decision, pinned: definition?.effort ?? null, agentType: e.subagentType,
500 source, fromModel: current, model, modelReason: routed ? reason : null,
501 })
502 return started
503 })
504 startingSpawns.add(startedPromise)
505 try {
506 return await startedPromise
507 } finally {
508 startingSpawns.delete(startedPromise)
509 }
510 })
511}
512
513// The debug log exists only under --debug, so an interactive session leaves
514// no trace of what was routed. This one does: one JSONL file per session in
515// `decisionLogDir` (default ~/.claude/jev-router), one line per decision.
516// `$.fs` has no append, so writes are chained and the file is rewritten
517// whole; a session is the only writer of its own file.
518type DecisionLog = { enabled: boolean; dir: string; queue: Promise<void>; path?: string }
519type SpawnDecision = {
520 decision: Decision | null
521 pinned: Effort | null
522 agentType: string
523 source: 'call' | 'definition' | 'parent'
524 fromModel: string
525 model: string | null
526 modelReason: string | null
527}
528type Engine = EngineInterface
529
530// An agent type's definition pins these values, or leaves them unset. Answers
531// are cached for a minute per type, so a burst of spawns scans folders once.
532const DEFINITION_TTL_MS = 60_000
533type AgentDefinition = { model: string | null; effort: Effort | null }
534const definitionCache = new Map<string, AgentDefinition & { at: number }>()
535
536async function definedAgent($: Engine, subagentType: string): Promise<AgentDefinition> {
537 const now = await $.clock.now()
538 const cached = definitionCache.get(subagentType)
539 if (cached && now - cached.at < DEFINITION_TTL_MS) return cached
540 const definition: AgentDefinition = { model: null, effort: null }
541 try {
542 const home = (await $.env.get('HOME')) ?? ''
543 const pluginDirs: string[] = []
544 const colon = subagentType.indexOf(':')
545 if (colon > 0) {
546 const registry = `${home}/.claude/plugins/installed_plugins.json`
547 if (await $.fs.exists(registry)) {
548 const plugins = JSON.parse(await $.fs.read(registry)).plugins ?? {}
549 const prefix = `${subagentType.slice(0, colon)}@`
550 const installs = Object.entries(plugins)
551 .filter(([key]) => key.startsWith(prefix))
552 .flatMap(([, entries]) => (entries as { installPath?: string }[]).map((entry) => entry.installPath ?? ''))
553 .filter(Boolean)
554 for (const install of installs) {
555 let manifest: unknown = {}
556 try {
557 const manifestPath = `${install}/.claude-plugin/plugin.json`
558 if (await $.fs.exists(manifestPath)) manifest = JSON.parse(await $.fs.read(manifestPath))
559 } catch {
560 // An unreadable manifest still leaves the default agents/ folder.
561 }
562 pluginDirs.push(...pluginAgentDirs(install, manifest))
563 }
564 }
565 }
566 const { agent, dirs } = definitionDirs(subagentType, await $.session.root(), home, pluginDirs)
567 const budget = { reads: DEFINITION_READ_BUDGET }
568 for (const dir of dirs) {
569 const found = await findDefinition($, dir, agent, budget)
570 if (found !== undefined) {
571 definition.model = definitionModel(found)
572 definition.effort = definitionEffort(found)
573 break
574 }
575 }
576 } catch (error) {
577 $.ui.log(`[jev-model-router] reading the ${subagentType} definition failed: ${String(error)}`)
578 }
579 definitionCache.set(subagentType, { ...definition, at: now })
580 return definition
581}
582
583// The text of the definition in `dir` (or up to two folders below it) that
584// defines `agent`, else undefined. In each folder the file named after the
585// agent is tried first, the common case. A file or folder that cannot be read
586// is skipped, not fatal. `budget` caps the files read across one lookup.
587const DEFINITION_READ_BUDGET = 300
588
589async function findDefinition(
590 $: Engine,
591 dir: string,
592 agent: string,
593 budget: { reads: number },
594 depth = 0,
595): Promise<string | undefined> {
596 const read = async (path: string): Promise<string | undefined> => {
597 if (budget.reads <= 0) return undefined
598 budget.reads--
599 try {
600 return await $.fs.read(path)
601 } catch {
602 return undefined
603 }
604 }
605 let entries: Awaited<ReturnType<Engine['fs']['list']>>
606 try {
607 if (!(await $.fs.exists(dir))) return undefined
608 entries = await $.fs.list(dir)
609 } catch {
610 return undefined
611 }
612 const guess = `${agent}.md`
613 const files = entries.filter((entry) => entry.kind === 'file' && entry.name.endsWith('.md'))
614 for (const entry of [...files.filter((f) => f.name === guess), ...files.filter((f) => f.name !== guess)]) {
615 const text = await read(`${dir}/${entry.name}`)
616 if (text !== undefined && definitionMatches(text, entry.name, agent)) return text
617 }
618 if (depth >= 2) return undefined
619 for (const entry of entries) {
620 if (entry.kind !== 'dir') continue
621 const found = await findDefinition($, `${dir}/${entry.name}`, agent, budget, depth + 1)
622 if (found !== undefined) return found
623 }
624 return undefined
625}
626
627// The session's starting model from `$.store`, recorded on first sight, one
628// key per session so two sessions starting at once cannot overwrite each
629// other. Entries unused for 30 days are dropped; each is re-read just before
630// deletion, so only a session idle that long and resuming in the same instant
631// could lose its entry. A map under the old shared key is still read, never
632// written. Any failure answers `current`.
633const STARTING_MODEL_PREFIX = 'start:'
634const STARTING_MODEL_TTL_MS = 30 * 24 * 60 * 60 * 1000
635const LEGACY_STARTING_MODELS = 'startingModels'
636type StartingModel = { model: string; at: number }
637
638function readStarting(value: unknown): StartingModel | null {
639 if (typeof value === 'string') return { model: value, at: 0 }
640 const entry = value as Partial<StartingModel> | null
641 return entry && typeof entry.model === 'string' ? { model: entry.model, at: Number(entry.at) || 0 } : null
642}
643
644async function startingModel($: Engine, current: string): Promise<string> {
645 try {
646 const now = await $.clock.now()
647 const sessionId = await $.session.id()
648 const key = `${STARTING_MODEL_PREFIX}${sessionId}`
649 const legacy = ((await $.store.get(LEGACY_STARTING_MODELS)) ?? {}) as Record<string, string>
650 const start = readStarting(await $.store.get(key))?.model ?? legacy[sessionId] ?? current
651 await $.store.set(key, { model: start, at: now } satisfies StartingModel)
652 for (const other of await $.store.keys()) {
653 if (other === key || !other.startsWith(STARTING_MODEL_PREFIX)) continue
654 const entry = readStarting(await $.store.get(other))
655 if (entry && entry.at > 0 && now - entry.at > STARTING_MODEL_TTL_MS) await $.store.delete(other)
656 }
657 return start
658 } catch (error) {
659 $.ui.log(`[jev-model-router] starting model not stored: ${String(error)}`)
660 }
661 return current
662}
663
664function record($: Engine, log: DecisionLog, entry: Record<string, unknown>): void {
665 if (!log.enabled) return
666 const ts = new Date().toISOString()
667 log.queue = log.queue
668 .then(async () => {
669 if (!log.path) {
670 const dir = log.dir || `${(await $.env.get('HOME')) ?? '.'}/.claude/jev-router`
671 log.path = `${dir}/${await $.session.id()}.jsonl`
672 }
673 const previous = (await $.fs.exists(log.path)) ? await $.fs.read(log.path) : ''
674 await $.fs.write(log.path, `${previous}${JSON.stringify({ ts, ...entry })}\n`)
675 })
676 .catch((error) => $.ui.log(`[jev-model-router] decision log failed: ${String(error)}`))
677}
678hooks/policy.ts 783 lines1/**
2 * jev-model-router — pure decision logic.
3 *
4 * No `$` and no I/O here: this module only builds the request the decision
5 * API takes, reads its answer, and turns that answer into a model id. The
6 * hooks module does every call on `$` at its own call site.
7 *
8 * Two backends speak to the same model with different wire shapes:
9 *
10 * typesafe POST https://api.typesafe.ai/v1/systemone
11 * `{ model, state, questions }`; a yes/no question is a `noul`
12 * and every answer carries its own `confidence`.
13 * gateway POST https://ai-gateway.vercel.sh/v4/ai/evaluation-model
14 * `{ state, questions }` with the model in a header; a yes/no
15 * question is a `boolean`, and there is no `confidence` field —
16 * it has to be derived from an optional distribution.
17 *
18 * The Gateway shape is not documented publicly; it was read from
19 * @ai-sdk/gateway and @ai-sdk/provider.
20 */
21
22export type Provider = 'typesafe' | 'gateway'
23
24export type Tier = 'fast' | 'balanced' | 'deep'
25
26export interface Tiers {
27 fast: string
28 balanced: string
29 deep: string
30}
31
32export interface Decision {
33 tier: Tier
34 /** Confidence in the tier, or null when the backend reported none. */
35 confidence: number | null
36 /** P(true) that carrying the task out would itself be costly or final. */
37 risky: number | null
38 /** 0..3 along the effort rubric, or null when absent. */
39 effort: number | null
40 /** Confidence in the effort, or null when the backend reported none. */
41 effortConfidence: number | null
42}
43
44/** The reasoning levels a turn can ask for, cheapest first. */
45export const EFFORT_ORDER = ['low', 'medium', 'high', 'xhigh'] as const
46
47export type RoutedEffort = (typeof EFFORT_ORDER)[number]
48export type Effort = RoutedEffort | 'max'
49
50export const TIER_ORDER: readonly Tier[] = ['fast', 'balanced', 'deep']
51
52/**
53 * How each tier is described to the decision model. Deliberately about the
54 * shape of the work, not about model names: the model never sees an id.
55 */
56const TIER_CRITERIA: Record<Tier, string> = {
57 fast: 'Mechanical and local: read or summarise a file, run one command, rename a symbol, answer something already in context.',
58 balanced:
59 'Ordinary engineering: implement a well-specified change across a few files, write tests, fix a clearly described bug, review a small diff.',
60 deep: 'Hard or high-stakes: architecture and design, debugging a failure whose cause is unknown, security, data migrations, concurrency, anything touching production or money.',
61}
62
63const EFFORT_RUBRIC = ['almost none', 'some', 'a lot', 'as much as possible'] as const
64
65export const DEFAULT_BASE_URL: Record<Provider, string> = {
66 typesafe: 'https://api.typesafe.ai',
67 gateway: 'https://ai-gateway.vercel.sh/v4/ai',
68}
69
70/**
71 * The Gateway's own protocol version, sent as `ai-gateway-protocol-version`.
72 * Tracks the `AI_GATEWAY_PROTOCOL_VERSION` of `@ai-sdk/gateway` (4.0.87).
73 */
74const AI_GATEWAY_PROTOCOL_VERSION = '0.0.1'
75
76export const DEFAULT_MODEL: Record<Provider, string> = {
77 typesafe: 'jev-latest',
78 gateway: 'typesafe-ai/jev',
79}
80
81/**
82 * Which backend a configuration asks for, or null for the built-in
83 * classifier. `auto` prefers TypeSafe, since it is the only one that reports
84 * a calibrated confidence; a forced backend whose key is missing resolves to
85 * null rather than falling through to the other one's key.
86 */
87export function selectProvider(
88 forced: string,
89 typesafeKey: string,
90 gatewayKey: string,
91): Provider | null {
92 if (forced === 'builtin') return null
93 if (forced === 'typesafe') return typesafeKey ? 'typesafe' : null
94 if (forced === 'gateway') return gatewayKey ? 'gateway' : null
95 if (typesafeKey) return 'typesafe'
96 if (gatewayKey) return 'gateway'
97 return null
98}
99
100/** The full endpoint a backend posts to. */
101export function endpoint(provider: Provider, baseUrl: string): string {
102 const root = baseUrl.replace(/\/+$/, '')
103 return provider === 'typesafe' ? `${root}/v1/systemone` : `${root}/evaluation-model`
104}
105
106/** The `questions` map, in the shape the backend's schema names. */
107export function questions(provider: Provider): Record<string, unknown> {
108 return {
109 tier: {
110 type: 'choice',
111 instructions: 'Which is the cheapest tier that can complete this coding task well?',
112 criteria: TIER_CRITERIA,
113 },
114 effort: {
115 type: 'score',
116 instructions: 'How much step-by-step reasoning does this task need?',
117 criteria: EFFORT_RUBRIC,
118 },
119 risky: {
120 // The same question under two names: `noul` on TypeSafe's own API,
121 // `boolean` in the AI SDK's evaluation schema.
122 type: provider === 'typesafe' ? 'noul' : 'boolean',
123 // Asked about the act, not the subject. The first wording ("the task
124 // touches production, money, credentials") scored 0.96 on "add a
125 // refund endpoint that calls Stripe" — ordinary code that happens to be
126 // about money — and would have escalated it past a 0.98-confidence
127 // answer of the balanced tier.
128 instructions:
129 'Carrying out this task would itself change production, move real money, or alter data that cannot be restored. Writing or testing code that deals with such things, without running it against the real system, does not count.',
130 },
131 }
132}
133
134/** The request body. The Gateway carries the model in a header instead. */
135export function requestBody(
136 provider: Provider,
137 state: Record<string, unknown>,
138 model: string,
139): string {
140 const body =
141 provider === 'typesafe'
142 ? { model, state, questions: questions(provider) }
143 : { state, questions: questions(provider) }
144 return JSON.stringify(body)
145}
146
147/** The request headers. */
148export function requestHeaders(
149 provider: Provider,
150 apiKey: string,
151 model: string,
152): Record<string, string> {
153 const common = { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` }
154 if (provider === 'typesafe') return common
155 return {
156 ...common,
157 'ai-gateway-auth-method': 'api-key',
158 'ai-model-id': model,
159 // The Gateway rejects any request that does not name the protocol it
160 // speaks: 400 "Unsupported gateway protocol version". Every other header
161 // here is accepted without it, so the omission fails the whole backend.
162 'ai-gateway-protocol-version': AI_GATEWAY_PROTOCOL_VERSION,
163 'ai-evaluation-model-specification-version': '4',
164 }
165}
166
167function isTier(value: unknown): value is Tier {
168 return value === 'fast' || value === 'balanced' || value === 'deep'
169}
170
171/**
172 * Reads a response from either backend.
173 *
174 * TypeSafe's own API reports a `confidence` per answer and a `noul` number
175 * for a yes/no question. The Gateway reports neither: confidence has to come
176 * from the highest probability of a distribution that is itself optional, and
177 * a yes/no answer arrives as `probability`. Both are handled, and a missing
178 * confidence reads as null rather than as a number the policy would trust.
179 */
180export function readDecision(responseText: string): Decision | null {
181 let parsed: unknown
182 try {
183 parsed = JSON.parse(responseText)
184 } catch {
185 return null
186 }
187 const answers = (parsed as { answers?: Record<string, Record<string, unknown>> }).answers
188 if (!answers) return null
189
190 const tierAnswer = answers.tier
191 if (!tierAnswer || !isTier(tierAnswer.choice)) return null
192
193 const effortAnswer = answers.effort
194 const riskyAnswer = answers.risky
195 const risky =
196 typeof riskyAnswer?.noul === 'number'
197 ? riskyAnswer.noul
198 : typeof riskyAnswer?.probability === 'number'
199 ? riskyAnswer.probability
200 : null
201
202 return {
203 tier: tierAnswer.choice,
204 confidence: confidenceOf(tierAnswer),
205 effort: typeof effortAnswer?.score === 'number' ? effortAnswer.score : null,
206 effortConfidence: effortAnswer ? confidenceOf(effortAnswer) : null,
207 risky,
208 }
209}
210
211/**
212 * How sure an answer is. TypeSafe reports it; the Gateway does not, so there
213 * it is the highest probability of a distribution that is itself optional.
214 */
215function confidenceOf(answer: Record<string, unknown>): number | null {
216 if (typeof answer.confidence === 'number') return answer.confidence
217 const probabilities = answer.probabilities as Record<string, number> | undefined
218 const values = probabilities ? Object.values(probabilities) : []
219 return values.length > 0 ? Math.max(...values) : null
220}
221
222/** The rubric score (0..3) as a reasoning level. */
223export function effortLevel(score: number): RoutedEffort {
224 const index = Math.min(EFFORT_ORDER.length - 1, Math.max(0, Math.round(score)))
225 return EFFORT_ORDER[index]!
226}
227
228/**
229 * Where a reasoning level sits on the ladder, or null when its place cannot
230 * be known. `max` is above every rung the rubric can produce, so it ranks
231 * above them without joining EFFORT_ORDER, which is also the set of values
232 * this router is allowed to ask for.
233 */
234export function effortRank(effort: string | number | undefined): number | null {
235 if (typeof effort !== 'string') return null
236 if (effort === 'max') return EFFORT_ORDER.length
237 const index = EFFORT_ORDER.findIndex((level) => level === effort)
238 return index === -1 ? null : index
239}
240
241/**
242 * Where a model id sits on the tier ladder, by matching it against the
243 * configured tier names first and then the family words. Null when it matches
244 * none, in which case the change is treated as an upgrade rather than guessed
245 * at: an unrecognised id gets the gentler threshold, never the strict one.
246 */
247export function rankOf(model: string, tiers: Tiers): number | null {
248 const lowered = model.toLowerCase()
249 for (let index = 0; index < TIER_ORDER.length; index++) {
250 const tier = TIER_ORDER[index] as Tier
251 const configured = tiers[tier].toLowerCase()
252 if (configured && lowered.includes(configured)) return index
253 }
254 if (lowered.includes('haiku')) return 0
255 if (lowered.includes('sonnet')) return 1
256 if (lowered.includes('opus')) return 2
257 return null
258}
259
260/** Recheck the effort ceiling against the model on this request. */
261export function effortForModel(effort: Effort | undefined, model: string, config: PolicyConfig): Effort | null {
262 if (effort === undefined) return null
263 const ceiling = config.effortCeiling
264 if (ceiling && rankOf(model, config.tiers) !== TIER_ORDER.indexOf('deep') && effortRank(effort)! > effortRank(ceiling)!) {
265 return ceiling
266 }
267 return effort
268}
269
270/** Reuse a turn's choice against the model the next step will actually send. */
271export function reuseForStep(
272 applied: { model?: string; effort?: Effort } | null,
273 step: { model: string; effort?: string | number },
274 config: PolicyConfig,
275): { model?: string; effort?: Effort } | null {
276 if (!applied) return null
277 const reused = step.effort === undefined ? (applied.model ? { model: applied.model } : null) : applied
278 if (!reused) return null
279 const effort = reused.effort ? effortForModel(reused.effort, reused.model ?? step.model, config) : null
280 return effort ? { ...reused, effort } : reused
281}
282
283/**
284 * The full id a family alias names on the main loop.
285 *
286 * `agent.spawn` takes an alias (`haiku`) the way the Agent tool does, but
287 * `turn.step`'s `model` is the id the engine already resolved for the request
288 * and goes to the API as written: an alias there is refused ("There's an
289 * issue with the selected model (haiku)"). So the tiers stay aliases in the
290 * options, and only a main-loop rewrite resolves them, here.
291 *
292 * A stale entry here is a silent downgrade: `opus` sat on `claude-opus-5`
293 * (Opus 5, not 5.5) for the whole time Opus 5.5 was current, so every
294 * `routeMainModel` upgrade to `deep` landed one point release behind. Bump
295 * this table whenever Anthropic ships a new point release in a family the
296 * tiers use — checked each time, not guessed: `claude --restricted -p
297 * --model <alias> --output-format json "OK"` names the id the engine itself
298 * resolves the alias to in `modelUsage` (`--restricted` keeps this router
299 * from rewriting the probe).
300 * Last checked 2026-09-28, against Opus 5.5 / Sonnet 5.5 / Haiku 4.5.
301 */
302const ALIAS_IDS: Record<string, string> = {
303 haiku: 'claude-haiku-4-5-20251001',
304 sonnet: 'claude-sonnet-5-5',
305 opus: 'claude-opus-5-5',
306}
307
308/**
309 * What to write into `turn.step`'s `model`: a full id as given, or the id
310 * behind a family alias. Anything else is returned unchanged for the engine
311 * to judge.
312 */
313export function requestModelId(model: string): string {
314 return ALIAS_IDS[model.trim().toLowerCase()] ?? model
315}
316
317export interface PolicyConfig {
318 tiers: Tiers
319 /** Whether the policy may change the request's model. Defaults to true. */
320 routeModel?: boolean
321 /**
322 * How sure the decision must be to spend more (a bigger model, more
323 * reasoning). Being wrong here costs money, so the bar is low.
324 */
325 minUpgradeConfidence: number
326 /**
327 * How sure it must be to spend less. Being wrong here means a task handled
328 * by too small a model or too little thought, so the bar is high.
329 */
330 minDowngradeConfidence: number
331 /**
332 * The lowest tier the model may be routed to; absent, any. Effort is not
333 * floored. The main loop sets `balanced`: a turn's prompt can read as
334 * mechanical ("проверь") while the work around it is not, and a haiku
335 * answer there was measured wrong (2026-09-26).
336 */
337 modelFloor?: Tier
338 /**
339 * The lowest reasoning level a turn may be routed to; absent, any. Unlike
340 * a moved level this is not a guess about the prompt, so a low reading is
341 * lifted to it without a confidence check. A downgrade stops here.
342 */
343 effortFloor?: Effort
344 /** A higher floor for a balanced reading; the larger of the two applies. */
345 balancedEffortFloor?: Effort
346 /**
347 * The highest level while the loop runs below the deep tier, so xhigh and
348 * max belong to opus. Like the floor it is policy: a level above it comes
349 * down without a confidence check.
350 */
351 effortCeiling?: Effort
352 /**
353 * The downgrade bar for the model alone, when it should differ from
354 * `minDowngradeConfidence`, which then still governs effort.
355 */
356 minModelDowngradeConfidence?: number
357}
358
359export interface Routing {
360 /** The model to run on, or null to leave the request as it is. */
361 model: string | null
362 /** The classified model even when this route cannot change it. */
363 wantedModel: string | null
364 /** The reasoning level to ask for, or null to leave it as it is. */
365 effort: Effort | null
366 /** Why, for the log line. */
367 reason: string
368 /**
369 * True when risk forced the deep tier past the thresholds. A caller's own
370 * gate (the main loop's context limit) must let such a change through too:
371 * risk is the one case that is not a cost question.
372 */
373 forced: boolean
374}
375
376const NOTHING: Routing = { model: null, wantedModel: null, effort: null, reason: 'no decision', forced: false }
377
378/**
379 * Whether a change of rank passes its threshold. Both directions are allowed;
380 * they just do not have to clear the same bar, because the two mistakes do not
381 * cost the same. A move whose direction cannot be told (an unrecognised
382 * current value) is treated as an upgrade.
383 */
384function allowed(
385 wanted: number,
386 current: number | null,
387 confidence: number | null,
388 config: PolicyConfig,
389 downgradeBar: number = config.minDowngradeConfidence,
390): boolean {
391 if (current !== null && wanted === current) return false
392 const isDowngrade = current !== null && wanted < current
393 const bar = isDowngrade ? downgradeBar : config.minUpgradeConfidence
394 // A backend that reports no confidence (the Gateway without a distribution,
395 // or the built-in classifier) clears the upgrade bar but never the
396 // downgrade one: spending less on an unmeasured hunch is the bad trade.
397 if (confidence === null) return !isDowngrade
398 return confidence >= bar
399}
400
401/**
402 * Turns a decision into a model and a reasoning level, either of which may be
403 * null to leave the request as it is. Both can move in either direction.
404 */
405export function route(
406 decision: Decision | null,
407 current: { model: string; effort?: string | number },
408 config: PolicyConfig,
409): Routing {
410 if (!decision) return NOTHING
411
412 let tier = decision.tier
413 let effortScore = decision.effort
414 let forced = false
415
416 // Carrying out something final is never worth the saving: take the deep
417 // tier and real reasoning, whatever the cheaper answer said, and skip the
418 // thresholds — this is the one case that is not a confidence question.
419 if (decision.risky !== null && decision.risky > 0.7) {
420 tier = 'deep'
421 effortScore = Math.max(effortScore ?? 0, 2)
422 forced = true
423 }
424
425 const floor = config.modelFloor ? TIER_ORDER.indexOf(config.modelFloor) : 0
426 const wantedTier = Math.max(TIER_ORDER.indexOf(tier), floor)
427 const currentTier = rankOf(current.model, config.tiers)
428 const wantedModel = config.tiers[TIER_ORDER[wantedTier] as Tier]
429 const modelBar = config.minModelDowngradeConfidence ?? config.minDowngradeConfidence
430
431 const model =
432 config.routeModel !== false &&
433 wantedModel &&
434 wantedModel !== current.model &&
435 (forced || allowed(wantedTier, currentTier, decision.confidence, config, modelBar))
436 ? wantedModel
437 : null
438
439 let effort: Effort | null = null
440 // An absent effort means this model cannot accept the API parameter.
441 if (effortScore !== null && current.effort !== undefined) {
442 const currentRank = effortRank(current.effort)
443 let wantedRank = EFFORT_ORDER.indexOf(effortLevel(effortScore))
444 const rankOfLevel = (level: Effort | undefined, none: number) => (level ? effortRank(level) ?? none : none)
445 const floorRank = Math.max(
446 rankOfLevel(config.effortFloor, 0),
447 tier === 'balanced' ? rankOfLevel(config.balancedEffortFloor, 0) : 0,
448 )
449 // An unknown model has no proof that it belongs to the deep tier.
450 const runsTier = model ? wantedTier : (currentTier ?? 0)
451 const ceilingRank = runsTier < TIER_ORDER.indexOf('deep') ? rankOfLevel(config.effortCeiling, Infinity) : Infinity
452 const lifted = wantedRank < floorRank && currentRank !== null && currentRank < floorRank
453 const capped = currentRank !== null && currentRank > ceilingRank
454 wantedRank = Math.min(Math.max(wantedRank, floorRank), ceilingRank)
455
456 // Risk raises the floor; it must never lower one. Forcing only skips the
457 // thresholds, so without this clamp a task already at `xhigh` or `max`
458 // and rated mechanically simple would be pulled down to `high` with no
459 // confidence check at all — the opposite of what the rule is for.
460 if (forced && currentRank !== null) wantedRank = Math.max(wantedRank, currentRank)
461
462 // A numeric effort is the caller's own scale, not this ladder; leave it.
463 const comparable = typeof current.effort !== 'number'
464 const wanted = EFFORT_ORDER[Math.min(EFFORT_ORDER.length - 1, wantedRank)]!
465 if (
466 comparable &&
467 wantedRank !== currentRank &&
468 (forced || lifted || capped || allowed(wantedRank, currentRank, decision.effortConfidence, config))
469 ) {
470 effort = wanted
471 }
472 }
473
474 const said = decision.confidence === null ? 'confidence n/d' : `confidence ${decision.confidence.toFixed(2)}`
475
476 if (!model && !effort) {
477 // Naming what it wanted and what it kept is the whole point of this line.
478 // Without it, a mod that classified and decided to leave the request alone
479 // is indistinguishable from one that never loaded.
480 const wantedEffort = effortScore === null ? null : effortLevel(effortScore)
481 const kept = `${current.model}${current.effort === undefined ? '' : `/${current.effort}`}`
482 const wanted = `${wantedModel}${wantedEffort ? `/${wantedEffort}` : ''}`
483 const floored = wantedTier > TIER_ORDER.indexOf(tier) ? `, ${tier} floored to ${TIER_ORDER[wantedTier]}` : ''
484 return { model: null, wantedModel, effort: null, reason: `kept ${kept}, wanted ${wanted}${floored} (${said})`, forced }
485 }
486
487 const label = wantedTier > TIER_ORDER.indexOf(tier) ? `${tier}, floored to ${TIER_ORDER[wantedTier]}` : tier
488 return { model, wantedModel, effort, reason: forced ? (config.routeModel === false ? 'effort forced by risk' : `${label}, forced by risk`) : `${label} (${said})`, forced }
489}
490
491/** Match model ids despite a context-window suffix such as `[1m]`. */
492export function sameModel(model: string, sessionModel?: string): boolean {
493 return sessionModel !== undefined && model.replace(/\[[^\]]+\]$/, '').toLowerCase() === sessionModel.replace(/\[[^\]]+\]$/, '').toLowerCase()
494}
495
496/** One Claude subagent's effort decision, including detectable definition pins. */
497export function subagentEffortRouting(
498 spawn: { decision: Decision | null; pinned: Effort | null },
499 current: { model: string; effort?: string | number },
500 config: PolicyConfig,
501 session?: { model?: string; effort?: string | number },
502): Pick<Routing, 'effort' | 'reason'> {
503 if (spawn.pinned) return { effort: null, reason: `effort pinned by definition (${spawn.pinned})` }
504 // A same-model pin equal to the session level is indistinguishable; pins on
505 // other models from --agents or AgentSpec are not detected by this heuristic.
506 if (sameModel(current.model, session?.model) && current.effort !== undefined && session?.effort !== undefined && current.effort !== session.effort) {
507 return { effort: null, reason: `effort set by agent definition (${current.effort}, session ${session.effort})` }
508 }
509 if (current.effort === undefined) return { effort: null, reason: 'model takes no effort' }
510 const { effort, reason } = route(spawn.decision, current, { ...config, routeModel: false })
511 return { effort, reason }
512}
513
514/**
515 * Whether a prompt is a slash command and nothing else (`/simplify`). The
516 * decision model sees only the name, never the skill or command it runs:
517 * measured on TypeSafe (three calls each), `/simplify`, `/run` and `/github`
518 * came back fast with effort 0.1 to 0.5, low effort for a multi-step skill,
519 * while `/code-review` and `/security-review` came back balanced and deep.
520 * With text after the name there is a task to read, and it is classified.
521 * A one-segment path alone (`/etc`) matches too; nobody sends one as a task.
522 */
523export function bareCommand(text: string): boolean {
524 return /^\/[^\s/]+$/.test(text.trim())
525}
526
527/**
528 * Holds a prompt's classification until the turn that reads that prompt
529 * starts.
530 *
531 * Nothing ties a decision to the turn it belongs to. Prompts can be queued
532 * while the model is busy, a peer session's message can be delivered inside a
533 * running turn, and `prompt.submit` carries no turn id at all while the
534 * session is idle. So when more than one prompt is waiting, `take` reports
535 * none: running a turn on another prompt's decision is a worse outcome than
536 * not routing it, and not routing is what every other failure path here does.
537 */
538export function pendingDecisions(): {
539 put(decision: Decision | null): void
540 take(): Decision | null
541} {
542 let held: Decision | null = null
543 let waiting = 0
544
545 return {
546 put(decision) {
547 waiting += 1
548 // Past the first, which prompt a turn will read is unknowable, so the
549 // slot is emptied instead of holding a decision that may not fit.
550 held = waiting === 1 ? decision : null
551 },
552 take() {
553 const decision = waiting === 1 ? held : null
554 held = null
555 waiting = 0
556 return decision
557 },
558 }
559}
560
561/** A number for the log, or `n/d` when the backend reported none. */
562function reported(value: number | null): string {
563 return value === null ? 'n/d' : value.toFixed(2)
564}
565
566/**
567 * The one-time line that says the router is alive, which backend answers it,
568 * and which routing switches are on.
569 *
570 * Without this, a router that loaded and a router that never loaded are told
571 * apart only by the absence of later lines, which is not evidence of anything.
572 */
573export function describeSetup(
574 provider: Provider | null,
575 url: string,
576 switches: { subagentModel: boolean; subagentEffort: boolean; mainEffort: boolean; mainModel: boolean },
577 // `provider: "builtin"` is a choice, not a missing key. Reporting it as a
578 // credential problem sends someone hunting for a key they meant to omit.
579 builtinByChoice = false,
580): string {
581 const backend = provider
582 ? `${provider} (${url})`
583 : builtinByChoice
584 ? 'the built-in classifier, by choice'
585 : 'the built-in classifier, no key set'
586 const on = [
587 switches.subagentModel && 'subagent model',
588 switches.subagentEffort && 'subagent effort',
589 switches.mainEffort && 'main effort',
590 switches.mainModel && 'main model',
591 ].filter(Boolean)
592 return `ready on ${backend}; routing ${on.length > 0 ? on.join(', ') : 'nothing, every switch is off'}`
593}
594
595/**
596 * What the decision model answered, before any policy touches it: the raw
597 * tier, effort and risk with their confidences, and how long it took.
598 *
599 * This is the line that shows the classification happened at all, separately
600 * from whether the policy then decided to act on it.
601 */
602export function describeDecision(decision: Decision | null, ms: number | null): string {
603 const took = ms === null ? '' : ` · ${Math.round(ms)}ms`
604 if (!decision) return `no answer${took}`
605
606 const parts = [`tier ${decision.tier} (${reported(decision.confidence)})`]
607 if (decision.effort !== null) {
608 parts.push(
609 `effort ${decision.effort.toFixed(1)} → ${effortLevel(decision.effort)} (${reported(decision.effortConfidence)})`,
610 )
611 }
612 if (decision.risky !== null) parts.push(`risky ${reported(decision.risky)}`)
613 return parts.join(' · ') + took
614}
615
616/**
617 * The persistent status line: the last thing the router did, short enough to
618 * sit on screen beside the engine's own notices.
619 */
620export function describeStatus(
621 decision: Decision | null,
622 change: { model?: string; effort?: Effort } | null,
623): string {
624 if (!decision) return 'jev · no answer'
625 const asked = `${decision.tier} ${reported(decision.confidence)}`
626 if (!change) return `jev · ${asked} · unchanged`
627 const to = [change.model, change.effort].filter(Boolean).join('/')
628 return `jev · ${asked} → ${to}`
629}
630
631/**
632 * The `model` an agent definition's frontmatter names, or null when it names
633 * none or `inherit`: both leave the parent's model to decide. `agent.spawn`
634 * reports only the Agent tool's `model` parameter and the parent's model, so
635 * without this a definition pinned to `sonnet` is measured as if it ran on
636 * the parent's `opus`, and a decision to route it up reads as a no-op.
637 */
638export function definitionModel(markdown: string): string | null {
639 const value = frontmatterField(markdown, 'model')
640 return value && value !== 'inherit' ? value : null
641}
642
643/** An agent definition's pinned reasoning effort, if it names a supported level. */
644export function definitionEffort(markdown: string): Effort | null {
645 const value = frontmatterField(markdown, 'effort')
646 return value && effortRank(value) !== null ? (value as Effort) : null
647}
648
649/** One scalar field of a Markdown file's YAML frontmatter, unquoted; null when absent. */
650function frontmatterField(markdown: string, key: string): string | null {
651 const frontmatter = /^---\r?\n([\s\S]*?)\r?\n---/.exec(markdown)
652 if (!frontmatter) return null
653 const line = new RegExp(`^${key}:[ \\t]*(.*)$`, 'm').exec(frontmatter[1])
654 const value = line?.[1]
655 .replace(/\s+#.*$/, '')
656 .trim()
657 .replace(/^(['"])(.*)\1$/, '$2')
658 return value || null
659}
660
661/**
662 * Whether a definition file defines `agent`. Claude Code names an agent by
663 * its frontmatter `name`, not its file name (`reviewer-config.md` may define
664 * `reviewer`); a file without a `name` falls back to its file name. A file
665 * without frontmatter defines no agent.
666 */
667export function definitionMatches(markdown: string, fileName: string, agent: string): boolean {
668 if (!/^---\r?\n/.test(markdown)) return false
669 const name = frontmatterField(markdown, 'name')
670 return name ? name === agent : fileName === `${agent}.md`
671}
672
673/**
674 * The folders a plugin keeps agents in: `agents/` by default, plus each path
675 * its manifest's `agents` field names (a string or a list, relative to the
676 * install). Both are scanned; a name match decides, so a path that turns out
677 * to replace the default rather than add to it costs only a wasted scan.
678 */
679export function pluginAgentDirs(installPath: string, manifest: unknown): string[] {
680 const declared = (manifest as { agents?: unknown } | null)?.agents
681 const extra = (Array.isArray(declared) ? declared : declared === undefined ? [] : [declared])
682 .filter((path): path is string => typeof path === 'string')
683 .map((path) => `${installPath}/${path.replace(/^\.\//, '').replace(/\/+$/, '')}`)
684 return [`${installPath}/agents`, ...extra.filter((dir) => dir !== `${installPath}/agents`)]
685}
686
687/**
688 * The agent name to match and the folders to search, most specific first: a
689 * `plugin:agent` type in that plugin's agent folders, any other in the
690 * project's `.claude/agents`, then the user's. The first definition that
691 * matches wins, whether or not it names a model. Managed and `--agents`
692 * definitions are not visible here; a miss falls back to the parent's model.
693 */
694export function definitionDirs(
695 subagentType: string,
696 projectRoot: string,
697 home: string,
698 pluginDirs: readonly string[],
699): { agent: string; dirs: string[] } {
700 const colon = subagentType.indexOf(':')
701 if (colon > 0) return { agent: subagentType.slice(colon + 1), dirs: [...pluginDirs] }
702 return { agent: subagentType, dirs: [`${projectRoot}/.claude/agents`, `${home}/.claude/agents`] }
703}
704
705/** A model id without its context-window suffix (`[1m]`). */
706function withoutWindow(model: string): string {
707 return model.replace(/\[[^\]]*\]$/, '')
708}
709
710/**
711 * Whether the main loop may move to `wanted` given how full the context is.
712 *
713 * A model switch forfeits the prompt cache: each model keeps its own, so the
714 * new one writes the conversation to its cache before it answers (measured
715 * 2026-09-26: a haiku → opus switch wrote 69k tokens and read none). Early in
716 * a session that is cheap and a mechanical turn on a small model can pay for
717 * it; late, the rewrite likely costs more than the smaller model saves. The
718 * limit is a heuristic, not a measured break-even: that depends on prices,
719 * the turns left and how warm each cache is. Above it the only move allowed
720 * is back to the session's starting model, so a session that dropped to a
721 * small model early is not stuck there for a hard turn later. Its cache is
722 * warm only if it actually answered recently; that is not tracked.
723 *
724 * The starting model is matched exactly, its `[1m]` suffix aside, and then
725 * sent as the starting id, so a 1M session is not cut to 200k by an alias
726 * resolving to the same model. Another version of the same family is a
727 * different model: it is sent as asked, and held above the limit.
728 * `contextTokens` is undefined before the first response of a fresh session:
729 * nothing to lose yet. A resumed one has its reading restored by the engine.
730 */
731export function gateMainModel(
732 wanted: string,
733 contextTokens: number | undefined,
734 maxContextTokens: number,
735 baseModel: string | undefined,
736): { model: string | null; reason?: string } {
737 const isBase = baseModel !== undefined && withoutWindow(wanted) === withoutWindow(baseModel)
738 const model = isBase ? (baseModel as string) : wanted
739 if (contextTokens === undefined || contextTokens <= maxContextTokens || isBase) return { model }
740 return {
741 model: null,
742 reason: `held model: context ${contextTokens} > ${maxContextTokens}, only the starting model ${baseModel ?? '(unknown)'} is allowed`,
743 }
744}
745
746/**
747 * The `--model` / `--effort` a `codex:codex-rescue` spawn should carry, from
748 * the same decision that routed the spawn itself.
749 *
750 * The rescue agent leaves both unset unless its prompt names them, so every
751 * delegation used to run on `~/.codex/config.toml`'s default. Calibrated
752 * 2026-09-27: the workhorse (sol) carries all hard work and effort carries
753 * the difficulty; the top model is not used by default (the ladder's `deep`
754 * rung is sol). Risk above 0.7 takes xhigh; deep takes high at the upgrade
755 * bar (0.3, or with no confidence), xhigh when the effort reads xhigh, medium
756 * below the bar; a balanced task whose effort reads high or more takes
757 * medium; the cheap model needs the downgrade bar (0.6), a measured
758 * confidence and low risk (≤ 0.4). Everything else, and no decision at all,
759 * is the default: sol at low. Null — prompt left alone — only when the
760 * prompt already names either flag: an explicit choice wins.
761 */
762export function codexFlags(
763 decision: Decision | null,
764 prompt: string,
765 ladder: Record<Tier, string>,
766): { model: string; effort: Effort } | null {
767 if (/(^|\s)--(model|effort)(\s|=|$)/.test(prompt)) return null
768 const fallback = { model: ladder.balanced, effort: 'low' as Effort }
769 if (!decision) return fallback
770 const { tier, confidence, risky } = decision
771 const level = decision.effort === null ? null : effortLevel(decision.effort)
772 if (risky !== null && risky > 0.7) return { model: ladder.deep, effort: 'xhigh' }
773 if (tier === 'deep') {
774 if (confidence === null || confidence >= 0.3) return { model: ladder.deep, effort: level === 'xhigh' ? 'xhigh' : 'high' }
775 return { model: ladder.balanced, effort: 'medium' }
776 }
777 if (tier === 'fast' && confidence !== null && confidence >= 0.6 && (risky === null || risky <= 0.4)) {
778 return { model: ladder.fast, effort: 'low' }
779 }
780 if (tier === 'balanced' && (level === 'high' || level === 'xhigh')) return { model: ladder.balanced, effort: 'medium' }
781 return fallback
782}
783