SLOPSHOPPER

jev-model-router

Picks the model and reasoning effort for each task with TypeSafe's Jev. Routes subagent models at agent.spawn and subagent and main-loop effort at turn.step…

newstatuspromptmodelnetworkagents
v0.5.3MITupdated 2026-10-04DmitryBMsk/thinkdial
A shopper browsing a rack in a slop shop
README

thinkdial

Per-task reasoning effort for Claude Code, decided by a typed decision model.

tests license: MIT Claude Code runtime: bun

thinkdial is a Claude Code mod (a plugin of function hooks). Before each turn it asks Jev, TypeSafe's System One decision model, three typed questions about the task: how hard is it, how much reasoning does it need, and is it risky. It then sets the reasoning effort of the request:

  • of the main conversation,
  • of every Claude subagent the Agent tool starts,
  • of every Codex delegation (codex:codex-rescue), as --effort.

It does not change which model answers, unless you turn that on (see why not). Every failure path is fail-open: if the classification is slow, errors out or makes no sense, the request goes out exactly as Claude Code built it.

The plugin id is still jev-model-router (that is what Claude Code and your pluginConfigs know it by). The project is thinkdial.


Contents


Why effort-only

The router started by routing models as well: haiku for mechanical turns, opus for hard ones. One user's week of real sessions, 1,174 routed main-loop turns across 133 sessions in late September 2026, showed why that loses money in a long Claude Code session:

What changed at the start of a turnTurnsPrompt-cache miss
nothing4331 %
effort only4951 %
model1681 %
model + effort2070 %
  • The prompt cache belongs to each model. A new model writes the whole conversation into its own cache before it answers, and in a long session that write is most of the request.
  • Cheaper models do not read the cache for less. Claude Opus 5.5 and Claude Sonnet 5.5 charge the same $0.20 per million cached input tokens, and cache reads were about half of the spend. Moving from opus to sonnet mid-session saved almost nothing and paid for a full cache rewrite.
  • Effort is free to move. Changing effort between turns did not invalidate the cache at all (1 % misses, the same as changing nothing). So effort is the dimension worth routing.

Model routing is still in the code, behind routeMainModel and routeSubagentModel, for anyone whose sessions look different. Measure your own before you turn it on: the decision log below has everything you need.

How it works

flowchart LR
  subgraph Main["Main conversation"]
    A[prompt.submit] -->|prompt text| J1{{Jev}}
    J1 -->|tier · effort · risky| P1[policy]
    P1 --> T1["turn.step<br/>effort for this turn"]
  end
  subgraph Sub["Claude subagent"]
    S[agent.spawn] -->|prompt · description · type| J2{{Jev}}
    J2 --> K[(decision by agentId)]
    K --> T2["subagent's turn.step<br/>effort for every request"]
  end
  subgraph Codex["Codex delegation"]
    C[agent.spawn<br/>codex:codex-rescue] --> J3{{Jev}}
    J3 --> F["prompt += --model gpt-6-sol --effort X"]
  end
  1. Classify once. The main-loop prompt is classified at prompt.submit, before the turn starts. A subagent's prompt is classified at agent.spawn.
  2. Apply at the first request. The decision is applied to the turn's first model request (turn.step) and reused by every later request of that turn, or of that subagent. Effort never changes inside a tool loop.
  3. Re-check every request. The effort ceiling is checked against the model actually sent on each request, so a fallback to a smaller model never carries an effort meant for a bigger one.

Quick start

Requirements: Claude Code 2.1.259+, function hooks enabled, and a TypeSafe API key. Without a key the router still runs on Claude Code's built-in classifier, which reports no confidence and so can only raise effort, never lower it.

git clone https://github.com/DmitryBMsk/thinkdial.git ~/src/thinkdial

# load it for every project: a user-level skills folder is auto-loaded as jev-model-router@skills-dir
ln -s ~/src/thinkdial ~/.claude/skills/jev-model-router

Add this to ~/.claude/settings.json. These are the recommended effort-only settings, and the defaults in the code differ from them, as noted below:

{
  "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" },
  "pluginConfigs": {
    "jev-model-router@skills-dir": {
      "options": {
        "typesafeApiKey": "<your TypeSafe key>",
        "provider": "typesafe",
        "timeoutMs": 1500,

        "routeMainEffort": true,
        "routeSubagentEffort": true,
        "routeMainModel": false,          // default false
        "routeSubagentModel": false,      // default TRUE in code: set false for effort-only
        "codexFastModel": "gpt-6-sol",    // default gpt-6-luna: pin Codex to one model

        "mainEffortFloor": "medium",
        "mainBalancedEffortFloor": "high",
        "mainEffortCeiling": "high"
      }
    }
  }
}

Start claude. The first routed turn prints [jev-model-router] ready on typesafe …; routing subagent effort, main effort.

CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir ~/src/thinkdial

Loaded this way the plugin's id is plain jev-model-router, so its options go under "pluginConfigs": { "jev-model-router": { … } }, for example in a file passed with --settings. For an isolated claude -p worker use --restricted, not --safe-mode: --safe-mode also switches off hooks loaded with --plugin-dir.

claude plugin validate ~/src/thinkdial prints every event the plugin hooks and every $ call it makes.

How it decides

Jev answers three questions in a single request:

QuestionShapeMeaning
tierchoice of 3mechanical and local · ordinary engineering · hard or high-stakes. The decision model never sees a model name.
effortscore 0–3how much step-by-step reasoning the task needs: low · medium · high · xhigh
riskyprobabilitydoes it touch production, money, credentials or state that cannot be undone?

The two kinds of mistake do not cost the same, so they do not clear the same bar:

  • Spending more needs minUpgradeConfidence (0.3). A wrong call costs money.
  • Spending less needs minDowngradeConfidence (0.6). A wrong call is a task done with too little thought.
  • risky above 0.7 forces at least high effort past both bars. It can raise effort but never lower it.
  • No confidence reported (the built-in classifier, or the Gateway without a distribution) means the router may only move effort up.

On top of the decision sit three policy limits. They apply without a confidence check:

OptionEffect
mainEffortFloorlowest effort a turn may get, so a short "check this" prompt is not sent at low
mainBalancedEffortFloorhigher floor for a task read as ordinary engineering
mainEffortCeilinghighest effort below the deep tier, so xhigh/max are kept for opus-class models

A slash command with nothing after it (/simplify) is not classified, because Jev would only see the command's name. That turn keeps the session's effort.

Subagents

The Agent tool has no effort parameter, so the router sets a subagent's effort from inside its loop:

  1. At agent.spawn it classifies the subagent's prompt, description and type.
  2. When the spawn resolves, it stores the decision by agentId.
  3. At the subagent's first turn.step it applies the effort, then reuses it for every later request and turn of that agent. If the first request overtakes the spawn, it waits for the spawn once, up to timeoutMs.

Subagents use the same floor and ceiling as the main loop. Some subagents are left alone:

CaseBehaviour
definition pins effort: in its .md filekept: effort pinned by definition (medium)
--agents / registered definition pins an effort on the session's modelkept, detected when its effort differs from the session's: effort set by agent definition (high, session medium)
model takes no effort (Haiku 4.5)nothing sent: model takes no effort
fork, Codex delegation, or a plugin's own $.agent.spawnuntouched

Known limits. The hook API does not say whether a subagent's effort was pinned. The router infers it, so a definition pinned to the same level as the session looks unpinned and may be routed. A pin on a different model given through --agents is not detected either. To make a pin certain, put effort: in the agent's .md file.

Codex delegations

A codex:codex-rescue spawn gets --model and --effort flags prepended to its prompt, taken from the same decision. The rescue agent passes them on to Codex. With codexFastModel set to the same model as the other rungs, only the effort varies:

DecisionCodex flags
no decision / ordinarygpt-6-sol · low
ordinary, high effortgpt-6-sol · medium
hardgpt-6-sol · high (xhigh when the effort reads xhigh)
risky > 0.7gpt-6-sol · xhigh

A prompt that already names --model or --effort is left alone: an explicit choice wins.

Seeing what it did

The effort the router sets is a parameter of each request, so Claude Code's status line and effort box never show it. The router reports its own work in three places.

Transcript lines (when logDecisions is on):

[jev-model-router] ready on typesafe (https://api.typesafe.ai/v1/systemone); routing subagent effort, main effort
[jev-model-router] jev: tier deep (0.98) · effort 2.1 → high (0.70) · risky 0.09 · 612ms
[jev-model-router] main loop → effort high: deep (confidence 0.98)
[jev-model-router] general-purpose → effort high: deep (confidence 0.98)
[jev-model-router] ra-impl-medium: effort pinned by definition (medium)

A status line, replaced as it goes: jev · deep 0.98 → high.

A decision log, one JSONL file per session in ~/.claude/jev-router/<session-id>.jsonl (no --debug needed, no prompt text):

{"ts":"2026-10-04T10:12:41.204Z","event":"subagent","agentId":"ad05ecf8…","agentType":"general-purpose",
 "tier":"deep","confidence":0.98,"effort":2.02,"effortConfidence":0.68,"risky":0.07,
 "from":{"model":"claude-sonnet-5-5","effort":"medium","source":"parent"},
 "applied":{"effort":"high"},"reason":"deep (confidence 0.98)"}

from is what the request would have used untouched. source says where its model came from: call, definition or parent. ran appears when the first request used a different model. applied is what the router changed. For cost analysis, join these records to the per-request usage in your session transcripts by timestamp.

Options

Set them under pluginConfigs.<id>.options in user settings (~/.claude/settings.json), with --settings <file>, in managed settings or through /config. The <id> depends on how the plugin was loaded: jev-model-router@skills-dir when it is auto-loaded from a skills folder, plain jev-model-router with --plugin-dir. Under the wrong key every option stays at its default, and the ready on line says no key set.

OptionDefault
typesafeApiKey—TypeSafe API key. Preferred, because it reports a calibrated confidence.
gatewayApiKey—Vercel AI Gateway key. Confidence is read from the optional distribution.
providerautoauto · typesafe · gateway · builtin
typesafeBaseUrl / typesafeModelhttps://api.typesafe.ai / jev-latest
gatewayBaseUrl / gatewayModelhttps://ai-gateway.vercel.sh/v4/ai / typesafe-ai/jev
timeoutMs800time limit for each classification. Past it the request goes out unchanged.
OptionDefault
routeMainEfforttrueeffort of the main conversation
routeSubagentEfforttrueeffort of each Claude subagent
routeCodexDelegationtrue--model / --effort for codex:codex-rescue spawns
routeMainModelfalsemodel of the main conversation (breaks the prompt cache, see above)
routeSubagentModeltruemodel of each subagent. Set false for effort-only.
OptionDefault
minUpgradeConfidence0.3confidence needed to spend more
minDowngradeConfidence0.6confidence needed to spend less
mainEffortFloor—lowest effort: low·medium·high·xhigh
mainBalancedEffortFloor—floor for a task read as ordinary engineering
mainEffortCeiling—highest effort below the deep tier
mainModelFloorbalancedlowest tier the main loop's model may go to (model routing only)
mainMinModelDowngradeConfidence0.9confidence needed for a main-loop model downgrade
mainModelMaxContextTokens80000above this context size, a main-loop model switch is held, except a return to the session's starting model
fastModel / balancedModel / deepModelhaiku / sonnet / opusthe three tiers, as aliases or full ids
OptionDefault
codexAgentTypescodex:codex-rescuecomma-separated agent types treated as a Codex delegation
codexFastModel / codexBalancedModel / codexDeepModelgpt-6-luna / gpt-6-sol / gpt-6-solthe Codex ladder. Make all three the same to route effort only.
logDecisionstruetranscript lines and the status line
decisionLogtruethe per-session JSONL decision log
decisionLogDir~/.claude/jev-routerwhere it is written

Troubleshooting

No [jev-model-router] lines at all. Check these in order:

  1. It is a headless run. claude -p and the SDK have no transcript, so every line goes to ~/.claude/debug/<session-id>.txt. The decision log is written either way.
  2. The plugin is not loaded. claude --debug should print hooks module jev-model-router@skills-dir loaded …; events: prompt.submit,turn.step,agent.spawn. A plugin in a project's .claude/skills/ is only read once that project is trusted; a user-level ~/.claude/skills/ link has no such step.
  3. Function hooks are off. The debug log says rollout flag (tengu_plugin_hooks_modules) is off. Set CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1.

ready on the built-in classifier, no key set when you did set a key: the key is under the wrong pluginConfigs id. See Options.

Effort never moves: check the reason in the decision log. kept …, wanted … means a floor or ceiling held it, or the confidence was below the bar. no decision means the classification timed out: raise timeoutMs.

Privacy and failure modes

  • What leaves the machine: with a key set, the main-loop prompt text, and for a subagent its prompt, description and agent type, all sent to the backend the key belongs to. Nothing else. With no key, nothing leaves the machine.
  • Fail-open: a timeout, a non-2xx response, a malformed body or a thrown error all leave the request exactly as Claude Code built it. The router never blocks a turn or a subagent.
  • Cost of a decision: one HTTP call per turn and per subagent spawn, bounded by timeoutMs.

Development

bun test            # pure policy tests: hooks/policy.ts
hooks/
  jev-model-router.ts   wiring: prompt.submit · turn.step · agent.spawn
  policy.ts             every decision as a pure function (route, effort ceiling, pins, Codex flags)
  hooks.json            module manifest
tests/policy.test.ts    bun:test
.claude-plugin/
  plugin.json           manifest and userConfig

All decision logic lives in policy.ts and is unit-tested. The hook wiring has no automated tests and is checked with live claude -p sessions against the decision log. .types/ and .claude-plugin/types/ are type declarations that Claude Code writes when it loads the plugin, so they are not in the repository, and tsconfig.json resolves only after a first load.

This is an early-access API: mods need Claude Code 2.1.259+, and the $ API may change between releases. The plugin is typed against Anthropic's declarations. It talks to both backends over $.http.fetch, because a mod runs without node_modules. The TypeSafe wire shape follows @typesafe-ai/sdk v0.6.0.

Credits and license

Based on the jev-model-router mod from davila7/claude-code-templates by Daniel (San) Ávila. This fork adds effort-only routing, subagent effort through turn.step, the Codex effort ladder, definition-pin detection and the measurements above.

MIT. See LICENSE.

Source 2 files
hooks/jev-model-router.ts 678 lines
1/**
2 * jev-model-router — Claude Mod (EARLY ACCESS)
3 *
4 * Picks the model each task runs on with TypeSafe's Jev, a System One
5 * decision model: unstructured state in, a typed choice with a probability
6 * distribution out.
7 *
8 * Jev is reached one of two ways, whichever key is configured: TypeSafe's
9 * own API (`typesafeApiKey`), which reports a calibrated confidence per
10 * answer, or the Vercel AI Gateway (`gatewayApiKey`), which does not. With
11 * neither, the engine's own `$.model.classify` stands in, so the mod is
12 * useful without any account.
13 *
14 * Four things it can set, each on its own switch:
15 *   agent.spawn  — the model of each subagent (on by default)
16 *   turn.step    — the reasoning effort of each Claude subagent (on by default)
17 *   turn.step    — the reasoning effort of the main loop (on by default)
18 *   turn.step    — the model of the main loop (off by default: switching
19 *                  models mid-session invalidates the prompt cache, which can
20 *                  cost more than the cheaper tier saves)
21 *
22 * Every one of them moves in both directions: a task the decision model reads
23 * as mechanical is routed down, one it reads as hard is routed up. The two
24 * mistakes do not cost the same, so they do not clear the same confidence bar
25 * (see `minUpgradeConfidence` / `minDowngradeConfidence` in policy.ts).
26 *
27 * The Agent tool has no effort parameter; a subagent's first turn.step sets it.
28 *
29 * The prompt is classified at `prompt.submit`, which runs before the turn
30 * starts, and the decision is applied at the turn's first request.
31 *
32 * Every failure path is fail-open: a classification that errors or runs past
33 * the latency budget leaves the request exactly as the engine built it.
34 *
35 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
36 * or "gatewayApiKey"). Never hardcode it in this file.
37 *
38 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (Claude Code >= 2.1.259). Typed
39 * against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
40 *
41 * Privacy: with a key set, the prompt text is sent to whichever backend the
42 * key belongs to.
43 */
44import type { EngineInterface, Register } from 'claude-code'
45import {
46  DEFAULT_BASE_URL,
47  DEFAULT_MODEL,
48  definitionDirs,
49  definitionMatches,
50  definitionModel,
51  definitionEffort,
52  pluginAgentDirs,
53  describeDecision,
54  describeSetup,
55  codexFlags,
56  describeStatus,
57  gateMainModel,
58  rankOf,
59  endpoint,
60  effortForModel,
61  reuseForStep,
62  pendingDecisions,
63  readDecision,
64  selectProvider,
65  requestBody,
66  requestHeaders,
67  requestModelId,
68  route,
69  subagentEffortRouting,
70  EFFORT_ORDER,
71  TIER_ORDER,
72  bareCommand,
73} from './policy.ts'
74import type { Decision, Effort, PolicyConfig, Provider, Tier } from './policy.ts'
75
76export const register: Register = (on, options) => {
77  const text = (key: string, fallback: string) =>
78    typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
79  const number = (key: string, fallback: number) =>
80    typeof options[key] === 'number' ? (options[key] as number) : fallback
81  const flag = (key: string, fallback: boolean) =>
82    typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
83
84  // TypeSafe's own API is preferred when both keys are set: it is the only
85  // one that reports a calibrated confidence, which the policy's threshold
86  // reads. `provider` forces one, including "builtin" to use neither.
87  const typesafeKey = text('typesafeApiKey', '')
88  const gatewayKey = text('gatewayApiKey', '')
89  const forced = text('provider', 'auto')
90  const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey)
91
92  // Each backend keeps its own URL and model, so an override written for one
93  // can never be sent to the other when `auto` picks differently than expected.
94  const apiKey = active === 'typesafe' ? typesafeKey : active === 'gateway' ? gatewayKey : ''
95  const modelId = !active
96    ? ''
97    : active === 'typesafe'
98      ? text('typesafeModel', DEFAULT_MODEL.typesafe)
99      : text('gatewayModel', DEFAULT_MODEL.gateway)
100  const url = !active
101    ? ''
102    : active === 'typesafe'
103      ? endpoint('typesafe', text('typesafeBaseUrl', DEFAULT_BASE_URL.typesafe))
104      : endpoint('gateway', text('gatewayBaseUrl', DEFAULT_BASE_URL.gateway))
105
106  // A backend named in the options but missing its key degrades to the
107  // built-in classifier, which is silent; say so once, when a hook first runs.
108  let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
109
110  const timeoutMs = number('timeoutMs', 800)
111  const routeSubagentModel = flag('routeSubagentModel', true)
112  const routeSubagentEffort = flag('routeSubagentEffort', true)
113  // A spawn of the Codex rescue agent gets `--model` / `--effort` for Codex
114  // itself from the same decision; the agent forwards them. Its own model is
115  // left alone: it only relays, so a bigger one would be money for nothing.
116  const routeCodexDelegation = flag('routeCodexDelegation', true)
117  const codexAgentTypes = text('codexAgentTypes', 'codex:codex-rescue').split(',').map((type) => type.trim())
118  const codexLadder: Record<Tier, string> = {
119    fast: text('codexFastModel', 'gpt-6-luna'),
120    balanced: text('codexBalancedModel', 'gpt-6-sol'),
121    deep: text('codexDeepModel', 'gpt-6-sol'),
122  }
123  const routeMainEffort = flag('routeMainEffort', true)
124  const routeMainModel = flag('routeMainModel', false)
125  // Above this many context tokens a main-loop model switch likely costs more
126  // than it saves (a heuristic, see gateMainModel); only a return to the
127  // starting model, a floor raise or a risk-forced switch passes.
128  const mainModelMaxContextTokens = number('mainModelMaxContextTokens', 80_000)
129  const routeMainLoop = routeMainEffort || routeMainModel
130  const switches = {
131    subagentModel: routeSubagentModel,
132    subagentEffort: routeSubagentEffort,
133    mainEffort: routeMainEffort,
134    mainModel: routeMainModel,
135  }
136  const logDecisions = flag('logDecisions', true)
137  // The debug log exists only under --debug, so an interactive session leaves
138  // no trace of what was routed. This one does: one JSONL file per session in
139  // `decisionLogDir` (default ~/.claude/jev-router), one line per decision.
140  // `$.fs` has no append, so writes are chained and the file is rewritten
141  // whole; a session is the only writer of its own file.
142  const decisionLog: DecisionLog = {
143    enabled: flag('decisionLog', true),
144    dir: text('decisionLogDir', ''),
145    queue: Promise.resolve(),
146  }
147
148  const policy: PolicyConfig = {
149    tiers: {
150      fast: text('fastModel', 'haiku'),
151      balanced: text('balancedModel', 'sonnet'),
152      deep: text('deepModel', 'opus'),
153    },
154    minUpgradeConfidence: number('minUpgradeConfidence', 0.3),
155    minDowngradeConfidence: number('minDowngradeConfidence', 0.6),
156  }
157  // The main loop answers the person directly and carries the whole session,
158  // so its model has a floor and a stricter bar to go down. Its effort has
159  // a floor and ceiling; subagent effort uses those same bounds.
160  const mainFloor = text('mainModelFloor', 'balanced')
161  const effortOption = (key: string): Effort | undefined => {
162    const value = text(key, '')
163    return (EFFORT_ORDER as readonly string[]).includes(value) ? (value as Effort) : undefined
164  }
165  const mainPolicy: PolicyConfig = {
166    ...policy,
167    routeModel: routeMainModel,
168    modelFloor: (TIER_ORDER as readonly string[]).includes(mainFloor) ? (mainFloor as Tier) : undefined,
169    minModelDowngradeConfidence: number('mainMinModelDowngradeConfidence', 0.9),
170    effortFloor: effortOption('mainEffortFloor'),
171    balancedEffortFloor: effortOption('mainBalancedEffortFloor'),
172    effortCeiling: effortOption('mainEffortCeiling'),
173  }
174  const subagentEffortPolicy: PolicyConfig = { ...mainPolicy, routeModel: false }
175  // A spawn's classification and origin wait for its first request, where
176  // the resolved model and runtime definition are finally available.
177  const pendingSpawns = new Map<string, SpawnDecision>()
178  // Null also records a miss, so unknown agents never wait a second time.
179  const subagentEffortById = new Map<string, Effort | null>()
180  let sessionEffort: string | number | undefined
181  let sessionModel: string | undefined
182  // A first step can overtake the agent.spawn continuation that gives us its id.
183  const startingSpawns = new Set<Promise<unknown>>()
184
185  // The classification waiting for the turn that reads its prompt, and what
186  // the current turn settled on. Both are single slots: main-loop turns run
187  // one at a time, so nothing accumulates over a long session. `pending`
188  // reports no decision when two prompts are waiting at once, rather than
189  // routing a turn on a decision made for a different prompt.
190  const pending = pendingDecisions()
191  // Said once, the first time a hook runs. A router that loaded and one that
192  // never loaded are otherwise told apart only by the absence of later lines,
193  // and absence is not evidence: the policy leaves most turns alone anyway.
194  let announced = false
195  let appliedTurnId: string | undefined
196  let applied: { model?: string; effort?: Effort } | null = null
197  // The main loop's model at its first routed turn, before any rewrite: the
198  // one a late switch may still return to. Kept in `$.store` by session id,
199  // since a resume or a reload starts this module over on whatever model the
200  // last turn was switched to.
201  let baseModel: string | undefined
202
203  on('prompt.submit', async ($, e, next) => {
204    // Before the routing guards: a module whose switches are all off has still
205    // loaded, and that is exactly when its silence is most misleading.
206    if (!announced) {
207      announced = true
208      if (logDecisions) {
209        $.ui.log(
210          `[jev-model-router] ${describeSetup(
211            active,
212            url,
213            switches,
214            forced === 'builtin',
215          )}`,
216        )
217      }
218    }
219    if (!routeMainLoop) return next(e)
220
221    if (!unusableReported) {
222      unusableReported = true
223      $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
224    }
225
226    // A slash command alone gives the decision model only the command's name.
227    // Its turn keeps the session's model and effort; the null put keeps a
228    // previous prompt's decision from reaching it.
229    if (bareCommand(e.text)) {
230      if (logDecisions) $.ui.log('[jev-model-router] a command with nothing after it; leaving the turn alone')
231      pending.put(null)
232      return next(e)
233    }
234
235    const startedAt = await $.clock.now()
236    let decision: Decision | null = null
237    if (active) {
238      try {
239        const response = await Promise.race([
240          $.http.fetch(url, {
241            method: 'POST',
242            headers: requestHeaders(active, apiKey, modelId),
243            body: requestBody(active, { prompt: e.text }, modelId),
244          }),
245          $.clock.sleep(timeoutMs),
246        ])
247        if (response && response.ok) decision = readDecision(response.text)
248        else if (response) $.ui.log(`[jev-model-router] ${active} responded ${response.status}`)
249        else $.ui.log(`[jev-model-router] classification passed ${timeoutMs}ms; leaving the turn alone`)
250      } catch (error) {
251        $.ui.log(`[jev-model-router] classification failed: ${String(error)}`)
252      }
253    } else {
254      // No backend: the engine's own small-model classifier answers the same
255      // question, without the confidence the policy's threshold reads.
256      try {
257        const label = await $.model.classify(e.text, TIER_ORDER)
258        if (label) {
259          decision = {
260            tier: label as Tier,
261            confidence: null,
262            risky: null,
263            effort: null,
264            effortConfidence: null,
265          }
266        }
267      } catch (error) {
268        $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
269      }
270    }
271
272    // What the decision model actually answered, whatever the policy then
273    // does with it. This is the line that proves the classification ran.
274    if (logDecisions) {
275      const ms = (await $.clock.now()) - startedAt
276      $.ui.log(`[jev-model-router] jev: ${describeDecision(decision, ms)}`)
277    }
278
279    pending.put(decision)
280    return next(e)
281  })
282
283  on('turn.step', async function* ($, e, next) {
284    if (e.agentId) {
285      if (!routeSubagentEffort) return yield* next(e)
286      let request = e
287      try {
288        if (!subagentEffortById.has(e.agentId)) {
289          if (!pendingSpawns.has(e.agentId) && startingSpawns.size) {
290            await Promise.race([Promise.allSettled([...startingSpawns]), $.clock.sleep(timeoutMs)])
291          }
292          const spawn = pendingSpawns.get(e.agentId)
293          // Consume once, including a missing spawn: later steps must be stable.
294          pendingSpawns.delete(e.agentId)
295          subagentEffortById.set(e.agentId, null)
296          if (spawn) {
297            const routing = subagentEffortRouting(spawn, { model: e.model, effort: e.effort }, subagentEffortPolicy, { model: sessionModel, effort: sessionEffort })
298            const effort = routing.effort
299            const applied = spawn.model || effort
300              ? { ...(spawn.model ? { model: spawn.model } : {}), ...(effort ? { effort } : {}) }
301              : null
302            const reason = spawn.modelReason
303              ? `model: ${spawn.modelReason}; effort: ${routing.reason}` : routing.reason
304            record($, decisionLog, {
305              event: 'subagent', agentId: e.agentId, agentType: spawn.agentType, ...spawn.decision,
306              from: { model: spawn.fromModel, effort: e.effort, source: spawn.source },
307              ...(e.model !== spawn.fromModel ? { ran: e.model } : {}),
308              applied, reason,
309            })
310            if (logDecisions) $.ui.log(
311              `[jev-model-router] ${spawn.agentType}${effort ? ` → effort ${effort}` : ''}: ${reason}`,
312            )
313            subagentEffortById.set(e.agentId, effort)
314          }
315        }
316        const effort = subagentEffortById.get(e.agentId)
317        if (effort && e.effort !== undefined) request = { ...e, effort: effortForModel(effort, e.model, subagentEffortPolicy)! }
318      } catch (error) {
319        subagentEffortById.set(e.agentId, null)
320        try {
321          $.ui.log(`[jev-model-router] subagent effort routing failed: ${String(error)}`)
322        } catch {
323          // A failed diagnostic must not block the agent's request.
324        }
325      }
326      return yield* next(request)
327    }
328    if (e.index === 0) {
329      sessionEffort = e.effort
330      sessionModel = e.model
331    }
332    if (!routeMainLoop) return yield* next(e)
333
334    // Every request after the first reuses what the turn settled on, so
335    // neither the model nor the effort changes under its own tool loop.
336    if (e.index > 0 && e.turnId === appliedTurnId) {
337      // A later request may have fallen back to a model without effort.
338      const reused = reuseForStep(applied, e, mainPolicy)
339      return yield* next(reused ? { ...e, ...reused } : e)
340    }
341
342    baseModel ??= await startingModel($, e.model)
343    const decision = pending.take()
344    const routing = route(decision, { model: e.model, effort: e.effort }, mainPolicy)
345    const change: { model?: string; effort?: Effort } = {}
346    let held: string | undefined
347    // The main loop's `model` is sent to the API as written, so an alias
348    // becomes its id here; a subagent's (agent.spawn) may stay an alias.
349    if (routeMainModel && routing.model) {
350      let contextTokens: number | undefined
351      try {
352        contextTokens = (await $.session.usage()).context.tokens
353      } catch {
354        // No reading: gate on nothing rather than block the turn.
355      }
356      // A loop already below its floor (switched there before the floor
357      // existed) is raised whatever the context: the floor is about quality.
358      const current = rankOf(e.model, policy.tiers)
359      const floor = mainPolicy.modelFloor ? TIER_ORDER.indexOf(mainPolicy.modelFloor) : 0
360      if (current !== null && current < floor) contextTokens = undefined
361      // Risk forcing the deep tier is not a cost question; the limit is.
362      if (routing.forced) contextTokens = undefined
363      const gated = gateMainModel(requestModelId(routing.model), contextTokens, mainModelMaxContextTokens, baseModel)
364      if (gated.model && gated.model !== e.model) change.model = gated.model
365      held = gated.reason
366    }
367    if (routeMainEffort && routing.effort) change.effort = routing.effort
368
369    appliedTurnId = e.turnId
370    applied = Object.keys(change).length > 0 ? change : null
371    // A row in the transcript scrolls away; this line stays on screen.
372    if (logDecisions) $.ui.status(describeStatus(decision, applied))
373    record($, decisionLog, { event: 'main', ...decision, from: { model: e.model, effort: e.effort }, applied, reason: held ? `${routing.reason}; ${held}` : routing.reason })
374
375    if (!applied) {
376      // A turn left alone is the common case, and it used to be silent, which
377      // made a working mod look like one that never loaded. Say what happened.
378      if (logDecisions) {
379        const suppressed = held ? ` (${held})` : routing.wantedModel && !routeMainModel ? ' (main-loop model routing off)' : ''
380        $.ui.log(`[jev-model-router] main loop: ${routing.reason}${suppressed}`)
381      }
382      return yield* next(e)
383    }
384    if (logDecisions) {
385      const what = [change.model, change.effort && `effort ${change.effort}`]
386        .filter(Boolean)
387        .join(', ')
388      $.ui.log(`[jev-model-router] main loop → ${what}: ${routing.reason}`)
389    }
390    return yield* next({ ...e, ...change })
391  })
392
393  on('agent.spawn', async ($, e, next) => {
394    // Before the routing guards: a module whose switches are all off has still
395    // loaded, and that is exactly when its silence is most misleading.
396    if (!announced) {
397      announced = true
398      if (logDecisions) {
399        $.ui.log(
400          `[jev-model-router] ${describeSetup(
401            active,
402            url,
403            switches,
404            forced === 'builtin',
405          )}`,
406        )
407      }
408    }
409
410    // A fork inherits its parent's model; `model` is ignored for it.
411    const codexType = codexAgentTypes.includes(e.subagentType)
412    const codexSpawn = routeCodexDelegation && codexType
413    const effortEnabled = routeSubagentEffort && !codexType
414    if ((!routeSubagentModel && !effortEnabled && !codexSpawn) || e.fork) return next(e)
415
416    if (!unusableReported) {
417      unusableReported = true
418      $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
419    }
420
421    const startedAt = await $.clock.now()
422    let decision: Decision | null = null
423    if (active) {
424      try {
425        const response = await Promise.race([
426          $.http.fetch(url, {
427            method: 'POST',
428            headers: requestHeaders(active, apiKey, modelId),
429            body: requestBody(
430              active,
431              { prompt: e.prompt, description: e.description, agentType: e.subagentType },
432              modelId,
433            ),
434          }),
435          $.clock.sleep(timeoutMs),
436        ])
437        if (response && response.ok) decision = readDecision(response.text)
438        else if (response) $.ui.log(`[jev-model-router] ${active} responded ${response.status}`)
439        else $.ui.log(`[jev-model-router] classification passed ${timeoutMs}ms; leaving the subagent alone`)
440      } catch (error) {
441        $.ui.log(`[jev-model-router] classification failed: ${String(error)}`)
442      }
443    } else {
444      try {
445        const label = await $.model.classify(e.prompt, TIER_ORDER)
446        if (label) {
447          decision = {
448            tier: label as Tier,
449            confidence: null,
450            risky: null,
451            effort: null,
452            effortConfidence: null,
453          }
454        }
455      } catch (error) {
456        $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
457      }
458    }
459
460    if (logDecisions) {
461      const ms = (await $.clock.now()) - startedAt
462      $.ui.log(`[jev-model-router] jev (${e.subagentType}): ${describeDecision(decision, ms)}`)
463    }
464
465    // The caller's model wins. Without one, the definition may pin a model;
466    // otherwise the agent inherits its parent's model. Keep that origin for
467    // the first-step record even if agent.spawn rewrites the request.
468    const definition = e.model && !effortEnabled ? null : await definedAgent($, e.subagentType)
469    const defined = definition?.model ?? null
470    const current = e.model ?? defined ?? e.parentModel
471    const source = e.model ? 'call' : defined ? 'definition' : 'parent'
472    const routed = codexSpawn || !routeSubagentModel ? null : route(decision, { model: current }, policy)
473    const model = routed?.model ?? null
474    const codex = codexSpawn ? codexFlags(decision, e.prompt, codexLadder) : null
475    const reason = codexSpawn
476      ? codex
477        ? `codex ${codex.model}/${codex.effort} (${decision?.tier ?? 'no decision, default'})`
478        : 'codex flags already in the prompt'
479      : (routed?.reason ?? 'no decision')
480    const applied = model || codex ? { ...(model ? { model } : {}), ...(codex ? { codex } : {}) } : null
481    if (!effortEnabled) {
482      record($, decisionLog, { event: 'subagent', agentType: e.subagentType, ...decision, from: { model: current, source }, applied, reason })
483    }
484    if (!applied) {
485      if (logDecisions && (!effortEnabled || routed)) $.ui.log(`[jev-model-router] ${e.subagentType}: ${reason}`)
486    } else if (logDecisions) {
487      $.ui.log(`[jev-model-router] ${e.subagentType} → ${[model, codex && `codex ${codex.model}/${codex.effort}`].filter(Boolean).join(', ')}: ${reason}`)
488    }
489    const changed = {
490      ...e,
491      ...(model ? { model } : {}),
492      ...(codex ? { prompt: `--model ${codex.model} --effort ${codex.effort} ${e.prompt}` } : {}),
493    }
494    if (!effortEnabled) return next(applied ? changed : e)
495    const startedPromise = next(changed).then((started) => {
496      // A timed-out first step has already cached a miss; do not leave a
497      // pending decision that no later step will consume.
498      if (started.agentId && !subagentEffortById.has(started.agentId)) pendingSpawns.set(started.agentId, {
499        decision, pinned: definition?.effort ?? null, agentType: e.subagentType,
500        source, fromModel: current, model, modelReason: routed ? reason : null,
501      })
502      return started
503    })
504    startingSpawns.add(startedPromise)
505    try {
506      return await startedPromise
507    } finally {
508      startingSpawns.delete(startedPromise)
509    }
510  })
511}
512
513// The debug log exists only under --debug, so an interactive session leaves
514// no trace of what was routed. This one does: one JSONL file per session in
515// `decisionLogDir` (default ~/.claude/jev-router), one line per decision.
516// `$.fs` has no append, so writes are chained and the file is rewritten
517// whole; a session is the only writer of its own file.
518type DecisionLog = { enabled: boolean; dir: string; queue: Promise<void>; path?: string }
519type SpawnDecision = {
520  decision: Decision | null
521  pinned: Effort | null
522  agentType: string
523  source: 'call' | 'definition' | 'parent'
524  fromModel: string
525  model: string | null
526  modelReason: string | null
527}
528type Engine = EngineInterface
529
530// An agent type's definition pins these values, or leaves them unset. Answers
531// are cached for a minute per type, so a burst of spawns scans folders once.
532const DEFINITION_TTL_MS = 60_000
533type AgentDefinition = { model: string | null; effort: Effort | null }
534const definitionCache = new Map<string, AgentDefinition & { at: number }>()
535
536async function definedAgent($: Engine, subagentType: string): Promise<AgentDefinition> {
537  const now = await $.clock.now()
538  const cached = definitionCache.get(subagentType)
539  if (cached && now - cached.at < DEFINITION_TTL_MS) return cached
540  const definition: AgentDefinition = { model: null, effort: null }
541  try {
542    const home = (await $.env.get('HOME')) ?? ''
543    const pluginDirs: string[] = []
544    const colon = subagentType.indexOf(':')
545    if (colon > 0) {
546      const registry = `${home}/.claude/plugins/installed_plugins.json`
547      if (await $.fs.exists(registry)) {
548        const plugins = JSON.parse(await $.fs.read(registry)).plugins ?? {}
549        const prefix = `${subagentType.slice(0, colon)}@`
550        const installs = Object.entries(plugins)
551          .filter(([key]) => key.startsWith(prefix))
552          .flatMap(([, entries]) => (entries as { installPath?: string }[]).map((entry) => entry.installPath ?? ''))
553          .filter(Boolean)
554        for (const install of installs) {
555          let manifest: unknown = {}
556          try {
557            const manifestPath = `${install}/.claude-plugin/plugin.json`
558            if (await $.fs.exists(manifestPath)) manifest = JSON.parse(await $.fs.read(manifestPath))
559          } catch {
560            // An unreadable manifest still leaves the default agents/ folder.
561          }
562          pluginDirs.push(...pluginAgentDirs(install, manifest))
563        }
564      }
565    }
566    const { agent, dirs } = definitionDirs(subagentType, await $.session.root(), home, pluginDirs)
567    const budget = { reads: DEFINITION_READ_BUDGET }
568    for (const dir of dirs) {
569      const found = await findDefinition($, dir, agent, budget)
570      if (found !== undefined) {
571        definition.model = definitionModel(found)
572        definition.effort = definitionEffort(found)
573        break
574      }
575    }
576  } catch (error) {
577    $.ui.log(`[jev-model-router] reading the ${subagentType} definition failed: ${String(error)}`)
578  }
579  definitionCache.set(subagentType, { ...definition, at: now })
580  return definition
581}
582
583// The text of the definition in `dir` (or up to two folders below it) that
584// defines `agent`, else undefined. In each folder the file named after the
585// agent is tried first, the common case. A file or folder that cannot be read
586// is skipped, not fatal. `budget` caps the files read across one lookup.
587const DEFINITION_READ_BUDGET = 300
588
589async function findDefinition(
590  $: Engine,
591  dir: string,
592  agent: string,
593  budget: { reads: number },
594  depth = 0,
595): Promise<string | undefined> {
596  const read = async (path: string): Promise<string | undefined> => {
597    if (budget.reads <= 0) return undefined
598    budget.reads--
599    try {
600      return await $.fs.read(path)
601    } catch {
602      return undefined
603    }
604  }
605  let entries: Awaited<ReturnType<Engine['fs']['list']>>
606  try {
607    if (!(await $.fs.exists(dir))) return undefined
608    entries = await $.fs.list(dir)
609  } catch {
610    return undefined
611  }
612  const guess = `${agent}.md`
613  const files = entries.filter((entry) => entry.kind === 'file' && entry.name.endsWith('.md'))
614  for (const entry of [...files.filter((f) => f.name === guess), ...files.filter((f) => f.name !== guess)]) {
615    const text = await read(`${dir}/${entry.name}`)
616    if (text !== undefined && definitionMatches(text, entry.name, agent)) return text
617  }
618  if (depth >= 2) return undefined
619  for (const entry of entries) {
620    if (entry.kind !== 'dir') continue
621    const found = await findDefinition($, `${dir}/${entry.name}`, agent, budget, depth + 1)
622    if (found !== undefined) return found
623  }
624  return undefined
625}
626
627// The session's starting model from `$.store`, recorded on first sight, one
628// key per session so two sessions starting at once cannot overwrite each
629// other. Entries unused for 30 days are dropped; each is re-read just before
630// deletion, so only a session idle that long and resuming in the same instant
631// could lose its entry. A map under the old shared key is still read, never
632// written. Any failure answers `current`.
633const STARTING_MODEL_PREFIX = 'start:'
634const STARTING_MODEL_TTL_MS = 30 * 24 * 60 * 60 * 1000
635const LEGACY_STARTING_MODELS = 'startingModels'
636type StartingModel = { model: string; at: number }
637
638function readStarting(value: unknown): StartingModel | null {
639  if (typeof value === 'string') return { model: value, at: 0 }
640  const entry = value as Partial<StartingModel> | null
641  return entry && typeof entry.model === 'string' ? { model: entry.model, at: Number(entry.at) || 0 } : null
642}
643
644async function startingModel($: Engine, current: string): Promise<string> {
645  try {
646    const now = await $.clock.now()
647    const sessionId = await $.session.id()
648    const key = `${STARTING_MODEL_PREFIX}${sessionId}`
649    const legacy = ((await $.store.get(LEGACY_STARTING_MODELS)) ?? {}) as Record<string, string>
650    const start = readStarting(await $.store.get(key))?.model ?? legacy[sessionId] ?? current
651    await $.store.set(key, { model: start, at: now } satisfies StartingModel)
652    for (const other of await $.store.keys()) {
653      if (other === key || !other.startsWith(STARTING_MODEL_PREFIX)) continue
654      const entry = readStarting(await $.store.get(other))
655      if (entry && entry.at > 0 && now - entry.at > STARTING_MODEL_TTL_MS) await $.store.delete(other)
656    }
657    return start
658  } catch (error) {
659    $.ui.log(`[jev-model-router] starting model not stored: ${String(error)}`)
660  }
661  return current
662}
663
664function record($: Engine, log: DecisionLog, entry: Record<string, unknown>): void {
665  if (!log.enabled) return
666  const ts = new Date().toISOString()
667  log.queue = log.queue
668    .then(async () => {
669      if (!log.path) {
670        const dir = log.dir || `${(await $.env.get('HOME')) ?? '.'}/.claude/jev-router`
671        log.path = `${dir}/${await $.session.id()}.jsonl`
672      }
673      const previous = (await $.fs.exists(log.path)) ? await $.fs.read(log.path) : ''
674      await $.fs.write(log.path, `${previous}${JSON.stringify({ ts, ...entry })}\n`)
675    })
676    .catch((error) => $.ui.log(`[jev-model-router] decision log failed: ${String(error)}`))
677}
678
hooks/policy.ts 783 lines
1/**
2 * jev-model-router — pure decision logic.
3 *
4 * No `$` and no I/O here: this module only builds the request the decision
5 * API takes, reads its answer, and turns that answer into a model id. The
6 * hooks module does every call on `$` at its own call site.
7 *
8 * Two backends speak to the same model with different wire shapes:
9 *
10 *   typesafe  POST https://api.typesafe.ai/v1/systemone
11 *             `{ model, state, questions }`; a yes/no question is a `noul`
12 *             and every answer carries its own `confidence`.
13 *   gateway   POST https://ai-gateway.vercel.sh/v4/ai/evaluation-model
14 *             `{ state, questions }` with the model in a header; a yes/no
15 *             question is a `boolean`, and there is no `confidence` field —
16 *             it has to be derived from an optional distribution.
17 *
18 * The Gateway shape is not documented publicly; it was read from
19 * @ai-sdk/gateway and @ai-sdk/provider.
20 */
21
22export type Provider = 'typesafe' | 'gateway'
23
24export type Tier = 'fast' | 'balanced' | 'deep'
25
26export interface Tiers {
27  fast: string
28  balanced: string
29  deep: string
30}
31
32export interface Decision {
33  tier: Tier
34  /** Confidence in the tier, or null when the backend reported none. */
35  confidence: number | null
36  /** P(true) that carrying the task out would itself be costly or final. */
37  risky: number | null
38  /** 0..3 along the effort rubric, or null when absent. */
39  effort: number | null
40  /** Confidence in the effort, or null when the backend reported none. */
41  effortConfidence: number | null
42}
43
44/** The reasoning levels a turn can ask for, cheapest first. */
45export const EFFORT_ORDER = ['low', 'medium', 'high', 'xhigh'] as const
46
47export type RoutedEffort = (typeof EFFORT_ORDER)[number]
48export type Effort = RoutedEffort | 'max'
49
50export const TIER_ORDER: readonly Tier[] = ['fast', 'balanced', 'deep']
51
52/**
53 * How each tier is described to the decision model. Deliberately about the
54 * shape of the work, not about model names: the model never sees an id.
55 */
56const TIER_CRITERIA: Record<Tier, string> = {
57  fast: 'Mechanical and local: read or summarise a file, run one command, rename a symbol, answer something already in context.',
58  balanced:
59    'Ordinary engineering: implement a well-specified change across a few files, write tests, fix a clearly described bug, review a small diff.',
60  deep: 'Hard or high-stakes: architecture and design, debugging a failure whose cause is unknown, security, data migrations, concurrency, anything touching production or money.',
61}
62
63const EFFORT_RUBRIC = ['almost none', 'some', 'a lot', 'as much as possible'] as const
64
65export const DEFAULT_BASE_URL: Record<Provider, string> = {
66  typesafe: 'https://api.typesafe.ai',
67  gateway: 'https://ai-gateway.vercel.sh/v4/ai',
68}
69
70/**
71 * The Gateway's own protocol version, sent as `ai-gateway-protocol-version`.
72 * Tracks the `AI_GATEWAY_PROTOCOL_VERSION` of `@ai-sdk/gateway` (4.0.87).
73 */
74const AI_GATEWAY_PROTOCOL_VERSION = '0.0.1'
75
76export const DEFAULT_MODEL: Record<Provider, string> = {
77  typesafe: 'jev-latest',
78  gateway: 'typesafe-ai/jev',
79}
80
81/**
82 * Which backend a configuration asks for, or null for the built-in
83 * classifier. `auto` prefers TypeSafe, since it is the only one that reports
84 * a calibrated confidence; a forced backend whose key is missing resolves to
85 * null rather than falling through to the other one's key.
86 */
87export function selectProvider(
88  forced: string,
89  typesafeKey: string,
90  gatewayKey: string,
91): Provider | null {
92  if (forced === 'builtin') return null
93  if (forced === 'typesafe') return typesafeKey ? 'typesafe' : null
94  if (forced === 'gateway') return gatewayKey ? 'gateway' : null
95  if (typesafeKey) return 'typesafe'
96  if (gatewayKey) return 'gateway'
97  return null
98}
99
100/** The full endpoint a backend posts to. */
101export function endpoint(provider: Provider, baseUrl: string): string {
102  const root = baseUrl.replace(/\/+$/, '')
103  return provider === 'typesafe' ? `${root}/v1/systemone` : `${root}/evaluation-model`
104}
105
106/** The `questions` map, in the shape the backend's schema names. */
107export function questions(provider: Provider): Record<string, unknown> {
108  return {
109    tier: {
110      type: 'choice',
111      instructions: 'Which is the cheapest tier that can complete this coding task well?',
112      criteria: TIER_CRITERIA,
113    },
114    effort: {
115      type: 'score',
116      instructions: 'How much step-by-step reasoning does this task need?',
117      criteria: EFFORT_RUBRIC,
118    },
119    risky: {
120      // The same question under two names: `noul` on TypeSafe's own API,
121      // `boolean` in the AI SDK's evaluation schema.
122      type: provider === 'typesafe' ? 'noul' : 'boolean',
123      // Asked about the act, not the subject. The first wording ("the task
124      // touches production, money, credentials") scored 0.96 on "add a
125      // refund endpoint that calls Stripe" — ordinary code that happens to be
126      // about money — and would have escalated it past a 0.98-confidence
127      // answer of the balanced tier.
128      instructions:
129        'Carrying out this task would itself change production, move real money, or alter data that cannot be restored. Writing or testing code that deals with such things, without running it against the real system, does not count.',
130    },
131  }
132}
133
134/** The request body. The Gateway carries the model in a header instead. */
135export function requestBody(
136  provider: Provider,
137  state: Record<string, unknown>,
138  model: string,
139): string {
140  const body =
141    provider === 'typesafe'
142      ? { model, state, questions: questions(provider) }
143      : { state, questions: questions(provider) }
144  return JSON.stringify(body)
145}
146
147/** The request headers. */
148export function requestHeaders(
149  provider: Provider,
150  apiKey: string,
151  model: string,
152): Record<string, string> {
153  const common = { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` }
154  if (provider === 'typesafe') return common
155  return {
156    ...common,
157    'ai-gateway-auth-method': 'api-key',
158    'ai-model-id': model,
159    // The Gateway rejects any request that does not name the protocol it
160    // speaks: 400 "Unsupported gateway protocol version". Every other header
161    // here is accepted without it, so the omission fails the whole backend.
162    'ai-gateway-protocol-version': AI_GATEWAY_PROTOCOL_VERSION,
163    'ai-evaluation-model-specification-version': '4',
164  }
165}
166
167function isTier(value: unknown): value is Tier {
168  return value === 'fast' || value === 'balanced' || value === 'deep'
169}
170
171/**
172 * Reads a response from either backend.
173 *
174 * TypeSafe's own API reports a `confidence` per answer and a `noul` number
175 * for a yes/no question. The Gateway reports neither: confidence has to come
176 * from the highest probability of a distribution that is itself optional, and
177 * a yes/no answer arrives as `probability`. Both are handled, and a missing
178 * confidence reads as null rather than as a number the policy would trust.
179 */
180export function readDecision(responseText: string): Decision | null {
181  let parsed: unknown
182  try {
183    parsed = JSON.parse(responseText)
184  } catch {
185    return null
186  }
187  const answers = (parsed as { answers?: Record<string, Record<string, unknown>> }).answers
188  if (!answers) return null
189
190  const tierAnswer = answers.tier
191  if (!tierAnswer || !isTier(tierAnswer.choice)) return null
192
193  const effortAnswer = answers.effort
194  const riskyAnswer = answers.risky
195  const risky =
196    typeof riskyAnswer?.noul === 'number'
197      ? riskyAnswer.noul
198      : typeof riskyAnswer?.probability === 'number'
199        ? riskyAnswer.probability
200        : null
201
202  return {
203    tier: tierAnswer.choice,
204    confidence: confidenceOf(tierAnswer),
205    effort: typeof effortAnswer?.score === 'number' ? effortAnswer.score : null,
206    effortConfidence: effortAnswer ? confidenceOf(effortAnswer) : null,
207    risky,
208  }
209}
210
211/**
212 * How sure an answer is. TypeSafe reports it; the Gateway does not, so there
213 * it is the highest probability of a distribution that is itself optional.
214 */
215function confidenceOf(answer: Record<string, unknown>): number | null {
216  if (typeof answer.confidence === 'number') return answer.confidence
217  const probabilities = answer.probabilities as Record<string, number> | undefined
218  const values = probabilities ? Object.values(probabilities) : []
219  return values.length > 0 ? Math.max(...values) : null
220}
221
222/** The rubric score (0..3) as a reasoning level. */
223export function effortLevel(score: number): RoutedEffort {
224  const index = Math.min(EFFORT_ORDER.length - 1, Math.max(0, Math.round(score)))
225  return EFFORT_ORDER[index]!
226}
227
228/**
229 * Where a reasoning level sits on the ladder, or null when its place cannot
230 * be known. `max` is above every rung the rubric can produce, so it ranks
231 * above them without joining EFFORT_ORDER, which is also the set of values
232 * this router is allowed to ask for.
233 */
234export function effortRank(effort: string | number | undefined): number | null {
235  if (typeof effort !== 'string') return null
236  if (effort === 'max') return EFFORT_ORDER.length
237  const index = EFFORT_ORDER.findIndex((level) => level === effort)
238  return index === -1 ? null : index
239}
240
241/**
242 * Where a model id sits on the tier ladder, by matching it against the
243 * configured tier names first and then the family words. Null when it matches
244 * none, in which case the change is treated as an upgrade rather than guessed
245 * at: an unrecognised id gets the gentler threshold, never the strict one.
246 */
247export function rankOf(model: string, tiers: Tiers): number | null {
248  const lowered = model.toLowerCase()
249  for (let index = 0; index < TIER_ORDER.length; index++) {
250    const tier = TIER_ORDER[index] as Tier
251    const configured = tiers[tier].toLowerCase()
252    if (configured && lowered.includes(configured)) return index
253  }
254  if (lowered.includes('haiku')) return 0
255  if (lowered.includes('sonnet')) return 1
256  if (lowered.includes('opus')) return 2
257  return null
258}
259
260/** Recheck the effort ceiling against the model on this request. */
261export function effortForModel(effort: Effort | undefined, model: string, config: PolicyConfig): Effort | null {
262  if (effort === undefined) return null
263  const ceiling = config.effortCeiling
264  if (ceiling && rankOf(model, config.tiers) !== TIER_ORDER.indexOf('deep') && effortRank(effort)! > effortRank(ceiling)!) {
265    return ceiling
266  }
267  return effort
268}
269
270/** Reuse a turn's choice against the model the next step will actually send. */
271export function reuseForStep(
272  applied: { model?: string; effort?: Effort } | null,
273  step: { model: string; effort?: string | number },
274  config: PolicyConfig,
275): { model?: string; effort?: Effort } | null {
276  if (!applied) return null
277  const reused = step.effort === undefined ? (applied.model ? { model: applied.model } : null) : applied
278  if (!reused) return null
279  const effort = reused.effort ? effortForModel(reused.effort, reused.model ?? step.model, config) : null
280  return effort ? { ...reused, effort } : reused
281}
282
283/**
284 * The full id a family alias names on the main loop.
285 *
286 * `agent.spawn` takes an alias (`haiku`) the way the Agent tool does, but
287 * `turn.step`'s `model` is the id the engine already resolved for the request
288 * and goes to the API as written: an alias there is refused ("There's an
289 * issue with the selected model (haiku)"). So the tiers stay aliases in the
290 * options, and only a main-loop rewrite resolves them, here.
291 *
292 * A stale entry here is a silent downgrade: `opus` sat on `claude-opus-5`
293 * (Opus 5, not 5.5) for the whole time Opus 5.5 was current, so every
294 * `routeMainModel` upgrade to `deep` landed one point release behind. Bump
295 * this table whenever Anthropic ships a new point release in a family the
296 * tiers use — checked each time, not guessed: `claude --restricted -p
297 * --model <alias> --output-format json "OK"` names the id the engine itself
298 * resolves the alias to in `modelUsage` (`--restricted` keeps this router
299 * from rewriting the probe).
300 * Last checked 2026-09-28, against Opus 5.5 / Sonnet 5.5 / Haiku 4.5.
301 */
302const ALIAS_IDS: Record<string, string> = {
303  haiku: 'claude-haiku-4-5-20251001',
304  sonnet: 'claude-sonnet-5-5',
305  opus: 'claude-opus-5-5',
306}
307
308/**
309 * What to write into `turn.step`'s `model`: a full id as given, or the id
310 * behind a family alias. Anything else is returned unchanged for the engine
311 * to judge.
312 */
313export function requestModelId(model: string): string {
314  return ALIAS_IDS[model.trim().toLowerCase()] ?? model
315}
316
317export interface PolicyConfig {
318  tiers: Tiers
319  /** Whether the policy may change the request's model. Defaults to true. */
320  routeModel?: boolean
321  /**
322   * How sure the decision must be to spend more (a bigger model, more
323   * reasoning). Being wrong here costs money, so the bar is low.
324   */
325  minUpgradeConfidence: number
326  /**
327   * How sure it must be to spend less. Being wrong here means a task handled
328   * by too small a model or too little thought, so the bar is high.
329   */
330  minDowngradeConfidence: number
331  /**
332   * The lowest tier the model may be routed to; absent, any. Effort is not
333   * floored. The main loop sets `balanced`: a turn's prompt can read as
334   * mechanical ("проверь") while the work around it is not, and a haiku
335   * answer there was measured wrong (2026-09-26).
336   */
337  modelFloor?: Tier
338  /**
339   * The lowest reasoning level a turn may be routed to; absent, any. Unlike
340   * a moved level this is not a guess about the prompt, so a low reading is
341   * lifted to it without a confidence check. A downgrade stops here.
342   */
343  effortFloor?: Effort
344  /** A higher floor for a balanced reading; the larger of the two applies. */
345  balancedEffortFloor?: Effort
346  /**
347   * The highest level while the loop runs below the deep tier, so xhigh and
348   * max belong to opus. Like the floor it is policy: a level above it comes
349   * down without a confidence check.
350   */
351  effortCeiling?: Effort
352  /**
353   * The downgrade bar for the model alone, when it should differ from
354   * `minDowngradeConfidence`, which then still governs effort.
355   */
356  minModelDowngradeConfidence?: number
357}
358
359export interface Routing {
360  /** The model to run on, or null to leave the request as it is. */
361  model: string | null
362  /** The classified model even when this route cannot change it. */
363  wantedModel: string | null
364  /** The reasoning level to ask for, or null to leave it as it is. */
365  effort: Effort | null
366  /** Why, for the log line. */
367  reason: string
368  /**
369   * True when risk forced the deep tier past the thresholds. A caller's own
370   * gate (the main loop's context limit) must let such a change through too:
371   * risk is the one case that is not a cost question.
372   */
373  forced: boolean
374}
375
376const NOTHING: Routing = { model: null, wantedModel: null, effort: null, reason: 'no decision', forced: false }
377
378/**
379 * Whether a change of rank passes its threshold. Both directions are allowed;
380 * they just do not have to clear the same bar, because the two mistakes do not
381 * cost the same. A move whose direction cannot be told (an unrecognised
382 * current value) is treated as an upgrade.
383 */
384function allowed(
385  wanted: number,
386  current: number | null,
387  confidence: number | null,
388  config: PolicyConfig,
389  downgradeBar: number = config.minDowngradeConfidence,
390): boolean {
391  if (current !== null && wanted === current) return false
392  const isDowngrade = current !== null && wanted < current
393  const bar = isDowngrade ? downgradeBar : config.minUpgradeConfidence
394  // A backend that reports no confidence (the Gateway without a distribution,
395  // or the built-in classifier) clears the upgrade bar but never the
396  // downgrade one: spending less on an unmeasured hunch is the bad trade.
397  if (confidence === null) return !isDowngrade
398  return confidence >= bar
399}
400
401/**
402 * Turns a decision into a model and a reasoning level, either of which may be
403 * null to leave the request as it is. Both can move in either direction.
404 */
405export function route(
406  decision: Decision | null,
407  current: { model: string; effort?: string | number },
408  config: PolicyConfig,
409): Routing {
410  if (!decision) return NOTHING
411
412  let tier = decision.tier
413  let effortScore = decision.effort
414  let forced = false
415
416  // Carrying out something final is never worth the saving: take the deep
417  // tier and real reasoning, whatever the cheaper answer said, and skip the
418  // thresholds — this is the one case that is not a confidence question.
419  if (decision.risky !== null && decision.risky > 0.7) {
420    tier = 'deep'
421    effortScore = Math.max(effortScore ?? 0, 2)
422    forced = true
423  }
424
425  const floor = config.modelFloor ? TIER_ORDER.indexOf(config.modelFloor) : 0
426  const wantedTier = Math.max(TIER_ORDER.indexOf(tier), floor)
427  const currentTier = rankOf(current.model, config.tiers)
428  const wantedModel = config.tiers[TIER_ORDER[wantedTier] as Tier]
429  const modelBar = config.minModelDowngradeConfidence ?? config.minDowngradeConfidence
430
431  const model =
432    config.routeModel !== false &&
433    wantedModel &&
434    wantedModel !== current.model &&
435    (forced || allowed(wantedTier, currentTier, decision.confidence, config, modelBar))
436      ? wantedModel
437      : null
438
439  let effort: Effort | null = null
440  // An absent effort means this model cannot accept the API parameter.
441  if (effortScore !== null && current.effort !== undefined) {
442    const currentRank = effortRank(current.effort)
443    let wantedRank = EFFORT_ORDER.indexOf(effortLevel(effortScore))
444    const rankOfLevel = (level: Effort | undefined, none: number) => (level ? effortRank(level) ?? none : none)
445    const floorRank = Math.max(
446      rankOfLevel(config.effortFloor, 0),
447      tier === 'balanced' ? rankOfLevel(config.balancedEffortFloor, 0) : 0,
448    )
449    // An unknown model has no proof that it belongs to the deep tier.
450    const runsTier = model ? wantedTier : (currentTier ?? 0)
451    const ceilingRank = runsTier < TIER_ORDER.indexOf('deep') ? rankOfLevel(config.effortCeiling, Infinity) : Infinity
452    const lifted = wantedRank < floorRank && currentRank !== null && currentRank < floorRank
453    const capped = currentRank !== null && currentRank > ceilingRank
454    wantedRank = Math.min(Math.max(wantedRank, floorRank), ceilingRank)
455
456    // Risk raises the floor; it must never lower one. Forcing only skips the
457    // thresholds, so without this clamp a task already at `xhigh` or `max`
458    // and rated mechanically simple would be pulled down to `high` with no
459    // confidence check at all — the opposite of what the rule is for.
460    if (forced && currentRank !== null) wantedRank = Math.max(wantedRank, currentRank)
461
462    // A numeric effort is the caller's own scale, not this ladder; leave it.
463    const comparable = typeof current.effort !== 'number'
464    const wanted = EFFORT_ORDER[Math.min(EFFORT_ORDER.length - 1, wantedRank)]!
465    if (
466      comparable &&
467      wantedRank !== currentRank &&
468      (forced || lifted || capped || allowed(wantedRank, currentRank, decision.effortConfidence, config))
469    ) {
470      effort = wanted
471    }
472  }
473
474  const said = decision.confidence === null ? 'confidence n/d' : `confidence ${decision.confidence.toFixed(2)}`
475
476  if (!model && !effort) {
477    // Naming what it wanted and what it kept is the whole point of this line.
478    // Without it, a mod that classified and decided to leave the request alone
479    // is indistinguishable from one that never loaded.
480    const wantedEffort = effortScore === null ? null : effortLevel(effortScore)
481    const kept = `${current.model}${current.effort === undefined ? '' : `/${current.effort}`}`
482    const wanted = `${wantedModel}${wantedEffort ? `/${wantedEffort}` : ''}`
483    const floored = wantedTier > TIER_ORDER.indexOf(tier) ? `, ${tier} floored to ${TIER_ORDER[wantedTier]}` : ''
484    return { model: null, wantedModel, effort: null, reason: `kept ${kept}, wanted ${wanted}${floored} (${said})`, forced }
485  }
486
487  const label = wantedTier > TIER_ORDER.indexOf(tier) ? `${tier}, floored to ${TIER_ORDER[wantedTier]}` : tier
488  return { model, wantedModel, effort, reason: forced ? (config.routeModel === false ? 'effort forced by risk' : `${label}, forced by risk`) : `${label} (${said})`, forced }
489}
490
491/** Match model ids despite a context-window suffix such as `[1m]`. */
492export function sameModel(model: string, sessionModel?: string): boolean {
493  return sessionModel !== undefined && model.replace(/\[[^\]]+\]$/, '').toLowerCase() === sessionModel.replace(/\[[^\]]+\]$/, '').toLowerCase()
494}
495
496/** One Claude subagent's effort decision, including detectable definition pins. */
497export function subagentEffortRouting(
498  spawn: { decision: Decision | null; pinned: Effort | null },
499  current: { model: string; effort?: string | number },
500  config: PolicyConfig,
501  session?: { model?: string; effort?: string | number },
502): Pick<Routing, 'effort' | 'reason'> {
503  if (spawn.pinned) return { effort: null, reason: `effort pinned by definition (${spawn.pinned})` }
504  // A same-model pin equal to the session level is indistinguishable; pins on
505  // other models from --agents or AgentSpec are not detected by this heuristic.
506  if (sameModel(current.model, session?.model) && current.effort !== undefined && session?.effort !== undefined && current.effort !== session.effort) {
507    return { effort: null, reason: `effort set by agent definition (${current.effort}, session ${session.effort})` }
508  }
509  if (current.effort === undefined) return { effort: null, reason: 'model takes no effort' }
510  const { effort, reason } = route(spawn.decision, current, { ...config, routeModel: false })
511  return { effort, reason }
512}
513
514/**
515 * Whether a prompt is a slash command and nothing else (`/simplify`). The
516 * decision model sees only the name, never the skill or command it runs:
517 * measured on TypeSafe (three calls each), `/simplify`, `/run` and `/github`
518 * came back fast with effort 0.1 to 0.5, low effort for a multi-step skill,
519 * while `/code-review` and `/security-review` came back balanced and deep.
520 * With text after the name there is a task to read, and it is classified.
521 * A one-segment path alone (`/etc`) matches too; nobody sends one as a task.
522 */
523export function bareCommand(text: string): boolean {
524  return /^\/[^\s/]+$/.test(text.trim())
525}
526
527/**
528 * Holds a prompt's classification until the turn that reads that prompt
529 * starts.
530 *
531 * Nothing ties a decision to the turn it belongs to. Prompts can be queued
532 * while the model is busy, a peer session's message can be delivered inside a
533 * running turn, and `prompt.submit` carries no turn id at all while the
534 * session is idle. So when more than one prompt is waiting, `take` reports
535 * none: running a turn on another prompt's decision is a worse outcome than
536 * not routing it, and not routing is what every other failure path here does.
537 */
538export function pendingDecisions(): {
539  put(decision: Decision | null): void
540  take(): Decision | null
541} {
542  let held: Decision | null = null
543  let waiting = 0
544
545  return {
546    put(decision) {
547      waiting += 1
548      // Past the first, which prompt a turn will read is unknowable, so the
549      // slot is emptied instead of holding a decision that may not fit.
550      held = waiting === 1 ? decision : null
551    },
552    take() {
553      const decision = waiting === 1 ? held : null
554      held = null
555      waiting = 0
556      return decision
557    },
558  }
559}
560
561/** A number for the log, or `n/d` when the backend reported none. */
562function reported(value: number | null): string {
563  return value === null ? 'n/d' : value.toFixed(2)
564}
565
566/**
567 * The one-time line that says the router is alive, which backend answers it,
568 * and which routing switches are on.
569 *
570 * Without this, a router that loaded and a router that never loaded are told
571 * apart only by the absence of later lines, which is not evidence of anything.
572 */
573export function describeSetup(
574  provider: Provider | null,
575  url: string,
576  switches: { subagentModel: boolean; subagentEffort: boolean; mainEffort: boolean; mainModel: boolean },
577  // `provider: "builtin"` is a choice, not a missing key. Reporting it as a
578  // credential problem sends someone hunting for a key they meant to omit.
579  builtinByChoice = false,
580): string {
581  const backend = provider
582    ? `${provider} (${url})`
583    : builtinByChoice
584      ? 'the built-in classifier, by choice'
585      : 'the built-in classifier, no key set'
586  const on = [
587    switches.subagentModel && 'subagent model',
588    switches.subagentEffort && 'subagent effort',
589    switches.mainEffort && 'main effort',
590    switches.mainModel && 'main model',
591  ].filter(Boolean)
592  return `ready on ${backend}; routing ${on.length > 0 ? on.join(', ') : 'nothing, every switch is off'}`
593}
594
595/**
596 * What the decision model answered, before any policy touches it: the raw
597 * tier, effort and risk with their confidences, and how long it took.
598 *
599 * This is the line that shows the classification happened at all, separately
600 * from whether the policy then decided to act on it.
601 */
602export function describeDecision(decision: Decision | null, ms: number | null): string {
603  const took = ms === null ? '' : ` · ${Math.round(ms)}ms`
604  if (!decision) return `no answer${took}`
605
606  const parts = [`tier ${decision.tier} (${reported(decision.confidence)})`]
607  if (decision.effort !== null) {
608    parts.push(
609      `effort ${decision.effort.toFixed(1)} → ${effortLevel(decision.effort)} (${reported(decision.effortConfidence)})`,
610    )
611  }
612  if (decision.risky !== null) parts.push(`risky ${reported(decision.risky)}`)
613  return parts.join(' · ') + took
614}
615
616/**
617 * The persistent status line: the last thing the router did, short enough to
618 * sit on screen beside the engine's own notices.
619 */
620export function describeStatus(
621  decision: Decision | null,
622  change: { model?: string; effort?: Effort } | null,
623): string {
624  if (!decision) return 'jev · no answer'
625  const asked = `${decision.tier} ${reported(decision.confidence)}`
626  if (!change) return `jev · ${asked} · unchanged`
627  const to = [change.model, change.effort].filter(Boolean).join('/')
628  return `jev · ${asked} → ${to}`
629}
630
631/**
632 * The `model` an agent definition's frontmatter names, or null when it names
633 * none or `inherit`: both leave the parent's model to decide. `agent.spawn`
634 * reports only the Agent tool's `model` parameter and the parent's model, so
635 * without this a definition pinned to `sonnet` is measured as if it ran on
636 * the parent's `opus`, and a decision to route it up reads as a no-op.
637 */
638export function definitionModel(markdown: string): string | null {
639  const value = frontmatterField(markdown, 'model')
640  return value && value !== 'inherit' ? value : null
641}
642
643/** An agent definition's pinned reasoning effort, if it names a supported level. */
644export function definitionEffort(markdown: string): Effort | null {
645  const value = frontmatterField(markdown, 'effort')
646  return value && effortRank(value) !== null ? (value as Effort) : null
647}
648
649/** One scalar field of a Markdown file's YAML frontmatter, unquoted; null when absent. */
650function frontmatterField(markdown: string, key: string): string | null {
651  const frontmatter = /^---\r?\n([\s\S]*?)\r?\n---/.exec(markdown)
652  if (!frontmatter) return null
653  const line = new RegExp(`^${key}:[ \\t]*(.*)$`, 'm').exec(frontmatter[1])
654  const value = line?.[1]
655    .replace(/\s+#.*$/, '')
656    .trim()
657    .replace(/^(['"])(.*)\1$/, '$2')
658  return value || null
659}
660
661/**
662 * Whether a definition file defines `agent`. Claude Code names an agent by
663 * its frontmatter `name`, not its file name (`reviewer-config.md` may define
664 * `reviewer`); a file without a `name` falls back to its file name. A file
665 * without frontmatter defines no agent.
666 */
667export function definitionMatches(markdown: string, fileName: string, agent: string): boolean {
668  if (!/^---\r?\n/.test(markdown)) return false
669  const name = frontmatterField(markdown, 'name')
670  return name ? name === agent : fileName === `${agent}.md`
671}
672
673/**
674 * The folders a plugin keeps agents in: `agents/` by default, plus each path
675 * its manifest's `agents` field names (a string or a list, relative to the
676 * install). Both are scanned; a name match decides, so a path that turns out
677 * to replace the default rather than add to it costs only a wasted scan.
678 */
679export function pluginAgentDirs(installPath: string, manifest: unknown): string[] {
680  const declared = (manifest as { agents?: unknown } | null)?.agents
681  const extra = (Array.isArray(declared) ? declared : declared === undefined ? [] : [declared])
682    .filter((path): path is string => typeof path === 'string')
683    .map((path) => `${installPath}/${path.replace(/^\.\//, '').replace(/\/+$/, '')}`)
684  return [`${installPath}/agents`, ...extra.filter((dir) => dir !== `${installPath}/agents`)]
685}
686
687/**
688 * The agent name to match and the folders to search, most specific first: a
689 * `plugin:agent` type in that plugin's agent folders, any other in the
690 * project's `.claude/agents`, then the user's. The first definition that
691 * matches wins, whether or not it names a model. Managed and `--agents`
692 * definitions are not visible here; a miss falls back to the parent's model.
693 */
694export function definitionDirs(
695  subagentType: string,
696  projectRoot: string,
697  home: string,
698  pluginDirs: readonly string[],
699): { agent: string; dirs: string[] } {
700  const colon = subagentType.indexOf(':')
701  if (colon > 0) return { agent: subagentType.slice(colon + 1), dirs: [...pluginDirs] }
702  return { agent: subagentType, dirs: [`${projectRoot}/.claude/agents`, `${home}/.claude/agents`] }
703}
704
705/** A model id without its context-window suffix (`[1m]`). */
706function withoutWindow(model: string): string {
707  return model.replace(/\[[^\]]*\]$/, '')
708}
709
710/**
711 * Whether the main loop may move to `wanted` given how full the context is.
712 *
713 * A model switch forfeits the prompt cache: each model keeps its own, so the
714 * new one writes the conversation to its cache before it answers (measured
715 * 2026-09-26: a haiku → opus switch wrote 69k tokens and read none). Early in
716 * a session that is cheap and a mechanical turn on a small model can pay for
717 * it; late, the rewrite likely costs more than the smaller model saves. The
718 * limit is a heuristic, not a measured break-even: that depends on prices,
719 * the turns left and how warm each cache is. Above it the only move allowed
720 * is back to the session's starting model, so a session that dropped to a
721 * small model early is not stuck there for a hard turn later. Its cache is
722 * warm only if it actually answered recently; that is not tracked.
723 *
724 * The starting model is matched exactly, its `[1m]` suffix aside, and then
725 * sent as the starting id, so a 1M session is not cut to 200k by an alias
726 * resolving to the same model. Another version of the same family is a
727 * different model: it is sent as asked, and held above the limit.
728 * `contextTokens` is undefined before the first response of a fresh session:
729 * nothing to lose yet. A resumed one has its reading restored by the engine.
730 */
731export function gateMainModel(
732  wanted: string,
733  contextTokens: number | undefined,
734  maxContextTokens: number,
735  baseModel: string | undefined,
736): { model: string | null; reason?: string } {
737  const isBase = baseModel !== undefined && withoutWindow(wanted) === withoutWindow(baseModel)
738  const model = isBase ? (baseModel as string) : wanted
739  if (contextTokens === undefined || contextTokens <= maxContextTokens || isBase) return { model }
740  return {
741    model: null,
742    reason: `held model: context ${contextTokens} > ${maxContextTokens}, only the starting model ${baseModel ?? '(unknown)'} is allowed`,
743  }
744}
745
746/**
747 * The `--model` / `--effort` a `codex:codex-rescue` spawn should carry, from
748 * the same decision that routed the spawn itself.
749 *
750 * The rescue agent leaves both unset unless its prompt names them, so every
751 * delegation used to run on `~/.codex/config.toml`'s default. Calibrated
752 * 2026-09-27: the workhorse (sol) carries all hard work and effort carries
753 * the difficulty; the top model is not used by default (the ladder's `deep`
754 * rung is sol). Risk above 0.7 takes xhigh; deep takes high at the upgrade
755 * bar (0.3, or with no confidence), xhigh when the effort reads xhigh, medium
756 * below the bar; a balanced task whose effort reads high or more takes
757 * medium; the cheap model needs the downgrade bar (0.6), a measured
758 * confidence and low risk (≤ 0.4). Everything else, and no decision at all,
759 * is the default: sol at low. Null — prompt left alone — only when the
760 * prompt already names either flag: an explicit choice wins.
761 */
762export function codexFlags(
763  decision: Decision | null,
764  prompt: string,
765  ladder: Record<Tier, string>,
766): { model: string; effort: Effort } | null {
767  if (/(^|\s)--(model|effort)(\s|=|$)/.test(prompt)) return null
768  const fallback = { model: ladder.balanced, effort: 'low' as Effort }
769  if (!decision) return fallback
770  const { tier, confidence, risky } = decision
771  const level = decision.effort === null ? null : effortLevel(decision.effort)
772  if (risky !== null && risky > 0.7) return { model: ladder.deep, effort: 'xhigh' }
773  if (tier === 'deep') {
774    if (confidence === null || confidence >= 0.3) return { model: ladder.deep, effort: level === 'xhigh' ? 'xhigh' : 'high' }
775    return { model: ladder.balanced, effort: 'medium' }
776  }
777  if (tier === 'fast' && confidence !== null && confidence >= 0.6 && (risky === null || risky <= 0.4)) {
778    return { model: ladder.fast, effort: 'low' }
779  }
780  if (tier === 'balanced' && (level === 'high' || level === 'xhigh')) return { model: ladder.balanced, effort: 'medium' }
781  return fallback
782}
783