Routes each Claude Code turn to the model and effort TypeSafe's Jev picks, weighed against what a switch costs, and compacts with Jev instead of a lossy…

A Claude Code plugin that picks the model and effort for every turn with Jev, TypeSafe's decision model, and weighs each switch against what it costs.
This is a fork. It is an extended version of satviksinha/jev-model-router by Satvik Sinha, with cost-aware routing, compaction by Jev (built on tamaratran/fast-jev-compaction), session persistence and a number of other additions. See What this fork adds and Credits.
MIT licensed (LICENSE).
For every prompt you send, Jev answers two questions in one request: which tier should handle it, and how hard that model should think. The plugin then checks the answer against the conversation's cost and size before any request goes out, and rewrites each model request in the turn to the result.
you type a prompt
↓
turn.start ask Jev → tier: fable, effort: max
↓ apply the checks → window, confidence, price, ceiling
turn.step each request → model: claude-fable-5-1, effort: xhigh
↓
reply > ✳️ fable · xhigh · Jev 97% · capped from max · 641ms
…your reply…
fable-5-1 ✓ xhigh · Jev 97% · $0.35 · 130k in (91% cached) · 2k out
The four tiers:
| Tier | For | Model |
|---|---|---|
haiku | Trivial: a lookup, a rename, a yes or no. | claude-haiku-4-5 |
sonnet | Straightforward and minor, no real decision to make. | claude-sonnet-5-5 |
opus | Plain implementation carrying some complexity. | claude-opus-5-5 |
fable | Planning, brainstorming, architecture, systematic debugging. | claude-fable-5-1 |
What Jev is told each tier is for lives in TIER_CRITERIA in hooks/policy.ts. Editing those strings is how you change the router's judgement; nothing else needs to change.
Why the checks matter: Claude's prompt cache is per model. Moving a long conversation to another model rewrites the whole context into that model's cache, which at a few hundred thousand tokens costs dollars. A router that follows every pick can cost more than it saves, so this one only switches when the switch is worth it, and says so when it is not.
Everything below goes to the provider you configured (TypeSafe, or the Vercel AI Gateway), and nowhere else:
Agent "reviewer" completed) when JEV_ROUTER_NOTIFY_CONTINUE=0 or there is no route to continue; never the task's result. A result can quote anything, so a prompt is cut at the first thing shaped like a task notification (<task-notification> followed by a tag), wherever it is: a prompt that quotes an example of one has only what comes before it graded (the model still gets all of it), and the same cut applies to what is kept on disk./jev compact off stops it.On disk, in Claude Code's plugin store, the router keeps each session's routing history, including the first 400 characters of each prompt, for the last 20 sessions.
~/.claude/settings.json, with function hooks enabled: {
"env": {
"TYPESAFE_API_KEY": "...",
"CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1"
}
}
Without CLAUDE_CODE_ENABLE_FUNCTION_HOOKS the plugin loads and silently does nothing.
~/.claude/skills/jev-claude-router/ to load it in every session, or run claude --plugin-dir /path/to/jev-claude-router for one session.npm run check-jev reports the provider and whether it answers. In a session, /jev shows the router's state.Everything else has a working default: the confidence bar at 75%, price checks on, a $1 upgrade limit, an effort ceiling of xhigh, and compaction by Jev on.
Each reply opens with one line saying what ran and why:
> ✳️ opus · high · Jev 98% · 555ms
The tier, the effort, Jev's confidence in the tier, and how long Jev took. When a check changed Jev's pick, the reason is written in plain words:
| The line says | Meaning |
|---|---|
kept fable: Jev 61% on haiku, needs 75% | Jev wanted haiku but was not sure enough to switch. |
kept fable: haiku costs $4.41 vs $0.13 | Moving down would have cost more than staying, cache included. |
kept opus: fable costs $5.03 vs $0.080, over the $1.00 limit | Moving up would have cost more than the upgrade limit over staying. |
kept fable: too long for haiku (310k) | The conversation does not fit haiku's window. |
haiku too long, moved up only to sonnet (Jev wanted fable) | The running tier outgrew its window; the turn went to the cheapest tier that fits. |
capped from max | Jev asked for more effort than the ceiling allows. |
your pick | You named the tier in your prompt. |
1st request runs medium as high | Fable runs medium as high on a conversation's first request, so that is what is sent. |
kept opus: Jev timed out after 1500ms | Jev did not answer in time; the turn stayed on the tier already running. |
> ⚠️ not routed: <reason> | Routing failed with nothing running to stay on; the turn ran on the session model. |
The line is part of the reply's text because that is the one channel every Claude Code surface draws: terminal, desktop app and IDE. The app's own model picker does not change, since the router rewrites individual requests, not the session. In the terminal the session-mode footer also shows the route (jev: opus, medium effort).
Under each finished reply, one block reports what the API says actually answered and what it cost at Anthropic's list price:
fable-5-1 ✓ xhigh · Jev 97% · $0.35 · 130k in (91% cached) · 2k out
✓ means the model that answered is the one the router asked for; a mismatch reads opus-5 ⚠ asked fable-5-1. 91% cached is the share of input read from the prompt cache.
A reply that launched background agents spans several turns and still gets one summary, written once every agent has finished:
3 turns: fable, fable, fable (2 woken by tasks) · $6.16 · 7.4M in (99% cached) · 61k out
agents: Explore haiku-4-5 $0.029, general-purpose opus-5-5 $0.36
turn 2: kept fable: haiku costs $4.41 vs $0.13
Lines after the first appear only when something did not run as Jev asked.
/jev quiet hides the line and the summary without stopping routing; /jev loud brings them back.
/jevjev-claude-router:
routing on
surface desktop
provider typesafe · TYPESAFE_API_KEY is set · jev-latest
budget 1500ms
sticky on, switch needs 75% (90% up past 100k)
price on, a downgrade has to pay, an upgrade may cost $1.00 over staying
ceiling xhigh (fable: medium)
compact on, Jev prunes tool calls · last: kept 41/87 messages, 63% smaller (12 calls kept, 9 cut, 30 dropped) · 2.1s
session claude-opus-5, running on fable
cache 1h writes · 201k context · fable→haiku pays below 3k
tiers haiku, sonnet, opus, fable
announce on, a line per turn
spent $4.12 this session
Recent turns, newest first:
0ms fable·medium [task finished] Agent "Review library-sync cluster" complet…
→ fable-5-1 ✓ · $0.076 · 45k in (98% cached) · 1k out
641ms fable·xhigh Jev 97%; capped from max help me plan the architecture
→ fable-5-1 ✓ · $0.35 · 130k in (91% cached) · 2k out
352ms fable·low kept fable: haiku costs $1.02 vs $0.020 what is 2+2
→ fable-5-1 ✓ · $0.034 · 47k in (99% cached) · 0k out
12ms not routed — gateway said HTTP 403 (customer_verification_required)
session is the model the session runs on, and the tier that is warm.cache is the context being priced, and the size below which the cheapest downgrade still pays.spent is the session's total at list price, routed turns or not.→) what actually answered. Turns you did not type are labelled [task finished], [continuing] or [type agent].The checks run in this order. Each one can only narrow what the one before allowed.
Naming a tier. A tier named with a routing verb (use opus, switch to fable, route to haiku, run this on sonnet, go with opus) skips Jev, runs at medium effort and shows your pick. The phrase has to be said to the model: at the start of a sentence or clause, or after "please", "just", "let's", "can you" and the like. Talk about a tier is not a route: "search for opus docs", "should I use opus or sonnet?", "make production use sonnet by default" and "I told you not to use haiku" are ordinary prompts, and so is a route under a condition ("if it runs long, switch to opus", "otherwise use opus"), or a tier naming something else ("use sonnet pricing"); "if you can, use opus" is still a request. A negation cancels the next route ("don't use haiku, use opus" goes to opus). Pasted content, code, quoted lines, text in double quotes (a short single-quoted phrase too) and background-task notifications are not read for this, so a pasted document that says "use opus" as an example does not route. JEV_ROUTER_ALLOW_OVERRIDE=0 turns this off, and a tier turned off with JEV_ROUTER_EXCLUDE or /jev tiers off cannot be named back in.
Go-aheads. Jev scores a bare "yes" as trivial, which is right about the text and wrong about the work. A prompt that is only a go-ahead (y, yes, ok, sure, go ahead, continue, do it, lgtm and similar) continues on the previous turn's tier and effort without asking Jev.
Wake-ups. A turn the engine starts itself, when a background task finishes or with its own "still working" nudge, continues the reply's route without a Jev call, so it adds no latency and cannot switch the model under a reply in progress. JEV_ROUTER_NOTIFY_CONTINUE=0 asks Jev about finished tasks anyway. Either way a finished task adds no route line to a reply that is still open or already summarised. The engine hands the router an empty prompt for its nudge and for a prompt with no words (an image on its own), and the two cannot be told apart, so a words-free prompt also continues the last route.
A turn is never sent to a tier whose context window it does not fit. Haiku 4.5 takes 200k tokens and the other tiers a million, less 16k of headroom. This applies whatever Jev said, and even to a tier you named: the turn stays on the tier already running.
When the running tier is the one that no longer fits (haiku past 184k), the turn moves up only as far as it must, to the cheapest offered tier that fits. That holds whether Jev picked haiku again, picked a higher tier the checks below held back, or the turn was a go-ahead. Because the running tier cannot take the turn at all, the step itself is not held back by the confidence bar or the price checks. With nothing known to be running, the session model keeps the turn, since its cache is the warm one.
A switch to a different tier than the one running needs Jev to be at least 75% sure. Moving up once the context is past 100k needs 90%, because it rewrites the whole context into a pricier cache. The bar follows the tier actually running, so a run of unsure picks cannot creep the session down one turn at a time. /jev sticky 0.6 moves the bar; /jev sticky off removes it.
JEV_ROUTER_UPGRADE_MAX over staying ($1 by default). With a typical turn that allows an upgrade up to about 48k of context from opus to fable, 126k from sonnet to opus, and 254k from haiku to sonnet./jev price off turns both off, independently of the confidence bar. npm run measure-switch-cost prints what a switch costs at each size.
The price follows the cache that is actually warm:
claude --resume, /model, or turning routing back on, the first turn is priced against the model that answered last./model alias names a tier, not a version: opus is held as Opus and a kept turn goes out as the engine's own Opus, but until the first response says which Opus answered, staying is priced at the current Opus's rates. opusplan and default name no single model, so nothing is held to them.claude-opus-5) is treated the same way: moving it to claude-opus-5-5 means a cold cache.A held turn still gets the effort Jev asked for on Haiku, Opus and Fable, since effort is sent per request and costs no cache. On Sonnet an effort change rewrites much of the cache, so the effort is held too unless Jev is sure enough of it. (Measured on Sonnet 5, to which Claude Code sent no effort; the hold stays for Sonnet 5.5 until it is measured there.)
Each tier has an effort ceiling, xhigh by default (matching the engine's own default). A turn Jev wanted higher runs at the ceiling and says capped from max. /jev ceiling changes it, per tier or for all of them: /jev ceiling xhigh then /jev ceiling medium fable runs everything at xhigh except Fable, held to medium. To drop a tier entirely instead of capping its effort, /jev tiers off fable takes it out of the question Jev is asked.
Fable 5.1 runs medium as high on the first request of a conversation, so the router sends high there and says so. Only a conversation's first request counts: after a compaction the effort asked for is the effort that runs, and /clear starts a new conversation.
Each spawned agent is routed on its own task, unless the call named a model or is a fork. A subagent starts with an empty context, so there is no cache to protect and no hold to the parent's tier; below 50% confidence it is left on its default model. Agents appear in /jev and in the reply's summary, never in the agent's own reply, which its parent reads as a tool result.
When a conversation fills its context, Claude Code compacts it into a summary and detail is lost. With this plugin, a compaction asks Jev about every tool call in the transcript instead — one request for a typical conversation, split into a few (at most two in flight at once) once there are enough calls to outgrow one request's budget: does this call still matter, and does its full output still need to be there?
| Jev's answer | What happens |
|---|---|
| The output still matters | Kept exactly as it was. |
| The call matters, its output does not | Kept, with the first 300 characters of the result and a note. |
| Neither | Removed, call and result together. |
Text messages are never changed, and the first message and the six most recent are never touched. What remains is the conversation itself, verbatim.
compact on, Jev prunes tool calls · last: kept 41/87 messages, 63% smaller (12 calls kept, 9 cut, 30 dropped) · 2.1s
The engine's summary runs instead, and /jev says why, when Jev removes less than 25%, takes longer than 8 seconds, fails, or the provider is the Vercel gateway (which does not answer the yes/no questions this uses). A /compact with instructions of its own is left to the engine's summary, which can follow them.
What Jev sees: the conversation's text and each tool call's input (up to 1,000 characters, so a Write or Edit call's content is included). Tool results are described only by their size and whether they errored; their contents are never sent. Claude Code compacts ahead of time and then for real a few messages later; the transcript is scored once.
/jev compact off restores the engine's summary, and /jev off turns compaction by Jev off along with routing.
It fails safe. When Jev is too slow or errors, the turn stays on the tier already running, so a hiccup never costs a cold cache and a switch back; the line says so. With nothing running yet, or no key at all, the turn runs exactly as it would without the plugin, and the line says why. The only cost is the wait, capped at the timeout. Prompts are cut to their first 12,000 characters before Jev sees them.
State survives a reload. The routing history, spend, the tier being held, the open reply and every /jev setting are saved in Claude Code's per-plugin store (~/.claude/plugins/store/jev-claude-router_*.json) and restored after an update or reload. The twenty most recently used sessions are kept.
One copy acts. More than one copy of the plugin can be loaded at once, after a reload or when the app continues a conversation under a new id. The newest copy acts and the others pass the turn through untouched: no Jev call, no model change, no line, no summary. Within one process this is decided in memory; across processes, by an owner record and a 60-second claim on each turn in the store.
Its own marks stay single. The model sees past lines and summaries in its own replies and can write look-alikes with invented figures. The plugin removes a route line at the start of the model's text and a summary at its end, as the text streams, before writing the real ones. One quoted mid-reply is left alone.
| Command | Effect |
|---|---|
/jev | Status, settings and recent turns. |
/jev on, /jev off | Turn routing (and compaction by Jev) on or off. |
/jev quiet, /jev loud | Hide or show the route line and summary. |
/jev sticky, /jev sticky 0.6, /jev sticky off | Turn the confidence bar on, set it, or turn it off. |
/jev price, /jev price on, /jev price off | Show or toggle the price checks. |
/jev ceiling | Show the effort ceiling. |
/jev ceiling xhigh, /jev xhigh | Raise every tier's ceiling. |
/jev ceiling xhigh fable, /jev xhigh fable | Raise one tier's ceiling. |
/jev ceiling off | Remove every cap. |
/jev tiers | Show which tiers are offered to Jev. |
/jev tiers off fable | Drop one or more tiers from the question entirely. |
/jev tiers on fable | Bring a dropped tier back. |
/jev compact, /jev compact on, /jev compact off | Show or toggle compaction by Jev. |
A command overrides the matching setting below for the rest of the session, and is kept across reloads.
All settings go in the env block of ~/.claude/settings.json.
Provider
| Variable | Default | Effect |
|---|---|---|
TYPESAFE_API_KEY | TypeSafe direct key. | |
AI_GATEWAY_API_KEY | Vercel AI Gateway key. | |
JEV_ROUTER_PROVIDER | TypeSafe if its key is set | typesafe (or direct) or gateway (or vercel); anything else chooses by the keys set. |
TYPESAFE_BASE_URL | https://api.typesafe.ai | Must be https. Only *.typesafe.ai is accepted unless JEV_ROUTER_ALLOW_CUSTOM_BASE=1. A trailing /v1/systemone is dropped, since the router adds it. |
JEV_ROUTER_ALLOW_CUSTOM_BASE | off | 1 lets TYPESAFE_BASE_URL name any https host. That host receives your key and prompts, so set both only in your own ~/.claude/settings.json, and check that a project's settings do not set them. |
JEV_ROUTER_JEV_MODEL | jev-latest | Pin a Jev version (e.g. jev-1.13.0) so confidences stay stable across releases. TypeSafe direct only; the gateway serves its own. |
JEV_ROUTER_TIMEOUT_MS | 1500 | How long a turn waits for Jev, in milliseconds. At least 100 (anything less is taken for a mistake and the default used), at most 8000. |
Routing
| Variable | Default | Effect |
|---|---|---|
JEV_ROUTER_STICKY | on | 0 removes the confidence bar. |
JEV_ROUTER_STICKY_CONFIDENCE | 0.75 | The confidence bar: 0.6, 60 or 60%. |
JEV_ROUTER_PRICE_CHECK | on | 0 turns both price checks off. |
JEV_ROUTER_UPGRADE_MAX | 1 | Dollars an upgrade may cost over staying (a plain number, $ allowed), or off. |
JEV_ROUTER_CACHE_TTL | 1h | Cache lifetime used for pricing: 1h (what Claude Code writes) or 5m. |
JEV_ROUTER_CEILING | xhigh | Effort ceiling: an effort for all tiers, or per tier, e.g. fable:medium,opus:high. |
JEV_ROUTER_EXCLUDE | Tiers not offered to Jev at session start, e.g. fable,haiku (commas, semicolons or spaces); /jev tiers toggles this per session. | |
JEV_ROUTER_ALLOW_OVERRIDE | on | 0 ignores tiers named in prompts. |
JEV_ROUTER_NOTIFY_CONTINUE | on | 0 asks Jev about finished-task turns. |
Compaction
| Variable | Default | Effect |
|---|---|---|
JEV_ROUTER_COMPACT | on | 0 leaves compaction to the engine's summary. |
JEV_ROUTER_COMPACT_TIMEOUT_MS | 8000 | How long scoring may take. At most 8000, under the hook's 10-second budget. |
JEV_ROUTER_COMPACT_MIN_REDUCTION | 0.25 | The share Jev must remove for its result to stand (40% works too). |
npm run check-jev # is the configured provider serving?
npm run try-prompts # Jev's tier, effort and confidence on sample prompts
npm run try-prompts -- "text" # the same for one prompt
npm run measure-switch-cost # what a switch costs at each context size
npm run bench-overhead # engine calls and plugin time per turn, with a fake engine
try-prompts is the tuning loop: edit TIER_CRITERIA, run it, check the picks moved the way you wanted. No script prints a key.
ms tier effort conf prompt
839 haiku low 1.00 what is 2+2
402 haiku medium 0.75 rename the variable foo to bar in utils.ts
482 sonnet medium 0.66 add a --verbose flag to the CLI
555 opus high 0.98 implement cursor pagination for the reports endpoint
641 fable xhigh 0.97 the e2e suite passes alone but fails with the others
734 fable xhigh 1.00 help me plan the architecture for multi-tenant billing
Tier and effort are separate questions and can disagree: a short question about unfamiliar code can be trivial to route but hard to answer.
hooks/register.ts the hooks, settings, turn history, saved state, which copy acts
hooks/policy.ts tiers, criteria, answers → model and effort, holds, the ceiling
hooks/prichooks/register.ts 2450 lines1import type { On } from "claude-code";
2
3import {
4 askJev,
5 timeoutOf,
6 type HttpInitLike,
7 type HttpResponseLike,
8 type StateSource,
9} from "./jev.ts";
10import { labelOf, withLabel } from "./label.ts";
11import {
12 TIER_ALIAS,
13 MODEL_ID,
14 asAsked,
15 capTo,
16 ceilingAt,
17 ceilingOf,
18 effortNamed,
19 excludedTiers,
20 firstTurnEffort,
21 isContinuation,
22 isEngineNudge,
23 notifyContinueOf,
24 priceCheckOf,
25 upgradeMaxOf,
26 offeredTiers,
27 overrideAllowedOf,
28 parseOverride,
29 sessionDecision,
30 stickyOf,
31 thresholdOf,
32 tierFilter,
33 DEFAULT_CEILING,
34 EFFORTS,
35 TIERS,
36 type Ceiling,
37 type Decision,
38 type Effort,
39 type Tier,
40} from "./policy.ts";
41import { tierOfModel, baseModel, sameModelAs, ttlOf, usageCost, type Ttl } from "./pricing.ts";
42import {
43 pack,
44 SNAPSHOT_PREFIX,
45 orphanOwnerKeys,
46 savedAtOf,
47 SNAPSHOTS_KEPT,
48 staleKeys,
49 unpack,
50 OVERRIDABLE,
51 type Overridable,
52 type State,
53} from "./persist.ts";
54import { providerOf, type ProviderResult } from "./provider.ts";
55import {
56 compactOnOf,
57 compactTimeoutOf,
58 minReductionOf,
59 pruneTranscript,
60 shortOf,
61 type Compaction,
62} from "./compactor.ts";
63import { messageChars } from "./compaction/compact.ts";
64import {
65 addUsage,
66 announceReply,
67 attemptOf,
68 carriedOf,
69 ceilingCommand,
70 normalUsage,
71 continuationOf,
72 continuationSkipped,
73 HISTORY_LIMIT,
74 kept,
75 liveLine,
76 FOOTER_SEPARATOR,
77 REPLY_SEPARATOR,
78 notificationOf,
79 notificationStateOf,
80 hasNotification,
81 notificationTaskOf,
82 replySummary,
83 spawnAttemptOf,
84 statusReport,
85 stickyCommand,
86 tiersCommand,
87 toggleReply,
88 TYPICAL_OUTPUT_TOKENS,
89 unknownCommandReply,
90 ImitationFilter,
91 isRouteLine,
92 type AgentTag,
93 type Attempt,
94} from "./status.ts";
95
96/** The slice of the engine a Jev call needs; every hook's `$` has it. */
97type Engine = {
98 env: { get: (key: string) => Promise<string | undefined> };
99 http: {
100 fetch: (url: string, init?: HttpInitLike) => Promise<HttpResponseLike>;
101 };
102 clock: { sleep: (ms: number, options?: { signal?: AbortSignal }) => Promise<unknown> };
103};
104
105/** The stamp of the copy that holds a session, from its owner record, or null. */
106async function ownerOf(
107 $: { store: { get: (key: string) => Promise<unknown> } },
108 key: string,
109): Promise<number | null> {
110 try {
111 return stampOf(await $.store.get(`${OWNER_PREFIX}${key}`));
112 } catch {
113 return null;
114 }
115}
116
117/** How much later than a copy's own last save a stored snapshot must be to count as another's. */
118const HANDOFF_SLACK_MS = 1000;
119
120/** Turns a reply keeps for its summary; past this the oldest go. */
121const REPLY_LIMIT = 64;
122
123/** Whether two ceilings cap every tier the same. */
124const sameCeiling = (a: Ceiling, b: Ceiling) => TIERS.every((t) => a[t] === b[t]);
125
126/** Turns kept in the decision cache before the oldest are dropped. */
127const CACHE_LIMIT = 32;
128
129/** A summary block, as `replySummary` writes it after `FOOTER_SEPARATOR`. */
130const SUMMARY = /^\n\n```\n[^\n]*\(\d+% cached\)/;
131
132/** Store key prefix for a claim on one turn, by its text. */
133const TURN_PREFIX = "turn:";
134
135/**
136 * How long a claim on a turn's text stands. Copies of the module that see
137 * the same turn see it within a second or two of each other; a prompt the
138 * person repeats minutes later is a turn of its own.
139 */
140const TURN_CLAIM_MS = 60_000;
141
142/** Turn claims kept in the store before the oldest are dropped. */
143const TURN_CLAIMS_KEPT = 50;
144
145/** A short stable hash of a turn's text, for its claim key. */
146function textHash(text: string): string {
147 let h = 5381;
148 for (let i = 0; i < text.length; i++) h = ((h << 5) + h + text.charCodeAt(i)) | 0;
149 return (h >>> 0).toString(36);
150}
151
152/**
153 * The claim key for a turn: its text, context, and session id, when known.
154 * Two different, unrelated warm sessions that happen to report the exact
155 * same token count for the same short prompt within the claim window no
156 * longer collide, since the session id is always folded in here (closed
157 * 2026-09-24; see the audit note in CHANGELOG-worthy commits for the
158 * measured collision).
159 *
160 * Why this stays safe for the case the id itself was chosen to solve: on
161 * 2026-09-23, one *live* copy kept routing a resumed conversation under its
162 * old session id after the app rotated it. `turn.start` re-reads
163 * `$.session.id()` on every turn (not only at `session.start`) and follows
164 * it when it changes, so that one copy's own claim key tracks the new id
165 * from its very next turn — nothing about folding the id in here breaks
166 * that, since it is still the *same* copy computing both the old and the
167 * new key over time, one after the other, not two copies racing on
168 * different ids at once.
169 *
170 * What remains open: two genuinely *separate* copies (a same-process
171 * module reload, or two processes) that each read a *different*, and
172 * non-converging, id for what is really one conversation. A same-process
173 * reload is already handled independently of this key, by `superseded` and
174 * `ownsSession` sharing `globalThis` — the newer copy wins outright and the
175 * older never even reaches a claim. A cross-process case with no shared
176 * `globalThis` has no such fallback: each copy's turn key now differs (it
177 * did not before this change either, once one of them reports a nonzero
178 * context, which a resumed conversation typically does immediately), so
179 * both may claim and both may write a line. This was not reproduced
180 * independently of the regression test that first covered it, and closing
181 * the far more easily reached collision — any two different sessions, same
182 * prompt, same reported context — was judged the higher-value fix.
183 */
184function turnKey(
185 text: string,
186 contextTokens: number | string | null,
187 sessionId: string | null = null,
188): string {
189 const base = `${textHash(text)}-${
190 typeof contextTokens === "string" ? textHash(contextTokens) : (contextTokens ?? 0)
191 }`;
192 return sessionId !== null ? `${base}-${textHash(sessionId)}` : base;
193}
194
195/** Store key prefix for which copy of the module owns a session. */
196const OWNER_PREFIX = "owner:";
197
198/**
199 * A claim unrefreshed for this long belongs to a copy that is gone. The
200 * holder refreshes it on every turn, so only a copy that died (a process
201 * killed without its session.end) lets it age this far.
202 */
203const OWNER_TTL_MS = 30 * 60 * 1000;
204
205/** How often the holder refreshes its claim, well inside `OWNER_TTL_MS`. */
206const OWNER_REFRESH_MS = 5 * 60 * 1000;
207
208/** Store key prefix for when a session's holder last claimed it. */
209const SEEN_PREFIX = "seen:";
210
211/**
212 * The holding copy's stamp from an owner record: a bare number, as every
213 * version writes it (and `{ birth }` from a short-lived one that did not).
214 */
215function stampOf(raw: unknown): number | null {
216 if (typeof raw === "number" && Number.isFinite(raw)) return raw;
217 if (typeof raw === "object" && raw !== null && typeof (raw as { birth?: unknown }).birth === "number")
218 return (raw as { birth: number }).birth;
219 return null;
220}
221
222/** A `seen:` record: which copy claimed (its stamp, and its id when written), and when. */
223function seenOf(raw: unknown): { birth: number; at: number; nonce?: string } | null {
224 if (
225 typeof raw === "object" && raw !== null &&
226 typeof (raw as { birth?: unknown }).birth === "number" &&
227 typeof (raw as { at?: unknown }).at === "number"
228 ) {
229 const nonce = (raw as { nonce?: unknown }).nonce;
230 return {
231 birth: (raw as { birth: number }).birth,
232 at: (raw as { at: number }).at,
233 ...(typeof nonce === "string" ? { nonce } : {}),
234 };
235 }
236 return null;
237}
238
239/** How old a claim on a session with no snapshot must be before it is dropped. */
240const ORPHAN_OWNER_MS = 24 * 60 * 60 * 1000;
241
242/** How often a turn still running saves its state. */
243const MID_TURN_SAVE_MS = 5_000;
244
245/**
246 * How many messages a transcript may have gained since it was scored for
247 * the scoring to be reused with them appended: within the newest messages
248 * pruning leaves alone (fast-jev-compaction's `preserveRecentMessages`).
249 */
250const PRUNE_TAIL_REUSED = 6;
251
252/**
253 * The decision a session is running on, from the model id the API reports
254 * answered: a dated id (`claude-opus-5-5-20260901`) is its undated model.
255 * Null for an id off the ladder or not a string.
256 */
257/**
258 * Whether a model the engine names is what `running` already runs: the
259 * same model, or a tier's own alias (`opus`) for a decision on that tier,
260 * whose version the alias does not say and so cannot contradict.
261 */
262function speaksFor(model: string, running: Decision): boolean {
263 const alias = TIER_ALIAS.exec(model);
264 return sameModelAs(running.model, model) || (alias !== null && alias[1]!.toLowerCase() === running.tier);
265}
266
267/** Marks the end of a step's stream, after its last chunk. */
268const STEP_END = Symbol("step-end");
269
270/** `source`, then `STEP_END`. */
271async function* withEnd<T>(source: AsyncIterable<T>): AsyncGenerator<T | typeof STEP_END> {
272 for await (const item of source) yield item;
273 yield STEP_END;
274}
275
276function warmDecision(model: unknown, effort: unknown): Decision | null {
277 if (typeof model !== "string" || model === "") return null;
278 const warm = sessionDecision(model.replace(/-\d{8}$/, ""));
279 if (warm === null) return null;
280 // The effort the step actually ran at, so a Sonnet effort hold does not
281 // bind to a placeholder; a numeric or absent effort leaves the default.
282 const name = typeof effort === "string" ? effort.trim().toLowerCase() : "";
283 const ran = (EFFORTS as readonly string[]).includes(name) ? (name as Effort) : null;
284 return ran === null ? warm : { ...warm, effort: ran };
285}
286
287/**
288 * Stop reasons that mean the turn continues: the engine will step again, so
289 * a summary would land in the middle of a reply. `tool_use` and `pause_turn`
290 * are the usual two; `compaction` is the engine compacting mid-turn and
291 * carrying on, which wrote a second summary under one reply when it was
292 * taken for an end (seen 2026-09-23). Every other reason ends the turn.
293 */
294const MID_TURN: ReadonlySet<string> = new Set([
295 "tool_use",
296 "pause_turn",
297 "compaction",
298]);
299
300/**
301 * Everything the router reads from the environment, read once. None of it
302 * changes within a session, and reading them all on every turn was an await
303 * each ahead of the Jev call. `sticky`, `ceiling`, `offered` and
304 * `excluded` start here and are then owned by `/jev sticky`, `/jev ceiling`
305 * and `/jev tiers`.
306 */
307type Settings = {
308 provider: ProviderResult;
309 timeoutMs: number;
310 offered: readonly Tier[];
311 excluded: readonly Tier[];
312 sticky: number | null;
313 ceiling: Ceiling;
314 ttl: Ttl;
315 allowOverride: boolean;
316 notifyContinue: boolean;
317 upgradeMax: number | null;
318 /** The downgrade and upgrade price checks; `/jev price` owns it after the environment. */
319 priceCheck: boolean;
320 /** Compaction by Jev: on, how long it may take, how much it must remove. */
321 compactOn: boolean;
322 compactTimeoutMs: number;
323 compactMinReduction: number;
324};
325
326/**
327 * Reads the settings from the environment on first use. A top-level
328 * function on purpose: the engine follows where `$` goes when it loads a
329 * module, and only lets it into a function declared here at the top, so a
330 * closure taking `$` inside `register` fails the whole module (measured
331 * 2026-09-23: it loaded nothing and every turn went unrouted).
332 */
333async function seedSettings(
334 $: Engine,
335 current: Settings | null,
336): Promise<Settings> {
337 if (current !== null) return current;
338 // Normalized together: JEV_ROUTER_EXCLUDE naming every tier falls back to
339 // the full ladder in `offered`, and `excluded` must agree with that or
340 // `/jev` can end up saying a tier is both offered and excluded.
341 const tiers = tierFilter(excludedTiers(await $.env.get("JEV_ROUTER_EXCLUDE")));
342 return {
343 provider: providerOf({
344 TYPESAFE_API_KEY: await $.env.get("TYPESAFE_API_KEY"),
345 AI_GATEWAY_API_KEY: await $.env.get("AI_GATEWAY_API_KEY"),
346 JEV_ROUTER_PROVIDER: await $.env.get("JEV_ROUTER_PROVIDER"),
347 TYPESAFE_BASE_URL: await $.env.get("TYPESAFE_BASE_URL"),
348 JEV_ROUTER_ALLOW_CUSTOM_BASE: await $.env.get(
349 "JEV_ROUTER_ALLOW_CUSTOM_BASE",
350 ),
351 JEV_ROUTER_JEV_MODEL: await $.env.get("JEV_ROUTER_JEV_MODEL"),
352 }),
353 timeoutMs: timeoutOf(await $.env.get("JEV_ROUTER_TIMEOUT_MS")),
354 offered: tiers.offered,
355 excluded: tiers.excluded,
356 sticky: stickyOf(await $.env.get("JEV_ROUTER_STICKY"))
357 ? thresholdOf(await $.env.get("JEV_ROUTER_STICKY_CONFIDENCE"))
358 : null,
359 ceiling: ceilingOf(await $.env.get("JEV_ROUTER_CEILING")),
360 ttl: ttlOf(await $.env.get("JEV_ROUTER_CACHE_TTL")),
361 allowOverride: overrideAllowedOf(
362 await $.env.get("JEV_ROUTER_ALLOW_OVERRIDE"),
363 ),
364 notifyContinue: notifyContinueOf(
365 await $.env.get("JEV_ROUTER_NOTIFY_CONTINUE"),
366 ),
367 upgradeMax: upgradeMaxOf(await $.env.get("JEV_ROUTER_UPGRADE_MAX")),
368 priceCheck: priceCheckOf(await $.env.get("JEV_ROUTER_PRICE_CHECK")),
369 compactOn: compactOnOf(await $.env.get("JEV_ROUTER_COMPACT")),
370 compactTimeoutMs: compactTimeoutOf(
371 await $.env.get("JEV_ROUTER_COMPACT_TIMEOUT_MS"),
372 ),
373 compactMinReduction: minReductionOf(
374 await $.env.get("JEV_ROUTER_COMPACT_MIN_REDUCTION"),
375 ),
376 };
377}
378
379/**
380 * Asks Jev about one piece of text. Top-level, for the same reason as above.
381 * `signal`, when given, is aborted by the caller if the turn is ceded to a
382 * newer copy before this resolves — a saving only where the fetch honours
383 * it (see askJev's own note); harmless to pass otherwise.
384 */
385async function classify(
386 $: Engine,
387 text: string,
388 offered: readonly Tier[],
389 settings: Settings,
390 signal?: AbortSignal,
391 source: StateSource = "prompt",
392) {
393 return askJev({
394 fetch: (url, init) => $.http.fetch(url, init),
395 sleep: (ms, options) => $.clock.sleep(ms, options),
396 provider: settings.provider,
397 state: text,
398 offered,
399 source,
400 timeoutMs: settings.timeoutMs,
401 signal,
402 });
403}
404
405/**
406 * The context the next turn will carry, in tokens, from the engine's own
407 * count of the last response, or null when it has none yet (a fresh
408 * session, or one just compacted, which is also when there is no cache to
409 * protect). Older engines have no `usage()`; that reads as null too.
410 */
411async function contextTokensOf($: {
412 session: {
413 usage: () => Promise<{ context?: { tokens?: number } } | undefined>;
414 };
415}): Promise<number | null> {
416 try {
417 const usage = await $.session.usage();
418 const tokens = usage?.context?.tokens;
419 return typeof tokens === "number" && tokens > 0 ? tokens : null;
420 } catch {
421 return null;
422 }
423}
424
425/** The store key for this session's snapshot, or null when the engine has no id. */
426async function snapshotKeyOf($: {
427 session: { id: () => Promise<string> };
428}): Promise<string | null> {
429 try {
430 const id = await $.session.id();
431 return typeof id === "string" && id !== "" ? `${SNAPSHOT_PREFIX}${id}` : null;
432 } catch {
433 return null;
434 }
435}
436
437/** The snapshot saved under `key`, or null. Never throws. */
438async function loadSnapshot(
439 $: { store: { get: (key: string) => Promise<unknown> } },
440 key: string,
441): Promise<State | null> {
442 try {
443 return unpack(await $.store.get(key));
444 } catch {
445 return null;
446 }
447}
448
449/**
450 * Saves `state` under `key`. A session's first save drops the oldest
451 * snapshots past `SNAPSHOTS_KEPT`. Never throws: losing a snapshot costs
452 * what a reload cost before, and must not cost the turn.
453 */
454async function saveSnapshot(
455 $: {
456 store: {
457 get: (key: string) => Promise<unknown>;
458 set: (key: string, value: unknown) => Promise<void>;
459 keys: () => Promise<string[]>;
460 delete: (key: string) => Promise<void>;
461 };
462 },
463 key: string,
464 state: State,
465 first: boolean,
466): Promise<void> {
467 try {
468 if (first) {
469 const keys = await $.store.keys();
470 // Only read when there is something to prune: one get per session.
471 const sessions = keys.filter((k) => k.startsWith(SNAPSHOT_PREFIX) && k !== key);
472 const savedAt = new Map<string, number>();
473 if (sessions.length + 1 > SNAPSHOTS_KEPT)
474 for (const k of sessions) {
475 const at = savedAtOf(await $.store.get(k));
476 if (at !== null) savedAt.set(k, at);
477 }
478 for (const stale of staleKeys(keys, key, savedAt)) {
479 await $.store.delete(stale);
480 await $.store.delete(`${OWNER_PREFIX}${stale}`);
481 await $.store.delete(`${SEEN_PREFIX}${stale}`);
482 }
483 // A claim on a session that never saved (left at once for a /resume
484 // or a /clear) has no snapshot to be pruned with; one a day old is
485 // nobody's live session any more.
486 // A seen: record whose owner record is gone (released by a copy of an
487 // earlier version, which does not know about seen:) is dropped once old.
488 for (const seenKey of keys.filter((k) => k.startsWith(SEEN_PREFIX) && !keys.includes(`${OWNER_PREFIX}${k.slice(SEEN_PREFIX.length)}`))) {
489 const seen = seenOf(await $.store.get(seenKey));
490 if (seen === null || Date.now() - seen.at > ORPHAN_OWNER_MS) await $.store.delete(seenKey);
491 }
492 for (const owner of orphanOwnerKeys(keys, OWNER_PREFIX, key)) {
493 const seenKey = `${SEEN_PREFIX}${owner.slice(OWNER_PREFIX.length)}`;
494 const stamp = stampOf(await $.store.get(owner));
495 const seen = seenOf(await $.store.get(seenKey));
496 const last = seen !== null && seen.birth === stamp ? seen.at : stamp;
497 if (last === null || Date.now() - last > ORPHAN_OWNER_MS) {
498 await $.store.delete(owner);
499 await $.store.delete(seenKey);
500 }
501 }
502 // Keys are kept in the order they were first written, and the oldest
503 // are dropped; moving this session to the end makes that the order
504 // of last use, so a long-lived session in use is never the one dropped.
505 // The old snapshot is put back if the new one cannot be written.
506 const previous = await $.store.get(key);
507 await $.store.delete(key);
508 try {
509 await $.store.set(key, pack(state));
510 } catch (error) {
511 if (previous !== undefined) await $.store.set(key, previous).catch(() => undefined);
512 throw error;
513 }
514 return;
515 }
516 await $.store.set(key, pack(state));
517 } catch {
518 // The store refused (over 4 MiB, a disk error): carry on unsaved.
519 }
520}
521
522/**
523 * Whether this copy of the module owns the session. When the plugin's files
524 * change, the engine loads a fresh copy without retiring the old one, and
525 * both then handle every turn: two Jev calls, two route lines, a summary
526 * from each (seen 2026-09-23, from 14:54 on in one session). The newest copy
527 * wins: each is stamped with its load time, the highest stamp is kept in the
528 * store under `owner:<session>`, and a copy that finds a newer stamp there
529 * stands aside. `claim` writes this copy's stamp when it is the newer one.
530 * A store that cannot be read leaves every copy in charge, as before.
531 */
532async function ownsSession(
533 $: {
534 store: {
535 get: (key: string) => Promise<unknown>;
536 set: (key: string, value: unknown) => Promise<void>;
537 };
538 },
539 key: string,
540 birth: number,
541 claim: boolean,
542 /** This copy's own id, which tells two copies with the same stamp apart. */
543 nonce?: string,
544): Promise<boolean> {
545 try {
546 const ownerKey = `${OWNER_PREFIX}${key}`;
547 const seenKey = `${SEEN_PREFIX}${key}`;
548 const owner = stampOf(await $.store.get(ownerKey));
549 const seen = seenOf(await $.store.get(seenKey));
550 // When the holder last claimed, if it says: a copy of an earlier version
551 // writes no `seen:` record, and its claim never goes stale here, as it
552 // never did before.
553 const lastSeen = owner !== null && seen !== null && seen.birth === owner ? seen.at : null;
554 // A newer copy that has not been seen for a while is gone (a process
555 // killed without its session.end): it no longer holds the session.
556 const gone = lastSeen !== null && Date.now() - lastSeen >= OWNER_TTL_MS;
557 if (owner !== null && owner > birth && !gone) return false;
558 // The same stamp from another copy (two processes loaded in the same
559 // millisecond: the stamp's fraction is too coarse at this size to keep
560 // them apart): the one whose id the seen: record holds keeps it.
561 if (owner === birth && seen !== null && seen.birth === birth && seen.nonce !== undefined && nonce !== undefined && seen.nonce !== nonce && !gone)
562 return false;
563 if (claim) {
564 // The owner record stays a bare stamp, which earlier versions read.
565 if (owner === null || owner < birth || gone) {
566 await $.store.set(ownerKey, birth);
567 await $.store.set(seenKey, { birth, at: Date.now(), ...(nonce !== undefined ? { nonce } : {}) });
568 } else if (owner === birth && (lastSeen === null || Date.now() - lastSeen >= OWNER_REFRESH_MS)) {
569 // The holder refreshes every few minutes, not on every step.
570 await $.store.set(seenKey, { birth, at: Date.now(), ...(nonce !== undefined ? { nonce } : {}) });
571 }
572 }
573 return true;
574 } catch {
575 return true;
576 }
577}
578
579/**
580 * Claims a turn for this copy, by the turn's text, context, and session id
581 * (see `turnKey` for what that key does and does not close). The newest
582 * copy wins: an older one that claimed first is overridden, and checks
583 * again before it writes. False means a newer copy holds the turn. A store
584 * that cannot be read lets the copy through, as before.
585 */
586async function claimTurn(
587 $: {
588 store: {
589 get: (key: string) => Promise<unknown>;
590 set: (key: string, value: unknown) => Promise<void>;
591 keys: () => Promise<string[]>;
592 delete: (key: string) => Promise<void>;
593 };
594 },
595 text: string,
596 contextTokens: number | string | null,
597 birth: number,
598 sessionId: string | null = null,
599): Promise<boolean> {
600 try {
601 const at = `${TURN_PREFIX}${turnKey(text, contextTokens, sessionId)}`;
602 const now = Date.now();
603 const held = (await $.store.get(at)) as { birth?: unknown; at?: unknown } | undefined;
604 if (
605 held &&
606 typeof held.at === "number" &&
607 typeof held.birth === "number" &&
608 now - held.at < TURN_CLAIM_MS &&
609 held.birth > birth
610 ) {
611 return false;
612 }
613 await $.store.set(at, { birth, at: now });
614 // Another process may have written between the read and the write;
615 // whichever claim the store holds now decides, newest winning.
616 const after = (await $.store.get(at)) as { birth?: unknown } | undefined;
617 if (typeof after?.birth === "number" && after.birth > birth) return false;
618 const claims = (await $.store.keys()).filter((k) => k.startsWith(TURN_PREFIX));
619 for (const old of claims.slice(0, Math.max(0, claims.length - TURN_CLAIMS_KEPT)))
620 await $.store.delete(old);
621 return true;
622 } catch {
623 return true;
624 }
625}
626
627/** Whether this copy still holds a turn it claimed. Never throws; true when unreadable. */
628async function holdsTurn(
629 $: { store: { get: (key: string) => Promise<unknown> } },
630 key: string,
631 birth: number,
632): Promise<boolean> {
633 try {
634 const held = (await $.store.get(`${TURN_PREFIX}${key}`)) as
635 | { birth?: unknown }
636 | undefined;
637 return typeof held?.birth !== "number" || held.birth === birth;
638 } catch {
639 return true;
640 }
641}
642
643/** Drops this copy's claim on the session, if it still holds it. Never throws. */
644async function releaseSession(
645 $: {
646 store: {
647 get: (key: string) => Promise<unknown>;
648 delete: (key: string) => Promise<void>;
649 };
650 },
651 key: string,
652 birth: number,
653 /** This copy's id: a copy with the same stamp that does not hold the claim leaves it. */
654 nonce?: string,
655): Promise<void> {
656 try {
657 const at = `${OWNER_PREFIX}${key}`;
658 const seen = seenOf(await $.store.get(`${SEEN_PREFIX}${key}`));
659 const theirs = seen !== null && seen.birth === birth && seen.nonce !== undefined && nonce !== undefined && seen.nonce !== nonce;
660 if (stampOf(await $.store.get(at)) === birth && !theirs) {
661 await $.store.delete(at);
662 await $.store.delete(`${SEEN_PREFIX}${key}`);
663 }
664 } catch {
665 // Nothing to release, or the store is unreadable: the next claim decides.
666 }
667}
668
669/** Where the session draws first (`terminal`, `desktop`, ...), or null in a plain -p run. */
670async function surfaceOf($: {
671 session: { surfaces: () => Promise<readonly string[]> };
672}): Promise<string | null> {
673 try {
674 return (await $.session.surfaces())[0] ?? null;
675 } catch {
676 return null;
677 }
678}
679
680/** The main loop's model as `/model` shows it, or null when the engine has none. */
681async function sessionModelOf($: {
682 session: { model: () => Promise<string> };
683}): Promise<string | null> {
684 try {
685 const model = await $.session.model();
686 return typeof model === "string" && model !== "" ? model : null;
687 } catch {
688 return null;
689 }
690}
691
692/**
693 * True while any of `ids` is still running: the reply that spawned them is
694 * not over. Only the reply's own agents count — one from an earlier reply,
695 * or a long-lived one, must not hold every later summary hostage.
696 */
697async function agentsRunning(
698 $: { agent: { list: () => Promise<readonly { id: string; status: string }[]> } },
699 ids: ReadonlySet<string>,
700): Promise<boolean> {
701 if (ids.size === 0) return false;
702 const rows = await $.agent.list().catch(() => []);
703 return rows.some((r) => ids.has(r.id) && r.status === "running");
704}
705
706/**
707 * Names the subagent a step runs in, from the session's agent list. A row may
708 * not be there yet for a loop that only just started; then the id stands in,
709 * which still says "not the main loop", the part that matters.
710 */
711async function agentTagOf(
712 $: {
713 agent: {
714 list: () => Promise<
715 readonly { id: string; type: string; description: string }[]
716 >;
717 };
718 },
719 agentId: string,
720): Promise<AgentTag> {
721 const rows = await $.agent.list().catch(() => []);
722 const row = rows.find((r) => r.id === agentId);
723 return row ? { type: row.type, label: row.description } : { label: agentId };
724}
725
726/**
727 * Registers the router: one Jev call per turn, applied to every model request
728 * that turn makes, and announced as it happens.
729 *
730 * The decision is made once in `turn.start`, where the person's text is, and
731 * read back in `turn.step`, which fires again after each tool result. Asking
732 * per step would pay Jev's latency several times over and could land two
733 * steps of one turn on different models.
734 *
735 * Every turn's outcome is kept, routed or not, because "did this do anything"
736 * is unanswerable otherwise: a router that fails open looks exactly like one
737 * that is not loaded.
738 *
739 * @param on the engine's registrar
740 */
741export function register(on: On) {
742 /** When this copy was loaded; the newest copy owns the session. */
743 // Copies loaded into one runtime share `globalThis`; the newest stands, and
744 // this holds even for a copy that cannot read a session id to claim with.
745 // A copy's stamp is always above every one already there, so two loaded
746 // in the same millisecond still have an order; the fraction keeps copies
747 // in different processes, which share only the store, from tying.
748 const runtime = globalThis as {
749 __jevRouterNewest?: number;
750 /** The newest copy's live state, for the copy that replaces it. */
751 __jevRouterLive?: () => { key: string; state: unknown; savedAt: number; birth: number } | null;
752 };
753 const birth = Math.max(
754 Date.now() + Math.random() * 0.001,
755 (runtime.__jevRouterNewest ?? 0) + 0.001,
756 );
757 runtime.__jevRouterNewest = birth;
758 const superseded = () => birth < (runtime.__jevRouterNewest ?? 0);
759 /** This copy's own id, for telling it from another loaded in the same millisecond. */
760 const nonce = `${Math.random().toString(36).slice(2)}${Math.random().toString(36).slice(2)}`;
761 // A reload mid-turn: the copy being replaced holds what it has not saved
762 // yet (mid-turn saves are throttled), so this copy takes its live state
763 // for the same session, once, over the store's older snapshot.
764 let previousLive = runtime.__jevRouterLive;
765 runtime.__jevRouterLive = () =>
766 snapshotKey ? { key: snapshotKey, state: pack(stateNow()), savedAt: lastSavedAt, birth } : null;
767 // Only from the copy that holds the session in this store (its claim is
768 // the owner record), so a copy for another store or a session that was
769 // never this one's is not taken for a reload; and only when the store's
770 // snapshot is no newer than that copy's own last save: one saved later
771 // came from another process that went on in the session.
772 // Whatever is restored, its save time is this copy's starting point, so
773 // the next reload can tell it from a newer one in turn.
774 const restoredFrom = (key: string, stored: State | null, owner: number | null): State | null => {
775 const previous = previousLive?.();
776 previousLive = undefined;
777 const fromStore = () => {
778 lastSavedAt = stored?.savedAt ?? lastSavedAt;
779 return stored;
780 };
781 if (!previous || previous.key !== key || owner !== previous.birth) return fromStore();
782 if (stored?.savedAt !== undefined && stored.savedAt > previous.savedAt + HANDOFF_SLACK_MS) return fromStore();
783 const live = unpack(JSON.parse(JSON.stringify(previous.state)));
784 if (live === null) return fromStore();
785 lastSavedAt = previous.savedAt;
786 return live;
787 };
788 /** True once a newer copy has claimed the session: this one stands aside. */
789 let inert = false;
790 /** The engine said the resumed session's cache has expired, and no response has written it since. */
791 let cacheExpired = false;
792 /** The model a resume event reported, until the resumed snapshot is restored against it. */
793 let resumedOn: string | null = null;
794 /**
795 * The model a placeholder guessed from a tier's alias (`opus` names the
796 * tier, not the version the engine resolves it to): a turn that holds to
797 * it goes out as the engine's own model of that tier, so "kept" is true.
798 */
799 let aliasGuess: string | null = null;
800 const placeholderOf = (model: string): Decision | null => {
801 const made = sessionDecision(model);
802 aliasGuess = made !== null && TIER_ALIAS.test(model) ? made.model : null;
803 return made;
804 };
805 /**
806 * A snapshot was restored since the last turn: the session model it held
807 * may be another process's, or from before a resume that did not name the
808 * model, so the next turn checks it against the engine's.
809 */
810 let liveModelDue: { model: string | null } | null = null;
811 /** A resume or fork event has said whether the cache expired: that outranks a snapshot's word. */
812 let resumeSpoke = false;
813 /**
814 * A response has been received in this conversation. The engine runs some
815 * efforts as others on a conversation's first request only
816 * (FIRST_TURN_EFFORT); measured 2026-09-23, the first request after a
817 * compaction runs the effort asked for, so a compaction does not reset
818 * this. A `/clear` does: it is a new conversation.
819 */
820 let answered = false;
821 /** The last compaction Jev was asked about, for /jev. */
822 let lastCompaction: Compaction | null = null;
823 /**
824 * The last transcript Jev pruned and what it kept: the engine compacts
825 * ahead of time (`precompute`) and then for real over the same messages,
826 * and each dispatch would otherwise be another scoring.
827 */
828 let prunedCache: {
829 handles: readonly string[];
830 messages: readonly unknown[];
831 reduction: number;
832 compaction: Compaction;
833 } | null = null;
834 /** Turns a newer copy claimed: this one passes them through untouched. */
835 const ceded = new Set<string>();
836 /** Agents whose reply's summary has been written: their wake-up joins no other. */
837 const summarisedAgents = new Set<string>();
838 /** Each claimed turn's text, to check the claim again before writing. */
839 const claimed = new Map<string, string>();
840 let settings: Settings | null = null;
841 const decisions = new Map<string, Decision>();
842 /**
843 * Turns whose reply has yet to open with its route line. The line itself is
844 * built at the first text chunk, not here: by then the step has said which
845 * loop the turn runs in, which the line names.
846 */
847 const pending = new Set<string>();
848 const attempts: Attempt[] = [];
849 /**
850 * turnId → its attempt, so each step's `stop` chunk can add what the API
851 * reported to the right turn. The same objects as in `attempts`.
852 */
853 const byTurn = new Map<string, Attempt>();
854 /**
855 * Every turn since the last one the person typed, agents included: what
856 * one reply took, written under it once, at the end. A reply that spawns
857 * background work is several turns — the typed one, then one per task
858 * that finished and woke the loop — and a block under each read as one
859 * reply changing model three times.
860 */
861 let reply: Attempt[] = [];
862 /** The agents the current reply spawned; its summary waits for them. */
863 let replyAgents = new Set<string>();
864 let latest: Decision | null = null;
865 /** The tier the last routed turn ran on; what a shaky switch is held to. */
866 let running: Decision | null = null;
867 /**
868 * What was running before this turn moved `running` to a new model, until
869 * a response on it confirms the new model's cache was written. A turn
870 * interrupted or failed before any response wrote nothing, so the next
871 * turn goes back to pricing against what is actually warm.
872 */
873 let unconfirmed: { was: Decision | null } | null = null;
874 /** What that turn carried and produced, for pricing the next switch. */
875 let lastUsage: { context: number; output: number } | null = null;
876 /** The main loop's model as `/model` shows it, read when first needed. */
877 let sessionModel: string | null = null;
878 /**
879 * What a bare go-ahead continues. Cleared on an unrouted turn: that turn
880 * ran on the session model, so re-applying the older routed decision would
881 * be wrong. Stickiness still holds to `running` (last routed).
882 */
883 let continueFrom: Decision | null = null;
884 let enabled = true;
885 let announce = true;
886 let surface: string | null = null;
887 /** Dollars across every turn seen this session, at list price. */
888 let spent = 0;
889 /** When the state was last saved mid-turn; end-of-turn saves are not throttled. */
890 let lastMidTurnSave = 0;
891 /**
892 * agentId → what its spawn settled on, for the subagent's own steps to
893 * apply and for /jev to show. Keyed by the id `next(e)` hands back from
894 * `agent.spawn`, which is the same id the loop's `turn.step` carries.
895 */
896 const spawned = new Map<string, Attempt>();
897 /** Agents the router left alone (forks, a named model), for one history row each. */
898 const unrouted = new Map<string, Attempt>();
899 /** Agents whose first step has run, so later steps get Jev's effort. */
900 const stepped = new Set<string>();
901
902 const trim = (map: Map<string, unknown>) => {
903 while (map.size > CACHE_LIMIT) {
904 const oldest = map.keys().next();
905 if (oldest.done) break;
906 map.delete(oldest.value);
907 }
908 };
909
910 /**
911 * Drop idle turn rows, but never an in-flight one still in `pending` or
912 * `decisions` — those still need the route line and usage fold-in. If every
913 * entry is protected, the map is allowed to grow past the limit.
914 */
915 // `keep` is the turn just added: it is not in `pending` or `decisions`
916 // yet, and with every other row protected it was the one evicted, so
917 // every 33rd turn lost its line, its summary and its usage.
918 const trimByTurn = (keep?: string) => {
919 let scanned = 0;
920 while (byTurn.size > CACHE_LIMIT && scanned < byTurn.size) {
921 const oldest = byTurn.keys().next();
922 if (oldest.done) break;
923 const key = oldest.value;
924 if (key === keep || pending.has(key) || decisions.has(key)) {
925 touch(byTurn, key, byTurn.get(key)!);
926 scanned++;
927 continue;
928 }
929 byTurn.delete(key);
930 scanned = 0;
931 }
932 };
933
934 /** Move a live entry to the end so FIFO trim drops idle keys first. */
935 const touch = <V>(map: Map<string, V>, key: string, value: V) => {
936 map.delete(key);
937 map.set(key, value);
938 };
939
940 const trimSet = (set: Set<string>) => {
941 while (set.size > CACHE_LIMIT) {
942 const oldest = set.values().next();
943 if (oldest.done) break;
944 set.delete(oldest.value);
945 }
946 };
947
948 /**
949 * Forget what the main loop was running on and the turns in flight. After
950 * `/jev off` the session model answers, and after `/clear` or a resume
951 * into another session the cache the hold was protecting is not this
952 * conversation's, so the next routed turn starts from Jev's word. (A
953 * compaction forgets only what was warm; see session.compact.)
954 */
955 const clearRouting = () => {
956 aliasGuess = null;
957 decisions.clear();
958 byTurn.clear();
959 pending.clear();
960 // spawned is kept: turn.step already ignores it while off, and clearing
961 // it made /jev on mid-agent invent "not routed at spawn" and drop effort.
962 latest = null;
963 continueFrom = null;
964 running = null;
965 unconfirmed = null;
966 lastUsage = null;
967 };
968
969 /**
970 * This session's snapshot key: undefined until looked up, null when the
971 * engine gives no session id (then nothing is saved or restored).
972 */
973 let snapshotKey: string | null | undefined = undefined;
974 let savedOnce = false;
975 /** False right after `/clear`: the next key lookup must not restore. */
976 let restoreOnKey = true;
977
978 /** When this copy last saved, for a replacing copy to tell its snapshot from a newer one. */
979 let lastSavedAt = 0;
980 /** The state to save, noting when. */
981 const stateToSave = (): State => {
982 lastSavedAt = Date.now();
983 return stateNow();
984 };
985
986 /** The settings a `/jev` command set this session, which outrank the environment. */
987 const overridden = new Set<Overridable>();
988
989 const stateNow = (): State => ({
990 attempts,
991 reply,
992 replyAgents: [...replyAgents],
993 spawned: [...spawned.entries()],
994 unrouted: [...unrouted.entries()],
995 turns: [...byTurn.entries()],
996 decisions: [...decisions.entries()],
997 pending: [...pending],
998 stepped: [...stepped],
999 running,
1000 continueFrom,
1001 latest,
1002 lastUsage,
1003 sessionModel,
1004 spent,
1005 enabled,
1006 announce,
1007 answered,
1008 sticky: settings?.sticky ?? null,
1009 ceiling: settings?.ceiling ?? ceilingAt(DEFAULT_CEILING),
1010 excludedTiers: [...(settings?.excluded ?? [])],
1011 compactOn: settings?.compactOn ?? true,
1012 priceCheck: settings?.priceCheck ?? true,
1013 overridden: [...overridden],
1014 summarisedAgents: [...summarisedAgents],
1015 compaction: lastCompaction,
1016 unconfirmed,
1017 cacheExpired,
1018 });
1019
1020 /** Puts a restored snapshot back, over what the environment seeded. */
1021 const applyState = (s: State | null) => {
1022 // Nothing to restore: a resume's reported model has nothing to correct.
1023 if (s === null) {
1024 resumedOn = null;
1025 liveModelDue = null;
1026 return;
1027 }
1028 liveModelDue = { model: s.sessionModel ?? null };
1029 // A placeholder made before the restore (a resume event first) is gone,
1030 // and its guess with it: the snapshot says what runs.
1031 aliasGuess = null;
1032 attempts.splice(0, attempts.length, ...s.attempts.slice(0, HISTORY_LIMIT));
1033 reply = s.reply;
1034 replyAgents = new Set(s.replyAgents);
1035 spawned.clear();
1036 for (const [id, a] of s.spawned) spawned.set(id, a);
1037 unrouted.clear();
1038 for (const [id, a] of s.unrouted) unrouted.set(id, a);
1039 byTurn.clear();
1040 for (const [id, a] of s.turns) byTurn.set(id, a);
1041 decisions.clear();
1042 for (const [id, d] of s.decisions) decisions.set(id, d);
1043 pending.clear();
1044 for (const id of s.pending) pending.add(id);
1045 stepped.clear();
1046 for (const id of s.stepped) stepped.add(id);
1047 running = s.running;
1048 continueFrom = s.continueFrom;
1049 latest = s.latest;
1050 lastUsage = s.lastUsage;
1051 sessionModel = s.sessionModel ?? sessionModel;
1052 spent = s.spent;
1053 enabled = s.enabled;
1054 announce = s.announce;
1055 answered = s.answered;
1056 lastCompaction = s.compaction;
1057 unconfirmed = s.unconfirmed;
1058 // A resume event this copy saw speaks for the session now; a snapshot
1059 // saved before it does not.
1060 if (!resumeSpoke) cacheExpired = s.cacheExpired;
1061 // A resume that reported another model than the snapshot's: the session
1062 // is on that one now, and what was warm under the snapshot's is not.
1063 if (resumedOn !== null) {
1064 if (running !== null && !speaksFor(resumedOn, running)) {
1065 running = placeholderOf(resumedOn);
1066 unconfirmed = null;
1067 }
1068 sessionModel = resumedOn;
1069 resumedOn = null;
1070 }
1071 summarisedAgents.clear();
1072 for (const id of s.summarisedAgents) summarisedAgents.add(id);
1073 // Only what a command set outranks the environment; the rest stays as
1074 // the environment seeded it, so a changed JEV_ROUTER_* holds on reload.
1075 // A snapshot from before `overridden` existed restores them all.
1076 const restore = new Set<Overridable>(s.overridden ?? OVERRIDABLE);
1077 overridden.clear();
1078 for (const k of restore) overridden.add(k);
1079 if (settings !== null) {
1080 if (restore.has("sticky")) settings.sticky = s.sticky;
1081 if (restore.has("ceiling")) settings.ceiling = s.ceiling;
1082 // undefined means the snapshot predates this field: leave the
1083 // environment's own JEV_ROUTER_EXCLUDE seeding in place rather than
1084 // overwrite it with "nothing excluded" (see State.excludedTiers).
1085 if (restore.has("excludedTiers") && s.excludedTiers !== undefined) {
1086 const tiers = tierFilter(s.excludedTiers as Tier[]);
1087 settings.excluded = tiers.excluded;
1088 settings.offered = tiers.offered;
1089 }
1090 if (restore.has("compactOn")) settings.compactOn = s.compactOn;
1091 if (restore.has("priceCheck")) settings.priceCheck = s.priceCheck;
1092 }
1093 };
1094
1095 /** Whether this save is the session's first, which prunes old sessions. */
1096 const firstSave = () => {
1097 const first = !savedOnce;
1098 savedOnce = true;
1099 return first;
1100 };
1101
1102 const record = (attempt: Attempt) => {
1103 attempts.unshift(attempt);
1104 attempts.length = Math.min(attempts.length, HISTORY_LIMIT);
1105 reply.push(attempt);
1106 // A reply that never closes (the engine's nudges alone, or quiet) must
1107 // not grow the snapshot without end: past the limit the oldest turns
1108 // after its first go. The first is kept: it is the one the person typed,
1109 // which lets the summary be written at all.
1110 if (reply.length > REPLY_LIMIT) reply.splice(1, reply.length - REPLY_LIMIT);
1111 };
1112
1113 on("session.start", async ($, e, next) => {
1114 await $.command.register({
1115 name: "jev",
1116 description: "Jev routing: status, on/off, sticky, price, ceiling, compact, quiet/loud.",
1117 });
1118 surface = await surfaceOf($);
1119 settings = await seedSettings($, settings);
1120 if (snapshotKey === undefined) {
1121 snapshotKey = await snapshotKeyOf($);
1122 if (snapshotKey !== null && restoreOnKey)
1123 applyState(restoredFrom(snapshotKey, await loadSnapshot($, snapshotKey), await ownerOf($, snapshotKey)));
1124 restoreOnKey = true;
1125 }
1126 // A reloaded copy gets its own session.start, so it claims the session
1127 // the moment it loads; an older copy then stands aside from the next
1128 // turn on, instead of both asking Jev on the first one.
1129 if (snapshotKey) inert = !(await ownsSession($, snapshotKey, birth, true, nonce));
1130 sessionModel = await sessionModelOf($);
1131 return next(e);
1132 });
1133
1134 // The session ending releases its claim, so a copy in another process
1135 // that resumes the same session later is not left standing aside behind
1136 // an owner that no longer exists.
1137 on("session.end", async ($, e, next) => {
1138 // What the throttled mid-turn saves have not written yet is saved first:
1139 // once the claim is gone, a later copy restores from the store alone.
1140 // Only the copy that still holds the session: one a reload replaced
1141 // may not have learnt it yet, and would write its stale state over.
1142 if (snapshotKey && settings && !inert && !superseded() && (await ownsSession($, snapshotKey, birth, false, nonce)))
1143 await saveSnapshot($, snapshotKey, stateToSave(), firstSave());
1144 if (snapshotKey && !inert) await releaseSession($, snapshotKey, birth, nonce);
1145 return next(e);
1146 });
1147
1148 // A resumed session is already running on something, with a cache the
1149 // first routed turn's switch should be priced against — unless the engine
1150 // says that cache has expired, in which case there is nothing to protect.
1151 // `/clear` starts a new conversation: nothing is running.
1152 on("classic.SessionStart", async ($, e, next) => {
1153 if (e.source === "clear") {
1154 // The old conversation is saved as it stands, and its claim goes; the
1155 // next lookup claims afresh.
1156 // Only the copy that still holds the session: one a reload replaced
1157 // may not have learnt it yet, and would write its stale state over.
1158 if (snapshotKey && settings && !inert && !superseded() && (await ownsSession($, snapshotKey, birth, false, nonce)))
1159 await saveSnapshot($, snapshotKey, stateToSave(), firstSave());
1160 if (snapshotKey && !inert) await releaseSession($, snapshotKey, birth, nonce);
1161 // A new conversation, and a new transcript id: its state is saved
1162 // under that, so a later resume of the old session restores the old
1163 // session's. The engine does not say when the id rotates, so the key
1164 // is looked up again on the next hook, when it has — and that lookup
1165 // must not restore, or it would undo the clear.
1166 clearRouting();
1167 // A restored model to check belongs to the conversation left.
1168 liveModelDue = null;
1169 reply = [];
1170 replyAgents = new Set();
1171 snapshotKey = undefined;
1172 restoreOnKey = false;
1173 attempts.length = 0;
1174 spent = 0;
1175 answered = false;
1176 savedOnce = false;
1177 // Compaction cache: the next turn starts fresh, so any previous prune
1178 // score is invalid.
1179 prunedCache = null;
1180 // A new conversation has no cache to have expired.
1181 cacheExpired = false;
1182 // Nor a compaction, or finished agents of its own. An agent still
1183 // running from before the clear keeps its routing.
1184 lastCompaction = null;
1185 summarisedAgents.clear();
1186 // An unreadable list keeps them all: it says nothing about which run.
1187 const rows = await $.agent.list().catch(() => null);
1188 if (rows !== null) {
1189 const live = new Set(rows.filter((a) => a.status === "running").map((a) => a.id));
1190 for (const id of [...spawned.keys()]) if (!live.has(id)) spawned.delete(id);
1191 for (const id of [...stepped]) if (!live.has(id)) stepped.delete(id);
1192 }
1193 unrouted.clear();
1194 }
1195 // A resume or fork into a different session, in a process already
1196 // running one: the old session's routing must not carry over. Its state
1197 // is dropped, and the next hook restores the resumed session's own.
1198 // Also right after a /clear, which left the key to be looked up again
1199 // and the restore off: a resume is a restore, whatever came before it.
1200 if (e.source === "resume" || e.source === "fork") {hooks/jev.ts 387 lines1/**
2 * Asking Jev, TypeSafe's decision model, through either the Vercel AI
3 * Gateway or TypeSafe's direct API.
4 *
5 * The gateway (POST /v1/evaluate) speaks its own vocabulary: question types
6 * are `choice` and `score` (never TypeSafe's native `noul`, which it rejects
7 * outright). Probabilities and confidences come back rounded to two decimal
8 * places.
9 *
10 * TypeSafe direct (POST /v1/systemone) supports choice, score, and noul and
11 * returns probabilities rounded to four decimal places.
12 *
13 * Both support the same `choice` and `score` question types and the same
14 * `answers` response shape, so the request/response handling is identical.
15 *
16 * `fetch` and `sleep` are arguments rather than imports so this file runs
17 * under plain `node` in tests, with no engine and no network.
18 */
19
20import { EFFORT_CRITERIA, PLAIN_DECIMAL, TIER_CRITERIA, type Tier } from "./policy.ts";
21import type { ProviderResult } from "./provider.ts";
22
23/**
24 * Measured against the live gateway on 2026-09-20: ten prompts ran 402ms to
25 * 839ms. An 800ms budget failed open on the slowest of them, so this leaves
26 * real headroom while still capping what a turn waits before giving up.
27 * `JEV_ROUTER_TIMEOUT_MS` overrides it.
28 */
29export const DEFAULT_TIMEOUT_MS = 1500;
30
31/**
32 * Jev takes 32k tokens of state and reads the whole of it; TypeSafe's own
33 * guidance is that accuracy falls as the state grows with content unrelated
34 * to the decision. A routing decision is made on how a request opens, so a
35 * long paste is cut here rather than sent whole and refused with a 422.
36 */
37export const MAX_STATE_CHARS = 12_000;
38
39/**
40 * The first `n` UTF-16 units of `text`, one fewer when the cut would split a
41 * surrogate pair: half an emoji is not valid Unicode, and a strict JSON
42 * parser on the other end refuses the whole body.
43 */
44export function headOf(text: string, n: number): string {
45 if (n <= 0) return "";
46 if (text.length <= n) return text;
47 const last = text.charCodeAt(n - 1);
48 return text.slice(0, last >= 0xd800 && last <= 0xdbff ? n - 1 : n);
49}
50
51/** The state Jev is sent: the prompt, cut at MAX_STATE_CHARS. */
52export function stateOf(text: string): string {
53 const trimmed = text.trim();
54 return trimmed.length <= MAX_STATE_CHARS
55 ? trimmed
56 : `${headOf(trimmed, MAX_STATE_CHARS)}…`;
57}
58
59/**
60 * An error's message as text, whatever was thrown: a fetch may reject with
61 * something that is not an Error, or one whose message is not a string, and
62 * turning that into text must not itself throw.
63 */
64export function messageOf(error: unknown): string {
65 try {
66 if (error instanceof Error && typeof error.message === "string") return error.message || "unknown error";
67 // A bare number or a blank string says nothing a person can use either.
68 if (typeof error === "string") return error.trim() || "unknown error";
69 // `undefined`, `[object Object]` and the like say nothing to a person.
70 return "unknown error";
71 } catch {
72 return "unknown error";
73 }
74}
75
76/**
77 * `text` with a `Bearer` value that reads as a credential taken out: one
78 * with a digit or an underscore, sixteen characters or more, or letters of
79 * both cases — not the words an error uses about one ("Bearer required",
80 * "Invalid Bearer token.", "Bearer auth/OAuth"). Each value is judged on
81 * its own, so the work is linear. The configured key is taken out by
82 * `withoutKey` wherever it is.
83 */
84export function withoutBearer(text: string): string {
85 // U+0085 counts as a gap too: `\s` leaves it out, and a line split later
86 // turns it into a space next to the token.
87 return text.replace(/\bBearer([\s\u0085]+)(["']?)([^\s\u0085"']+)(["']?)/gi, (all, space: string, _open: string, value: string) => {
88 // Trailing punctuation found by hand: `[…]+$` rescanned a long run of
89 // it from each position.
90 let cut = value.length;
91 while (cut > 0 && ".,;:!?)]".includes(value[cut - 1]!)) cut--;
92 const trail = value.slice(cut);
93 const bare = value.slice(0, value.length - trail.length);
94 const credential =
95 /[\d_]/.test(bare) || bare.length >= 16 || (/^[A-Za-z]{8,}$/.test(bare) && /[a-z]/.test(bare) && /[A-Z]/.test(bare));
96 return credential ? `${all.slice(0, 6)}${space}…${trail}` : all;
97 });
98}
99
100/** `text` with every occurrence of `key` taken out: an error can quote it anywhere. */
101export function withoutKey(text: string, key: string | undefined): string {
102 return key && key.length >= 4 ? text.split(key).join("…") : text;
103}
104
105/**
106 * An engine fetch error as a few words for the route line: without the
107 * engine's "<plugin>: $.http.fetch(<url>) failed:" preamble, repeats and
108 * advice, e.g. "ECONNREFUSED" or "getaddrinfo ENOTFOUND host".
109 */
110export function shortError(detail: string): string {
111 // Cut first, and newlines folded by splitting: `\s*\n\s*` is quadratic
112 // on a long run of spaces with no newline in it.
113 // Control characters become spaces and credentials go before the cut, so
114 // none can hide a token from the Bearer rule or split one at the cut.
115 const bare = withoutBearer(
116 // eslint-disable-next-line no-control-regex
117 detail.replace(/[\x00-\x08\x0e-\x1f\x7f-\x84\x86-\x9f]/g, " "),
118 )
119 .slice(0, 2000)
120 .split(/[\n\r\v\f\u0085\u2028\u2029]/)
121 .map((t) => t.trim())
122 .filter((t) => t !== "")
123 .join(" ")
124 .replace(/^[\w.-]+: \$\.http\.fetch\([^)]*\) failed: /, "");
125 const said = bare
126 .split(/[.?!]\s/)[0]!
127 .replace(/^(\w+): \1\b:?\s*/, "$1: ")
128 .replace(/:\s*$/, "")
129 .trim();
130 return said.length > 60 ? `${said.slice(0, 57)}…` : said;
131}
132
133/**
134 * Below this a budget is taken for a mistake, most likely seconds written
135 * where milliseconds are read (`1.5`), which would time every turn out.
136 */
137export const MIN_TIMEOUT_MS = 100;
138
139/** A timeout from the environment, or the default when it is unusable. */
140export function timeoutOf(raw: string | undefined): number {
141 const v = (raw ?? "").trim();
142 const parsed = Number(v);
143 if (!PLAIN_DECIMAL.test(v) || parsed < MIN_TIMEOUT_MS) return DEFAULT_TIMEOUT_MS;
144 return Math.min(parsed, MAX_TIMEOUT_MS);
145}
146
147/**
148 * The longest a turn may wait for Jev. The wait runs on `$.clock`, which
149 * counts against the hook's 10-second budget; a hook over it is skipped as
150 * absent and its turn goes unrecorded. So the budget from the environment is
151 * held under it, with room for the hook's own work.
152 */
153export const MAX_TIMEOUT_MS = 8000;
154
155export type HttpResponseLike = {
156 ok: boolean;
157 status: number;
158 text: string;
159};
160
161/**
162 * What one attempt at Jev came to. A failure carries its reason so the
163 * session can say why a turn went unrouted instead of going quiet.
164 */
165export type JevResult =
166 | { ok: true; answers: unknown; ms: number }
167 | { ok: false; reason: string; ms: number };
168
169export type AskArgs = {
170 fetch: (url: string, init?: HttpInitLike) => Promise<HttpResponseLike>;
171 /** The engine's `$.clock.sleep`; `signal` ends the wait early, so no timer outlives the call. */
172 sleep: (ms: number, options?: { signal?: AbortSignal }) => Promise<unknown>;
173 provider: ProviderResult;
174 state: string;
175 offered: readonly Tier[];
176 /** Who wrote `state`, which the question to Jev says; `prompt` when absent. */
177 source?: StateSource;
178 timeoutMs?: number;
179 /**
180 * Aborted by the caller when the answer is no longer wanted (the turn was
181 * ceded to a newer copy before this resolved). Wired to the same fetch
182 * signal as the timeout abort; whether it actually stops the request
183 * depends on the fetch implementation honouring it — see the note below.
184 */
185 signal?: AbortSignal;
186 /** Injected so tests can measure without a real clock. */
187 now?: () => number;
188};
189
190export type HttpInitLike = {
191 method?: string;
192 headers?: Record<string, string>;
193 body?: string;
194 signal?: AbortSignal;
195};
196
197/** What the text Jev grades is: a typed prompt, a subagent's task, a finished task's report. */
198export type StateSource = "prompt" | "task" | "notification";
199
200/**
201 * How the text is introduced to Jev. A subagent's task is written by the
202 * model, and a finished task's report is a line about work done, not the
203 * work to do next: graded as a developer's request, "Agent X completed"
204 * reads as trivial when the turn is about to work through its results.
205 */
206const TIER_QUESTION: Record<StateSource, string> = {
207 prompt: "A developer typed this request to a coding agent. Which model tier should answer it?",
208 task:
209 "A coding agent handed this task to a subagent of its own. Which model tier " +
210 "should the subagent run on?",
211 notification:
212 "A background task the coding agent started has finished and reported back; " +
213 "this is its report. The agent now works through the result and decides what " +
214 "to do next. Which model tier should do that?",
215};
216
217/**
218 * The request body for one routing decision: two questions Jev answers in
219 * parallel, the tier as a Choice and the effort as a Score.
220 *
221 * The model field is added by askJev depending on which provider is used.
222 */
223export function requestBodyOf(state: string, offered: readonly Tier[], source: StateSource = "prompt") {
224 const criteria: Record<string, string> = {};
225 for (const tier of offered) criteria[tier] = TIER_CRITERIA[tier];
226
227 return {
228 state,
229 questions: {
230 tier: {
231 type: "choice",
232 instructions: TIER_QUESTION[source],
233 criteria,
234 },
235 effort: {
236 type: "score",
237 instructions: "How much thinking does answering this request take?",
238 criteria: [...EFFORT_CRITERIA],
239 },
240 },
241 };
242}
243
244/**
245 * Asks Jev and answers with the response's `answers` object, or a reason.
246 *
247 * Every failure is still a pass for the turn, but it is a named one: the
248 * caller reports the reason rather than leaving the person guessing whether
249 * the router ran at all.
250 *
251 * On timeout, or on the caller's own `signal` aborting (the turn was ceded
252 * before this resolved), the turn moves on without the answer. The engine's
253 * `$.http.fetch` takes no abort signal, so the request itself runs to
254 * completion and is billed (about $0.00003) either way; the signal is
255 * passed for a plain `fetch`, as the scripts use, which does honour it.
256 */
257export async function askJev(args: AskArgs): Promise<JevResult> {
258 const {
259 fetch,
260 sleep,
261 provider,
262 state,
263 offered,
264 now = () => Date.now(),
265 } = args;
266 const timeoutMs = args.timeoutMs ?? DEFAULT_TIMEOUT_MS;
267 const started = now();
268 const since = () => now() - started;
269
270 if (!provider.ok) return { ok: false, reason: provider.reason, ms: 0 };
271 if (state.trim() === "") return { ok: false, reason: "empty prompt", ms: 0 };
272 if (offered.length === 0)
273 return { ok: false, reason: "no tiers offered", ms: 0 };
274 if (args.signal?.aborted) return { ok: false, reason: "ceded", ms: 0 };
275
276 const TIMED_OUT = Symbol("timed-out");
277 const CEDED = Symbol("ceded");
278 const controller = new AbortController();
279 // The caller's signal cancels the same in-flight request the timeout does,
280 // and resolves the race below the moment it fires.
281 let onCeded: (() => void) | undefined;
282 const ceded = new Promise<typeof CEDED>((resolve) => {
283 onCeded = () => {
284 controller.abort();
285 resolve(CEDED);
286 };
287 args.signal?.addEventListener("abort", onCeded);
288 });
289
290 const body = {
291 ...requestBodyOf(stateOf(state), offered, args.source),
292 model: provider.model,
293 };
294
295 // Called inside an async function, so a fetch that throws at once is a
296 // rejection handled below, not an escape past the finally.
297 const call = (async () =>
298 fetch(provider.endpoint, {
299 method: "POST",
300 headers: {
301 authorization: `Bearer ${provider.apiKey}`,
302 "content-type": "application/json",
303 },
304 body: JSON.stringify(body),
305 signal: controller.signal,
306 }))();
307 // Ended in the finally, so the timeout does not keep running (and a
308 // script's process alive) after the answer is in.
309 const timer = new AbortController();
310
311 let response: HttpResponseLike;
312 try {
313 const raced = await Promise.race([
314 call,
315 sleep(timeoutMs, { signal: timer.signal }).then(
316 () => TIMED_OUT,
317 () => new Promise<never>(() => {}),
318 ),
319 ceded,
320 ]);
321 if (raced === CEDED) {
322 void call.catch(() => undefined);
323 return { ok: false, reason: "ceded", ms: since() };
324 }
325 if (raced === TIMED_OUT) {
326 controller.abort();
327 void call.catch(() => undefined);
328 return {
329 ok: false,
330 reason: `timed out after ${timeoutMs}ms`,
331 ms: since(),
332 };
333 }
334 response = raced as HttpResponseLike;
335 } catch (error) {
336 if (args.signal?.aborted) {
337 return { ok: false, reason: "ceded", ms: since() };
338 }
339 if (controller.signal.aborted) {
340 return {
341 ok: false,
342 reason: `timed out after ${timeoutMs}ms`,
343 ms: since(),
344 };
345 }
346 return { ok: false, reason: `request failed: ${shortError(withoutKey(messageOf(error), provider.apiKey))}`, ms: since() };
347 } finally {
348 timer.abort();
349 if (onCeded) args.signal?.removeEventListener("abort", onCeded);
350 }
351
352 if (!response) return { ok: false, reason: "no response", ms: since() };
353
354 if (!response.ok) {
355 const who = provider.name;
356 return {
357 ok: false,
358 reason: `${who} said HTTP ${response.status}${providerNoteOf(response)}`,
359 ms: since(),
360 };
361 }
362
363 try {
364 const parsed = JSON.parse(response.text) as { answers?: unknown };
365 if (typeof parsed !== "object" || parsed === null || !parsed.answers) {
366 return { ok: false, reason: "response carried no answers", ms: since() };
367 }
368 return { ok: true, answers: parsed.answers, ms: since() };
369 } catch {
370 return { ok: false, reason: "response was not JSON", ms: since() };
371 }
372}
373
374/** The provider's own error type, when it sent one, for the status line. */
375export function providerNoteOf(response: HttpResponseLike): string {
376 try {
377 const body = JSON.parse(response.text) as { error?: { type?: string } };
378 const type = body?.error?.type;
379 // It lands in the reply's route line: a short, plain word or nothing.
380 if (typeof type !== "string") return "";
381 const plain = type.replace(/[^\w.-]/g, "").slice(0, 40);
382 return plain === "" ? "" : ` (${plain})`;
383 } catch {
384 return "";
385 }
386}
387hooks/label.ts 63 lines1/**
2 * The footer label. `SessionMode` draws the strings it is handed, so this
3 * file's only job is to turn a decision into one of them.
4 */
5
6import type { Decision } from "./policy.ts";
7
8/** Below this, the pick is marked so a bad route is visible rather than silent. */
9export const LOW_CONFIDENCE = 0.5;
10
11/** A previous router label left in SessionMode's modes list, old style or new. */
12const JEV_MODE = /^jev( →|:| off)/;
13
14/**
15 * The label for the footer, or null to add nothing.
16 *
17 * Null rather than a placeholder before the first turn: a footer that says
18 * nothing reads better than one that says the router has not run yet.
19 */
20export function labelOf(
21 decision: Decision | null,
22 enabled: boolean,
23 /** How the turn got its decision, as the attempt's `kind` says. */
24 kind?: "notify" | "agent" | "continue" | "nudge",
25 /** The decision was carried on from an earlier turn without asking Jev. */
26 continued?: boolean,
27): string | null {
28 if (!enabled) return "jev off";
29 if (!decision) return null;
30
31 // A named tier and a failed Jev call carry 0 as a placeholder, not a
32 // score: Jev was not asked, or did not answer, so there is no doubt to show.
33 // A held turn's confidence is Jev's in the tier it did not move to, and a
34 // continued turn's was an earlier turn's: neither is doubt about this one.
35 const scored =
36 !(decision.forced && decision.confidence === 0) &&
37 decision.jevFailed === undefined &&
38 decision.held === undefined &&
39 kind !== "continue" &&
40 kind !== "nudge" &&
41 continued !== true;
42 const doubt =
43 scored && decision.confidence < LOW_CONFIDENCE
44 ? `, only ${Math.round(decision.confidence * 100)}% sure`
45 : "";
46 return `jev: ${decision.tier}, ${decision.effort} effort${doubt}`;
47}
48
49/**
50 * The modes array `SessionMode` should draw, with our label on the end.
51 * Prior `jev → …` / `jev off` entries are stripped so crumbs do not
52 * accumulate across tier changes.
53 */
54export function withLabel(
55 modes: readonly string[],
56 label: string | null,
57): readonly string[] {
58 const cleared = modes.filter((m) => !JEV_MODE.test(m));
59 if (label === null) return cleared;
60 if (cleared.includes(label)) return cleared;
61 return [...cleared, label];
62}
63hooks/policy.ts 1095 lines1/**
2 * The routing policy: which tiers exist, what Jev is told each one is for,
3 * and how Jev's answers become a model and an effort level.
4 *
5 * Nothing here touches the engine or the network, so it runs under plain
6 * `node` in tests.
7 */
8
9import { sameModelAs,
10 baseModel,
11 fitsWindow,
12 tierOfModel,
13 type SwitchVerdict,
14} from "./pricing.ts";
15
16export type Tier = "haiku" | "sonnet" | "opus" | "fable";
17
18/** The efforts the engine accepts, low to high. There is no rung above max. */
19export type Effort = "low" | "medium" | "high" | "xhigh" | "max";
20
21export type Decision = {
22 tier: Tier;
23 model: string;
24 effort: Effort;
25 /** Jev's confidence in the tier, 0 to 1. The gateway rounds to 2 places. */
26 confidence: number;
27 /**
28 * Jev's confidence in the effort score, separately. Measured 2026-09-22:
29 * the two move independently (tier 0.81 with effort 0.49 on the same
30 * prompt), and effort confidence is lowest on terse follow-ups, which is
31 * where an effort flip is least worth paying for. 0 when the answer
32 * carried none.
33 */
34 effortConfidence?: number;
35 /**
36 * The tier Jev named, when stickiness kept the turn on the previous one
37 * instead. Absent on a turn that went where Jev pointed. Kept so the route
38 * line can say a hold happened; a hold nobody can see is indistinguishable
39 * from a router that is not running.
40 */
41 held?: Tier;
42 /**
43 * The model Jev's tier would have run on, when held. Usually implied by
44 * `held`; it differs when the session runs a model off the ladder
45 * (`claude-opus-5`) and Jev named the same tier (`claude-opus-5-5`): the
46 * same rung, a different cache.
47 */
48 heldModel?: string;
49 /**
50 * The two prices a held downgrade was decided between, when it was the
51 * cost of the switch and not Jev's doubt that held it. Absent otherwise.
52 */
53 heldCost?: { stay: number; go: number; limit?: number };
54 /**
55 * The tier the turn had to leave because the context no longer fits it,
56 * when the move went only as far up as it had to (not to Jev's pick).
57 */
58 outgrew?: Tier;
59 /** Jev's pick, when the turn moved up only part of the way to it. */
60 wanted?: Tier;
61 /**
62 * Why Jev gave no answer (a timeout, an error), when the turn stayed on the
63 * tier already running instead of dropping to the session model.
64 */
65 jevFailed?: string;
66 /**
67 * The context this turn carries, when that is what held it: the tier Jev
68 * named cannot take a prompt this long at all. Absent otherwise.
69 */
70 heldWindow?: number;
71 /** The confidence the switch needed, when Jev's doubt is what held it. */
72 heldBar?: number;
73 /**
74 * The effort Jev named, when a turn staying on Sonnet kept the previous
75 * turn's effort instead (see `holdsSonnetEffort`). Absent otherwise.
76 */
77 heldEffort?: Effort;
78 /**
79 * The tier was named in the prompt itself ("use opus"), so Jev's tier
80 * answer was set aside and stickiness did not get a vote. Jev is not asked
81 * at all, so it runs at medium effort. Shown on the route line, since a forced turn at 43%
82 * would otherwise read as a low-confidence pick.
83 */
84 forced?: true;
85 /**
86 * The effort Jev named, when the tier's ceiling capped this turn. Absent
87 * when no cap applied.
88 */
89 cappedEffort?: Effort;
90 /**
91 * Jev's probability for each offered tier, when the answer carried them.
92 * They sum to one; `confidence` is derived from them when the provider
93 * sends none (the Vercel gateway does not).
94 */
95 probabilities?: Partial<Record<Tier, number>>;
96 /**
97 * The effort Jev asked for, when `effort` is instead what the engine runs
98 * on a conversation's first turn (see `FIRST_TURN_EFFORT`). Absent when
99 * the two agree.
100 */
101 askedEffort?: Effort;
102};
103
104export const TIERS: readonly Tier[] = ["haiku", "sonnet", "opus", "fable"];
105
106export const EFFORTS: readonly Effort[] = [
107 "low",
108 "medium",
109 "high",
110 "xhigh",
111 "max",
112];
113
114/** Model ids as the engine names them. */
115export const MODEL_OF: Record<Tier, string> = {
116 haiku: "claude-haiku-4-5",
117 sonnet: "claude-sonnet-5-5",
118 opus: "claude-opus-5-5",
119 fable: "claude-fable-5-1",
120};
121
122/**
123 * What Jev is told each tier is for. This is the policy: edit these lines to
124 * change how the router behaves, and nothing else.
125 */
126export const TIER_CRITERIA: Record<Tier, string> = {
127 haiku:
128 "Trivial. A lookup, a rename, a yes or no question, reading one short file, " +
129 "restating something already on screen.",
130 sonnet:
131 "Straightforward and minor. A small edit whose shape is already obvious from " +
132 "the request, with no real decision to make.",
133 opus:
134 "Plain implementation carrying some complexity. Writing or changing real code, " +
135 "possibly across a few files, where the approach is known but the work is not " +
136 "mechanical.",
137 fable:
138 "High complexity needing higher-order reasoning. Planning, brainstorming, " +
139 "architecture, systematic debugging, weighing trade-offs, research. Anything " +
140 "where working out the approach is itself the hard part.",
141};
142
143/**
144 * Ordered low to high; the index Jev scores is the effort level. Written as
145 * situations rather than degrees, which is what TypeSafe's guidance for a
146 * score question asks for ("with numbers only, the model has nothing to
147 * match against and splits the probability"). Five levels, one per effort
148 * the engine accepts.
149 */
150export const EFFORT_CRITERIA: readonly string[] = [
151 "The answer is already known or on screen: a lookup, a rename, a yes or " +
152 "no, restating something.",
153 "One or two obvious steps: a small edit whose shape the request already " +
154 "gives, a short explanation.",
155 "Several steps that have to fit together, or a choice worth weighing: " +
156 "real code across a file or two, a bug with a likely cause.",
157 "Many interacting parts, or a subtle failure to chase down: a change " +
158 "across several files, a bug with no obvious cause, a design with " +
159 "trade-offs.",
160 "Open-ended or ambiguous, or the cost of being wrong is high: " +
161 "architecture, a systematic debugging campaign, a migration plan, " +
162 "anything where the approach itself is the hard part.",
163];
164
165/** Tiers dropped from the question entirely, lowercase, from the env var. */
166export function excludedTiers(raw: string | undefined): Set<Tier> {
167 // Commas, semicolons or spaces between the names.
168 const names = (raw ?? "")
169 .split(/[\s,;]+/)
170 .map((s) => s.trim().toLowerCase())
171 .filter(Boolean);
172 return new Set(
173 names.filter((n): n is Tier => (TIERS as string[]).includes(n)),
174 );
175}
176
177/** The tiers offered to Jev, in ladder order, never empty. */
178export function offeredTiers(excluded: Set<Tier>): Tier[] {
179 const kept = TIERS.filter((t) => !excluded.has(t));
180 return kept.length > 0 ? [...kept] : [...TIERS];
181}
182
183/**
184 * `offered` paired with the `excluded` list that actually matches it. Naming
185 * every tier excluded (a misconfigured `JEV_ROUTER_EXCLUDE`, or a matching
186 * combination of `/jev tiers off` calls before the last-tier guard existed)
187 * makes `offeredTiers` fall back to the full ladder rather than nothing —
188 * every caller that stores or displays `excluded` alongside `offered` needs
189 * the two to agree, or `/jev` can end up saying a tier is both offered and
190 * excluded.
191 */
192export function tierFilter(excluded: Iterable<Tier>): {
193 offered: Tier[];
194 excluded: Tier[];
195} {
196 const set = excluded instanceof Set ? excluded : new Set(excluded);
197 const offered = offeredTiers(set);
198 return { offered, excluded: offered.length === TIERS.length ? [] : [...set] };
199}
200
201type ChoiceAnswer = {
202 type: "choice";
203 choice?: unknown;
204 confidence?: unknown;
205 probabilities?: unknown;
206};
207type ScoreAnswer = { type: "score"; score?: unknown; confidence?: unknown };
208
209function isRecord(v: unknown): v is Record<string, unknown> {
210 return typeof v === "object" && v !== null;
211}
212
213/** Jev's score across EFFORT_CRITERIA to the nearest effort level. */
214export function effortOf(score: unknown): Effort {
215 if (typeof score !== "number" || !Number.isFinite(score)) return "medium";
216 const i = Math.min(Math.max(Math.round(score), 0), EFFORTS.length - 1);
217 return EFFORTS[i] ?? "medium";
218}
219
220/**
221 * Turns the `answers` object of a Jev response into a decision.
222 *
223 * Returns null whenever the answer is missing, malformed, or names a tier
224 * that was not offered: the caller then leaves the turn alone.
225 */
226export function decisionOf(
227 answers: unknown,
228 offered: readonly Tier[] = TIERS,
229): Decision | null {
230 if (!isRecord(answers)) return null;
231
232 const tier = answers.tier as ChoiceAnswer | undefined;
233 if (!isRecord(tier) || tier.type !== "choice") return null;
234
235 const choice = tier.choice;
236 if (typeof choice !== "string") return null;
237 if (!offered.includes(choice as Tier)) return null;
238
239 const effort = answers.effort as ScoreAnswer | undefined;
240 // Clamped to 0..1: a Decision's confidence is trusted as a probability
241 // everywhere it is read, and `persist.ts`'s unpack validation now rejects
242 // one that is not — a provider that ever sends something outside that
243 // range (measured possible, not measured live) would otherwise route on
244 // it live and then have the whole Decision silently dropped on restore.
245 const confidenceOf = (v: unknown) =>
246 typeof v === "number" && Number.isFinite(v)
247 ? Math.min(1, Math.max(0, v))
248 : 0;
249
250 const probabilities = probabilitiesOf(tier.probabilities, offered);
251 const confidence =
252 typeof tier.confidence === "number" && Number.isFinite(tier.confidence)
253 ? Math.min(1, Math.max(0, tier.confidence))
254 : confidenceFrom(probabilities, offered.length, choice as Tier);
255
256 return {
257 tier: choice as Tier,
258 model: MODEL_OF[choice as Tier],
259 effort: effortOf(isRecord(effort) ? effort.score : undefined),
260 confidence,
261 effortConfidence: confidenceOf(isRecord(effort) ? effort.confidence : 0),
262 ...(probabilities !== undefined ? { probabilities } : {}),
263 };
264}
265
266/** The per-tier probabilities from a choice answer, offered tiers only. */
267function probabilitiesOf(
268 raw: unknown,
269 offered: readonly Tier[],
270): Partial<Record<Tier, number>> | undefined {
271 if (!isRecord(raw)) return undefined;
272 const out: Partial<Record<Tier, number>> = {};
273 let any = false;
274 for (const tier of offered) {
275 const p = raw[tier];
276 if (typeof p === "number" && Number.isFinite(p)) {
277 out[tier] = Math.min(1, Math.max(0, p));
278 any = true;
279 }
280 }
281 return any ? out : undefined;
282}
283
284/**
285 * TypeSafe's confidence, from the probabilities, for a provider that sends
286 * none. Their documented measure is how far the mass sits on one option:
287 * all of it gives 1, an even spread gives 0, and their worked example
288 * (0.85 / 0.15 / 0 → 0.78) is `(n·p_max − 1) / (n − 1)` for n options. Same
289 * scale as the confidence the direct API sends, so the sticky bar means the
290 * same thing on either provider.
291 *
292 * With `chosen`, the mass is the chosen tier's own, not the largest: an
293 * answer whose choice and probabilities disagree (haiku chosen at 0.05,
294 * fable at 0.95) is not sure of haiku, and must not clear a bar as if it were.
295 */
296export function confidenceFrom(
297 probabilities: Partial<Record<Tier, number>> | undefined,
298 options: number,
299 chosen?: Tier,
300): number {
301 if (probabilities === undefined) return 0;
302 const values = Object.values(probabilities).filter(
303 (v): v is number => typeof v === "number",
304 );
305 if (values.length === 0) return 0;
306 const p = chosen === undefined ? Math.max(...values) : (probabilities[chosen] ?? 0);
307 const clamp = (n: number) => Math.min(1, Math.max(0, n));
308 if (options <= 1) return clamp(p);
309 return clamp((options * p - 1) / (options - 1));
310}
311
312/**
313 * The confidence a switch must clear before the model moves, when stickiness
314 * is on. Jev's confidence is its top probability, normalised so an even
315 * spread reads 0: with four tiers, 0.75 means the named tier holds about 81%
316 * of the mass. 0.75 is a starting point, not a measured optimum; retune with
317 * `/jev sticky` or `npm run try-prompts`.
318 *
319 * The bar is one of two things a shaky switch has to clear. The other is the
320 * price: a downgrade whose cold cache write costs more than the turn would
321 * cost on the tier already warm is held whatever Jev's confidence
322 * (`switchVerdict` in pricing.ts, with the context size from the engine).
323 */
324export const DEFAULT_STICKY_CONFIDENCE = 0.75;
325
326/**
327 * Whether stickiness is on. **On by default** (unset/empty). Opt out with
328 * `0`/`false`/`off`/`no`/`none`; opt in explicitly with `1`/`true`/`yes`/`on`.
329 */
330export function stickyOf(raw: string | undefined): boolean {
331 // Unset, explicit on, or anything else → on (session default).
332 return !flagOff(raw);
333}
334
335/**
336 * The bar from the environment, or the default when it is unusable.
337 *
338 * A value above 1 is read as a percentage, since `JEV_ROUTER_STICKY_CONFIDENCE=80`
339 * is the likelier intent than a bar no turn can ever clear. 0 and 1 are both
340 * refused: one would hold every switch forever, the other would hold none,
341 * and each is better said by leaving the flag off.
342 */
343export function thresholdOf(raw: string | undefined): number {
344 return confidenceShareOf(raw) ?? DEFAULT_STICKY_CONFIDENCE;
345}
346
347/**
348 * A confidence bar as a share strictly between 0 and 1, or null. `0.6`,
349 * `60` and `60%` are the same bar; anything written with `%` is a
350 * percentage, so `0.5%` is half a percent, not half. Without `%`, a number
351 * past 1 is a percentage, but one between 1 and 10 must be whole: `1.5` is
352 * refused rather than read as 1.5%, a bar so low it is as good as none.
353 * Plain decimals only (`Number` alone reads `0x40` as 64).
354 */
355export function confidenceShareOf(raw: string | undefined): number | null {
356 const trimmed = (raw ?? "").trim();
357 const percent = trimmed.endsWith("%");
358 const v = percent ? trimmed.slice(0, -1).trim() : trimmed;
359 if (!PLAIN_DECIMAL.test(v)) return null;
360 const n = Number(v);
361 // `1.5` could be 1.5% or a slip for 0.15; `60.5` can only be a percentage.
362 if (!percent && n > 1 && n < 10 && !Number.isInteger(n)) return null;
363 const ratio = percent || n > 1 ? n / 100 : n;
364 return ratio > 0 && ratio < 1 ? ratio : null;
365}
366
367/**
368 * Holds a switch on the tier the last turn used when Jev is not sure enough
369 * of it, or when the move costs more than it is worth (`verdict`, computed
370 * by the caller from the context size: for a downgrade, whether it saves
371 * anything; for an upgrade, whether it costs more than the upgrade limit
372 * over staying; null when price checks are off or nothing is known).
373 *
374 * Only the model is held here. The effort Jev asked for is applied either
375 * way: on Opus and Haiku it is sent per request and costs no cache, so a
376 * held turn still gets to think harder or less hard than the one before it.
377 * Sonnet is the exception, and `holdsSonnetEffort` handles it separately.
378 *
379 * `previous` is the tier the last routed turn ran on, or null on the first
380 * turn of a session, which has nothing to hold to.
381 *
382 * `offered` refuses to hold on a `previous` that has since been turned off
383 * with `/jev tiers off` — but only when `previous` was itself a real routed
384 * decision (`effortConfidence` set). A `previous` seeded only as a
385 * placeholder from the session model (nothing ever routed there) is exempt:
386 * an unrouted turn runs on that same placeholder anyway, so refusing to
387 * weigh it against a switch's real cost does not stop the plugin from
388 * "using" the tier — the tier is not being used *by a choice this plugin
389 * made* either way — and it does force a switch whose cache-write cost can
390 * run many times what staying would have, for no benefit.
391 */
392export function stickyDecision(
393 fresh: Decision,
394 previous: Decision | null,
395 threshold: number,
396 verdict: SwitchVerdict | null = null,
397 offered: readonly Tier[] = TIERS,
398): Decision {
399 if (previous === null) return fresh;
400 // The model, not the tier: a session on `claude-opus-5` that Jev keeps on
401 // opus is still a switch, to `claude-opus-5-5` and a cold cache. The
402 // engine's `[1m]` suffix is not a different model, and the session's own
403 // spelling is what is sent back, so nothing changes under it.
404 // A provider's spelling (`…@date`, `us.anthropic.…`) is the same model too.
405 if (sameModelAs(fresh.model, previous.model))
406 return fresh.model === previous.model
407 ? fresh
408 : { ...fresh, model: previous.model };
409 const shaky = fresh.confidence < threshold;
410 const unprofitable = verdict !== null && verdict.hold;
411 const droppedTier =
412 !offered.includes(previous.tier) && previous.effortConfidence !== undefined;
413 if ((!shaky && !unprofitable) || droppedTier) return fresh;
414 return {
415 tier: previous.tier,
416 model: previous.model,
417 effort: fresh.effort,
418 confidence: fresh.confidence,
419 // Sonnet effort gating reads this next; dropping it made every held
420 // Sonnet turn look like effort confidence 0 and always hold effort.
421 effortConfidence: fresh.effortConfidence,
422 ...(fresh.probabilities !== undefined
423 ? { probabilities: fresh.probabilities }
424 : {}),
425 held: fresh.tier,
426 heldModel: fresh.model,
427 ...(shaky && !unprofitable ? { heldBar: threshold } : {}),
428 ...(unprofitable
429 ? {
430 heldCost: {
431 stay: verdict.stay,
432 go: verdict.go,
433 ...(verdict.limit !== undefined ? { limit: verdict.limit } : {}),
434 },
435 }
436 : {}),
437 };
438}
439
440/**
441 * Keeps a turn off a tier whose window it does not fit: on the tier already
442 * running when that one takes it, otherwise nowhere (null), so the caller's
443 * own step-up runs instead. A `use haiku` at 300k is refused the same way;
444 * the API would refuse it with "Prompt is too long", and did, three times in
445 * a week. `offered` refuses `previous` as a landing spot when it has been
446 * turned off since it started running, even though it still fits — the
447 * caller still sees `previous` was non-null (nothing here nulls it), so its
448 * own step-up runs from `decision.tier` rather than giving up outright.
449 */
450export function withinWindow(
451 decision: Decision,
452 previous: Decision | null,
453 contextTokens: number,
454 offered: readonly Tier[] = TIERS,
455): Decision | null {
456 if (fitsWindow(decision.tier, contextTokens)) return decision;
457 if (
458 previous === null ||
459 !fitsWindow(previous.tier, contextTokens) ||
460 !offered.includes(previous.tier)
461 )
462 return null;
463 return {
464 tier: previous.tier,
465 model: previous.model,
466 effort: decision.effort,
467 confidence: decision.confidence,
468 effortConfidence: decision.effortConfidence,
469 ...(decision.probabilities !== undefined
470 ? { probabilities: decision.probabilities }
471 : {}),
472 held: decision.tier,
473 heldModel: decision.model,
474 heldWindow: contextTokens,
475 // A named tier that does not fit is still the person's pick.
476 ...(decision.forced ? { forced: true as const } : {}),
477 };
478}
479
480/**
481 * The decision a session is already running on, for the turns the router
482 * did not route: a resumed session, `/jev on` after a stretch off, a plugin
483 * loaded into a live session. Its cache is what the first routed turn's
484 * switch is priced against. Null for a model off the ladder.
485 */
486export function sessionDecision(model: string): Decision | null {
487 // A spelling a snapshot could not hold (a control character, markdown, a
488 // runaway length) is not adopted: saved, it would lose the whole state.
489 if (!MODEL_ID.test(model)) return null;
490 // An alias that names no one model (`opusplan` runs Sonnet outside plan
491 // mode; `default` is whatever the account gets) makes no placeholder, so
492 // nothing is held to a model not running. A tier's own alias (`opus`,
493 // `sonnet[1m]`) runs that tier's model, and stands for it.
494 if (AMBIGUOUS_ALIAS.test(model)) return null;
495 const plainAlias = TIER_ALIAS.exec(model);
496 if (plainAlias) {
497 const tier = plainAlias[1]!.toLowerCase() as Tier;
498 return { tier, model: MODEL_OF[tier], effort: "medium", confidence: 1 };
499 }
500 const tier = tierOfModel(model);
501 if (tier === null) return null;
502 return { tier, model, effort: "medium", confidence: 1 };
503}
504
505/**
506 * A model id as the engine spells one — `claude-opus-5-5[1m]`, a Bedrock id
507 * or ARN, a Vertex path, `+build` — and nothing the route line (a rendered
508 * blockquote) would read as markdown: no spaces, parentheses or emphasis.
509 */
510export const MODEL_ID = /^[\w.:/@+\[\]-]{1,512}$/;
511
512/** A tier's own alias (`opus`, `sonnet[1m]`): its tier is known, its version is not. */
513export const TIER_ALIAS = /^(opus|sonnet|haiku|fable)(?:\[1m\])?$/i;
514
515/** A model alias that names no one model: a placeholder cannot be made of it. */
516export const AMBIGUOUS_ALIAS = /^(?:opusplan|default|best)(?:\[1m\])?$/i;
517
518/**
519 * A turn the engine started, not the person: its "say what you are doing,
520 * then continue" nudge when a turn has run long without a reply. Jev would
521 * grade the nudge's text (opus at 46%, measured 2026-09-23) and move the
522 * model under a task that is mid-flight; the turn continues instead.
523 */
524const NUDGE = /^\s*The user hasn't heard from you in a while/i;
525
526export function isEngineNudge(text: string): boolean {
527 return NUDGE.test(normalizeQuotes(text));
528}
529
530/**
531 * A bare go-ahead: the person is answering the previous turn, not starting a
532 * task. Jev reads these as trivial with near-total confidence ("yes" 1.00,
533 * "y" 0.98, "go ahead" 0.79, measured 2026-09-22), which is right about the
534 * text and wrong about the work. Stickiness cannot catch this, since its bar
535 * is a confidence and these clear any bar. The list is intentionally narrow
536 * — bare `k`/`go`/`next`/`approved` used to false-positive on real tasks.
537 * Trailing punctuation (`.`, `!`, `?`, `,`) is tolerated; anything longer is
538 * a real prompt and goes to Jev.
539 */
540const CONTINUATION =
541 /^(?:y|yes|yep|yeah|yup|ok|okay|sure|go ahead|go on|go for it|proceed|continue|carry on|do it|ok do it|let'?s do it|please do|yes please|sounds good|lgtm)[\s.!,?]*$/i;
542
543export function isContinuation(text: string): boolean {
544 return CONTINUATION.test(normalizeQuotes(text).trim());
545}
546
547/**
548 * Whether natural-language tier overrides ("use opus") are honored.
549 * On by default; `JEV_ROUTER_ALLOW_OVERRIDE=0` disables them.
550 */
551export function overrideAllowedOf(raw: string | undefined): boolean {
552 return !flagOff(raw);
553}
554
555/**
556 * Whether a task-notification turn continues the previous route instead of
557 * asking Jev. On by default: the turn's text is the engine's XML about a
558 * finished background task, not work to grade, and the reply it wakes is
559 * the one already under way. Skipping Jev there saves a round trip on the
560 * critical path of every task that finishes. `JEV_ROUTER_NOTIFY_CONTINUE=0`
561 * asks Jev anyway.
562 */
563export function notifyContinueOf(raw: string | undefined): boolean {
564 return !flagOff(raw);
565}
566
567/** The words every on/off setting reads as off: `0`, `false`, `no`, `off`, `none`. */
568export function flagOff(raw: string | undefined): boolean {
569 const flag = (raw ?? "").trim().toLowerCase();
570 return flag === "0" || flag === "false" || flag === "no" || flag === "off" || flag === "none";
571}
572
573/** A plain decimal (`12`, `0.5`, `1000.`, `.5`): no sign, hex or exponent, the same for every setting. */
574export const PLAIN_DECIMAL = /^(?:\d+(?:\.\d*)?|\.\d+)$/;
575
576/**
577 * A tier named in the prompt: "use opus", "go with fable", "switch to haiku",
578 * "run this on sonnet", "do it using opus". Only verbs that actually mean
579 * "run on" are accepted; bare "on"/"for"/"with"/"using" are not (they turned
580 * "happy with opus" and "I'm using opus for comparison" into routes). A bare
581 * tier glued to another word ("sonnet-level") is not a name either. Negations
582 * skip only the first run-on *or* bare `using <tier>` after them, so
583 * "stop using haiku and use opus" still forces opus. A model id names its
584 * tier too. Returns the tier, or null when none is named or not offered.
585 */
586/** The verbs that route, shared by OVERRIDE and the backticked-tier unwrap in ownWords. */
587const ROUTE_VERB =
588 "use|do (?:it |this )?using|switch(?:ing)?(?: over| back)?(?: (?:the )?model)? to|route to|run (?:it |this )?on|go with";
589
590const OVERRIDE = new RegExp(
591 `\\b(?:${ROUTE_VERB})\\s+(?:claude-)?(haiku|sonnet|opus|fable)(?:-\\d+)*(?![\\w-])`,
592 "gi",
593);
594
595/**
596 * Bare `using <tier>` is not an override, but it can absorb a negation so a
597 * later affirmative is not wrongly skipped ("stop using haiku and use opus").
598 */
599const USING_SINK =
600 /\busing\s+(?:claude-)?(?:haiku|sonnet|opus|fable)(?:-\d+)*(?![\w-])/gi;
601
602/** Bare `<tier>` after `avoid`/`stop` only ("avoid haiku and use opus"). */
603const BARE_TIER_SINK =
604 /\b(?:claude-)?(?:haiku|sonnet|opus|fable)(?:-\d+)*(?![\w-])/gi;
605
606/**
607 * Negation starters. Bare `\bnot` is omitted: "why not use opus" is
608 * affirmative. `never mind` is omitted (`never(?!\s+mind)`). Every modal's
609 * contracted AND spaced form is included ("shouldn't"/"should not",
610 * "won't"/"will not", ...) — a prior version had the contractions but
611 * missed the spaced form for should/would/could/will, so "we should not use
612 * haiku" read as affirmative and forced the very tier it refused.
613 */
614const OVERRIDE_NEGATION_AT =
615 /\b(?:do\s*n'?t|doesn'?t|didn'?t|won'?t|will\s+not|wouldn'?t|would\s+not|shouldn'?t|should\s+not|mustn'?t|couldn'?t|could\s+not|can(?:'?t|not|\s+not)|never(?!\s+mind)|avoid|stop|do\s+not|must\s+not|may\s+not)\b/gi;
616
617/**
618 * Words allowed between a negation and its target. Anything else (you, what,
619 * doing, and, …) means the negation is discourse/rhetorical, not "don't use".
620 */
621const NEGATION_BRIDGE =
622 /^(?:\s+(?:want|to|try|ever|really|please|just|even|still|actually|also|need|have|you\s+to))*\s*$/i;
623
624/** Fold typographic apostrophes so iOS/macOS quotes match the ASCII forms. */
625function normalizeQuotes(text: string): string {
626 return text.replace(/[‘’ʼ]/g, "'");
627}
628
629/**
630 * The words of a prompt that are the person's own: without pasted content
631 * (the engine wraps it in `<pasted_content>` tags), code blocks and spans,
632 * and quoted lines. A handoff or log pasted in can say "use opus" as an
633 * example; that is not an instruction to route there.
634 */
635/**
636 * `text` without its `<tag …>…</tag>` blocks, found by hand: the pattern
637 * this replaces rescanned to the end from every unclosed opening (a paste
638 * of 100k `<pasted_content ` took half a second).
639 */
640function withoutBlocks(text: string, tag: string): string {
641 const OPEN = `<${tag}`;
642 const CLOSE = `</${tag}`;
643 let out = "";
644 let pos = 0;
645 let at = text.indexOf(OPEN);
646 while (at !== -1) {
647 const next = text[at + OPEN.length];
648 if (next !== undefined && /\w/.test(next)) {
649 at = text.indexOf(OPEN, at + 1);
650 continue;
651 }
652 const openEnd = text.indexOf(">", at);
653 if (openEnd === -1) break;
654 const close = text.indexOf(CLOSE, openEnd + 1);
655 if (close === -1) break;
656 const closeEnd = text.indexOf(">", close);
657 if (closeEnd === -1) break;
658 out += `${text.slice(pos, at)} `;
659 pos = closeEnd + 1;
660 at = text.indexOf(OPEN, pos);
661 }
662 return out + text.slice(pos);
663}
664
665export function ownWords(text: string): string {
666 return withoutBlocks(withoutBlocks(text, "pasted_content"), "task-notification")
667 .replace(/```[\s\S]*?(?:```|$)/g, " ")
668 // A tier alone in backticks after a route verb is the person's own ask
669 // ("use `opus`"), not code: unwrapped before code spans go.
670 .replace(new RegExp(`\\b(${ROUTE_VERB})\\s+\`((?:claude-)?(?:haiku|sonnet|opus|fable)(?:-\\d+)*)\``, "gi"), "$1 $2")
671 .replace(/`[^`\n]*`/g, " ")
672 // A code comment on a line of its own, outside a fence: `// use opus`.
673 .replace(/^[ \t]*\/\/.*$/gm, " ")
674 // A phrase in double quotes is being quoted, not said: "use opus".
675 .replace(/"[^"\n]{1,200}"/g, " ")
676 .replace(/\u201c[^\u201d\n]{1,200}\u201d/g, " ")
677 // Single quotes too, around a short phrase with no punctuation inside:
678 // the README says 'use opus'. An apostrophe (don't, 'em, users') is not
679 // a quote: it does not both open after a space and close before one
680 // around a phrase that short and plain.
681 .replace(/(^|[\s(])['\u2018][^'\u2018\u2019\n.,;:!?]{1,60}['\u2019](?=[\s.,;:!?)]|$)/g, "$1 ")
682 .replace(/^[ \t]*>.*$/gm, " ");
683}
684
685/** How much of the person's own words a named tier is looked for in. */
686const OWN_WORDS_MAX = 20_000;
687
688export function parseOverride(
689 text: string,
690 offered: readonly Tier[] = TIERS,
691): Tier | null {
692 // A tier named in a prompt past this many characters is in a paste the
693 // engine did not mark; the person's own ask is at the start or the end.
694 const own = normalizeQuotes(ownWords(text));
695 // Joined with a sentence break, so a phrase cannot form across the cut
696 // ("…use" + "opus…" from "user" and "octopus").
697 const normalized = own.length <= OWN_WORDS_MAX ? own : `${own.slice(0, OWN_WORDS_MAX / 2)}\n.\n${own.slice(-OWN_WORDS_MAX / 2)}`;
698 const matches = [...normalized.matchAll(OVERRIDE)];
699 const negated = negatedAt(normalized, matches);
700 let named: Tier | null = null;
701 for (const [i, match] of matches.entries()) {
702 const at = match.index ?? 0;
703 if (negated.has(at)) continue;
704 const previous = matches[i - 1];
705 const from = previous ? (previous.index ?? 0) + previous[0].length : 0;
706 if (!addressedAt(normalized, from, at)) continue;
707 if (TIER_AS_NAME.test(normalized.slice(at + match[0].length, at + match[0].length + 40))) continue;
708 const tier = match[1]?.toLowerCase() as Tier | undefined;
709 if (tier !== undefined && offered.includes(tier)) named = tier;
710 }
711 return named;
712}
713
714/**
715 * What may stand ahead of the verb, between the start of its clause and the
716 * verb, for the phrase to be said to the model. Measured against a labelled
717 * set of prompts (tests/fixtures/override-corpus.ts): refusing only what
718 * looks like talk (a deny-list) let through prose with any subject not
719 * listed ("anyone can use opus", "they want to use opus", "the job will
720 * switch to haiku"), so this lists what a request opens with instead —
721 * softeners, acknowledgements, scope ("for the migration", "this time"),
722 * and the ways of asking — and anything else is talk about a tier. The cost
723 * of a miss is a turn left to Jev; of a false match, a forced switch with no
724 * checks, which is the one to avoid.
725 */
726const OPENER = new RegExp(
727 "^(?:" +
728 [
729 // Softeners and acknowledgements.
730 "please|pls|plz|pleae|kindly|pretty please|just|now|then|so|ok|okay|kk|cool|oh|hey|hi|yes|yeah|yep|yup|sure|hmm+|um+|well|again",
731 "maybe|perhaps|actually|instead|also|and|but|or|rather|here|claude|nope|no|alright|right",
732 "fine|anyway|honestly|really|definitely|ideally|tbh|always|only|probably|better",
733 // Addressing the model by name: `@claude`.
734 "@[\\w-]+",
735 // Scope.
736 // One word after "for the": a second is the clause's own subject
737 // ("for these tasks people use haiku" is talk).
738 "this time|for this one|for this|for now|from now on|going forward|for the rest of (?:the|this) [\\w-]+|for (?:the|this|that|these|those|each|every|all) [\\w-]+",
739 // Asking.
740 "let'?s|let us|let me|go ahead and|i want you to|i want to|we want to|i'?d like (?:you )?to|i would like (?:you )?to",
741 "i need you to|we need to|you need to|i'?d rather you|i would rather you|i'?d prefer (?:(?:that |if )?you)?|i think (?:you|we) should",
742 "you should|u should|we should|you can|you may|you could|can you|can u|could you|would you|will you|can we|could we|shall we",
743 "feel free to|you'?re free to|make sure (?:to|you)|remember to|be sure to|try to|time to|it'?s time to",
744 "(?:please )?don'?t hesitate to|i'?m going to ask you to|i'?m asking you to|i said(?: to)?|wouldn'?t hurt to",
745 "you might as well|might as well|you might want to",
746 // Tag questions that ask for it.
747 "why not|why don'?t you|can'?t you|won'?t you|couldn'?t you|wouldn'?t you",
748 ].join("|") +
749 ")(?: |$)",
750);
751
752/** A list marker opening the clause: `-`, `*`, `+`, `•`, `- [ ]`, `1.`, `1)`, `(1)`, `a)`. */
753const BULLET = /^(?:[-*+•](?:\s*\[[ x]?\])?|\[[ x]?\]|\(?(?:\d+|[a-z])[.)])\s*/;
754
755/** A clause break: sentence ends, commas, dashes, ellipses, a new line, a joining and/then/but. */
756const CLAUSE_BREAK = /[.!?,;:\n—–…]|\s-\s|\s(?:and|then|but)\s/gi;
757
758/**
759 * Whether the verb at `at` asks the model to run on the tier: the text from
760 * the last clause break (or the end of the previous route phrase, `from`) up
761 * to the verb is nothing but `OPENER`s.
762 */
763function addressedAt(text: string, from: number, at: number): boolean {
764 let start = from;
765 for (const brk of text.slice(from, at).matchAll(CLAUSE_BREAK))
766 start = from + (brk.index ?? 0) + brk[0].length;
767 // A clause under a condition describes what happens then, not what to do
768 // now: "if it runs long, switch to opus", "otherwise use opus".
769 let sentence = from;
770 for (const brk of text.slice(from, start).matchAll(/[.!?\n]/g)) sentence = from + (brk.index ?? 0) + 1;
771 // Any clause of the sentence so far: "Add a fallback: if it times out, …".
772 for (const part of text.slice(sentence, start).toLowerCase().split(/[,;:—–]|\s-\s/)) {
773 const clause = part.trim().replace(BULLET, "");
774 if (CONDITION.test(clause) && !POLITE_CONDITION.test(clause) && !SET_PHRASE.test(clause))
775 return false;
776 }
777 let lead = text.slice(start, at).trim().toLowerCase().replace(/\s+/g, " ").replace(BULLET, "");
778 for (let guard = 0; lead !== "" && guard < 12; guard++) {
779 const m = lead.match(OPENER);
780 if (m === null) return false;
781 lead = lead.slice(m[0].length).trimStart();
782 }
783 return lead === "";
784}
785
786/** A clause that sets a condition, ahead of the one naming the tier. */
787const CONDITION = /^(?:if|when|whenever|unless|once|until|in case|otherwise|else)\b/;
788
789/** Set phrases that are not conditions on anything: "once again", "if needed". */
790const SET_PHRASE =
791 /^(?:once (?:again|more)|if (?:so|not)|(?:if|when) in doubt|if that'?s the case|whenever|until the end of (?:this|the) (?:session|conversation|task|chat)|if so|if (?:needed|necessary|possible|required|appropriate|applicable)|when(?:ever)? (?:done|ready|finished|possible)|until further notice|if (?:that|this|it)(?:'?s| is)? (?:ok|okay|fine|alright|all right|not too much trouble)(?: with \w+)?)\s*(?:then)?\s*$/;
792
793/**
794 * A condition that is only manners, or the person's say-so, not a state of
795 * things: "if you can", "if you want", "whenever you're ready", "unless you
796 * disagree", "until I say otherwise". A listed phrase and nothing more: "if
797 * you get a 429" or "if my repo is large" describes behaviour.
798 */
799const POLITE_CONDITION = new RegExp(
800 "^(?:if|when|whenever|unless|until)\\s+(?:" +
801 [
802 "(?:you|u|ya)\\s+(?:can|could|would|will|may|might|want(?: to)?|like|wish|prefer|please|must|disagree|object|agree|think so|think otherwise|see fit|are able|are ready|are free|are willing|feel like it|don'?t mind|do not mind|wouldn'?t mind|would not mind|could please|would please|get (?:a|the) chance|have (?:a )?(?:sec|second|moment|minute|chance|time)|have a better idea|think (?:it'?s|it is) (?:needed|necessary|worth it|best|better|wise))",
803 "you'?d (?:like|prefer|be so kind|rather|be willing|not mind)",
804 "you'?re (?:ready|able|free|ok|okay|happy|willing|up for it|good)(?: with (?:it|that|this))?",
805 "you are (?:ok|okay|happy|fine|good) with (?:it|that|this)",
806 "i (?:say|tell you) (?:otherwise|so|to stop)",
807 "i (?:change my mind|change it|switch (?:it )?back)",
808 "it'?s all the same to you",
809 // Acceptable to the model, not an outcome: "if it works" alone is one.
810 "(?:that|it|this) works for you",
811 "possible",
812 "(?:that|it)(?:'?s| is) (?:ok|okay|fine|alright|all right|not too much(?: trouble)?)(?: with you)?",
813 ].join("|") +
814 ")\\s*(?:then)?\\s*$",
815);
816
817/**
818 * Words after the tier that make it a name for something else: "use sonnet
819 * pricing" is about a price table, "use haiku ids in the test" about ids.
820 */
821const TIER_AS_NAME =
822 /^[ \t]+(?:pricing|prices?|rates?|ids?|names?|constants?|strings?|labels?|values?|entr(?:y|ies)|fields?|columns?|tables?|fixtures?|mocks?|stubs?|numbers?|figures?|tokens?|limits?|costs?|windows?)\b/i;
823
824/** True when the gap is only light bridge words and no clause break. */
825function proximityOk(gap: string): boolean {
826 if (/[.!?,;:—–…]/.test(gap)) return false;
827 return NEGATION_BRIDGE.test(gap);
828}
829
830/**
831 * The route phrases a negation binds: each negation binds the first run-on
832 * (or sink) attached after it, and only that one. Discourse ("Stop what
833 * you're doing and use opus") and tags ("why don't you use opus") do not
834 * bind. Worked out once per prompt, in one pass over the negations with the
835 * sinks found once: re-finding every sink for every phrase and negation made
836 * a long prompt cubic (a 20k-character one took a minute).
837 */
838function negatedAt(text: string, matches: readonly RegExpMatchArray[]): Set<number> {
839 const at = (m: RegExpMatchArray) => m.index ?? -1;
840 const sorted = (xs: number[]) => [...new Set(xs.filter((i) => i >= 0))].sort((a, b) => a - b);
841 const sinks = sorted([...matches.map(at), ...[...text.matchAll(USING_SINK)].map(at)]);
842 const withBare = sorted([...sinks, ...[...text.matchAll(BARE_TIER_SINK)].map(at)]);
843 // The first position in `xs` at or after `from`.
844 const firstFrom = (xs: readonly number[], from: number) => {
845 let lo = 0;
846 let hi = xs.length;
847 while (lo < hi) {
848 const mid = (lo + hi) >> 1;
849 if (xs[mid]! < from) lo = mid + 1;
850 else hi = mid;
851 }
852 return xs[lo];
853 };
854 const bound = new Set<number>();
855 for (const neg of text.matchAll(OVERRIDE_NEGATION_AT)) {
856 const negEnd = (neg.index ?? 0) + neg[0].length;
857 const negWord = neg[0].toLowerCase().replace(/\s+/g, " ");
858 const sink = firstFrom(negWord === "avoid" || negWord === "stop" ? withBare : sinks, negEnd);
859 if (sink !== undefined && proximityOk(text.slice(negEnd, sink))) bound.add(sink);
860 }
861 return bound;
862}
863
864/**
865 * A decision forced to a named tier. The router does not ask Jev for one
866 * (`fresh` is null), so it runs at medium; given an answer, its effort would
867 * be kept and its tier set aside.
868 */
869export function forcedDecision(tier: Tier, fresh: Decision | null): Decision {
870 return {
871 tier,
872 model: MODEL_OF[tier],
873 effort: fresh?.effort ?? "medium",
874 confidence: fresh?.confidence ?? 0,
875 effortConfidence: fresh?.effortConfidence ?? 0,
876 forced: true,
877 };
878}
879
880/**
881 * Whether a turn staying on Sonnet should keep the previous turn's effort.
882 *
883 * Measured 2026-09-22 on one session at ~58k context: an effort change on
884 * Opus 5.5 and Haiku 4.5 costs nothing (the engine sends it per turn), but
885 * on Sonnet 5 it rewrites everything after the system block, about half the
886 * prefix, $0.12 at that size. So on Sonnet an effort flip is a cache miss
887 * and gets the same treatment as a model switch: it has to clear the bar.
888 *
889 * The bar is read against Jev's confidence in the effort score, not the
890 * tier, because the two are separate answers and the effort one is the
891 * shakier (0.00 to 0.81 across ten prompts; lowest on the short follow-ups
892 * where a flip is least worth $0.12). Symmetric on purpose: letting rises
893 * through freely ratchets a Sonnet stretch up to xhigh and holds it there.
894 *
895 * With `ceiling`, an effort the ceiling no longer allows is not held: it
896 * would be cut to the cap anyway, so the effort changes and the cache is
897 * rewritten whatever the hold does, and Jev's own pick should run instead.
898 */
899export function holdsSonnetEffort(
900 fresh: Decision,
901 previous: Decision | null,
902 threshold: number,
903 ceiling?: Ceiling,
904): boolean {
905 if (previous === null) return false;
906 if (fresh.tier !== "sonnet" || previous.tier !== "sonnet") return false;
907 if (fresh.effort === previous.effort) return false;
908 if (ceiling !== undefined && effortRank(previous.effort) > effortRank(ceiling.sonnet)) return false;
909 return (fresh.effortConfidence ?? 0) < threshold;
910}
911
912/**
913 * The confidence a subagent's classification must reach before its model is
914 * set, up or down. A subagent starts with an empty conversation, so there is
915 * no cache to protect and stickiness does not apply; what the bar guards is
916 * a guess. Calibrated 2026-09-22 on six subagent-style prompts: the
917 * well-specified ones scored 0.72 to 0.98, the one vague audit 0.22, so 0.5
918 * splits them. Below it the subagent runs on what it would have anyway.
919 */
920export const SUBAGENT_CONFIDENCE = 0.5;
921
922/**
923 * Jev's decision for a spawned subagent, or null to leave the spawn alone.
924 * Nothing is held to: the parent's tier is only what "alone" resolves to.
925 */
926export function subagentDecision(
927 fresh: Decision | null,
928 threshold: number = SUBAGENT_CONFIDENCE,
929): Decision | null {
930 if (fresh === null) return null;
931 return fresh.confidence >= threshold ? fresh : null;
932}
933
934/**
935 * The most effort each tier may be asked for, xhigh on every tier by
936 * default — matching the engine's own default — and adjustable per session
937 * with `JEV_ROUTER_CEILING` or `/jev ceiling`. One effort per tier is the
938 * whole policy: what Jev asks for above it is capped to it, and
939 * `cappedEffort` keeps what Jev wanted so the route line can say so.
940 */
941export type Ceiling = Record<Tier, Effort>;
942
943export const DEFAULT_CEILING: Effort = "xhigh";
944
945/** Ladder position of an effort, low to high. */
946function effortRank(effort: Effort): number {
947 return EFFORTS.indexOf(effort);
948}
949
950/** An effort by name, or null. `off` and `none` mean no cap, which is max. */
951export function effortNamed(raw: string): Effort | null {
952 const name = raw.trim().toLowerCase();
953 if (name === "off" || name === "none") return "max";
954 return (EFFORTS as string[]).includes(name) ? (name as Effort) : null;
955}
956
957/** The same ceiling on every tier. */
958export function ceilingAt(effort: Effort): Ceiling {
959 return { haiku: effort, sonnet: effort, opus: effort, fable: effort };
960}
961
962/**
963 * Reads `JEV_ROUTER_CEILING`: one effort for every tier (`xhigh`), or a
964 * comma list of `tier:effort` pairs for some (`fable:xhigh,opus:high`) with
965 * the rest at the default. Anything unreadable is ignored, so a typo leaves
966 * the default in place rather than opening the ceiling.
967 */
968export function ceilingOf(raw: string | undefined): Ceiling {
969 const ceiling = ceilingAt(DEFAULT_CEILING);
970 const text = (raw ?? "").trim().toLowerCase();
971 if (!text) return ceiling;
972 const whole = effortNamed(text);
973 if (whole !== null) return ceilingAt(whole);
974 // Commas, semicolons or spaces between the parts, as JEV_ROUTER_EXCLUDE.
975 // Spaces around a colon folded by splitting, not `\s*:\s*`, which
976 // rescanned a long run of spaces from every position.
977 const joined = text
978 .split(":")
979 .map((s) => s.trim())
980 .join(":");
981 for (const part of joined.split(/[\s,;]+/)) {
982 const [tierName, effortName] = part.split(":").map((s) => s.trim());
983 if (tierName === undefined || effortName === undefined) continue;
984 const effort = effortNamed(effortName);
985 if (effort === null || !(TIERS as string[]).includes(tierName)) continue;
986 ceiling[tierName as Tier] = effort;
987 }
988 return ceiling;
989}
990
991/** A copy of `ceiling` with `effort` set on `tiers`, or on every tier. */
992export function withCeiling(
993 ceiling: Ceiling,
994 effort: Effort,
995 tiers: readonly Tier[] = TIERS,
996): Ceiling {
997 const next = { ...ceiling };
998 for (const tier of tiers) next[tier] = effort;
999 return next;
1000}
1001
1002/** Caps a decision's effort at its tier's ceiling, keeping what Jev named. */
1003export function capTo(decision: Decision, ceiling: Ceiling): Decision {
1004 const cap = ceiling[decision.tier];
1005 if (effortRank(decision.effort) <= effortRank(cap)) return decision;
1006 return { ...decision, effort: cap, cappedEffort: decision.effort };
1007}
1008
1009/**
1010 * What the engine runs on the first request of a conversation when asked
1011 * for an effort it does not honour there. Measured 2026-09-23 on Claude
1012 * Code 2.1.280, by the transcript's `perTurnEffort`: Fable 5.1 runs
1013 * `medium` as `high` on the first turn of a session (five of five), and
1014 * honours it from the second turn on (three of three); `low` and `high` go
1015 * through on every turn, and Opus 5.5 honours all five. The router sends
1016 * what will run, so the route line does not claim an effort the engine did
1017 * not use. The cost is the same either way. Remove an entry once the engine
1018 * honours it, and the request goes back to what Jev asked for.
1019 */
1020export const FIRST_TURN_EFFORT: Partial<
1021 Record<Tier, Partial<Record<Effort, Effort>>>
1022> = {
1023 fable: { medium: "high" },
1024};
1025
1026/**
1027 * The decision as the engine will run it on a conversation's first turn,
1028 * with what Jev asked kept in `askedEffort` so the next turn, where the
1029 * engine honours it, starts from Jev's word and not the quirk.
1030 */
1031export function firstTurnEffort(decision: Decision): Decision {
1032 const ran = FIRST_TURN_EFFORT[decision.tier]?.[decision.effort];
1033 if (ran === undefined) return decision;
1034 return { ...decision, effort: ran, askedEffort: decision.effort };
1035}
1036
1037/** The decision as Jev asked for it, for the turns that hold to or continue it. */
1038export function asAsked(decision: Decision): Decision {
1039 if (decision.askedEffort === undefined) return decision;
1040 const { askedEffort, ...rest } = decision;
1041 return { ...rest, effort: askedEffort };
1042}
1043
1044/**
1045 * The context size from which an upgrade has to be surer than the bar.
1046 *
1047 * An upgrade writes the whole context to the dearer tier's cache: at 250k,
1048 * five dollars for fable. Over a week of transcripts (2026-09-23), 54 of 72
1049 * routed upgrades ran under 75% confidence and 7 more under 90%, every one
1050 * of those past 100k context, while the prompts that are plainly planning
1051 * work measure 0.97 to 1.00. So past this size an upgrade needs
1052 * `UPGRADE_CONFIDENCE`, or the bar if that is higher. A tier the prompt
1053 * names is not an upgrade in this sense and is never held.
1054 */
1055export const UPGRADE_CONTEXT_TOKENS = 100_000;
1056export const UPGRADE_CONFIDENCE = 0.9;
1057
1058/**
1059 * The most an upgrade may cost this turn over staying, in dollars. Writing a
1060 * large context to a dearer tier's cache is the one cost a confident Jev
1061 * does not see: opus to fable at 250k is about $5 before any output. At $1
1062 * and a typical turn, an upgrade goes through up to about 48k of context
1063 * from opus to fable, 126k from sonnet to opus, 254k from haiku to sonnet.
1064 */
1065export const UPGRADE_MAX_USD = 1;
1066
1067/**
1068 * `JEV_ROUTER_UPGRADE_MAX`: dollars an upgrade may cost over staying, or
1069 * `off` for no limit (the confidence bar still applies). Anything else is
1070 * the default.
1071 */
1072export function upgradeMaxOf(raw: string | undefined): number | null {
1073 const v = (raw ?? "").trim().toLowerCase().replace(/^\$/, "");
1074 if (v === "off" || v === "none") return null;
1075 // Plain dollars only: `-1` or `0x10` is a mistake, not a limit.
1076 if (!PLAIN_DECIMAL.test(v)) return UPGRADE_MAX_USD;
1077 return Number(v);
1078}
1079
1080/**
1081 * `JEV_ROUTER_PRICE_CHECK`: the downgrade and upgrade price checks, on
1082 * unless `0`, `false`, `no`, `off` or `none`. Separate from sticky, which is the
1083 * confidence bar alone.
1084 */
1085export function priceCheckOf(raw: string | undefined): boolean {
1086 return !flagOff(raw);
1087}
1088
1089/** The bar an upgrade must clear, given the context it would write. */
1090export function upgradeBar(bar: number, contextTokens: number): number {
1091 return contextTokens >= UPGRADE_CONTEXT_TOKENS
1092 ? Math.max(bar, UPGRADE_CONFIDENCE)
1093 : bar;
1094}
1095hooks/pricing.ts 300 lines1/**
2 * What a turn costs, and what a switch would cost, in dollars.
3 *
4 * Prices are Anthropic's public list, per million tokens, read from
5 * platform.claude.com/docs/en/about-claude/pricing on PRICE_DATE. A cache
6 * read is a tenth of input on most models, a fortieth on Fable 5.1 and a
7 * twentieth on Opus 5.5; a cache write is 1.25× input for the five-minute
8 * cache and 2× for the one-hour one. Claude Code writes the one-hour cache
9 * (every one of 18,204 writes in a week of this machine's transcripts), so
10 * that is the default here.
11 *
12 * Nothing here touches the engine or the network.
13 */
14
15import type { Tier } from "./policy.ts";
16
17export const PRICE_DATE = "2026-09-23";
18
19/** Dollars per million tokens. */
20export type Price = {
21 input: number;
22 write5m: number;
23 write1h: number;
24 read: number;
25 output: number;
26};
27
28export type Ttl = "5m" | "1h";
29
30/** The ladder's models. */
31export const PRICE: Record<Tier, Price> = {
32 haiku: { input: 1, write5m: 1.25, write1h: 2, read: 0.1, output: 5 },
33 sonnet: { input: 2, write5m: 2.5, write1h: 4, read: 0.2, output: 10 },
34 opus: { input: 4, write5m: 5, write1h: 8, read: 0.2, output: 20 },
35 fable: { input: 10, write5m: 12.5, write1h: 20, read: 0.25, output: 50 },
36};
37
38/**
39 * Models the session may run on that are not on the ladder, so an unrouted
40 * turn's cost is still right. Matched by prefix of the id the API reports.
41 */
42const OPUS_4_0: Price = { input: 15, write5m: 18.75, write1h: 30, read: 1.5, output: 75 };
43/** Fable 5 and Mythos 5: Fable 5.1's price, but four times its cache read. */
44const FABLE_5_0: Price = { ...PRICE.fable, read: 1 };
45
46// First match wins, so a longer id goes above the prefix it starts with.
47const OTHER_PRICE: readonly (readonly [string, Price])[] = [
48 ["claude-opus-5-5", PRICE.opus],
49 ["claude-opus-5", { input: 5, write5m: 6.25, write1h: 10, read: 0.5, output: 25 }],
50 // Opus 4 and 4.1, and Opus 4's dated id: three times the later 4.x price.
51 ["claude-opus-4-0", OPUS_4_0],
52 ["claude-opus-4-1", OPUS_4_0],
53 ["claude-opus-4-2025", OPUS_4_0],
54 ["claude-opus-4", { input: 5, write5m: 6.25, write1h: 10, read: 0.5, output: 25 }],
55 ["claude-sonnet-5", PRICE.sonnet],
56 ["claude-sonnet-4", { input: 3, write5m: 3.75, write1h: 6, read: 0.3, output: 15 }],
57 ["claude-haiku-4", PRICE.haiku],
58 ["claude-fable-5-1", PRICE.fable],
59 ["claude-mythos-5-1", PRICE.fable],
60 ["claude-fable-5", FABLE_5_0],
61 ["claude-mythos-5", FABLE_5_0],
62];
63
64/** The cache TTL from the environment; `1h` unless told `5m`. */
65export function ttlOf(raw: string | undefined): Ttl {
66 return (raw ?? "").trim().toLowerCase() === "5m" ? "5m" : "1h";
67}
68
69/** The price of the model the API named, or null for one we do not know. */
70export function priceOfModel(model: string): Price | null {
71 if (typeof model !== "string") return null;
72 // Bedrock spells an id `us.anthropic.claude-…-v1:0`, Vertex `claude-…@date`:
73 // the same model, priced by the id inside.
74 const id = model
75 .toLowerCase()
76 .replace(/^(?:[a-z]{2,}\.)*anthropic\./, "")
77 .replace(/-v\d+(?::\d+)?$/, "")
78 .replace(/@\d{8}$/, "");
79 // Bare, it is Opus 4.0 (Vertex's `claude-opus-4@date` comes to this).
80 if (id === "claude-opus-4") return OPUS_4_0;
81 for (const [prefix, price] of OTHER_PRICE)
82 if (id.startsWith(prefix)) return price;
83 return null;
84}
85
86/**
87 * The rung a model id or alias sits on: `claude-opus-5` and `opus` are both
88 * opus-class, whatever the session runs. Null for a model off the ladder.
89 */
90export function tierOfModel(model: string): Tier | null {
91 if (typeof model !== "string") return null;
92 const id = model.toLowerCase();
93 if (id.includes("haiku")) return "haiku";
94 if (id.includes("sonnet")) return "sonnet";
95 if (id.includes("opus")) return "opus";
96 if (id.includes("fable") || id.includes("mythos")) return "fable";
97 return null;
98}
99
100/** Token counts as the engine's `TurnUsage` carries them. */
101type Tokens = {
102 input_tokens: number;
103 output_tokens: number;
104 cache_read_input_tokens: number;
105 cache_creation_input_tokens: number;
106};
107
108/**
109 * What one turn's requests cost, from the API's own counts. Null when the
110 * model is one we have no price for, which is better than a wrong number.
111 */
112export function usageCost(
113 model: string,
114 usage: Tokens,
115 ttl: Ttl = "1h",
116): number | null {
117 const p = priceOfModel(model);
118 if (p === null) return null;
119 const write = ttl === "1h" ? p.write1h : p.write5m;
120 return (
121 (usage.input_tokens * p.input +
122 usage.cache_creation_input_tokens * write +
123 usage.cache_read_input_tokens * p.read +
124 usage.output_tokens * p.output) /
125 1e6
126 );
127}
128
129/**
130 * The context window of each tier's model, in tokens: Haiku 4.5 takes 200K,
131 * the rest 1M (platform.claude.com/docs/en/about-claude/models, 2026-09-23).
132 * A request past it is refused with "Prompt is too long", and a week of
133 * transcripts holds three of those, each right after a turn at 358k–605k
134 * was routed to haiku. Nothing on the ladder is smaller than a session
135 * with a `[1m]` model: the plain ids the router sends were answered at
136 * 737k on fable and 344k on sonnet.
137 */
138export const WINDOW_TOKENS: Record<Tier, number> = {
139 haiku: 200_000,
140 sonnet: 1_000_000,
141 opus: 1_000_000,
142 fable: 1_000_000,
143};
144
145/** Room left for the prompt and the reply when a turn is judged to fit. */
146const WINDOW_HEADROOM_TOKENS = 16_000;
147
148/** True when a turn carrying `contextTokens` can be sent to `tier` at all. */
149export function fitsWindow(tier: Tier, contextTokens: number): boolean {
150 return contextTokens + WINDOW_HEADROOM_TOKENS <= WINDOW_TOKENS[tier];
151}
152
153/** `claude-opus-5-5[1m]` and `claude-opus-5-5` are one model: the suffix is the engine's. */
154export function baseModel(model: string): string {
155 // By hand, not `/\[[^\]]*\]$/`, which rescans to the end from every `[`.
156 if (!model.endsWith("]")) return model;
157 const open = model.indexOf("[", model.lastIndexOf("]", model.length - 2) + 1);
158 return open === -1 || open === model.length - 1 ? model : model.slice(0, open);
159}
160
161/** The model a spelling names: without `[1m]` or a date suffix, so both read as one. */
162export function sameModelAs(a: string, b: string): boolean {
163 return modelKey(a) === modelKey(b);
164}
165
166/**
167 * A model's name without its spelling: `[1m]`, a date, and the Bedrock and
168 * Vertex wrappings (`us.anthropic.…-v1:0`, an ARN, `…@20260901`, a path)
169 * all name the same model as the plain id.
170 */
171function modelKey(model: string): string {
172 let m = baseModel(model).toLowerCase();
173 m = m.slice(m.lastIndexOf("/") + 1);
174 m = m.replace(/^(?:us|eu|apac|global|au|jp|ca)\./, "").replace(/^anthropic\./, "");
175 m = m.replace(/@.*$/, "").replace(/-v\d+(?::\d+)?$/, "").replace(/-\d{8}$/, "");
176 return m;
177}
178
179/** Ladder order, low to high, for telling a downgrade from an upgrade. */
180const RANK: Record<Tier, number> = { haiku: 0, sonnet: 1, opus: 2, fable: 3 };
181
182export function isDowngrade(from: Tier, to: Tier): boolean {
183 return RANK[to] < RANK[from];
184}
185
186/**
187 * The two prices a shaky downgrade is decided between.
188 *
189 * `stay` is the next turn on the tier already running, warm: its context
190 * read from cache plus its output. `go` is the same turn on the cheaper
191 * tier, cold: the whole context written to that tier's cache, its output,
192 * and then the write that comes due when the session returns to the tier it
193 * left, whose cache the detour let go cold (measured over a week of
194 * transcripts: 31 of 38 returns from haiku to fable paid it in full).
195 *
196 * A downgrade that costs more than it saves is held. Output tokens are the
197 * only term where the cheaper tier wins, so the balance tips with context:
198 * at a few thousand tokens the cheaper output carries it; at the sizes a
199 * working session actually runs (150k–330k at the median, measured) the
200 * writes are tens of times the output and no downgrade pays.
201 */
202export type SwitchVerdict = {
203 /** Dollars for this turn on the running tier, cache warm. */
204 stay: number;
205 /** Dollars for this turn on the new tier, cache cold, return write included. */
206 go: number;
207 /** True when going costs at least as much as staying. */
208 hold: boolean;
209 /**
210 * For an upgrade: the most the move may cost over staying. Absent for a
211 * downgrade, which has to pay for itself.
212 */
213 limit?: number;
214};
215
216export function switchVerdict(
217 from: Tier,
218 to: Tier,
219 contextTokens: number,
220 outputTokens: number,
221 ttl: Ttl = "1h",
222 /**
223 * The running model's own price when it is not the ladder's model for its
224 * tier: a session on `claude-opus-5` reads its cache at $0.50, not the
225 * $0.20 of the opus tier's `claude-opus-5-5`.
226 */
227 fromPrice: Price = PRICE[from],
228 /** The running model's cache has expired: staying writes it too. */
229 fromCold = false,
230): SwitchVerdict {
231 const write = (p: Price) => (ttl === "1h" ? p.write1h : p.write5m);
232 const ctx = contextTokens / 1e6;
233 const out = outputTokens / 1e6;
234 const stay =
235 ctx * (fromCold ? write(fromPrice) : fromPrice.read) + out * fromPrice.output;
236 const go =
237 ctx * write(PRICE[to]) + out * PRICE[to].output + ctx * write(fromPrice);
238 return { stay, go, hold: go >= stay };
239}
240
241/**
242 * What moving up from `from` to `to` costs this turn, against staying. Going
243 * writes the whole context to the dearer tier's cache and pays its output
244 * price; staying reads the warm cache. The way back is not counted: it may
245 * never happen, and if it does, the downgrade is priced then. The move is
246 * held when it costs more than `limit` over staying.
247 */
248export function upgradeVerdict(
249 from: Tier,
250 to: Tier,
251 contextTokens: number,
252 outputTokens: number,
253 limit: number,
254 ttl: Ttl = "1h",
255 fromPrice: Price = PRICE[from],
256 /** The running model's cache has expired: staying writes it too. */
257 fromCold = false,
258): SwitchVerdict {
259 const write = (p: Price) => (ttl === "1h" ? p.write1h : p.write5m);
260 const ctx = contextTokens / 1e6;
261 const out = outputTokens / 1e6;
262 const stay =
263 ctx * (fromCold ? write(fromPrice) : fromPrice.read) + out * fromPrice.output;
264 const go = ctx * write(PRICE[to]) + out * PRICE[to].output;
265 return { stay, go, hold: go - stay > limit, limit };
266}
267
268/**
269 * The context size below which a downgrade from `from` to `to` still pays,
270 * for a turn of `outputTokens`. Shown in the status report so the bar is
271 * visible; zero when no context is small enough.
272 */
273export function breakEvenTokens(
274 from: Tier,
275 to: Tier,
276 outputTokens: number,
277 ttl: Ttl = "1h",
278 /** The running model's own price, as in `switchVerdict`. */
279 fromPrice: Price = PRICE[from],
280 /** The running model's cache has expired, as in `switchVerdict`. */
281 fromCold = false,
282): number {
283 const write = (p: Price) => (ttl === "1h" ? p.write1h : p.write5m);
284 // stay = ctx·read(from) + out·output(from); go = ctx·(write(to)+write(from)) + out·output(to)
285 // go < stay ⇔ ctx·(write(to)+write(from)−read(from)) < out·(output(from)−output(to))
286 // Cold, staying writes too: read(from) becomes write(from).
287 const perCtx = write(PRICE[to]) + write(fromPrice) - (fromCold ? write(fromPrice) : fromPrice.read);
288 const perOut = fromPrice.output - PRICE[to].output;
289 if (perOut <= 0 || perCtx <= 0) return 0;
290 return Math.floor((outputTokens * perOut) / perCtx);
291}
292
293/** `$4.41`, `$0.36`, `$0.024`, `$0.0035`: enough places to show a small turn. */
294export function usd(n: number): string {
295 // Decided on the rounded figure: $0.09999 is $0.10, not $0.100.
296 if (Number(n.toFixed(3)) >= 0.1) return `$${n.toFixed(2)}`;
297 if (Number(n.toFixed(4)) >= 0.01) return `$${n.toFixed(3)}`;
298 return `$${n.toFixed(4)}`;
299}
300hooks/persist.ts 410 lines1/**
2 * What of a session's routing survives a reload of this module.
3 *
4 * The engine reloads a hooks module when its files change (an update, a
5 * `git pull`), and every `let` in `register` starts over: the history empty,
6 * `spent` at zero, and — the part that costs money — nothing held, so the
7 * first switch after a reload was priced against the session model instead
8 * of the tier actually warm. `$.store` is the engine's own JSON store for
9 * the plugin, kept across reloads and sessions; this file turns the state
10 * into data for it and back. Nothing here touches the engine.
11 *
12 * Attempts are shared: one object sits in the history, in the open reply
13 * and in the agent map at once, and usage folded into it must show in all
14 * three. JSON would copy each, so they are written once, as a pool, and
15 * referred to by index.
16 */
17
18import type { Compaction } from "./compactor.ts";
19import { EFFORTS, MODEL_ID, TIERS, type Ceiling, type Decision } from "./policy.ts";
20import { tierOfModel } from "./pricing.ts";
21
22/** A model id as the engine spells one: `claude-opus-5-5[1m]`, a Bedrock or Vertex id. */
23
24/**
25 * The most entries any list in a snapshot can hold: the router keeps far
26 * fewer (a few turns of history, 64 of a reply, 32-odd agents). A larger
27 * one did not come from it, and restoring it could overflow a spread.
28 */
29const LIST_MAX = 1_000;
30import type { Attempt } from "./status.ts";
31
32export const SNAPSHOT_VERSION = 1;
33
34/** Keys under which snapshots are kept, one per session. */
35export const SNAPSHOT_PREFIX = "session:";
36
37/** Sessions whose snapshots are kept; older ones are dropped on save. */
38export const SNAPSHOTS_KEPT = 20;
39
40/** The settings a `/jev` command can set, which then outrank the environment. */
41export const OVERRIDABLE = ["sticky", "ceiling", "excludedTiers", "compactOn", "priceCheck"] as const;
42export type Overridable = (typeof OVERRIDABLE)[number];
43
44export type State = {
45 attempts: Attempt[];
46 reply: Attempt[];
47 replyAgents: string[];
48 spawned: [string, Attempt][];
49 /** Agents the router left alone, one history row each. */
50 unrouted: [string, Attempt][];
51 /**
52 * The turns in flight: their attempts, their decisions, and which still
53 * await their route line. A reload mid-turn used to leave the rest of
54 * that turn unrouted, since its next step found nothing under its id.
55 */
56 turns: [string, Attempt][];
57 decisions: [string, Decision][];
58 pending: string[];
59 /** Agents whose first request has run, so a resumed one is not re-snapped. */
60 stepped: string[];
61 running: Decision | null;
62 continueFrom: Decision | null;
63 latest: Decision | null;
64 lastUsage: { context: number; output: number } | null;
65 sessionModel: string | null;
66 spent: number;
67 enabled: boolean;
68 announce: boolean;
69 /** A response has been received in this conversation: no request is its first any more. */
70 answered: boolean;
71 sticky: number | null;
72 ceiling: Ceiling;
73 /**
74 * Tiers `/jev tiers off` dropped from the question Jev is asked.
75 * `undefined` only comes back from `unpack` on a snapshot from before this
76 * field existed — never from `pack`, which always writes the live array —
77 * and means "this snapshot has no opinion", not "nothing is excluded": the
78 * caller should leave the environment's own `JEV_ROUTER_EXCLUDE` seeding
79 * in place rather than overwrite it with an empty array.
80 */
81 excludedTiers: string[] | undefined;
82 /** Compaction by Jev is on. */
83 compactOn: boolean;
84 /** The downgrade and upgrade price checks are on. */
85 priceCheck: boolean;
86 /**
87 * The settings above that a `/jev` command set this session; only these
88 * are restored over the environment, so a reload or resume still follows
89 * a changed `JEV_ROUTER_*` for everything no command touched.
90 * `undefined` only from `unpack` on a snapshot from before this field
91 * existed, which restores every setting, as those snapshots always did.
92 */
93 overridden: Overridable[] | undefined;
94 /** Agents whose reply's summary was written: their late wake-up joins no block. */
95 summarisedAgents: string[];
96 /** The last compaction Jev was asked about, for /jev. */
97 compaction: Compaction | null;
98 /**
99 * What was running before the last turn switched, while no response has
100 * confirmed the switch; null when there is nothing to take back.
101 */
102 unconfirmed: { was: Decision | null } | null;
103 /** The engine said the resumed session's cache expired, and no response has written it since. */
104 cacheExpired: boolean;
105 /** When the snapshot was written; only from `unpack`, since `pack` stamps its own. */
106 savedAt?: number;
107};
108
109type Packed = Omit<State, "attempts" | "reply" | "spawned" | "unrouted" | "turns"> & {
110 v: number;
111 /** When it was written, for pruning the least recently used first. */
112 savedAt: number;
113 pool: Attempt[];
114 attempts: number[];
115 reply: number[];
116 spawned: [string, number][];
117 unrouted: [string, number][];
118 turns: [string, number][];
119};
120
121/** The state as JSON data, attempts written once each. */
122export function pack(state: State): Packed {
123 const pool: Attempt[] = [];
124 const index = new Map<Attempt, number>();
125 const ref = (a: Attempt) => {
126 let i = index.get(a);
127 if (i === undefined) {
128 i = pool.length;
129 pool.push(a);
130 index.set(a, i);
131 }
132 return i;
133 };
134 return {
135 v: SNAPSHOT_VERSION,
136 savedAt: Date.now(),
137 pool,
138 attempts: state.attempts.map(ref),
139 reply: state.reply.map(ref),
140 replyAgents: [...state.replyAgents],
141 spawned: state.spawned.map(([id, a]) => [id, ref(a)]),
142 unrouted: (state.unrouted ?? []).map(([id, a]) => [id, ref(a)]),
143 turns: state.turns.map(([id, a]) => [id, ref(a)]),
144 decisions: state.decisions,
145 pending: [...state.pending],
146 stepped: [...state.stepped],
147 running: state.running,
148 continueFrom: state.continueFrom,
149 latest: state.latest,
150 lastUsage: state.lastUsage,
151 sessionModel: state.sessionModel,
152 spent: state.spent,
153 enabled: state.enabled,
154 announce: state.announce,
155 answered: state.answered,
156 sticky: state.sticky,
157 ceiling: state.ceiling,
158 excludedTiers: state.excludedTiers,
159 compactOn: state.compactOn,
160 priceCheck: state.priceCheck,
161 overridden: state.overridden,
162 summarisedAgents: state.summarisedAgents,
163 compaction: state.compaction,
164 unconfirmed: state.unconfirmed,
165 cacheExpired: state.cacheExpired,
166 };
167}
168
169const isRecord = (v: unknown): v is Record<string, unknown> =>
170 typeof v === "object" && v !== null && !Array.isArray(v);
171
172/** A finite number no smaller than 0. */
173const isCount = (v: unknown): v is number => typeof v === "number" && Number.isFinite(v) && v >= 0;
174
175/**
176 * A decision as the router writes one: a tier and effort it knows, a model
177 * id, a confidence 0–1.
178 */
179function isValidDecision(v: unknown): v is Decision {
180 return (
181 isRecord(v) &&
182 (TIERS as readonly unknown[]).includes(v.tier) &&
183 typeof v.model === "string" &&
184 // A model id of the tier it names: the route line says the tier and the
185 // request sends the model, so the two cannot be allowed to differ.
186 MODEL_ID.test(v.model) &&
187 tierOfModel(v.model) === v.tier &&
188 (EFFORTS as readonly unknown[]).includes(v.effort) &&
189 typeof v.confidence === "number" &&
190 Number.isFinite(v.confidence) &&
191 v.confidence >= 0 &&
192 v.confidence <= 1 &&
193 // What `/jev` and the route line print from: a number where one is read.
194 [v.effortConfidence, v.heldWindow, v.heldBar].every((n) => n === undefined || isCount(n)) &&
195 (v.heldCost === undefined ||
196 (isRecord(v.heldCost) &&
197 isCount(v.heldCost.stay) &&
198 isCount(v.heldCost.go) &&
199 (v.heldCost.limit === undefined || isCount(v.heldCost.limit)))) &&
200 [v.jevFailed, v.heldModel].every((t) => t === undefined || typeof t === "string") &&
201 [v.held, v.outgrew, v.wanted].every((t) => t === undefined || (TIERS as readonly unknown[]).includes(t)) &&
202 [v.heldEffort, v.cappedEffort].every((t) => t === undefined || (EFFORTS as readonly unknown[]).includes(t))
203 );
204}
205
206/** An attempt as the router writes one: a prompt, a time, a decision or a reason, usage in numbers. */
207function isValidAttempt(v: unknown): boolean {
208 if (!isRecord(v) || typeof v.prompt !== "string" || typeof v.ms !== "number" || !Number.isFinite(v.ms)) return false;
209 if ("decision" in v ? !isValidDecision(v.decision) : typeof v.skipped !== "string") return false;
210 if (v.cost !== undefined && !isCount(v.cost)) return false;
211 if (
212 v.agent !== undefined &&
213 !(isRecord(v.agent) && typeof v.agent.label === "string" && (v.agent.type === undefined || typeof v.agent.type === "string"))
214 )
215 return false;
216 if (v.usage !== undefined) {
217 const u = v.usage;
218 if (
219 !isRecord(u) ||
220 typeof u.model !== "string" ||
221 ![u.input_tokens, u.output_tokens, u.cache_read_input_tokens, u.cache_creation_input_tokens].every(isCount)
222 )
223 return false;
224 }
225 return true;
226}
227
228/**
229 * The state back from what the store returned, or null for anything that is
230 * not a snapshot this version wrote. A bad snapshot is ignored, never
231 * trusted: the router then starts over, which is what it did before this
232 * file existed.
233 */
234export function unpack(raw: unknown): State | null {
235 if (!isRecord(raw) || raw.v !== SNAPSHOT_VERSION) return null;
236 const pool = raw.pool;
237 // Every attempt is checked as a decision is: a corrupt one reaches the
238 // route line, the summary and the spend total, which trust its fields.
239 if (!Array.isArray(pool) || pool.length > LIST_MAX || !pool.every(isValidAttempt)) return null;
240 for (const list of [raw.attempts, raw.reply, raw.spawned, raw.unrouted, raw.turns, raw.decisions, raw.pending, raw.stepped, raw.replyAgents])
241 if (Array.isArray(list) && list.length > LIST_MAX) return null;
242 const at = (i: unknown): Attempt | null =>
243 typeof i === "number" && Number.isInteger(i) && i >= 0 && i < pool.length
244 ? (pool[i] as Attempt)
245 : null;
246 const refs = (v: unknown): Attempt[] | null => {
247 if (!Array.isArray(v)) return null;
248 const out = v.map(at);
249 return out.every((a) => a !== null) ? (out as Attempt[]) : null;
250 };
251 const attempts = refs(raw.attempts);
252 const reply = refs(raw.reply);
253 if (attempts === null || reply === null) return null;
254 if (!Array.isArray(raw.spawned) || !Array.isArray(raw.replyAgents))
255 return null;
256 const pairs = (v: unknown): [string, Attempt][] | null => {
257 // Absent in a snapshot from before the field existed: nothing in flight.
258 if (v === undefined) return [];
259 if (!Array.isArray(v)) return null;
260 const out: [string, Attempt][] = [];
261 for (const pair of v) {
262 if (!Array.isArray(pair) || typeof pair[0] !== "string") return null;
263 const a = at(pair[1]);
264 if (a === null) return null;
265 out.push([pair[0], a]);
266 }
267 return out;
268 };
269 const spawned = pairs(raw.spawned);
270 const unrouted = pairs(raw.unrouted);
271 const turns = pairs(raw.turns);
272 if (spawned === null || unrouted === null || turns === null) return null;
273 const strings = (v: unknown): string[] =>
274 Array.isArray(v) ? v.filter((x): x is string => typeof x === "string") : [];
275 // Invalid decisions are dropped rather than trusted to their detriment.
276 const decisions: [string, Decision][] = Array.isArray(raw.decisions)
277 ? raw.decisions.filter(
278 (p): p is [string, Decision] =>
279 Array.isArray(p) && typeof p[0] === "string" && isValidDecision(p[1]),
280 )
281 : [];
282 const decision = (v: unknown) => (isValidDecision(v) ? v : null);
283 const lastUsage =
284 isRecord(raw.lastUsage) &&
285 isCount(raw.lastUsage.context) &&
286 isCount(raw.lastUsage.output)
287 ? { context: raw.lastUsage.context, output: raw.lastUsage.output }
288 : null;
289 // Every tier's cap must be an effort: one missing or misspelled would
290 // cap that tier to nothing, and the request would go out with no effort.
291 const ceiling = raw.ceiling;
292 if (
293 !isRecord(ceiling) ||
294 !TIERS.every((t) => (EFFORTS as readonly unknown[]).includes(ceiling[t]))
295 )
296 return null;
297 return {
298 attempts,
299 reply,
300 replyAgents: raw.replyAgents.filter(
301 (id): id is string => typeof id === "string",
302 ),
303 spawned,
304 unrouted,
305 turns,
306 decisions,
307 pending: strings(raw.pending),
308 stepped: strings(raw.stepped),
309 running: decision(raw.running),
310 continueFrom: decision(raw.continueFrom),
311 latest: decision(raw.latest),
312 lastUsage,
313 sessionModel:
314 typeof raw.sessionModel === "string" && MODEL_ID.test(raw.sessionModel) ? raw.sessionModel : null,
315 spent: isCount(raw.spent) ? raw.spent : 0,
316 enabled: raw.enabled !== false,
317 announce: raw.announce !== false,
318 // A snapshot from before this field exists has turns behind it.
319 answered: raw.answered !== false,
320 // A bar outside 0–1 would hold every switch, or none.
321 sticky: typeof raw.sticky === "number" && raw.sticky > 0 && raw.sticky < 1 ? raw.sticky : null,
322 ceiling: Object.fromEntries(TIERS.map((t) => [t, ceiling[t]])) as Ceiling,
323 // Absent (a snapshot from before this field existed) is left undefined
324 // — a signal to leave the environment's own JEV_ROUTER_EXCLUDE seeding
325 // alone — rather than defaulted to an empty array, which used to
326 // silently clear an env-seeded exclusion the moment such a snapshot was
327 // restored (the field never existed to preserve it).
328 excludedTiers:
329 raw.excludedTiers === undefined
330 ? undefined
331 : strings(raw.excludedTiers).filter((t) =>
332 (TIERS as readonly string[]).includes(t),
333 ),
334 compactOn: raw.compactOn !== false,
335 priceCheck: raw.priceCheck !== false,
336 overridden:
337 raw.overridden === undefined
338 ? undefined
339 : strings(raw.overridden).filter((k): k is Overridable =>
340 (OVERRIDABLE as readonly string[]).includes(k),
341 ),
342 summarisedAgents: Array.isArray(raw.summarisedAgents)
343 ? raw.summarisedAgents.filter((a): a is string => typeof a === "string")
344 : [],
345 compaction: compactionOf(raw.compaction),
346 unconfirmed:
347 isRecord(raw.unconfirmed) && (raw.unconfirmed.was === null || isValidDecision(raw.unconfirmed.was))
348 ? { was: raw.unconfirmed.was as Decision | null }
349 : null,
350 cacheExpired: raw.cacheExpired === true,
351 ...(typeof raw.savedAt === "number" && Number.isFinite(raw.savedAt) ? { savedAt: raw.savedAt } : {}),
352 };
353}
354
355/** A saved compaction with every field it needs, or null. */
356function compactionOf(raw: unknown): Compaction | null {
357 if (!isRecord(raw) || !isRecord(raw.calls)) return null;
358 const n = (v: unknown) => typeof v === "number" && Number.isFinite(v);
359 const c = raw.calls;
360 if (![raw.at, raw.kept, raw.of, raw.reduction, raw.ms, c.kept, c.cut, c.dropped].every(n)) return null;
361 return {
362 at: raw.at as number,
363 kept: raw.kept as number,
364 of: raw.of as number,
365 reduction: raw.reduction as number,
366 ms: raw.ms as number,
367 calls: { kept: c.kept as number, cut: c.cut as number, dropped: c.dropped as number },
368 ...(typeof raw.fallback === "string" ? { fallback: raw.fallback } : {}),
369 };
370}
371
372/**
373 * Claims (`<ownerPrefix><snapshot key>`) on a session that has no snapshot
374 * and is not the current one: candidates to drop once they are old.
375 */
376export function orphanOwnerKeys(keys: readonly string[], ownerPrefix: string, current: string): string[] {
377 const snapshots = new Set(keys.filter((k) => k.startsWith(SNAPSHOT_PREFIX)));
378 return keys.filter((k) => {
379 if (!k.startsWith(ownerPrefix)) return false;
380 const target = k.slice(ownerPrefix.length);
381 return target.startsWith(SNAPSHOT_PREFIX) && target !== current && !snapshots.has(target);
382 });
383}
384
385/** When a stored snapshot was written, or null for one from before the field or not a snapshot. */
386export function savedAtOf(raw: unknown): number | null {
387 return isRecord(raw) && typeof raw.savedAt === "number" && Number.isFinite(raw.savedAt) ? raw.savedAt : null;
388}
389
390/**
391 * The snapshot keys to drop so `SNAPSHOTS_KEPT` remain, least recently saved
392 * first when `savedAt` says (a session saved on every turn is never the one
393 * dropped, however long ago it started), else in the store's key order.
394 */
395export function staleKeys(
396 keys: readonly string[],
397 current: string,
398 savedAt?: ReadonlyMap<string, number>,
399): string[] {
400 const sessions = keys.filter(
401 (k) => k.startsWith(SNAPSHOT_PREFIX) && k !== current,
402 );
403 const excess = sessions.length + 1 - SNAPSHOTS_KEPT;
404 if (excess <= 0) return [];
405 const order = sessions
406 .map((k, i) => ({ k, i, at: savedAt?.get(k) ?? -Infinity }))
407 .sort((a, b) => a.at - b.at || a.i - b.i);
408 return order.slice(0, excess).map((x) => x.k);
409}
410hooks/provider.ts 197 lines1/**
2 * Provider resolution: which backend (TypeSafe direct or Vercel AI Gateway)
3 * should handle this request, based on available keys and user overrides.
4 *
5 * Precedence:
6 * 1. JEV_ROUTER_PROVIDER=typesafe|gateway forces one (and errors if its key is missing)
7 * 2. TypeSafe direct if TYPESAFE_API_KEY is set
8 * 3. Gateway if AI_GATEWAY_API_KEY is set
9 * 4. Error if neither is set
10 *
11 * TYPESAFE_BASE_URL overrides the TypeSafe endpoint base (defaults to
12 * https://api.typesafe.ai). Only https://api.typesafe.ai and hosts under
13 * *.typesafe.ai are accepted unless JEV_ROUTER_ALLOW_CUSTOM_BASE=1.
14 *
15 * JEV_ROUTER_JEV_MODEL pins the Jev version on the direct API. The default
16 * `jev-latest` is an alias that moves when TypeSafe ships a release, and the
17 * confidences the sticky bar is tuned against can move with it; TypeSafe's
18 * own advice is to pin (`jev-1.13.0` on api.typesafe.ai; a passthrough such
19 * as OpenRouter spells it `jev-1.13`, measured 2026-09-23).
20 */
21
22export type ProviderResult =
23 | {
24 ok: true;
25 name: "typesafe" | "gateway";
26 endpoint: string;
27 model: string;
28 apiKey: string;
29 }
30 | {
31 ok: false;
32 reason: string;
33 };
34
35export type ProviderEnv = {
36 TYPESAFE_API_KEY: string | undefined;
37 AI_GATEWAY_API_KEY: string | undefined;
38 JEV_ROUTER_PROVIDER: string | undefined;
39 TYPESAFE_BASE_URL: string | undefined;
40 JEV_ROUTER_ALLOW_CUSTOM_BASE?: string | undefined;
41 JEV_ROUTER_JEV_MODEL?: string | undefined;
42};
43
44const TYPESAFE_MODEL_DEFAULT = "jev-latest";
45
46const TYPESAFE_BASE_DEFAULT = "https://api.typesafe.ai";
47const GATEWAY_BASE = "https://ai-gateway.vercel.sh";
48
49/**
50 * Resolves a TypeSafe API base URL. Rejects non-https and unknown hosts
51 * unless custom bases are explicitly allowed — otherwise a mistyped or
52 * malicious settings value would send the Bearer key elsewhere.
53 */
54export function typesafeBaseOf(
55 raw: string | undefined,
56 allowCustom: string | undefined,
57): { ok: true; base: string } | { ok: false; reason: string } {
58 const trimmed = (raw ?? "").trim();
59 if (!trimmed) return { ok: true, base: TYPESAFE_BASE_DEFAULT };
60
61 let url: URL;
62 try {
63 url = new URL(trimmed.replace(/\/$/, ""));
64 } catch {
65 return { ok: false, reason: "TYPESAFE_BASE_URL is not a valid URL" };
66 }
67 if (url.protocol !== "https:") {
68 return { ok: false, reason: "TYPESAFE_BASE_URL must use https" };
69 }
70
71 const host = url.hostname.toLowerCase();
72 const allowed =
73 host === "api.typesafe.ai" || host.endsWith(".typesafe.ai");
74 const customOk = flagOn(allowCustom);
75 if (!allowed && !customOk) {
76 return {
77 ok: false,
78 reason:
79 "TYPESAFE_BASE_URL host is not allowlisted; set " +
80 "JEV_ROUTER_ALLOW_CUSTOM_BASE=1 to permit it",
81 };
82 }
83 // The endpoint's own path is added after the base; a base that already
84 // carries it (copied from the docs' full URL) would otherwise double it.
85 // Trailing slashes trimmed by hand: `\/+$` rescanned a long run of them
86 // from every position.
87 let end = url.pathname.length;
88 while (end > 0 && url.pathname[end - 1] === "/") end--;
89 const path = url.pathname.slice(0, end).replace(/\/v1(?:\/systemone)?$/, "");
90 return { ok: true, base: `${url.origin}${path}` };
91}
92
93function flagOn(raw: string | undefined): boolean {
94 const flag = (raw ?? "").trim().toLowerCase();
95 return flag === "1" || flag === "true" || flag === "yes" || flag === "on";
96}
97
98function typesafeProvider(
99 apiKey: string,
100 env: ProviderEnv,
101): ProviderResult {
102 const base = typesafeBaseOf(
103 env.TYPESAFE_BASE_URL,
104 env.JEV_ROUTER_ALLOW_CUSTOM_BASE,
105 );
106 if (!base.ok) return base;
107 const pinned = (env.JEV_ROUTER_JEV_MODEL ?? "").trim();
108 return {
109 ok: true,
110 name: "typesafe",
111 endpoint: `${base.base}/v1/systemone`,
112 model: pinned || TYPESAFE_MODEL_DEFAULT,
113 apiKey,
114 };
115}
116
117export function providerOf(raw: ProviderEnv): ProviderResult {
118 // A key pasted with a newline or spaces is the key without them, and one
119 // that is only whitespace is no key: it must not win over a real one.
120 const keyOf = (v: string | undefined) => (v ?? "").trim() || undefined;
121 const env = {
122 ...raw,
123 TYPESAFE_API_KEY: keyOf(raw.TYPESAFE_API_KEY),
124 AI_GATEWAY_API_KEY: keyOf(raw.AI_GATEWAY_API_KEY),
125 };
126 const chosen = chooseProvider(env);
127 if (!chosen.ok) return chosen;
128 // Only what the chosen provider uses is checked: a stale key for the
129 // other one, or a pinned model the gateway ignores, blocks nothing.
130 // A key with a line break or other character inside it cannot go in a
131 // header: the request would fail with an error quoting the header, key
132 // and all, into the route line and the store. Refused up front instead.
133 const keyName = chosen.name === "typesafe" ? "TYPESAFE_API_KEY" : "AI_GATEWAY_API_KEY";
134 if (!/^[\x21-\x7e]+$/.test(chosen.apiKey))
135 return { ok: false, reason: `${keyName} has a line break, space or other character a key cannot have` };
136 if (chosen.name === "typesafe" && !/^[\w.:/@+-]{1,100}$/.test(chosen.model))
137 return { ok: false, reason: "JEV_ROUTER_JEV_MODEL is not a model id" };
138 return chosen;
139}
140
141function chooseProvider(env: ProviderEnv): ProviderResult {
142 const named = (env.JEV_ROUTER_PROVIDER ?? "").toLowerCase().trim();
143 // The gateway's other names are read as the gateway; anything else
144 // unknown falls back to choosing by the keys set, as before.
145 const forced =
146 named === "vercel" || named === "ai-gateway" || named === "ai_gateway"
147 ? "gateway"
148 : named === "direct"
149 ? "typesafe"
150 : named;
151
152 if (forced === "typesafe") {
153 if (!env.TYPESAFE_API_KEY) {
154 return {
155 ok: false,
156 reason: "JEV_ROUTER_PROVIDER=typesafe but TYPESAFE_API_KEY is not set",
157 };
158 }
159 return typesafeProvider(env.TYPESAFE_API_KEY, env);
160 }
161
162 if (forced === "gateway") {
163 if (!env.AI_GATEWAY_API_KEY) {
164 return {
165 ok: false,
166 reason: "JEV_ROUTER_PROVIDER=gateway but AI_GATEWAY_API_KEY is not set",
167 };
168 }
169 return {
170 ok: true,
171 name: "gateway",
172 endpoint: `${GATEWAY_BASE}/v1/evaluate`,
173 model: "typesafe-ai/jev",
174 apiKey: env.AI_GATEWAY_API_KEY,
175 };
176 }
177
178 if (env.TYPESAFE_API_KEY) {
179 return typesafeProvider(env.TYPESAFE_API_KEY, env);
180 }
181
182 if (env.AI_GATEWAY_API_KEY) {
183 return {
184 ok: true,
185 name: "gateway",
186 endpoint: `${GATEWAY_BASE}/v1/evaluate`,
187 model: "typesafe-ai/jev",
188 apiKey: env.AI_GATEWAY_API_KEY,
189 };
190 }
191
192 return {
193 ok: false,
194 reason: "no TYPESAFE_API_KEY or AI_GATEWAY_API_KEY",
195 };
196}
197hooks/compactor.ts 304 lines1/**
2 * Compaction by Jev: instead of the engine's summary, every tool call in the
3 * transcript is scored — one Jev request for a typical conversation, split
4 * into a few (capped at 2 in flight at once) once there are enough calls to
5 * outgrow one request's token budget — and the stale ones are dropped or
6 * cut, so what stays is the conversation itself, verbatim. The scoring is
7 * the vendored fast-jev-compaction library (hooks/compaction/); this file is
8 * what ties it to the plugin's provider, its settings and `/jev`.
9 *
10 * Only TypeSafe direct serves it: the questions are `noul` (a probability
11 * for a yes/no), which the Vercel gateway rejects. Anything that goes wrong
12 * — no key, the gateway, a timeout, too little removed — leaves the engine's
13 * own compaction to run, and `/jev` says why.
14 */
15
16import { compact, reductionRatio, resolveOptions } from "./compaction/compact.ts";
17import { buildJevRequest, parseJevResponse } from "./compaction/request.ts";
18import type {
19 CompactOptions,
20 CompactResult,
21 JevAsker,
22 Message,
23 ToolResult,
24 ToolUse,
25} from "./compaction/types.ts";
26import { withoutBearer, messageOf, providerNoteOf, withoutKey, type HttpInitLike, type HttpResponseLike } from "./jev.ts";
27import { flagOff, PLAIN_DECIMAL } from "./policy.ts";
28import { hasNotification, notificationOf, notificationStateOf, words } from "./status.ts";
29import type { ProviderResult } from "./provider.ts";
30
31/** Below this share removed, the engine's summary does better; its default. */
32export const MIN_REDUCTION = 0.25;
33
34/**
35 * How long the whole scoring may take before the engine's summary runs
36 * instead. The hook itself has a 10-second budget that a wait counts
37 * against, so the cap stays under it: a hook the engine kills leaves no
38 * record and resets nothing.
39 */
40export const DEFAULT_COMPACT_TIMEOUT_MS = 8_000;
41const MAX_COMPACT_TIMEOUT_MS = 8_000;
42
43/** `JEV_ROUTER_COMPACT`: on unless `0`, `false`, `no`, `off` or `none`. */
44export function compactOnOf(raw: string | undefined): boolean {
45 return !flagOff(raw);
46}
47
48/**
49 * Below this a scoring budget is taken for a mistake (seconds written as
50 * `8`), as for JEV_ROUTER_TIMEOUT_MS; a small budget above it is honoured.
51 */
52export const MIN_COMPACT_TIMEOUT_MS = 100;
53
54/** `JEV_ROUTER_COMPACT_TIMEOUT_MS`, clamped; the default when unset or bad. */
55export function compactTimeoutOf(raw: string | undefined): number {
56 const v = (raw ?? "").trim();
57 const n = Number(v);
58 if (!PLAIN_DECIMAL.test(v) || n < MIN_COMPACT_TIMEOUT_MS)
59 return DEFAULT_COMPACT_TIMEOUT_MS;
60 return Math.min(n, MAX_COMPACT_TIMEOUT_MS);
61}
62
63/**
64 * `JEV_ROUTER_COMPACT_MIN_REDUCTION`: a share 0–1; the default when unset or
65 * bad. A number past 1 is a percentage, and so is anything written with `%`,
66 * whatever its size: `1%` is a hundredth, not all of it.
67 */
68export function minReductionOf(raw: string | undefined): number {
69 const trimmed = (raw ?? "").trim();
70 const percent = trimmed.endsWith("%");
71 const v = percent ? trimmed.slice(0, -1).trim() : trimmed;
72 // Plain decimals, as the other settings: `0x19` or `1e1` is a mistake.
73 if (!PLAIN_DECIMAL.test(v)) return MIN_REDUCTION;
74 const n = Number(v);
75 return percent || n > 1 ? Math.min(n / 100, 1) : n;
76}
77
78/** Why a pruning that removed too little does not stand; undefined when it does. */
79export function shortOf(reduction: number, minReduction: number): string | undefined {
80 if (reduction >= minReduction) return undefined;
81 // The removed share rounded down and the needed one up, so they never
82 // read "only 25% removed, needs 25%" (a bar of 25.4% needs 26%); the
83 // epsilons keep 0.29 × 100 = 28.999… at 29 and 0.25 × 100 at 25.
84 const needs = Math.ceil(minReduction * 100 - 1e-9);
85 const removed = Math.min(Math.floor(reduction * 100 + 1e-9), needs - 1);
86 return `only ${Math.max(0, removed)}% removed, needs ${needs}%`;
87}
88
89/** A transcript message as the engine hands it to `session.compact`. */
90export type EngineMessage = Message & { handle?: string };
91
92/** What one compaction came to, for `/jev` and the store. */
93export type Compaction = {
94 at: number;
95 /** Messages after and before. */
96 kept: number;
97 of: number;
98 /** Share of characters removed, 0–1. */
99 reduction: number;
100 /** Tool calls kept whole, cut to their head, and removed. */
101 calls: { kept: number; cut: number; dropped: number };
102 ms: number;
103 /** Why the engine's own summary ran instead; absent when Jev's stood. */
104 fallback?: string;
105};
106
107export type PruneResult =
108 | { ok: true; messages: EngineMessage[]; compaction: Compaction }
109 | { ok: false; compaction: Compaction };
110
111/** A message's text as the scoring shows it to Jev: a notification by its summary alone. */
112function stateTextOf(message: Message): string {
113 const notice = notificationOf(message.text) !== null || hasNotification(message.text);
114 return message.role === "user" && notice ? notificationStateOf(message.text) : message.text;
115}
116
117/** A `JevAsker` over the plugin's provider and the engine's fetch. */
118function askerOf(
119 provider: Extract<ProviderResult, { ok: true }>,
120 fetch: (url: string, init?: HttpInitLike) => Promise<HttpResponseLike>,
121 signal?: AbortSignal,
122): JevAsker {
123 return {
124 async ask(state, questions) {
125 const request = buildJevRequest(
126 { apiKey: provider.apiKey, model: provider.model, baseUrl: provider.endpoint },
127 state,
128 questions,
129 );
130 const response = await fetch(request.url, {
131 method: request.method,
132 headers: request.headers,
133 body: request.body,
134 // Passed for a fetch that honours it; the engine's does not yet,
135 // so a timed-out request still runs to completion (at Jev's flat
136 // per-call price).
137 ...(signal !== undefined ? { signal } : {}),
138 });
139 // The provider's body is not shown: it can echo the key or the state.
140 // Its status and error type say enough, as on a routed turn.
141 if (!response.ok) throw new Error(`${provider.name} said HTTP ${response.status}${providerNoteOf(response)}`);
142 return parseJevResponse(response.status, response.ok, response.text);
143 },
144 };
145}
146
147/**
148 * Maps the library's output back onto the engine's messages. A message the
149 * library left alone is the engine's own object, handle and all, so the
150 * engine keeps it whole; one it rebuilt has no handle, and the engine
151 * takes its edited content.
152 */
153export function toEngineMessages(
154 input: readonly EngineMessage[],
155 output: readonly Message[],
156): EngineMessage[] {
157 const own = new Set<Message>(input);
158 const uses = new Set<ToolUse>();
159 const results = new Set<ToolResult>();
160 for (const m of input) {
161 for (const t of m.toolUses) uses.add(t);
162 for (const r of m.toolResults ?? []) results.add(r);
163 }
164 return output.map((m) => {
165 if (own.has(m)) return m as EngineMessage;
166 const rebuilt: EngineMessage = {
167 role: m.role,
168 text: m.text,
169 toolUses: m.toolUses.map((t) => (uses.has(t) ? t : withoutFalse(t))),
170 };
171 if (m.toolResults && m.toolResults.length > 0)
172 // A result's isError is a plain boolean to the engine, false included.
173 rebuilt.toolResults = m.toolResults.map((r) => (results.has(r) ? r : { ...r, isError: r.isError === true }));
174 return rebuilt;
175 });
176}
177
178/** A rebuilt tool use without `isError: false`: the engine spells a use's `true | undefined`. */
179function withoutFalse<T extends { isError?: boolean }>(block: T): T {
180 if (block.isError) return { ...block };
181 const { isError: _, ...rest } = block;
182 void _;
183 return rest as T;
184}
185
186function compactionOf(result: CompactResult, ms: number): Compaction {
187 return {
188 at: Date.now(),
189 kept: result.stats.messagesAfter,
190 of: result.stats.messagesBefore,
191 reduction: reductionRatio(result),
192 calls: {
193 // Pinned calls (the recent zone) are kept whole too.
194 kept: result.stats.kept + result.stats.pinned,
195 cut: result.stats.resultsDropped,
196 dropped: result.stats.callsDropped,
197 },
198 ms,
199 };
200}
201
202/**
203 * Scores the transcript with Jev and returns what to keep, or why the
204 * engine's summary should run instead. Never throws.
205 */
206export async function pruneTranscript(args: {
207 messages: readonly EngineMessage[];
208 provider: ProviderResult;
209 fetch: (url: string, init?: HttpInitLike) => Promise<HttpResponseLike>;
210 sleep: (ms: number, options?: { signal?: AbortSignal }) => Promise<unknown>;
211 timeoutMs: number;
212 minReduction: number;
213 options?: CompactOptions;
214 now?: () => number;
215}): Promise<PruneResult> {
216 const now = args.now ?? (() => Date.now());
217 const started = now();
218 const none = (fallback: string): PruneResult => ({
219 ok: false,
220 compaction: {
221 at: Date.now(),
222 kept: args.messages.length,
223 of: args.messages.length,
224 reduction: 0,
225 calls: { kept: 0, cut: 0, dropped: 0 },
226 ms: now() - started,
227 fallback,
228 },
229 });
230 if (!args.provider.ok) return none(args.provider.reason);
231 if (args.provider.name !== "typesafe")
232 return none("the gateway does not answer yes/no questions; needs TYPESAFE_API_KEY");
233
234 const TIMED_OUT = Symbol("timed-out");
235 const controller = new AbortController();
236 // Ended in the finally, so the timeout does not run on after the scoring.
237 const timer = new AbortController();
238 try {
239 const work = compact(
240 args.messages,
241 askerOf(args.provider, args.fetch, controller.signal),
242 // A task's notification in the transcript is shown by its summary,
243 // never its result, as a notification turn is.
244 resolveOptions({ ...(args.options ?? {}), textOf: stateTextOf }),
245 controller.signal,
246 );
247 const raced = await Promise.race([
248 work,
249 args.sleep(args.timeoutMs, { signal: timer.signal }).then(
250 () => TIMED_OUT,
251 () => new Promise<never>(() => {}),
252 ),
253 ]);
254 if (raced === TIMED_OUT) {
255 controller.abort();
256 void work.catch(() => undefined);
257 return none(`timed out after ${args.timeoutMs}ms`);
258 }
259 const result = raced as CompactResult;
260 const compaction = compactionOf(result, now() - started);
261 // A result the library marked for cutting is left whole when it is
262 // already short: count what actually changed, not what was marked.
263 const originals = new Set<ToolResult>(args.messages.flatMap((m) => m.toolResults ?? []));
264 const cut = result.messages.flatMap((m) => m.toolResults ?? []).filter((r) => !originals.has(r)).length;
265 compaction.calls = {
266 kept: compaction.calls.kept + compaction.calls.cut - cut,
267 cut,
268 dropped: compaction.calls.dropped,
269 };
270 const short = shortOf(compaction.reduction, args.minReduction);
271 if (short !== undefined) return { ok: false, compaction: { ...compaction, fallback: short } };
272 return { ok: true, messages: toEngineMessages(args.messages, result.messages), compaction };
273 } catch (error) {
274 // Control characters go first, so none can sit between "Bearer" and a
275 // token and hide it; the token is judged whole, before any cut.
276 const detail = withoutBearer(
277 withoutKey(messageOf(error), args.provider.ok ? args.provider.apiKey : undefined)
278 // eslint-disable-next-line no-control-regex
279 .replace(/[\x00-\x08\x0e-\x1f\x7f-\x9f]/g, " "),
280 );
281 // Shown in /jev and saved: plain words only, whatever the provider sent.
282 return none(
283 detail
284 .replace(/\s+/g, " ")
285 .replace(/[`*_#<>\[\]()|]/g, "")
286 .slice(0, 120),
287 );
288 } finally {
289 timer.abort();
290 }
291}
292
293/** `kept 41/87 messages, 63% smaller (12 calls kept, 9 cut, 30 dropped) · 2.1s`, or why not. */
294export function compactionLine(c: Compaction): string {
295 const when = c.ms >= 1000 ? `${(c.ms / 1000).toFixed(1)}s` : `${Math.round(c.ms)}ms`;
296 // Plain words whatever the store held: a fallback restored from a snapshot
297 // is printed as it was saved.
298 if (c.fallback !== undefined) return `engine summary: ${words(c.fallback)} · ${when}`;
299 return (
300 `kept ${c.kept}/${c.of} messages, ${Math.round(c.reduction * 100)}% smaller ` +
301 `(${c.calls.kept} calls kept, ${c.calls.cut} cut, ${c.calls.dropped} dropped) · ${when}`
302 );
303}
304hooks/compaction/compact.ts 390 lines1// Vendored from fast-jev-compaction (https://github.com/tamaratran/fast-jev-compaction)
2// commit e3f262a7f4d4, MIT licensed; see LICENSE-fast-jev-compaction. Imports
3// renamed to .ts. Changed on top of the vendored source: `concurrentMap`
4// (below) was added, and `compact()`'s batch `Promise.all` replaced by it, to
5// cap batch concurrency; `compact()` takes an optional `signal` to stop
6// starting batches. Everything else in this file is unchanged.
7
8import { noulAnswer } from "./request.ts";
9import { collectToolCalls, estimateTokens, fitState } from "./state.ts";
10import type {
11 CallAnswer,
12 CallDecision,
13 CompactOptions,
14 CompactResult,
15 CompactionState,
16 JevAsker,
17 JevQuestions,
18 Message,
19 ResolvedCompactOptions,
20 ToolCall,
21 ToolUse,
22} from "./types.ts";
23
24export const DEFAULT_OPTIONS: ResolvedCompactOptions = {
25 goal: '',
26 keepThreshold: 0.5,
27 preserveRecentMessages: 6,
28 maxStateTokens: 25_000,
29 maxRequestTokens: 30_000,
30 truncateHeadChars: 300,
31};
32
33/** Tokens the request envelope (`model`, key names) adds around state and questions. */
34const REQUEST_OVERHEAD_TOKENS = 20;
35
36function finite(value: number | undefined, fallback: number): number {
37 return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
38}
39
40export function resolveOptions(options: CompactOptions = {}): ResolvedCompactOptions {
41 return {
42 goal: options.goal ?? DEFAULT_OPTIONS.goal,
43 keepThreshold: finite(options.keepThreshold, DEFAULT_OPTIONS.keepThreshold),
44 preserveRecentMessages: Math.max(
45 0,
46 Math.floor(
47 finite(options.preserveRecentMessages, DEFAULT_OPTIONS.preserveRecentMessages),
48 ),
49 ),
50 maxStateTokens: Math.max(1, finite(options.maxStateTokens, DEFAULT_OPTIONS.maxStateTokens)),
51 maxRequestTokens: Math.max(
52 1,
53 finite(options.maxRequestTokens, DEFAULT_OPTIONS.maxRequestTokens),
54 ),
55 truncateHeadChars: Math.max(
56 0,
57 Math.floor(finite(options.truncateHeadChars, DEFAULT_OPTIONS.truncateHeadChars)),
58 ),
59 ...(options.textOf ? { textOf: options.textOf } : {}),
60 };
61}
62
63/** The two `noul` questions asked about one call: keep the call, keep its result. */
64export function questionsFor(call: ToolCall): JevQuestions {
65 return {
66 [`call_${call.id}`]: {
67 type: 'noul',
68 instructions: `Tool call ${call.id} (${call.tool}) should stay in the history: knowing this call was made, with its input, still matters for what the assistant does next`,
69 },
70 [`result_${call.id}`]: {
71 type: 'noul',
72 instructions: `The full output of tool call ${call.id} (${call.tool}, ${call.resultChars} chars) should stay in the history verbatim: the assistant still needs its contents and re-running the tool would not do`,
73 },
74 };
75}
76
77/**
78 * Splits the candidate calls into batches whose questions, together with the
79 * (always complete) state, fit one request.
80 */
81export function batchCalls(
82 calls: readonly ToolCall[],
83 stateTokens: number,
84 options: Pick<ResolvedCompactOptions, 'maxRequestTokens'>,
85): ToolCall[][] {
86 const budget = options.maxRequestTokens - stateTokens - REQUEST_OVERHEAD_TOKENS;
87 const batches: ToolCall[][] = [];
88 let current: ToolCall[] = [];
89 let currentTokens = 0;
90 for (const call of calls) {
91 const tokens = estimateTokens(JSON.stringify(questionsFor(call)));
92 if (current.length > 0 && currentTokens + tokens > budget) {
93 batches.push(current);
94 current = [];
95 currentTokens = 0;
96 }
97 if (current.length === 0 && tokens > budget) {
98 throw new Error(
99 `state leaves no room for questions (~${stateTokens} of ${options.maxRequestTokens} tokens)`,
100 );
101 }
102 current.push(call);
103 currentTokens += tokens;
104 }
105 if (current.length > 0) batches.push(current);
106 return batches;
107}
108
109export function decideCall(
110 call: Pick<ToolCall, 'id' | 'tool' | 'pinned'>,
111 answer: CallAnswer,
112 options: Pick<ResolvedCompactOptions, 'keepThreshold'>,
113): CallDecision {
114 const base = { id: call.id, tool: call.tool, ...answer };
115 if (call.pinned) return { ...base, action: 'keep', reason: 'pinned' };
116 if (answer.keepResult >= options.keepThreshold) {
117 return { ...base, action: 'keep', reason: 'kept' };
118 }
119 if (answer.keepCall >= options.keepThreshold) {
120 return { ...base, action: 'drop_result', reason: 'result_dropped' };
121 }
122 return { ...base, action: 'drop_call', reason: 'call_dropped' };
123}
124
125async function askBatch(
126 asker: JevAsker,
127 state: CompactionState,
128 batch: readonly ToolCall[],
129): Promise<Map<string, CallAnswer>> {
130 const questions: JevQuestions = Object.assign({}, ...batch.map(questionsFor));
131 const { answers } = await asker.ask(state, questions);
132 return new Map(
133 batch.map((call) => [
134 call.id,
135 {
136 keepCall: noulAnswer(answers, `call_${call.id}`),
137 keepResult: noulAnswer(answers, `result_${call.id}`),
138 },
139 ]),
140 );
141}
142
143/**
144 * Runs `fn` over `items` with at most `concurrency` in flight at once, in
145 * item order (results are collected by index, not completion order, so
146 * this is a true `map`, not a fire-and-forget pool). Useful for controlling
147 * resource use when scoring large transcript fragments in many batches.
148 *
149 * If `signal` aborts, no new work starts and the promise rejects straight
150 * away, with calls already started left to settle on their own. None of
151 * them becomes an unhandled rejection, since each already has handlers
152 * attached; calls already sent to the network complete regardless, since aborting the
153 * signal only cancels a fetch that itself honours it (see askJev's note —
154 * the engine's own fetch today does not).
155 */
156async function concurrentMap<T, U>(
157 items: readonly T[],
158 fn: (item: T) => Promise<U>,
159 concurrency: number = 2,
160 signal?: AbortSignal,
161): Promise<U[]> {
162 if (signal?.aborted) throw new Error("aborted");
163 const results: U[] = new Array(items.length);
164 const active = new Set<Promise<void>>();
165 let aborted = false;
166 // Resolved (not rejected — nothing here should reach an unhandled state)
167 // the moment `signal` aborts, so a wait for a free slot wakes up right
168 // away instead of only noticing abort on its next loop iteration.
169 let wakeAborted: () => void = () => {};
170 const abortedWake = new Promise<void>((resolve) => { wakeAborted = resolve; });
171 const onAbort = () => { aborted = true; wakeAborted(); };
172 signal?.addEventListener("abort", onAbort);
173 try {
174 for (const [i, item] of items.entries()) {
175 if (aborted) throw new Error("aborted");
176 // `p` closes over itself so its own settlement removes itself from
177 // `active` — adding the `.finally()` wrapper instead (a *different*
178 // promise) while deleting the original left `active` growing forever,
179 // so the concurrency cap silently stopped limiting after the first
180 // batch settled (measured: 67 batches → 66 in flight at once).
181 const p: Promise<void> = fn(item)
182 .then((result) => {
183 results[i] = result;
184 })
185 .finally(() => {
186 active.delete(p);
187 });
188 active.add(p);
189 if (active.size >= concurrency) await Promise.race([...active, abortedWake]);
190 }
191 await Promise.all(active);
192 } finally {
193 signal?.removeEventListener("abort", onAbort);
194 }
195 if (aborted) {
196 // Let whatever is still in flight settle before returning, so none of
197 // it becomes an unhandled rejection after this function has returned.
198 await Promise.all(active).catch(() => undefined);
199 throw new Error("aborted");
200 }
201 return results;
202}
203
204function truncatedResultText(text: string, isError: boolean, headChars: number): string {
205 if (text.length <= headChars + 120) return text;
206 const head = headChars > 0 ? `${text.slice(0, headChars)}\n` : '';
207 return `${head}[fast-jev-compaction truncated ${text.length - headChars} chars of this tool result${
208 isError ? ' (error)' : ''
209 }; re-run the tool if needed]`;
210}
211
212/**
213 * Rebuilds the conversation from the decisions. A dropped call disappears
214 * together with its result; a dropped result keeps a bounded head and note.
215 * Messages that lose all their content are removed; untouched messages are
216 * returned as the same objects they came in as.
217 */
218export function applyDecisions(
219 messages: readonly Message[],
220 decisions: readonly CallDecision[],
221 calls: readonly ToolCall[],
222 headChars: number,
223): Message[] {
224 const byId = new Map(calls.map((call) => [call.id, call]));
225 const actions = new Map<string, CallDecision['action']>();
226 for (const decision of decisions) {
227 const call = byId.get(decision.id);
228 if (call && decision.action !== 'keep') actions.set(call.tool_use_id, decision.action);
229 }
230 const kept: Message[] = [];
231 for (const message of messages) {
232 const touched =
233 message.toolUses.some((tool) => actions.has(tool.tool_use_id)) ||
234 (message.toolResults ?? []).some((result) => actions.has(result.tool_use_id));
235 if (!touched) {
236 kept.push(message);
237 continue;
238 }
239 const toolUses = message.toolUses
240 .filter((tool) => actions.get(tool.tool_use_id) !== 'drop_call')
241 .map((tool) => {
242 if (actions.get(tool.tool_use_id) !== 'drop_result') return tool;
243 const text = truncatedResultText(
244 tool.text ?? '',
245 tool.isError ?? false,
246 headChars,
247 );
248 if ((tool.text ?? '') === text) return tool;
249 const copy: ToolUse = {
250 tool_use_id: tool.tool_use_id,
251 tool: tool.tool,
252 input: tool.input,
253 text,
254 };
255 if (tool.isError) copy.isError = true;
256 return copy;
257 });
258 const toolResults = (message.toolResults ?? [])
259 .filter((result) => actions.get(result.tool_use_id) !== 'drop_call')
260 .map((result) => {
261 if (actions.get(result.tool_use_id) !== 'drop_result') return result;
262 const text = truncatedResultText(result.text, result.isError ?? false, headChars);
263 return text === result.text
264 ? result
265 : {
266 tool_use_id: result.tool_use_id,
267 text,
268 isError: result.isError,
269 };
270 });
271 if (
272 !message.toolUses.some(
273 (tool) => actions.get(tool.tool_use_id) === 'drop_call',
274 ) &&
275 !(message.toolResults ?? []).some(
276 (result) => actions.get(result.tool_use_id) === 'drop_call',
277 ) &&
278 toolUses.every((tool, index) => tool === message.toolUses[index]) &&
279 toolResults.every(
280 (result, index) => result === message.toolResults?.[index],
281 )
282 ) {
283 kept.push(message);
284 continue;
285 }
286 if (message.text.trim().length === 0 && toolUses.length === 0 && toolResults.length === 0) {
287 continue;
288 }
289 const rebuilt: Message = { role: message.role, text: message.text, toolUses };
290 if (toolResults.length > 0) rebuilt.toolResults = toolResults;
291 kept.push(rebuilt);
292 }
293 return kept;
294}
295
296/** Characters of text, tool input and tool output a message holds. */
297export function messageChars(message: Message): number {
298 let total = message.text.length;
299 for (const tool of message.toolUses) {
300 try {
301 total += JSON.stringify(tool.input).length;
302 } catch {
303 total += 20;
304 }
305 }
306 for (const result of message.toolResults ?? []) total += result.text.length;
307 return total;
308}
309
310export function reductionRatio(result: Pick<CompactResult, 'stats'>): number {
311 const { charsBefore, charsAfter } = result.stats;
312 return charsBefore === 0 ? 0 : (charsBefore - charsAfter) / charsBefore;
313}
314
315function count(decisions: readonly CallDecision[], reason: CallDecision['reason']): number {
316 return decisions.filter((decision) => decision.reason === reason).length;
317}
318
319/**
320 * Compacts a transcript by asking Jev, for every tool call outside the pinned
321 * first and newest messages, whether the call and whether its result must
322 * stay. The whole history (results omitted, fitted into `maxStateTokens`) is
323 * sent as state with every batch of questions. Throws when Jev fails or the
324 * history cannot be fitted; the caller decides whether to fall back.
325 */
326export async function compact(
327 messages: readonly Message[],
328 asker: JevAsker,
329 options: CompactOptions = {},
330 signal?: AbortSignal,
331): Promise<CompactResult> {
332 const started = Date.now();
333 const resolved = resolveOptions(options);
334 const calls = collectToolCalls(messages, resolved.preserveRecentMessages);
335 const candidates = calls.filter((call) => !call.pinned);
336 const charsBefore = messages.reduce((sum, message) => sum + messageChars(message), 0);
337
338 let fitted: { tokens: number; stage: string } = { tokens: 0, stage: '' };
339 let batches: ToolCall[][] = [];
340 const answers = new Map<string, CallAnswer>();
341 if (candidates.length > 0) {
342 const state = fitState(messages, calls, resolved);
343 fitted = state;
344 batches = batchCalls(candidates, state.tokens, resolved);
345 // Limit batch concurrency to 2 to avoid overwhelming the provider or local
346 // resources when scoring large transcripts. The caller's signal (its own
347 // timeout, enforced outside this function) stops any further batches from
348 // being launched once it fires; batches already in flight are cancelled at
349 // the network level only where the asker's own fetch honours the signal
350 // baked into it (see askerOf in compactor.ts) — otherwise they still run
351 // to completion and their answers are discarded.
352 const answered = await concurrentMap(
353 batches,
354 (batch) => askBatch(asker, state.state, batch),
355 2,
356 signal,
357 );
358 for (const map of answered) for (const [id, answer] of map) answers.set(id, answer);
359 }
360
361 const decisions = calls.map((call) =>
362 decideCall(call, answers.get(call.id) ?? { keepCall: 1, keepResult: 1 }, resolved),
363 );
364 const kept = applyDecisions(
365 messages,
366 decisions,
367 calls,
368 resolved.truncateHeadChars,
369 );
370 return {
371 messages: kept,
372 decisions,
373 stats: {
374 messagesBefore: messages.length,
375 messagesAfter: kept.length,
376 charsBefore,
377 charsAfter: kept.reduce((sum, message) => sum + messageChars(message), 0),
378 calls: calls.length,
379 kept: count(decisions, 'kept'),
380 resultsDropped: count(decisions, 'result_dropped'),
381 callsDropped: count(decisions, 'call_dropped'),
382 pinned: count(decisions, 'pinned'),
383 stateTokens: fitted.tokens,
384 stateStage: fitted.stage,
385 requests: batches.length,
386 ms: Date.now() - started,
387 },
388 };
389}
390hooks/status.ts 1857 lines1/**
2 * What `/jev` prints, the line at the top of a reply, and the summary under
3 * it. `/jev` is the router's only guaranteed-visible surface: a command's
4 * output row draws on every surface, where a footer label may not, so
5 * anything you need to be sure of belongs here.
6 *
7 * Everything shown to the person is in short plain words: `kept fable:
8 * haiku costs $4.41 vs $0.13`, not `held:haiku·$4.41>$0.125`.
9 */
10
11import type { JevResult } from "./jev.ts";
12import {
13 capTo,
14 ceilingAt,
15 confidenceShareOf,
16 MODEL_OF,
17 decisionOf,
18 DEFAULT_STICKY_CONFIDENCE,
19 effortNamed,
20 EFFORTS,
21 forcedDecision,
22 holdsSonnetEffort,
23 offeredTiers,
24 stickyDecision,
25 SUBAGENT_CONFIDENCE,
26 subagentDecision,
27 TIERS,
28 UPGRADE_CONTEXT_TOKENS,
29 upgradeBar,
30 withCeiling,
31 withinWindow,
32 type Ceiling,
33 type Decision,
34 type Tier,
35} from "./policy.ts";
36import { sameModelAs,
37 baseModel,
38 breakEvenTokens,
39 isDowngrade,
40 PRICE,
41 upgradeVerdict,
42 priceOfModel,
43 switchVerdict,
44 usageCost,
45 usd,
46 fitsWindow,
47 WINDOW_TOKENS,
48 type Ttl,
49} from "./pricing.ts";
50import type { ProviderResult } from "./provider.ts";
51import { compactionLine, type Compaction } from "./compactor.ts";
52
53/**
54 * What the API said a turn cost, and which model it says answered. The
55 * shape of the engine's `TurnUsage`, spelled out here so this file stays
56 * free of engine types and runs under plain `node`.
57 */
58export type Usage = {
59 /** The model that answered, by the id the API reports. */
60 model: string;
61 input_tokens: number;
62 output_tokens: number;
63 cache_read_input_tokens: number;
64 cache_creation_input_tokens: number;
65};
66
67/**
68 * One turn's outcome, kept for the status report. `usage` arrives after the
69 * decision, from the `stop` chunk of each step, so it is filled in later and
70 * is absent for a turn still running or one whose response never came.
71 */
72export type Attempt = {
73 prompt: string;
74 ms: number;
75 usage?: Usage;
76 /** Dollars for `usage` at list price, or absent for a model without one. */
77 cost?: number;
78 /**
79 * What started the turn, when it was not the person typing a task. Absent
80 * for a typed prompt. `notify`: the main loop woke because a background
81 * task finished, and the engine's `<task-notification>` was the turn's
82 * text. `agent`: a subagent's own loop, which no `turn.start` announces;
83 * its model was settled at `agent.spawn`, and its steps carry that. Without
84 * this, one prompt that spawned three reviewers read as one reply that
85 * changed model three times. `continue`: a bare go-ahead ("yes"), which
86 * ran on the previous turn's decision without asking Jev. `nudge`: the
87 * engine's own "say what you are doing, then continue", the same way, and
88 * announced nowhere but here.
89 */
90 kind?: "notify" | "agent" | "continue" | "nudge";
91 /**
92 * Carried on from an earlier turn's decision without asking Jev: a
93 * go-ahead, the engine's nudge, or a task's notification that continued
94 * the reply's route. Its decision's confidence is that earlier turn's.
95 */
96 continued?: true;
97 /** For `kind: 'agent'`: which subagent, as `$.agent.list()` describes it. */
98 agent?: AgentTag;
99} & ({ decision: Decision } | { skipped: string });
100
101/**
102 * Which subagent a turn ran in. `type` is the definition (`general-purpose`,
103 * `Explore`); `label` its row's description (`Review library-sync cluster`),
104 * or the id when the list has no row for it yet, in which case `type` is
105 * absent too.
106 */
107export type AgentTag = {
108 type?: string;
109 label: string;
110};
111
112/** Percent, rounded, for a confidence. */
113const pct = (n: number) => `${Math.round(n * 100)}%`;
114
115/**
116 * Why a turn did not run exactly as Jev asked, in plain words: held on its
117 * previous tier for doubt or for price, held on its previous Sonnet effort,
118 * forced to a tier the prompt named, capped at the ceiling, or sent the
119 * effort a first request runs. Empty when it ran as asked.
120 */
121export function reasonsOf(attempt: Attempt): string[] {
122 if (!("decision" in attempt)) return [];
123 const d = attempt.decision;
124 const out: string[] = [];
125 if (d.held !== undefined) {
126 // Same rung, different model: a session model off the ladder.
127 const wanted =
128 d.heldModel !== undefined && d.held === d.tier ? plain(d.heldModel) : d.held;
129 const kept = d.held === d.tier ? plain(d.model) : d.tier;
130 out.push(
131 d.heldWindow !== undefined
132 ? `kept ${kept}: too long for ${wanted} (${kOf(d.heldWindow)})`
133 : d.heldCost !== undefined
134 ? `kept ${kept}: ${wanted} costs ${usd(d.heldCost.go)} vs ${usd(d.heldCost.stay)}` +
135 (d.heldCost.limit !== undefined
136 ? `, over the ${usd(d.heldCost.limit)} limit`
137 : "")
138 : `kept ${kept}: Jev ${pct(d.confidence)} on ${wanted}` +
139 (d.heldBar !== undefined ? `, needs ${pct(d.heldBar)}` : ""),
140 );
141 }
142 // Kept, unless the tier it was on had outgrown its window: then the
143 // step-up below says where it went.
144 if (d.jevFailed !== undefined)
145 out.push(d.outgrew !== undefined ? `Jev ${words(d.jevFailed)}` : `kept ${d.tier}: Jev ${words(d.jevFailed)}`);
146 if (d.outgrew !== undefined)
147 out.push(
148 `${d.outgrew} too long, moved up only to ${d.tier}` +
149 (d.wanted !== undefined && d.wanted !== d.tier ? ` (Jev wanted ${d.wanted})` : ""),
150 );
151 // Sonnet only: an effort change there re-caches half the prefix.
152 if (d.heldEffort !== undefined) {
153 out.push(
154 `kept ${d.effort}: Jev ${pct(d.effortConfidence ?? 0)} on ${d.heldEffort}`,
155 );
156 }
157 if (d.forced) out.push(pickOf(d));
158 // The ceiling, and the engine running the capped effort higher on a
159 // conversation's first request, read as one fact: what Jev wanted, what
160 // the ceiling allowed, what actually ran. A first request that ran what
161 // Jev wanted anyway was not capped in any way that matters.
162 const capped = d.cappedEffort !== undefined && d.cappedEffort !== d.effort;
163 if (capped && d.askedEffort !== undefined)
164 out.push(
165 `capped ${d.cappedEffort}→${d.askedEffort}; 1st request runs it as ${d.effort}`,
166 );
167 else if (capped) out.push(`capped from ${d.cappedEffort}`);
168 else if (d.askedEffort !== undefined)
169 out.push(`1st request runs ${d.askedEffort} as ${d.effort}`);
170 return out;
171}
172
173/**
174 * The cheapest tier above `from`, up to `to`, that takes `contextTokens`:
175 * where a turn goes when what it would run on is too small. `strictlyAbove`
176 * false lets `from` itself count. Null when none fits.
177 */
178function stepUp(
179 from: Tier,
180 to: Tier,
181 offered: readonly Tier[],
182 contextTokens: number,
183 strictlyAbove = true,
184): Tier | null {
185 const lo = TIERS.indexOf(from);
186 const hi = TIERS.indexOf(to);
187 return (
188 TIERS.find(
189 (t, i) =>
190 (strictlyAbove ? i > lo : i >= lo) &&
191 i <= hi &&
192 offered.includes(t) &&
193 fitsWindow(t, contextTokens),
194 ) ?? null
195 );
196}
197
198/**
199 * How a named tier reads: `your pick` when the turn runs on it, `you picked
200 * haiku` when it did not fit and the turn ran elsewhere (kept on the running
201 * tier, or stepped up), so the tier shown is never called the person's pick
202 * when it was not.
203 */
204function pickOf(d: Decision): string {
205 const named = d.held ?? d.wanted ?? d.outgrew ?? d.tier;
206 return named === d.tier ? "your pick" : `you picked ${named}`;
207}
208
209/** The tier a step-up says was too long: the running one if it was, else Jev's pick. */
210function outgrownOf(running: Decision | null, decision: Decision, contextTokens: number): Tier {
211 return running !== null && !fitsWindow(running.tier, contextTokens) ? running.tier : decision.tier;
212}
213
214/**
215 * A name from outside the plugin (an agent's type, a model id) as plain
216 * words for the route line, the summary's fence and `/jev`: no backticks or
217 * newlines to close the fence or start a heading, and not too long.
218 */
219export function plain(text: string): string {
220 // Only what can close the fence or open a tag: inside a fence and on one
221 // line, `#`, `|`, `_` and `[1m]` are plain text and stay.
222 const flat = oneLine(String(text).replace(/\t/g, " ")).replace(/[`<>]/g, "");
223 return flat.length > 60 ? `${flat.slice(0, 59)}…` : flat;
224}
225
226/** What started a turn nobody typed, in plain words; null for a typed prompt. */
227export function originOf(
228 attempt: Pick<Attempt, "kind" | "agent">,
229): string | null {
230 if (attempt.kind === "notify") return "task finished";
231 if (attempt.kind === "continue") return "continuing";
232 if (attempt.kind === "nudge") return "continuing";
233 if (attempt.kind === "agent")
234 return attempt.agent?.type ? `${plain(attempt.agent.type)} agent` : "agent";
235 return null;
236}
237
238/**
239 * A turn that is a task's notification: it opens with the envelope, an
240 * element after the tag (prose that opens with the tag is a prompt).
241 */
242const NOTIFICATION = { test: (text: string) => opensWithEnvelope(text) };
243/**
244 * A tag's text, found by hand: the lazy pattern it replaces rescanned to the
245 * end from every opening tag, quadratic on a run of openings with no close.
246 */
247const tagOf = (text: string, tag: string): string | undefined => {
248 const open = text.indexOf(`<${tag}>`);
249 if (open === -1) return undefined;
250 const from = open + tag.length + 2;
251 const close = text.indexOf(`</${tag}>`, from);
252 return close === -1 ? undefined : text.slice(from, close).trim();
253};
254
255/**
256 * Reads the engine's task notification, when the turn's text is one: what
257 * the row should say instead of the XML envelope.
258 */
259export function notificationOf(text: string): string | null {
260 if (!NOTIFICATION.test(text)) return null;
261 // Read before the result, as what Jev is sent is: a result can quote a
262 // summary or a task id of its own.
263 const head = text.split(/<result\b/i)[0]!;
264 return plain(tagOf(head, "summary") ?? `task ${tagOf(head, "task-id") ?? "?"}`);
265}
266
267/**
268 * What Jev is told about a notification turn: the task's one-line summary,
269 * and any text typed before the envelope, never the task's result, which
270 * can quote whatever the agent read.
271 */
272export function notificationStateOf(text: string): string {
273 // Only what cannot be a result: the text before the notification, and
274 // its summary, read before its result starts. Nothing after the result's
275 // opening is sent: a result can quote a whole envelope,
276 // `</task-notification>` and all (an agent reading this very code), so
277 // nothing after it can be told from the result. A turn that is a
278 // notification is cut at its opening tag, whatever shape the rest has.
279 const at = NOTIFICATION.test(text) ? text.search(/<task-notification\b/i) : envelopeAt(text);
280 if (at === -1) return text.trim();
281 const before = text.slice(0, at).trim();
282 const head = text.slice(at).split(/<result\b/i)[0]!;
283 const summary = plain(tagOf(head, "summary") ?? `task ${tagOf(head, "task-id") ?? "?"}`);
284 return [before, summary].filter((p) => p !== "").join("\n");
285}
286
287/**
288 * A notification's envelope, not a mention of the tag: an element follows
289 * the opening tag (the engine's fields, in any order, or a comment).
290 */
291const TAG = "<task-notification";
292
293/**
294 * Whether an envelope opens at `at`: the tag (attributes of any length, on
295 * one line), then an element or a comment. By hand, with the next `>` and
296 * newline found once and reused: a pattern either rescanned a long line
297 * of unclosed openings from each one, or, bounded, missed a tag with long
298 * attributes and let its result through.
299 */
300function envelopeOpensAt(
301 text: string,
302 at: number,
303 next: { gt: number; nl: number; seen?: number; ok?: boolean },
304): boolean {
305 const after = text.charCodeAt(at + TAG.length);
306 if (after === after && /\w/.test(String.fromCharCode(after))) return false;
307 if (next.gt !== -1 && next.gt < at) next.gt = text.indexOf(">", at);
308 if (next.nl !== -1 && next.nl < at) next.nl = text.indexOf("\n", at);
309 if (next.gt === -1 || (next.nl !== -1 && next.nl < next.gt)) return false;
310 // Openings that share a `>` share what follows it: read once.
311 if (next.seen === next.gt) return next.ok!;
312 next.seen = next.gt;
313 next.ok = elementAfter(text, next.gt + 1);
314 return next.ok;
315}
316
317/** Whether an element or a comment opens at `i`, after blank space. */
318function elementAfter(text: string, i: number): boolean {
319 while (i < text.length && /\s/.test(text[i]!)) i++;
320 if (text[i] !== "<") return false;
321 if (text.startsWith("!--", i + 1)) return true;
322 let j = i + 1;
323 if (!/[a-z]/i.test(text[j] ?? "")) return false;
324 while (j < text.length && /[\w-]/.test(text[j]!)) j++;
325 return j < text.length && /[\s/>]/.test(text[j]!);
326}
327
328/** Where the first envelope in `text` opens, or -1. Linear. */
329function firstEnvelope(text: string): number {
330 // Found in `text` itself: lowercasing it first can change its length
331 // ("İ" becomes two characters) and so every position after it. ASCII
332 // case only, as the pattern's `i` without `u` folds (no Kelvin sign).
333 const next = { gt: text.indexOf(">"), nl: text.indexOf("\n") };
334 for (const m of text.matchAll(/<task-notification/gi)) if (envelopeOpensAt(text, m.index!, next)) return m.index!;
335 return -1;
336}
337
338/** Whether `text` opens (after blank space) with an envelope. */
339function opensWithEnvelope(text: string): boolean {
340 const at = text.search(/\S/);
341 if (at === -1 || !/^<task-notification/i.test(text.slice(at, at + TAG.length))) return false;
342 return envelopeOpensAt(text, at, { gt: text.indexOf(">", at), nl: text.indexOf("\n", at) });
343}
344
345/**
346 * Where a notification starts in text the person typed, or -1: the first
347 * envelope, wherever it is and whatever quotes it. Nothing after it is
348 * trusted, since a task's result can quote anything — fences, closing
349 * tags, pastes — and whatever follows the envelope (a trailer, queued
350 * prompts) cannot be told from the result. Decided for privacy over
351 * routing: a prompt that quotes an example envelope has what follows the
352 * quote left out of what Jev grades (the model still gets all of it).
353 */
354function envelopeAt(text: string): number {
355 return firstEnvelope(text);
356}
357
358/** Whether `text` carries a task's notification anywhere, typed text before it or not. */
359export function hasNotification(text: string): boolean {
360 return envelopeAt(text) !== -1;
361}
362
363/** The task a notification is about: the agent's id, as `$.agent.list()` names it. */
364export function notificationTaskOf(text: string): string | null {
365 if (!NOTIFICATION.test(text)) return null;
366 return tagOf(text.split(/<result\b/i)[0]!, "task-id") ?? null;
367}
368
369/**
370 * Folds one step's usage into its turn: counts sum, the model is the last
371 * step's, as the engine defines a turn's usage, and the dollars are re-priced
372 * from the sum. Mutates, because the same object sits in the history and in
373 * the by-turn lookup. Returns this one step's own cost (0 when it cannot be
374 * priced), for the caller's running session total — which must count every
375 * step's real cost regardless of whether the *row's* total stays presentable
376 * (see below).
377 *
378 * `attempt.cost` is the row's own field, for display, and is deliberately
379 * `undefined` — not "however much we could price" — the moment any one of
380 * the turn's steps cannot be priced (a synthetic or unknown model): a partial
381 * dollar figure with a $ sign in front of it reads as the whole turn's cost,
382 * which it is not. That is a display choice; it must not double as the
383 * accounting for money actually spent. An earlier version conflated the two
384 * by having the caller diff `attempt.cost` before and after this call: the
385 * moment `attempt.cost` was cleared (this step or an earlier one lacked a
386 * price), the diff went negative and silently subtracted a step already
387 * billed, or if the *first* step was unpriced, `attempt.cost` stayed
388 * `undefined` forever and every later step's real cost added `0 - 0` —
389 * missing the whole turn from the session total.
390 */
391export function addUsage(
392 attempt: Attempt,
393 usage: Usage,
394 ttl: Ttl = "1h",
395): number {
396 const prior = attempt.usage;
397 attempt.usage = {
398 model: usage.model,
399 input_tokens: (prior?.input_tokens ?? 0) + usage.input_tokens,
400 output_tokens: (prior?.output_tokens ?? 0) + usage.output_tokens,
401 cache_read_input_tokens:
402 (prior?.cache_read_input_tokens ?? 0) + usage.cache_read_input_tokens,
403 cache_creation_input_tokens:
404 (prior?.cache_creation_input_tokens ?? 0) +
405 usage.cache_creation_input_tokens,
406 };
407 // Each step at the model that answered it: a turn whose steps ran on
408 // different models (a fallback) is not priced wholly at the last one's.
409 const step = usageCost(usage.model, usage, ttl);
410 if (step === null || (prior !== undefined && attempt.cost === undefined))
411 delete attempt.cost;
412 else attempt.cost = (attempt.cost ?? 0) + step;
413 return step ?? 0;
414}
415
416/**
417 * How much of what the turn's requests carried was read from cache, 0 to 1.
418 * Everything carried is uncached input plus cache reads plus cache writes;
419 * this is the cost-relevant measure, since reads bill at a tenth.
420 */
421export function cacheRatio(usage: Usage): number {
422 const carried = carriedOf(usage);
423 return carried === 0 ? 0 : usage.cache_read_input_tokens / carried;
424}
425
426/** What the last turn carried into the model: the context size, in tokens. */
427export function carriedOf(usage: Usage): number {
428 return (
429 usage.input_tokens +
430 usage.cache_read_input_tokens +
431 usage.cache_creation_input_tokens
432 );
433}
434
435/** No usage count past this: ten times the largest window. */
436const MAX_COUNT = 10_000_000;
437
438/**
439 * A usage record with every count a number. The API omits the cache fields
440 * on some paths; summed unchecked they made NaN of the turn's cost, the
441 * session's `spent` and the context every price hold reads.
442 */
443export function normalUsage(usage: {
444 model?: unknown;
445 input_tokens?: unknown;
446 output_tokens?: unknown;
447 cache_read_input_tokens?: unknown;
448 cache_creation_input_tokens?: unknown;
449}): Usage {
450 // A count the API could never mean (negative, NaN) is none: a negative
451 // one would make a cost negative and the snapshot unloadable.
452 // One larger than any window could carry is capped, so a cost or a
453 // snapshot cannot overflow to Infinity (saved as null, the snapshot lost).
454 const n = (v: unknown) =>
455 typeof v === "number" && Number.isFinite(v) && v > 0 ? Math.min(v, MAX_COUNT) : 0;
456 return {
457 model: typeof usage.model === "string" ? usage.model : "",
458 input_tokens: n(usage.input_tokens),
459 output_tokens: n(usage.output_tokens),
460 cache_read_input_tokens: n(usage.cache_read_input_tokens),
461 cache_creation_input_tokens: n(usage.cache_creation_input_tokens),
462 };
463}
464
465/**
466 * How much of a prompt is kept in the history. Only its opening is shown,
467 * and a session of long pastes kept whole outran the store's size cap, so
468 * nothing saved.
469 */
470export const PROMPT_KEPT = 400;
471export const kept = (text: string) => {
472 // Kept in the history and the store: a task's result, typed ahead of or
473 // not, is no more kept than it is sent to Jev.
474 const own = NOTIFICATION.test(text) || hasNotification(text) ? notificationStateOf(text) : text;
475 return own.length > PROMPT_KEPT ? own.slice(0, PROMPT_KEPT) : own;
476};
477
478export type Status = {
479 enabled: boolean;
480 /** The resumed session's cache expired and nothing has written it since. */
481 cold?: boolean;
482 surface: string | null;
483 provider: ProviderResult;
484 timeoutMs: number;
485 /** The confidence a switch must clear, or null when stickiness is off. */
486 sticky: number | null;
487 /** The most an upgrade may cost over staying, or null for no limit. */
488 upgradeMax?: number | null;
489 /** The price checks: on, off, or absent on older callers. */
490 price?: boolean;
491 /** The most effort each tier may be asked for. */
492 ceiling: Ceiling;
493 /** Compaction by Jev: on, and the last one; absent on older callers. */
494 compactOn?: boolean;
495 compaction?: Compaction | null;
496 /** Which prompt cache the session writes; the price of a switch depends on it. */
497 ttl: Ttl;
498 /** The context size the next turn would carry, or null before the first reply. */
499 contextTokens: number | null;
500 /** The main loop's model as `/model` shows it, or null when unknown. */
501 sessionModel: string | null;
502 /** What the main loop is running on, as the router last saw it; null when nothing yet. */
503 running: Decision | null;
504 offered: readonly Tier[];
505 excluded: readonly Tier[];
506 announce: boolean;
507 attempts: readonly Attempt[];
508 /** Dollars across every turn the router saw this session, at list price. */
509 spent: number;
510};
511
512/**
513 * What a switch is priced against: the context the turn carries, the output
514 * it is likely to produce (the last turn's, or a typical one), and which
515 * cache the session writes.
516 */
517type Economics = {
518 contextTokens: number;
519 outputTokens: number;
520 ttl: Ttl;
521 /** The running model's cache has expired (a resume after the TTL): staying is a write too. */
522 cold?: boolean;
523};
524
525/**
526 * What settles a main-loop turn beyond Jev's answer. `sticky` is the bar a
527 * switch must clear, or null when switches are free; `running` what the last
528 * routed turn ran on; `forced` a tier the prompt itself named, which takes
529 * the tier question away from Jev and from stickiness both; `ceiling` the
530 * most effort each tier may be asked for; `economics` what a downgrade is
531 * priced against, absent when nothing is known about the context yet.
532 */
533type Hold = {
534 sticky: number | null;
535 running: Decision | null;
536 forced?: Tier | null;
537 ceiling?: Ceiling;
538 economics?: Economics;
539 /** The most an upgrade may cost over staying; null or absent for no limit. */
540 upgradeMax?: number | null;
541 /** The downgrade and upgrade price checks; absent reads as off. */
542 price?: boolean;
543};
544
545/** A typical turn's output when the session has not produced one yet. */
546export const TYPICAL_OUTPUT_TOKENS = 1500;
547
548/**
549 * One turn's outcome from Jev's answer, so the three ways a turn can fail to
550 * route all land in one place and all get announced the same way.
551 *
552 * This is the only place a main-loop decision is settled: the route the
553 * engine applies and the line /jev shows are the same object, so the two
554 * cannot disagree. The order is the policy: a named tier first (it needs no
555 * answer from Jev at all — effort defaults to medium), then stickiness on
556 * the tier (Jev's doubt, or the price of a downgrade), then, for a turn that
557 * stays on Sonnet, stickiness on the effort, then the tier's ceiling.
558 */
559export function attemptOf(
560 text: string,
561 result: JevResult,
562 offered: readonly Tier[],
563 hold: Hold = { sticky: null, running: null },
564): Attempt {
565 const summary = notificationOf(text);
566 const head =
567 summary === null
568 ? { prompt: kept(text) }
569 : { prompt: kept(summary), kind: "notify" as const };
570 const forced = hold.forced ?? null;
571 // A tier `/jev tiers off` dropped since it started running must not be
572 // actively chosen or held to any more — but simply erasing `hold.running`
573 // here for every such case went too far: it also disabled the window
574 // guard's step-up (nothing to step *from*, so a turn too long for Jev's
575 // pick fell to fully unrouted instead of the next tier up) and, for a
576 // `running` that is only a placeholder seeded from the session model (no
577 // key, nothing ever routed there), stripped the price check's own real
578 // pricing data for no reason — an unrouted turn runs on that same
579 // placeholder anyway, so refusing to weigh it costs money without
580 // stopping anything. `withinWindow` and `stickyDecision` below are given
581 // `offered` instead, and refuse to *land on or stay on* an excluded tier
582 // at their own single decision points, while `hold.running` stays intact
583 // everywhere else — including as the signal that there is something to
584 // step up from, and as the real cache a switch away is priced against.
585
586 if (!result.ok && forced === null) {
587 // No answer from Jev: stay on the tier already running rather than drop
588 // to the session model, which could be a cold cache and a switch back
589 // afterwards. With nothing running, or nothing that fits, the turn is
590 // left to the session model as before.
591 // Only a tier Jev actually routed: a placeholder seeded from the
592 // session model (no key, nothing routed yet) is the session model.
593 if (hold.running !== null && hold.running.effortConfidence !== undefined) {
594 const stay = continuationOf(
595 text,
596 hold.running,
597 hold.ceiling ?? ceilingAt("max"),
598 hold.economics?.contextTokens ?? null,
599 offered,
600 );
601 if ("decision" in stay)
602 return {
603 ...head,
604 ms: result.ms,
605 decision: { ...stay.decision, confidence: 0, jevFailed: result.reason },
606 };
607 }
608 return { ...head, ms: result.ms, skipped: result.reason };
609 }
610
611 const fresh = result.ok ? decisionOf(result.answers, offered) : null;
612 let decision = forced !== null ? forcedDecision(forced, fresh) : fresh;
613 if (!decision) {
614 return {
615 ...head,
616 ms: result.ms,
617 skipped: "Jev answered but named no tier we offered",
618 };
619 }
620 // First of all, can the tier take a prompt this long? Haiku's window is
621 // 200k; a turn carrying more is refused by the API, forced or not.
622 if (hold.economics !== undefined) {
623 const fits = withinWindow(
624 decision,
625 hold.running,
626 hold.economics.contextTokens,
627 offered,
628 );
629 if (fits === null) {
630 // Neither Jev's pick nor what is running takes a context this long
631 // (haiku past its window, Jev saying haiku again). Rather than leave
632 // the turn to whatever the session model is, go up only as far as the
633 // context needs.
634 // With nothing known to be running, the session model holds the warm
635 // cache, and staying there is cheaper than a cold write anywhere.
636 const step =
637 hold.running === null
638 ? null
639 : stepUp(decision.tier, TIERS.at(-1)!, offered, hold.economics.contextTokens);
640 if (step === null) {
641 return {
642 ...head,
643 ms: result.ms,
644 skipped:
645 `too long for ${decision.tier} (${kOf(hold.economics.contextTokens)})`,
646 };
647 }
648 return {
649 ...head,
650 ms: result.ms,
651 decision: capTo(
652 {
653 tier: step,
654 model: MODEL_OF[step],
655 effort: decision.effort,
656 confidence: decision.confidence,
657 ...(decision.effortConfidence !== undefined
658 ? { effortConfidence: decision.effortConfidence }
659 : {}),
660 ...(decision.probabilities !== undefined
661 ? { probabilities: decision.probabilities }
662 : {}),
663 // What was too long: the running tier when it was (a running
664 // tier that fits but was turned off was not outgrown), else
665 // Jev's pick.
666 outgrew: outgrownOf(hold.running, decision, hold.economics.contextTokens),
667 ...(decision.tier !== outgrownOf(hold.running, decision, hold.economics.contextTokens)
668 ? { wanted: decision.tier }
669 : {}),
670 ...(decision.forced ? { forced: true as const } : {}),
671 },
672 hold.ceiling ?? ceilingAt("max"),
673 ),
674 };
675 }
676 if (fits.heldWindow !== undefined) {
677 // Held on the running tier: on Sonnet an effort change still rewrites
678 // much of the cache, so the effort gate applies here as well.
679 const held =
680 !fits.forced &&
681 hold.sticky !== null &&
682 hold.running !== null &&
683 hold.running.effortConfidence !== undefined &&
684 holdsSonnetEffort(fits, hold.running, hold.sticky, hold.ceiling ?? undefined)
685 ? { ...fits, effort: hold.running.effort, heldEffort: fits.effort }
686 : fits;
687 return {
688 ...head,
689 ms: result.ms,
690 decision: capTo(held, hold.ceiling ?? ceilingAt("max")),
691 };
692 }
693 }
694 // What Jev (or the prompt) picked, before any hold: what a hold that does
695 // not fit gives way to.
696 const picked: Decision | null = decision;
697 const priced = hold.price === true;
698 if ((hold.sticky !== null || priced) && !decision.forced) {
699 const running = hold.running;
700 const downgrade =
701 running !== null && isDowngrade(running.tier, decision.tier);
702 // The same rung on a different model — a session on `claude-opus-5`
703 // that Jev keeps on opus — is a switch to a cold cache too, and priced
704 // like a downgrade; there is no doubt to weigh, Jev agreed on the tier.
705 const lateral =
706 running !== null &&
707 running.tier === decision.tier &&
708 running.model !== decision.model;
709 const fromPrice =
710 running !== null ? (priceOfModel(running.model) ?? PRICE[running.tier]) : null;
711 const upgrade =
712 running !== null && !downgrade && !lateral && running.tier !== decision.tier;
713 const verdict =
714 !priced || running === null || fromPrice === null || hold.economics === undefined
715 ? null
716 : downgrade || lateral
717 ? switchVerdict(
718 running.tier,
719 decision.tier,
720 hold.economics.contextTokens,
721 hold.economics.outputTokens,
722 hold.economics.ttl,
723 fromPrice,
724 hold.economics.cold === true,
725 )
726 : upgrade && hold.upgradeMax != null
727 ? upgradeVerdict(
728 running.tier,
729 decision.tier,
730 hold.economics.contextTokens,
731 hold.economics.outputTokens,
732 hold.upgradeMax,
733 hold.economics.ttl,
734 fromPrice,
735 hold.economics.cold === true,
736 )
737 : null;
738 // An upgrade writes the whole context to the dearer tier; past 100k it
739 // has to be surer than the bar.
740 // Sticky off: no confidence bar, the price checks stand on their own.
741 const bar =
742 lateral || hold.sticky === null
743 ? 0
744 : !downgrade && hold.economics !== undefined
745 ? upgradeBar(hold.sticky, hold.economics.contextTokens)
746 : hold.sticky;
747 decision = stickyDecision(decision, running, bar, verdict, offered);
748 }
749 // A forced turn named its tier and runs at medium (Jev is not asked).
750 // The Sonnet effort gate does not get a vote here.
751 if (
752 !decision.forced &&
753 hold.sticky !== null &&
754 hold.running !== null &&
755 // A placeholder for what a session runs on carries no effort Jev
756 // chose; there is nothing to hold to.
757 hold.running.effortConfidence !== undefined &&
758 holdsSonnetEffort(decision, hold.running, hold.sticky, hold.ceiling ?? undefined)
759 ) {
760 decision = {
761 ...decision,
762 effort: hold.running.effort,
763 heldEffort: decision.effort,
764 };
765 }
766 // A hold must fit too: holding to haiku at 190k would send the turn where
767 // the API refuses it. Jev's pick passed the check above, so the hold gives
768 // way to it.
769 if (
770 hold.economics !== undefined &&
771 decision.held !== undefined &&
772 !fitsWindow(decision.tier, hold.economics.contextTokens) &&
773 picked !== null
774 ) {
775 // The tier the hold kept is outgrown, but what held the move (doubt,
776 // or an upgrade over the limit) still stands: go up only as far as the
777 // context needs, the cheapest tier that fits, not all the way to Jev's
778 // pick. When that is Jev's pick, it runs as picked.
779 const held = decision;
780 const step = stepUp(held.tier, picked.tier, offered, hold.economics.contextTokens, true);
781 if (step === null || step === picked.tier) {
782 decision = picked;
783 } else {
784 // What held the move priced haiku against Jev's pick; neither figure
785 // describes the step, so the line says what happened instead.
786 const { heldCost: _c, heldBar: _b, heldWindow: _w, held: _h, heldModel: _m, ...rest } = held;
787 void _c, _b, _w, _h, _m;
788 decision = { ...rest, tier: step, model: MODEL_OF[step], outgrew: held.tier, wanted: picked.tier };
789 }
790 }
791 decision = capTo(decision, hold.ceiling ?? ceilingAt("max"));
792 return { ...head, ms: result.ms, decision };
793}
794
795/**
796 * A bare go-ahead's outcome: the previous turn's decision, carried over as
797 * is. `held` and the like are dropped, since they described that turn's
798 * choice, not this one's; the `continue` tag says what happened here. The
799 * ceiling is re-applied so a change mid-session still binds.
800 */
801export function continuationOf(
802 text: string,
803 running: Decision,
804 ceiling: Ceiling = ceilingAt("max"),
805 contextTokens: number | null = null,
806 offered: readonly Tier[] = TIERS,
807): Attempt {
808 const { tier, model, effort, confidence, effortConfidence } = running;
809 // A tier `/jev tiers off` dropped since it started running is nothing
810 // safe to continue: same outcome as having nothing to continue at all
811 // (unrouted, the session model answers) rather than a switch quietly
812 // biased toward a tier the person just turned off.
813 if (!offered.includes(tier)) return continuationSkipped(text);
814 if (contextTokens !== null && !fitsWindow(tier, contextTokens)) {
815 const step = stepUp(tier, TIERS.at(-1)!, offered, contextTokens);
816 if (step === null)
817 return {
818 prompt: kept(text),
819 ms: 0,
820 kind: "continue",
821 skipped: `too long for ${tier} (${kOf(contextTokens)})`,
822 };
823 return {
824 prompt: kept(text),
825 ms: 0,
826 kind: "continue",
827 continued: true,
828 decision: capTo(
829 { tier: step, model: MODEL_OF[step], effort, confidence, effortConfidence, outgrew: tier },
830 ceiling,
831 ),
832 };
833 }
834 return {
835 prompt: kept(text),
836 ms: 0,
837 kind: "continue",
838 continued: true,
839 decision: capTo(
840 { tier, model, effort, confidence, effortConfidence },
841 ceiling,
842 ),
843 };
844}
845
846/**
847 * A bare go-ahead when there is nothing safe to continue (first turn, or the
848 * previous turn was unrouted / routing was off). Jev must not be asked: it
849 * grades these as trivial at ~1.00 and would clear any sticky bar. The turn
850 * stays on the session model.
851 */
852export function continuationSkipped(text: string): Attempt {
853 return {
854 prompt: kept(text),
855 ms: 0,
856 kind: "continue",
857 skipped: "nothing to continue",
858 };
859}
860
861/**
862 * A spawned subagent's outcome from Jev's answer to its task. No stickiness
863 * and no forcing: a subagent starts with an empty conversation, so there is
864 * no cache to hold to, and the tier named in the person's prompt was for the
865 * main loop. What there is instead is a confidence floor (`SUBAGENT_CONFIDENCE`),
866 * below which the spawn is left on the model it would have had anyway.
867 */
868export function spawnAttemptOf(
869 description: string,
870 result: JevResult,
871 offered: readonly Tier[],
872 agent: AgentTag,
873 ceiling: Ceiling = ceilingAt("max"),
874): Attempt {
875 const head = { prompt: kept(description), kind: "agent" as const, agent };
876 if (!result.ok) return { ...head, ms: result.ms, skipped: result.reason };
877 const fresh = decisionOf(result.answers, offered);
878 if (!fresh) {
879 return {
880 ...head,
881 ms: result.ms,
882 skipped: "Jev answered but named no tier we offered",
883 };
884 }
885 const decision = subagentDecision(fresh);
886 if (!decision) {
887 return {
888 ...head,
889 ms: result.ms,
890 skipped: `Jev ${pct(fresh.confidence)} on ${fresh.tier}, needs ${pct(SUBAGENT_CONFIDENCE)}`,
891 };
892 }
893 return { ...head, ms: result.ms, decision: capTo(decision, ceiling) };
894}
895
896/** The last few turns, newest first, so the report stays one screen. */
897export const HISTORY_LIMIT = 5;
898
899function shorten(text: string, width = 44): string {
900 // A prompt or an agent's description can carry a terminal's escape
901 // sequences (a model wrote it): dropped, with every line break folded.
902 const flat = text.replace(CONTROL, "").replace(/[\s\u0085\u2028\u2029]+/g, " ").trim();
903 return flat.length > width ? `${flat.slice(0, width - 1)}…` : flat;
904}
905
906/**
907 * `Jev 57%`: how sure Jev was of the tier. Empty when Jev was not asked this
908 * turn: a named tier, a go-ahead, or the engine's nudge.
909 */
910function sureOf(d: Decision, kind?: Attempt["kind"], continued?: boolean): string {
911 // A notification that carried the reply's route on was not asked about.
912 if (continued) return "";
913 if (kind === "continue" || kind === "nudge") return "";
914 if (d.forced && d.confidence === 0) return "";
915 if (d.jevFailed !== undefined) return "";
916 return `Jev ${pct(d.confidence)}`;
917}
918
919/** `opus-5-5` for `claude-opus-5-5-20260901`: the id without its prefix and date. */
920function shortModel(model: string): string {
921 // A usage record without a model id must not throw inside turn.step.
922 if (typeof model !== "string" || model === "") return "unknown model";
923 return plain(model.replace(/^claude-/, "").replace(/-\d{8}$/, ""));
924}
925
926function attemptLine(attempt: Attempt): string {
927 const when = `${String(attempt.ms).padStart(4)}ms`;
928 const origin = originOf(attempt);
929 const what = `${origin ? `[${origin}] ` : ""}${shorten(attempt.prompt)}`;
930 if ("skipped" in attempt) {
931 // A subagent's row names the agent first, then why: "under the bar"
932 // and "Jev timed out" are different stories.
933 return attempt.kind === "agent"
934 ? ` ${when} not routed — ${what} · ${words(attempt.skipped)}`
935 : ` ${when} not routed — ${words(attempt.skipped)}`;
936 }
937 const d = attempt.decision;
938 // A held turn's reason carries the confidence; saying it twice is noise.
939 const notes = [
940 ...(d.held === undefined ? [sureOf(d, attempt.kind, attempt.continued)] : []),
941 ...reasonsOf(attempt),
942 ].filter((n) => n !== "");
943 return ` ${when} ${d.tier}·${d.effort} ${notes.join("; ")} ${what}`;
944}
945
946/** Thousands rounded down, for a limit that must not be overstated: 4.1k, 45k. */
947function kOfDown(n: number): string {
948 const k = n / 1000;
949 return k < 10 ? `${Math.floor(k * 10) / 10}k` : `${Math.floor(k)}k`;
950}
951
952/** Thousands or millions, rounded, for token counts: 130k, 2k, 3.3M. */
953function kOf(n: number): string {
954 const k = Math.round(n / 1000);
955 // 999,500 rounds to a thousand k, which is a million.
956 if (k >= 1000) return `${(n / 1_000_000).toFixed(1)}M`;
957 return `${k}k`;
958}
959
960/**
961 * `fable-5-1 ✓`: the model the API says answered, and whether it is the one
962 * the router asked for (a dated id still counts). A different model is the
963 * one case worth looking at: `sonnet-5 ⚠ asked opus-5-5`.
964 */
965function answeredBy(attempt: Attempt): string {
966 const usage = attempt.usage!;
967 const got = shortModel(usage.model);
968 if (!("decision" in attempt)) return got;
969 const asked = attempt.decision.model;
970 return sameModel(asked, usage.model)
971 ? `${got} ✓`
972 : `${got} ⚠ asked ${shortModel(asked)}`;
973}
974
975/**
976 * Whether the model the API reports is the one asked for: the same id, give
977 * or take the engine's `[1m]` and a trailing date. `claude-opus-5-5` does not
978 * confirm `claude-opus-5`, although one begins with the other.
979 */
980function sameModel(asked: string, got: unknown): boolean {
981 // `[1m]`, a date, and a provider's spelling (`…@date`, `us.anthropic.…`)
982 // all answer for the model asked.
983 return typeof got === "string" && sameModelAs(asked, got);
984}
985
986/** `$0.50 · 47k in (49% cached) · 0k out`. */
987function costPhrase(usage: Usage, cost: number | undefined): string {
988 const parts = [];
989 if (cost !== undefined) parts.push(usd(cost));
990 parts.push(
991 `${kOf(carriedOf(usage))} in (${Math.round(cacheRatio(usage) * 100)}% cached)`,
992 `${kOf(usage.output_tokens)} out`,
993 );
994 return parts.join(" · ");
995}
996
997/**
998 * The line under a turn in /jev saying what the API reports actually
999 * answered, and what the requests carried. This is the intrinsic check: the
1000 * route line is what we asked for; this is what we got.
1001 */
1002function usageLine(attempt: Attempt): string | null {
1003 if (!attempt.usage) return null;
1004 return ` → ${answeredBy(attempt)} · ${costPhrase(attempt.usage, attempt.cost)}`;
1005}
1006
1007/**
1008 * The summary under a finished reply: one block for everything the reply
1009 * took, however many turns it spanned. A reply that spawns background work
1010 * is several turns — the one you typed, then one per task that finished and
1011 * woke the loop — and a block under each read as one reply changing model
1012 * three times. So the turns are gathered and written once, at the end.
1013 *
1014 * Two lines as a rule: what answered and what it cost, then why it did not
1015 * run as Jev asked, when it did not.
1016 *
1017 * fable-5-1 ✓ xhigh · Jev 97% · $0.14 · 130k in (91% cached) · 2k out
1018 * kept fable: haiku costs $4.41 vs $0.13
1019 *
1020 * Fenced, because markdown collapses leading whitespace and joins
1021 * consecutive lines into one paragraph: unfenced, the rows would render as
1022 * a single run-on.
1023 */
1024export function replySummary(turns: readonly Attempt[]): string | null {
1025 const main = turns.filter((t) => t.kind !== "agent");
1026 const agents = turns.filter((t) => t.kind === "agent");
1027 const priced = turns.filter((t) => t.usage !== undefined);
1028 if (main.length === 0 || priced.length === 0) return null;
1029
1030 const rows: string[] = [];
1031 const head: string[] = [];
1032
1033 // What answered.
1034 if (main.length === 1) {
1035 const only = main[0]!;
1036 if ("decision" in only) {
1037 const d = only.decision;
1038 head.push(
1039 `${only.usage ? answeredBy(only) : shortModel(d.model)} ${d.effort}`,
1040 );
1041 // How the tier was settled, always in this spot: Jev's confidence, the
1042 // prompt's own pick, or a go-ahead carrying the last one on.
1043 // A held turn's confidence is in the tier it did not move to; the
1044 // note under this line says so, so none is shown here.
1045 const how = d.forced
1046 ? pickOf(d)
1047 : only.kind === "continue" || only.kind === "nudge"
1048 ? "continuing"
1049 : d.held !== undefined
1050 ? ""
1051 : sureOf(d, only.kind, only.continued);
1052 if (how !== "") head.push(how);
1053 } else {
1054 head.push(only.usage ? answeredBy(only) : "session model");
1055 head.push(`not routed: ${words(only.skipped)}`);
1056 }
1057 } else {
1058 // A tier the person named is marked, as the single-turn line says "your pick".
1059 const legs = main.map((t) =>
1060 "decision" in t
1061 ? `${t.decision.tier}${t.decision.forced ? ` (${pickOf(t.decision)})` : ""}${t.usage && !answeredBy(t).endsWith("✓") ? " ⚠" : ""}`
1062 : "session",
1063 );
1064 const woken = main.filter((t) => t.kind === "notify").length;
1065 const nudged = main.filter((t) => t.kind === "nudge").length;
1066 const because = [
1067 ...(woken > 0 ? [`${woken} woken by tasks`] : []),
1068 ...(nudged > 0 ? [`${nudged} nudged`] : []),
1069 ];
1070 head.push(
1071 `${main.length} turns: ${legs.join(", ")}` +
1072 (because.length > 0 ? ` (${because.join(", ")})` : ""),
1073 );
1074 }
1075
1076 // The whole reply's cost.
1077 const sum = priced.reduce(
1078 (acc, t) => {
1079 const u = t.usage!;
1080 acc.input += u.input_tokens;
1081 acc.read += u.cache_read_input_tokens;
1082 acc.write += u.cache_creation_input_tokens;
1083 acc.output += u.output_tokens;
1084 if (t.cost !== undefined) acc.cost += t.cost;
1085 else acc.unpriced = true;
1086 return acc;
1087 },
1088 { input: 0, read: 0, write: 0, output: 0, cost: 0, unpriced: false },
1089 );
1090 const carried = sum.input + sum.read + sum.write;
1091 const cached = carried === 0 ? 0 : Math.round((100 * sum.read) / carried);
1092 if (!sum.unpriced) head.push(usd(sum.cost));
1093 head.push(`${kOf(carried)} in (${cached}% cached)`, `${kOf(sum.output)} out`);
1094 rows.push(head.join(" · "));
1095
1096 // What its agents ran on and cost.
1097 if (agents.length > 0) {
1098 const legs = agents.map((a) => {
1099 const name = plain(a.agent?.type ?? "agent");
1100 // What ran it: the model the API reported, else the tier routed to.
1101 const on = a.usage
1102 ? shortModel(a.usage.model)
1103 : "decision" in a
1104 ? a.decision.tier
1105 : "its own model";
1106 return `${name} ${on}${a.cost !== undefined ? ` ${usd(a.cost)}` : ""}`;
1107 });
1108 rows.push(`agents: ${legs.join(", ")}`);
1109 }
1110
1111 // Why a turn did not run exactly as Jev asked.
1112 // (How the tier was settled is on the first line already; a multi-turn
1113 // reply counts its wake-ups and nudges there.)
1114 main.forEach((t, i) => {
1115 const why = reasonsOf(t).filter((r) => !r.startsWith("your pick") && !r.startsWith("you picked"));
1116 if (why.length === 0) return;
1117 rows.push(`${main.length > 1 ? `turn ${i + 1}: ` : ""}${why.join("; ")}`);
1118 });
1119
1120 return ["```", ...rows, "```"].join("\n");
1121}
1122
1123/** `medium for all`, or `medium (opus: xhigh, fable: xhigh)`. */
1124export function ceilingLine(ceiling: Ceiling): string {
1125 const counts = new Map<string, number>();
1126 for (const tier of TIERS)
1127 counts.set(ceiling[tier], (counts.get(ceiling[tier]) ?? 0) + 1);
1128 let common = ceiling.haiku;
1129 for (const tier of TIERS) {
1130 const effort = ceiling[tier];
1131 if ((counts.get(effort) ?? 0) > (counts.get(common) ?? 0)) common = effort;
1132 }
1133 const rest = TIERS.filter((t) => ceiling[t] !== common).map(
1134 (t) => `${t}: ${ceiling[t]}`,
1135 );
1136 return rest.length === 0
1137 ? `${common} for all`
1138 : `${common} (${rest.join(", ")})`;
1139}
1140
1141/**
1142 * The report, as plain lines. Written so the first three tell you whether
1143 * the thing is on at all, which is the question that brings people here.
1144 */
1145export function statusReport(status: Status): string {
1146 // The engine prefixes the plugin's name; a header here said it twice.
1147 const lines: string[] = [""];
1148
1149 lines.push(` routing ${status.enabled ? "on" : "off (/jev on)"}`);
1150 lines.push(` surface ${status.surface === null ? "unknown" : plain(status.surface)}`);
1151
1152 if (status.provider.ok) {
1153 const key =
1154 status.provider.name === "typesafe"
1155 ? "TYPESAFE_API_KEY"
1156 : "AI_GATEWAY_API_KEY";
1157 lines.push(
1158 ` provider ${status.provider.name} · ${key} is set · ${plain(status.provider.model)}`,
1159 );
1160 } else {
1161 // The reason names the fix: a missing key, a bad base URL, a provider
1162 // forced without its key. "No keys" alone sent people after the wrong one.
1163 lines.push(` provider NOT SET UP — ${words(status.provider.reason)}; nothing will route`);
1164 }
1165
1166 lines.push(` budget ${status.timeoutMs}ms`);
1167 lines.push(
1168 ` sticky ${
1169 status.sticky === null
1170 ? "off (/jev sticky on)"
1171 : `on, switch needs ${pct(status.sticky)} ` +
1172 `(${pct(upgradeBar(status.sticky, UPGRADE_CONTEXT_TOKENS))} up past ` +
1173 `${kOf(UPGRADE_CONTEXT_TOKENS)})`
1174 }`,
1175 );
1176 if (status.price !== undefined)
1177 lines.push(
1178 ` price ${
1179 status.price
1180 ? "on, a downgrade has to pay" +
1181 (status.upgradeMax === 0
1182 ? ", an upgrade may not cost more than staying"
1183 : status.upgradeMax != null
1184 ? `, an upgrade may cost ${usd(status.upgradeMax)} over staying`
1185 : "")
1186 : "off (/jev price on)"
1187 }`,
1188 );
1189 lines.push(` ceiling ${ceilingLine(status.ceiling)}`);
1190 if (status.compactOn !== undefined)
1191 lines.push(
1192 ` compact ${
1193 status.compactOn
1194 ? `${
1195 status.provider.ok && status.provider.name !== "typesafe"
1196 ? "on, but the gateway cannot score tool calls (needs TypeSafe direct): the engine summarises"
1197 : "on, Jev prunes tool calls"
1198 }${status.compaction ? ` · last: ${compactionLine(status.compaction)}` : ""}`
1199 : "off (/jev compact on)"
1200 }`,hooks/compaction/request.ts 85 lines1// Vendored from fast-jev-compaction (https://github.com/tamaratran/fast-jev-compaction)
2// commit e3f262a7f4d4, MIT licensed; see LICENSE-fast-jev-compaction. Imports
3// renamed to .ts; otherwise unchanged.
4
5import type { JevAnswer, JevQuestions, JevResponse, JevState } from "./types.ts";
6
7export const SYSTEM_ONE_URL = 'https://api.typesafe.ai/v1/systemone';
8export const DEFAULT_MODEL = 'jev-latest';
9
10export interface JevRequest {
11 url: string;
12 method: 'POST';
13 headers: Record<string, string>;
14 body: string;
15}
16
17/** The HTTP request for one Jev call, for any fetch-like transport. */
18export function buildJevRequest(
19 params: {
20 apiKey: string;
21 model?: string;
22 baseUrl?: string;
23 },
24 state: JevState,
25 questions: JevQuestions,
26): JevRequest {
27 return {
28 url: params.baseUrl ?? SYSTEM_ONE_URL,
29 method: 'POST',
30 headers: {
31 authorization: `Bearer ${params.apiKey}`,
32 'content-type': 'application/json',
33 },
34 body: JSON.stringify({
35 model: params.model ?? DEFAULT_MODEL,
36 state,
37 questions,
38 }),
39 };
40}
41
42/** Validates a Jev response body; throws on anything but an `answers` object. */
43export function parseJevResponse(
44 status: number,
45 ok: boolean,
46 text: string,
47): JevResponse {
48 if (!ok) {
49 throw new Error(`Jev request failed (${status}): ${text.slice(0, 200)}`);
50 }
51 let parsed: unknown;
52 try {
53 parsed = JSON.parse(text);
54 } catch {
55 throw new Error('Jev returned malformed JSON');
56 }
57 if (
58 parsed === null ||
59 typeof parsed !== 'object' ||
60 !('answers' in parsed) ||
61 parsed.answers === null ||
62 typeof parsed.answers !== 'object'
63 ) {
64 throw new Error('Jev response is missing answers');
65 }
66 return parsed as JevResponse;
67}
68
69/** The `noul` probability of one answer; throws when it is not there. */
70export function noulAnswer(
71 answers: Record<string, JevAnswer>,
72 name: string,
73): number {
74 const answer = answers[name];
75 if (
76 !answer ||
77 !('noul' in answer) ||
78 typeof answer.noul !== 'number' ||
79 !Number.isFinite(answer.noul)
80 ) {
81 throw new Error(`Invalid Jev answer for ${name}`);
82 }
83 return answer.noul;
84}
85hooks/compaction/types.ts 213 lines1// Vendored from fast-jev-compaction (https://github.com/tamaratran/fast-jev-compaction)
2// commit e3f262a7f4d4, MIT licensed; see LICENSE-fast-jev-compaction. Imports
3// renamed to .ts; otherwise unchanged.
4
5export type Role = 'user' | 'assistant';
6
7/**
8 * A tool_use block of an assistant message. `text` and `isError` mirror the
9 * outcome once the transcript holds it (Claude Code attaches them).
10 */
11export interface ToolUse {
12 tool_use_id: string;
13 tool: string;
14 input: Record<string, unknown>;
15 text?: string;
16 isError?: boolean;
17}
18
19/** A tool_result block of a user message. */
20export interface ToolResult {
21 tool_use_id: string;
22 text: string;
23 isError?: boolean;
24}
25
26/**
27 * One transcript message. The shape is a subset of Claude Code's
28 * `SessionMessage`, so a session transcript can be passed in as is.
29 */
30export interface Message {
31 role: Role;
32 text: string;
33 toolUses: ToolUse[];
34 toolResults?: ToolResult[];
35}
36
37/** A tool call paired with its result by `tool_use_id`. */
38export interface ToolCall {
39 /** Short id used in the Jev state and question names (`t1`, `t2`, ...). */
40 id: string;
41 tool_use_id: string;
42 tool: string;
43 input: Record<string, unknown>;
44 /** Index of the message holding the tool_use block. */
45 callIndex: number;
46 /** Index of the message holding the tool_result block. */
47 resultIndex: number;
48 resultChars: number;
49 isError: boolean;
50 /** In the first or the newest preserved messages; never a candidate. */
51 pinned: boolean;
52}
53
54export interface CallAnswer {
55 /** Jev's probability that the call itself still matters. */
56 keepCall: number;
57 /** Jev's probability that the full result still needs to stay verbatim. */
58 keepResult: number;
59}
60
61export type CallAction = 'keep' | 'drop_result' | 'drop_call';
62
63export interface CallDecision extends CallAnswer {
64 id: string;
65 tool: string;
66 action: CallAction;
67 reason: 'pinned' | 'kept' | 'result_dropped' | 'call_dropped';
68}
69
70export interface HistoryToolCall {
71 id: string;
72 tool: string;
73 input: string;
74 result: string;
75}
76
77export interface HistoryEntry {
78 i: number;
79 role: Role;
80 text: string;
81 /** Structured per call, or one compact line per call once the state has to shrink. */
82 tool_calls?: HistoryToolCall[] | string[];
83}
84
85/** The state sent with every Jev request: the whole history, results omitted. */
86export interface CompactionState {
87 context: string;
88 goal: string;
89 history: HistoryEntry[];
90}
91
92export interface FittedState {
93 state: CompactionState;
94 tokens: number;
95 /** Which fitting stage produced the state, for diagnostics. */
96 stage: string;
97}
98
99export interface CompactOptions {
100 /** Ongoing task description; defaults to the last few user prompts. */
101 goal?: string;
102 /** Minimum keep probability for a call or result to stay. Default 0.5. */
103 keepThreshold?: number;
104 /** Newest messages never touched (the first message is always kept). Default 6. */
105 preserveRecentMessages?: number;
106 /** Estimated token ceiling for the state. Default 25000. */
107 maxStateTokens?: number;
108 /** Estimated token ceiling for state plus one batch of questions. Default 30000. */
109 maxRequestTokens?: number;
110 /** Characters of a dropped tool result to retain. Default 300. */
111 truncateHeadChars?: number;
112 /**
113 * A message's text as Jev is shown it in the state (the output keeps the
114 * message's own); the text itself when absent.
115 */
116 textOf?: (message: Message) => string;
117}
118
119export interface ResolvedCompactOptions {
120 goal: string;
121 keepThreshold: number;
122 preserveRecentMessages: number;
123 maxStateTokens: number;
124 maxRequestTokens: number;
125 truncateHeadChars: number;
126 textOf?: (message: Message) => string;
127}
128
129export interface CompactResult {
130 /** The compacted transcript; untouched messages are the input objects. */
131 messages: Message[];
132 decisions: CallDecision[];
133 stats: {
134 messagesBefore: number;
135 messagesAfter: number;
136 charsBefore: number;
137 charsAfter: number;
138 calls: number;
139 kept: number;
140 resultsDropped: number;
141 callsDropped: number;
142 pinned: number;
143 stateTokens: number;
144 /** Which fitting stage the state needed, '' when no request was made. */
145 stateStage: string;
146 requests: number;
147 ms: number;
148 };
149}
150
151/** The `state` of a Jev request: a string or any JSON-serialisable object. */
152export type JevState = string | object;
153
154export interface NoulQuestion {
155 type: 'noul';
156 instructions: string;
157 criteria?: {
158 true?: string;
159 false?: string;
160 };
161}
162
163export interface ChoiceQuestion {
164 type: 'choice';
165 instructions: string;
166 criteria: Record<string, string | null>;
167}
168
169export interface ScoreQuestion {
170 type: 'score';
171 instructions: string;
172 criteria: string[];
173}
174
175export type JevQuestion = NoulQuestion | ChoiceQuestion | ScoreQuestion;
176export type JevQuestions = Record<string, JevQuestion>;
177
178export interface NoulAnswer {
179 type?: 'noul';
180 noul: number;
181}
182
183export interface ChoiceAnswer {
184 type?: 'choice';
185 choice: string;
186 confidence: number;
187 probabilities: Record<string, number>;
188}
189
190export interface ScoreAnswer {
191 type?: 'score';
192 score: number;
193 confidence: number;
194 probabilities: Record<string, number>;
195}
196
197export type JevAnswer = NoulAnswer | ChoiceAnswer | ScoreAnswer;
198
199export interface JevResponse {
200 model?: string;
201 answers: Record<string, JevAnswer>;
202 usage?: {
203 input_tokens?: number;
204 output_tokens?: number;
205 };
206 [key: string]: unknown;
207}
208
209/** Anything that can answer Jev questions: `JevClient`, or a host-provided adapter. */
210export interface JevAsker {
211 ask(state: JevState, questions: JevQuestions): Promise<JevResponse>;
212}
213