SLOPSHOPPER

jev-router

Routes each turn to a Claude model chosen by TypeSafe's Jev, asked through the Vercel AI Gateway.

newspinnercommandnetwork
★ 4v0.1.0MITupdated 2026-09-23satviksinha/jev-model-router
A shopper browsing a rack in a slop shop
README

jev-router

Picks the model for each turn with Jev, TypeSafe's decision model. Supports both TypeSafe's direct API and the Vercel AI Gateway.

MIT licensed.

You type a prompt. Before the turn runs, Jev is asked two questions at once: which tier should answer this, and how hard should it think. Every model request in that turn then goes to the model Jev named, and a line above the reply says which one.

Context    you type a prompt
              ↓
turn.start    ask Jev  →  tier: fable   effort: 3
              ↓
turn.step     next({ ...e, model: 'claude-fable-5-1', effort: 'xhigh' })
              first text chunk ← 'jev → fable·xhigh  0.97 · 641ms\n\n' + text
              ↓
/jev          the full history, with reasons for anything unrouted

The ladder

TierForModel
haikuTrivial. A lookup, a rename, a yes or no.claude-haiku-4-5
sonnetStraightforward and minor, no real decision to make.claude-sonnet-5
opusPlain implementation carrying some complexity.claude-opus-5
fablePlanning, brainstorming, architecture, systematic debugging.claude-fable-5-1

The policy lives in TIER_CRITERIA in hooks/policy.ts. Those strings are what Jev is told each tier is for, so editing them is how you change the router's behaviour. Nothing else needs to change.

Setup

Provider: TypeSafe direct or Vercel AI Gateway

Get either a TypeSafe API key or an AI Gateway key and put it in the env block of ~/.claude/settings.json. The router prefers TypeSafe direct when both keys are set:

{
  "env": {
    "TYPESAFE_API_KEY": "...",
    "AI_GATEWAY_API_KEY": "...",
    "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1"
  }
}

For TypeSafe direct: Get a key from your TypeSafe account.

For the Vercel AI Gateway: Get a key from your Vercel dashboard and note that the gateway needs a card on the account, not just a key. A valid key on an account with no payment method gets:

HTTP 403  customer_verification_required
"AI Gateway requires a valid credit card on file to service requests."

A personal-scope gateway key needs a card before it serves anything, even on free credits. A team-scope key reportedly does not.

Forcing a provider: If both keys are set and you want to use the gateway, set JEV_ROUTER_PROVIDER=gateway. Similarly, JEV_ROUTER_PROVIDER=typesafe forces TypeSafe direct.

CLAUDE_CODE_ENABLE_FUNCTION_HOOKS is not optional. Without it the mod loads and silently does nothing, with no warning.

Because the router fails open, setup issues show up as the mod doing nothing at all rather than as an error. Run npm run check-jev when nothing seems to route to see which provider is configured and whether it's serving.

Then run Claude Code with the mod:

claude --plugin-dir ~/Desktop/jev-router

To keep it on permanently, move the folder to ~/.claude/skills/jev-router/, where it loads on its own next session.

Knowing whether it is working

Four signals, in order of how much you can trust them.

/jev prints the full state. A command's output row draws on every surface, so this always works:

jev-router
  routing   on
  surface   desktop
  provider  typesafe · TYPESAFE_API_KEY is set
  budget    1500ms
  sticky    on, switch needs 75%
  tiers     haiku, sonnet, opus, fable

  Recent turns, newest first:
   653ms  fable·xhigh 0.61  [notify] Agent "Review library-sync cluster" com…
          answered claude-fable-5-1 ✓  cache 98%  45k in  1k out
     0ms  unrouted — [agent:general-purpose] Review library-sync cluster
          answered claude-opus-5  cache 82%  22k in  0k out
   641ms  fable·xhigh 0.97  help me plan the architecture
          answered claude-fable-5-1 ✓  cache 91%  130k in  2k out
   402ms  haiku·medium 0.75  rename the variable foo to bar
          answered claude-haiku-4-5 ✓  cache 4%  128k in  0k out
    12ms  unrouted — gateway said HTTP 403 (customer_verification_required)

Not every turn is you typing, and the history says which are not. A prompt that spawns background agents produces more turns than replies: each agent that finishes wakes the main loop with a <task-notification>, and the engine starts a fresh turn with that XML as its text. Those are tagged [notify], with the notification's summary in place of the envelope, and the same tag rides in the route line and footer of the reply they produce (· notify ·). Without it, one prompt that dispatched three reviewers reads as one reply that changed model three times.

The agents themselves are the [agent:type] rows. A subagent's loop gets no turn.start, so the router never sees a prompt to ask Jev about, and it runs on whatever the Agent tool resolved (the session model, usually). It is listed as unrouted with the model that answered it, so the requests one prompt really caused are all on the screen. Nothing is written into a subagent's reply: that text is a tool result its parent reads.

The same answered information is under each reply as it happens, in the footer below; /jev is where you go to see it across turns.

An unrouted turn says why. That matters because the router fails open, so a dead provider and a missing plugin look identical from the outside.

The answered line under each turn is the API's own report, taken from the usage on each step's stop chunk: which model actually answered, and what the turn's requests carried. The route line above it is what the mod asked for; this is what it got. ✓ means they agree (a dated id such as claude-opus-5-20260901 still counts); ≠ claude-opus-5 means something else answered, which is the one case worth looking into. There is no need to proxy traffic or force a bogus model id to check the rewrite lands.

cache is the share of the turn's input read from the prompt cache. The cache is per model, so the turn after a switch runs cold: cache 4% on the haiku turn above is the price of leaving fable. Cache reads bill at a tenth of uncached input, so a switch on a large context costs roughly ten times what staying would have, once. Watch this number to see whether Jev's switching is eating what the cheaper tiers save.

The first line of every reply. The route is written into the reply's own text, as the first text chunk streams through turn.step:

> ✳️ `opus` · high · 98% · 555ms

---

Here is the implementation...

Markdown, because the line rides in the reply's text and that is what the transcript renders, so it is the only styling available. The blockquote sets it off from prose with a rail and dimmer text; the tier is inline code, which the theme colours. The blank line before --- is load-bearing: a rule on the line directly after text is a setext heading underline, and the route would render as a heading.

An unrouted turn opens with > ⚠️ \unrouted\ · reason. A ? after the percentage means Jev was under 50% sure.

The footer under every finished reply, which is the same information settled. The top line is what the router asked for, before the reply exists; the footer is what the API says it got, and it can only be written once the response is whole:

──────────────────────────────────────────────────────
jev  fable·xhigh · 97% · 641ms
api  claude-fable-5-1 ✓ · cache 90% · 130k in · 1k out

api is read off the usage on the step's stop chunk, so ✓ is the API's own confirmation that the model rewrite landed — no proxy, no bogus model id. A dated id such as claude-fable-5-1-20260901 still counts as a match; a real mismatch reads claude-opus-5 ≠ claude-fable-5-1.

cache is the share of the turn's input read from the prompt cache. The cache is per model, so the turn after a switch runs cold:

─────────────────────────────────────────────────────
jev  haiku·medium · 75% · 402ms
api  claude-haiku-4-5 ✓ · cache 4% · 128k in · 0k out

That 4% is the price of leaving fable. Cache reads bill at a tenth of uncached input, so a switch on a large context costs roughly ten times what staying would have, once. Watch it to see whether the switching is eating what the cheaper tiers save.

The footer is fenced because markdown collapses leading whitespace and joins consecutive lines: unfenced, the rule and the two rows render as one run-on paragraph. <details> was tried first, for a fold; the desktop app renders it as raw tags.

It is emitted as a chunk the hook built rather than one the engine streamed, at one past the last text block's index, and only on a step whose stop reason ends the turn — a tool_use step is mid-reply. The index is load-bearing: a chunk yielded at an index the engine has already streamed is dropped silently. Probed live, a chunk at lastTextIndex never reached the transcript and one at lastTextIndex + 1 did, so the footer opens a block of its own and the reply above it is untouched.

Note that claude -p shows only the last text block in its result, so the reply looks like it vanished when the footer lands. It has not: --output-format stream-json --verbose shows both blocks whole.

/jev quiet drops both the line and the footer without turning routing off; /jev loud brings them back. Both ride in the reply's recorded text, so the model sees them on its own past replies; that is the standing cost of a marker on a surface that draws neither render sites nor ui.log.

The line is part of the recorded message, so the model sees its own past replies open with it. That is the cost of a marker that reaches the desktop app: two cleaner mechanisms were tried first and neither drew there. An AssistantMessage render rewrite was correct against the generated types and drew nothing; $.ui.log, documented as a dim transcript row, also drew nothing. That tab reports $.session.surface() as unknown and appears to be an SDK host that drops both. Reply text and command output are the two channels that reach it.

The footer, via SessionMode. Terminal only in practice, so treat its absence as meaning nothing.

The app's own model indicator will never change. It shows the session model, which this mod does not touch — the rewrite happens per request, in turn.step.

Commands and switches

  • /jev prints the status above. /jev on and /jev off set routing explicitly rather than toggling blind.
  • /jev quiet and /jev loud control the line at the top of each reply. Quiet still routes and still records, so /jev shows what you missed.
  • JEV_ROUTER_PROVIDER=typesafe|gateway forces a specific provider. If both keys are set, the default is TypeSafe direct; use this to force the gateway.
  • TYPESAFE_BASE_URL=https://api.example.com overrides the TypeSafe endpoint base (defaults to https://api.typesafe.ai). Useful for custom deployments.
  • JEV_ROUTER_EXCLUDE=fable,haiku drops those tiers from the question entirely, so Jev is never offered them. Excluding all four is ignored.
  • JEV_ROUTER_TIMEOUT_MS=2500 changes how long a turn waits for Jev before giving up and running unrouted. The default is 1500ms. Ten live calls on 2026-09-20 ran 402ms to 839ms, so an earlier 800ms default was failing open on the slowest of them.
  • /jev sticky makes a tier switch clear a confidence bar before the model moves, /jev sticky 0.6 sets that bar, /jev sticky off stops. --sticky works too. See below.
  • JEV_ROUTER_STICKY=1 and JEV_ROUTER_STICKY_CONFIDENCE=0.6 set the same thing for a session before it starts, for a project that always wants it. The command overrides them from then on.

Holding a shaky switch

The prompt cache is per model. A session cached under fable is cold for haiku, so the turn that switches pays full input tokens and a slower first token. A router that flips tier on a 51% hunch can pick the cheaper model every time and still cost more than staying put.

Run /jev sticky and a turn that names a different tier than the last one has to clear the bar to move. Below it, the turn runs on the tier already loaded, and says so:

> ✳️ `fable` · low · 61% · held:haiku · 512ms

Jev wanted haiku, was 61% sure, and the bar is 75%, so the turn stayed on fable. The same held:haiku appears in the footer and in /jev, because a hold nobody can see is indistinguishable from a router that is not running.

Only the model is held. The effort Jev asked for is applied either way, since effort does not change the model and so costs no cache: a held turn still thinks harder or less hard than the one before it.

What the next turn holds to is the tier actually running, not the one Jev named. Three shaky haiku calls in a row will not creep the session onto haiku one turn at a time. An unrouted turn changes nothing, since nothing ran.

The bar starts at 0.75, which is a starting point rather than a measured optimum. Retune it in place with /jev sticky 0.6 and watch the next few turns; npm run try-prompts prints Jev's confidence across a set of prompts, which is the other input to picking a number. /jev sticky on its own keeps a bar you have already set, so turning it off and on again does not lose it.

A session that should always be sticky can say so before it starts, with JEV_ROUTER_STICKY=1 in the env block of settings.json. The command wins after that.

Checking and tuning

npm run check-jev        # is the configured provider serving?
npm run try-prompts      # what tier does Jev give a spread of prompts?
npm run try-prompts -- "your prompt"

check-jev detects which provider is configured and runs the appropriate diagnostic. For the gateway it checks credits; for TypeSafe direct it makes a live call. try-prompts is the tuning loop: edit TIER_CRITERIA, run it, and see whether the picks moved the way you wanted. Neither script ever prints the key.

Measured on 2026-09-20 against the shipped criteria:

  ms  tier    effort  conf  prompt
 839  haiku   low     1.00  what is 2+2
 508  haiku   xhigh   0.86  what does this function return
 402  haiku   medium  0.75  rename the variable foo to bar in utils.ts
 482  sonnet  medium  0.66  add a --verbose flag to the CLI
 426  opus    high    0.66  write a test for the pagination helper
 555  opus    high    0.98  implement cursor pagination for the reports endpoint
 483  opus    xhigh   0.97  refactor the auth module to use the new session interface
 641  fable   xhigh   0.97  the e2e suite passes alone but fails with the others
 734  fable   xhigh   1.00  help me plan the architecture for multi-tenant billing
 479  fable   xhigh   1.00  should we use event sourcing here or is that overkill

Note row two. Tier and effort are separate questions, so they can disagree: an out-of-context question reads as trivial to route but hard to answer. The engine silently downgrades an effort the chosen model does not support, so this is harmless, but it is why a haiku·xhigh label is possible.

When it does nothing

The router fails open at every step, and a turn it cannot decide runs exactly as it would without the mod:

  • no TYPESAFE_API_KEY or AI_GATEWAY_API_KEY, no request is made at all
  • the provider takes longer than the timeout (1500ms by default)
  • the provider refuses the key, errors, or returns a body we cannot read
  • Jev names a tier that was not offered

The only cost of a failure is the latency spent waiting, capped at the timeout.

Low confidence is not a failure. The pick is used and the line marks it, so jev → opus·high? means Jev was under 50% sure of the tier.

An unrouted turn announces itself too, with the reason:

jev → unrouted (gateway said HTTP 403 (customer_verification_required))

Layout

hooks/register.ts   the five hooks, the per-turn cache, the turn history
hooks/jev.ts        the request shape, timeout, named failures
hooks/provider.ts   which backend (TypeSafe direct or gateway) to use
hooks/policy.ts     the tiers, the criteria, answers → model and effort
hooks/label.ts      decision → footer string
hooks/status.ts     the per-turn line, what /jev prints, usage per turn
tests/              node:test suites; register.test.ts drives the real hooks
                    with a fake engine and asserts the stream transform
scripts/            check-jev and try-prompts, for setup and tuning

jev.ts takes fetch and sleep as arguments rather than importing them, so the tests run with no engine and no network. provider.ts is pure: it takes environment variables and returns which provider to use, with all credentials and endpoints already resolved.

Working on it

node --test 'tests/*.test.ts'
claude plugin validate .

plugin validate is worth running on every change. It does static analysis and prints every event the module hooks, everything it calls on $, and every environment variable it reads or writes, without executing anything. It also catches shape errors that are easy to get wrong, such as turn.step needing to be an async function* because it streams.

There is no claude plugin test in Claude Code 2.1.275, so the engine-level test kit described in the upstream mods/README.md is not available yet.

Backends

Both backends speak the same choice and score question types and the same answers response shape, so the request/response handling is identical.

TypeSafe direct endpoint: POST https://api.typesafe.ai/v1/systemone. Bearer auth, body is { model: "jev-latest", state, questions }. Probabilities and confidences come back rounded to four decimal places. Supports all three question types: choice, score, and noul.

Vercel AI Gateway endpoint: POST https://ai-gateway.vercel.sh/v1/evaluate. Bearer auth, body is { model, state, questions } (the gateway ignores the model field). Probabilities and confidences come back rounded to two decimal places. Only supports choice and score — the gateway rejects noul outright.

Source 6 files
hooks/register.ts 326 lines
1import type { On } from "claude-code";
2
3import { askJev, timeoutOf } from "./jev.ts";
4import { labelOf, withLabel } from "./label.ts";
5import {
6  excludedTiers,
7  offeredTiers,
8  stickyDecision,
9  stickyOf,
10  thresholdOf,
11  type Decision,
12} from "./policy.ts";
13import { providerOf } from "./provider.ts";
14import {
15  addUsage,
16  announceReply,
17  attemptOf,
18  HISTORY_LIMIT,
19  liveLine,
20  FOOTER_SEPARATOR,
21  REPLY_SEPARATOR,
22  statusReport,
23  toggleReply,
24  stickyCommand,
25  usageFooter,
26  type AgentTag,
27  type Attempt,
28} from "./status.ts";
29
30/** Turns kept in the decision cache before the oldest are dropped. */
31const CACHE_LIMIT = 32;
32
33/**
34 * Stop reasons that mean the turn continues: the engine will step again, so
35 * the footer would land in the middle of a reply. Every other reason ends it.
36 */
37const MID_TURN: ReadonlySet<string> = new Set(["tool_use", "pause_turn"]);
38
39/**
40 * Names the subagent a step runs in, from the session's agent list. A row may
41 * not be there yet for a loop that only just started; then the id stands in,
42 * which still says "not the main loop", the part that matters.
43 */
44async function agentTagOf(
45  $: {
46    agent: {
47      list: () => Promise<
48        readonly { id: string; type: string; description: string }[]
49      >;
50    };
51  },
52  agentId: string,
53): Promise<AgentTag> {
54  const rows = await $.agent.list().catch(() => []);
55  const row = rows.find((r) => r.id === agentId);
56  return row ? { type: row.type, label: row.description } : { label: agentId };
57}
58
59/**
60 * Registers the router: one Jev call per turn, applied to every model request
61 * that turn makes, and announced as it happens.
62 *
63 * The decision is made once in `turn.start`, where the person's text is, and
64 * read back in `turn.step`, which fires again after each tool result. Asking
65 * per step would pay Jev's latency several times over and could land two
66 * steps of one turn on different models.
67 *
68 * Every turn's outcome is kept, routed or not, because "did this do anything"
69 * is unanswerable otherwise: a router that fails open looks exactly like one
70 * that is not loaded.
71 *
72 * @param on the engine's registrar
73 */
74export function register(on: On) {
75  const decisions = new Map<string, Decision>();
76  /**
77   * Turns whose reply has yet to open with its route line. The line itself is
78   * built at the first text chunk, not here: by then the step has said which
79   * loop the turn runs in, which the line names.
80   */
81  const pending = new Set<string>();
82  const attempts: Attempt[] = [];
83  /**
84   * turnId → its attempt, so each step's `stop` chunk can add what the API
85   * reported to the right turn. The same objects as in `attempts`.
86   */
87  const byTurn = new Map<string, Attempt>();
88  let latest: Decision | null = null;
89  /**
90   * The confidence a switch must clear, or null when switches are free. The
91   * env vars are the session's starting value; `/jev sticky` overrides them
92   * from then on, so retuning does not mean restarting the session.
93   */
94  let sticky: number | null = null;
95  /** The tier the last routed turn ran on; what a shaky switch is held to. */
96  let running: Decision | null = null;
97  let enabled = true;
98  let announce = true;
99  let surface: string | null = null;
100
101  const trim = (map: Map<string, unknown>) => {
102    while (map.size > CACHE_LIMIT) {
103      const oldest = map.keys().next();
104      if (oldest.done) break;
105      map.delete(oldest.value);
106    }
107  };
108
109  const trimSet = (set: Set<string>) => {
110    while (set.size > CACHE_LIMIT) {
111      const oldest = set.values().next();
112      if (oldest.done) break;
113      set.delete(oldest.value);
114    }
115  };
116
117  const record = (attempt: Attempt) => {
118    attempts.unshift(attempt);
119    attempts.length = Math.min(attempts.length, HISTORY_LIMIT);
120  };
121
122  on("session.start", async ($, e, next) => {
123    await $.command.register({
124      name: "jev",
125      description: "Jev routing: status, or `on` / `off`.",
126    });
127    surface = await $.session.surface();
128    sticky = stickyOf(await $.env.get("JEV_ROUTER_STICKY"))
129      ? thresholdOf(await $.env.get("JEV_ROUTER_STICKY_CONFIDENCE"))
130      : null;
131    return next(e);
132  });
133
134  on("command.run", { command: "jev" }, async ($, e) => {
135    const arg = e.args.trim().toLowerCase();
136
137    if (arg === "on" || arg === "off") {
138      enabled = arg === "on";
139      if (!enabled) latest = null;
140      return { text: toggleReply(enabled) };
141    }
142
143    if (arg === "quiet" || arg === "loud") {
144      announce = arg === "loud";
145      return { text: announceReply(announce) };
146    }
147
148    // `--sticky` as well as `sticky`: the flag spelling is what people reach
149    // for, and refusing it would teach nothing.
150    const sub = arg.replace(/^-+/, "");
151    if (sub === "sticky" || sub.startsWith("sticky ")) {
152      const result = stickyCommand(sub.slice("sticky".length), sticky);
153      sticky = result.sticky;
154      return { text: result.text };
155    }
156
157    const excluded = excludedTiers(await $.env.get("JEV_ROUTER_EXCLUDE"));
158    const provider = providerOf({
159      TYPESAFE_API_KEY: await $.env.get("TYPESAFE_API_KEY"),
160      AI_GATEWAY_API_KEY: await $.env.get("AI_GATEWAY_API_KEY"),
161      JEV_ROUTER_PROVIDER: await $.env.get("JEV_ROUTER_PROVIDER"),
162      TYPESAFE_BASE_URL: await $.env.get("TYPESAFE_BASE_URL"),
163    });
164    return {
165      text: statusReport({
166        enabled,
167        surface: surface ?? (await $.session.surface()),
168        provider,
169        timeoutMs: timeoutOf(await $.env.get("JEV_ROUTER_TIMEOUT_MS")),
170        sticky,
171        offered: offeredTiers(excluded),
172        excluded: [...excluded],
173        announce,
174        attempts,
175      }),
176    };
177  });
178
179  on("turn.start", async ($, e, next) => {
180    if (!enabled) return next(e);
181
182    const offered = offeredTiers(
183      excludedTiers(await $.env.get("JEV_ROUTER_EXCLUDE")),
184    );
185
186    const provider = providerOf({
187      TYPESAFE_API_KEY: await $.env.get("TYPESAFE_API_KEY"),
188      AI_GATEWAY_API_KEY: await $.env.get("AI_GATEWAY_API_KEY"),
189      JEV_ROUTER_PROVIDER: await $.env.get("JEV_ROUTER_PROVIDER"),
190      TYPESAFE_BASE_URL: await $.env.get("TYPESAFE_BASE_URL"),
191    });
192
193    const result = await askJev({
194      fetch: (url, init) => $.http.fetch(url, init),
195      sleep: (ms) => $.clock.sleep(ms),
196      provider,
197      state: e.text,
198      offered,
199      timeoutMs: timeoutOf(await $.env.get("JEV_ROUTER_TIMEOUT_MS")),
200    });
201
202    // One place where the turn's outcome is settled, so the report and the
203    // announcement can never disagree about what happened.
204    const attempt = attemptOf(e.text, result, offered, { sticky, running });
205    record(attempt);
206    byTurn.set(e.turnId, attempt);
207    trim(byTurn);
208
209    // The line goes into the reply's own text, in turn.step below. Render
210    // hooks and $.ui.log both drew nothing in the desktop app; the model's
211    // text is the one channel that reaches every surface.
212    if (announce) {
213      pending.add(e.turnId);
214      trimSet(pending);
215    }
216
217    if ("decision" in attempt) {
218      decisions.set(e.turnId, attempt.decision);
219      trim(decisions);
220      latest = attempt.decision;
221      // What the next turn holds to is the tier actually running, which on a
222      // held turn is the previous one, not the one Jev named.
223      running = attempt.decision;
224    }
225
226    return next(e);
227  });
228
229  // turn.step streams, so it is an async generator. The model rewrite goes
230  // down in `e`; the label comes back up in the first text chunk of the turn,
231  // and the `stop` chunk's usage, which names the model the API says answered,
232  // is kept on the turn. That is the check on the rewrite: the route line is
233  // what was asked for, /jev shows what was got.
234  //
235  // Text chunks concatenate per block, so prefixing the first one puts the
236  // line at the top of the reply. This is the recorded text too, so the model
237  // sees its past replies open with the line; that is the price of a marker
238  // that reaches a surface which draws neither render sites nor ui.log.
239  on("turn.step", async function* ($, e, next) {
240    const decision = decisions.get(e.turnId);
241
242    // A subagent's loop gets no turn.start (probed live: its steps arrive
243    // with agentId set and nothing in byTurn), so its turn is first seen
244    // here. It is not routed: there is no prompt to ask Jev about, and a
245    // line in its reply would land in the tool result its parent reads. It
246    // is recorded, though, with the model that ran it, or /jev would show
247    // one prompt and hide the four requests it caused.
248    let attempt = byTurn.get(e.turnId);
249    if (attempt === undefined && e.agentId !== undefined) {
250      const agent = await agentTagOf($, e.agentId);
251      attempt = {
252        prompt: agent.label,
253        ms: 0,
254        skipped: "subagent runs on the session model",
255        kind: "agent",
256        agent,
257      };
258      record(attempt);
259      byTurn.set(e.turnId, attempt);
260      trim(byTurn);
261    }
262    const step = decision
263      ? next({ ...e, model: decision.model, effort: decision.effort })
264      : next(e);
265
266    // The block the footer joins, so it lands at the end of the reply's text
267    // rather than opening a block of its own.
268    let lastTextIndex = 0;
269
270    for await (const chunk of step) {
271      if (chunk.kind === "text") {
272        lastTextIndex = chunk.index;
273        if (attempt && pending.has(e.turnId)) {
274          pending.delete(e.turnId);
275          yield {
276            ...chunk,
277            text: `${liveLine(attempt)}${REPLY_SEPARATOR}${chunk.text}`,
278          };
279          continue;
280        }
281      }
282
283      if (chunk.kind === "stop") {
284        if (attempt && chunk.usage) addUsage(attempt, chunk.usage);
285
286        // The index must be one past the last text block, and this is
287        // load-bearing. A chunk yielded at an index the engine already
288        // streamed is dropped on the floor, silently: probed live, a chunk
289        // at `lastTextIndex` never reached the transcript, one at
290        // `lastTextIndex + 1` did. It opens a block of its own, which is
291        // what a footer wants anyway — the reply above it stays untouched.
292        //
293        // No `ref`, because the engine's handle belongs to a chunk the
294        // engine streamed; one a hook built has none and is taken at its
295        // word. It goes before the stop chunk, the last thing the engine
296        // expects to see.
297        // Not in a subagent's reply: that is a tool result its parent reads.
298        if (
299          attempt &&
300          announce &&
301          attempt.kind !== "agent" &&
302          !MID_TURN.has(chunk.stopReason ?? "")
303        ) {
304          const footer = usageFooter(attempt);
305          if (footer !== null) {
306            yield {
307              kind: "text" as const,
308              index: lastTextIndex + 1,
309              text: `${FOOTER_SEPARATOR}${footer}`,
310            };
311          }
312        }
313      }
314
315      yield chunk;
316    }
317  });
318
319  // The footer, where a surface draws one. The announcement above is what
320  // carries on surfaces that draw no footer, which is most of them.
321  on("ui.render", { component: "SessionMode" }, async ($, e, next) => {
322    const modes = withLabel(e.props.modes, labelOf(latest, enabled));
323    return next({ ...e, props: { ...e.props, modes } });
324  });
325}
326
hooks/jev.ts 193 lines
1/**
2 * Asking Jev, TypeSafe's decision model, through either the Vercel AI
3 * Gateway or TypeSafe's direct API.
4 *
5 * The gateway (POST /v1/evaluate) speaks its own vocabulary: question types
6 * are `choice`, `score` and `boolean`, never TypeSafe's native `noul`, which
7 * it rejects outright. Probabilities and confidences come back rounded to two
8 * decimal places.
9 *
10 * TypeSafe direct (POST /v1/systemone) supports all three question types
11 * (choice, score, noul) and returns probabilities rounded to four decimal
12 * places.
13 *
14 * Both support the same `choice` and `score` question types and the same
15 * `answers` response shape, so the request/response handling is identical.
16 *
17 * `fetch` and `sleep` are arguments rather than imports so this file runs
18 * under plain `node` in tests, with no engine and no network.
19 */
20
21import { EFFORT_CRITERIA, TIER_CRITERIA, type Tier } from "./policy.ts";
22import type { ProviderResult } from "./provider.ts";
23
24/**
25 * Measured against the live gateway on 2026-09-20: ten prompts ran 402ms to
26 * 839ms. An 800ms budget failed open on the slowest of them, so this leaves
27 * real headroom while still capping what a turn waits before giving up.
28 * `JEV_ROUTER_TIMEOUT_MS` overrides it.
29 */
30export const DEFAULT_TIMEOUT_MS = 1500;
31
32/** A timeout from the environment, or the default when it is unusable. */
33export function timeoutOf(raw: string | undefined): number {
34  const parsed = Number(raw);
35  if (!Number.isFinite(parsed) || parsed <= 0) return DEFAULT_TIMEOUT_MS;
36  return parsed;
37}
38
39export type HttpResponseLike = {
40  ok: boolean;
41  status: number;
42  text: string;
43};
44
45/**
46 * What one attempt at Jev came to. A failure carries its reason so the
47 * session can say why a turn went unrouted instead of going quiet.
48 */
49export type JevResult =
50  | { ok: true; answers: unknown; ms: number }
51  | { ok: false; reason: string; ms: number };
52
53export type AskArgs = {
54  fetch: (url: string, init?: HttpInitLike) => Promise<HttpResponseLike>;
55  sleep: (ms: number) => Promise<unknown>;
56  provider: ProviderResult;
57  state: string;
58  offered: readonly Tier[];
59  timeoutMs?: number;
60  /** Injected so tests can measure without a real clock. */
61  now?: () => number;
62};
63
64export type HttpInitLike = {
65  method?: string;
66  headers?: Record<string, string>;
67  body?: string;
68};
69
70/**
71 * The request body for one routing decision: two questions Jev answers in
72 * parallel, the tier as a Choice and the effort as a Score.
73 *
74 * The model field is added by askJev depending on which provider is used.
75 */
76export function requestBodyOf(state: string, offered: readonly Tier[]) {
77  const criteria: Record<string, string> = {};
78  for (const tier of offered) criteria[tier] = TIER_CRITERIA[tier];
79
80  return {
81    state,
82    questions: {
83      tier: {
84        type: "choice",
85        instructions:
86          "A developer typed this request to a coding agent. Which model tier " +
87          "should answer it?",
88        criteria,
89      },
90      effort: {
91        type: "score",
92        instructions: "How much thinking does answering this request take?",
93        criteria: [...EFFORT_CRITERIA],
94      },
95    },
96  };
97}
98
99/**
100 * Asks Jev and answers with the response's `answers` object, or a reason.
101 *
102 * Every failure is still a pass for the turn, but it is a named one: the
103 * caller reports the reason rather than leaving the person guessing whether
104 * the router ran at all.
105 */
106export async function askJev(args: AskArgs): Promise<JevResult> {
107  const {
108    fetch,
109    sleep,
110    provider,
111    state,
112    offered,
113    now = () => Date.now(),
114  } = args;
115  const timeoutMs = args.timeoutMs ?? DEFAULT_TIMEOUT_MS;
116  const started = now();
117  const since = () => now() - started;
118
119  if (!provider.ok) return { ok: false, reason: provider.reason, ms: 0 };
120  if (state.trim() === "") return { ok: false, reason: "empty prompt", ms: 0 };
121  if (offered.length === 0)
122    return { ok: false, reason: "no tiers offered", ms: 0 };
123
124  const TIMED_OUT = Symbol("timed-out");
125
126  // Build the request body. For TypeSafe direct, we use the model name directly.
127  // For the gateway, we still ask for it but the gateway ignores our model field
128  // and uses typesafe-ai/jev regardless.
129  const body = {
130    ...requestBodyOf(state, offered),
131    model: provider.model,
132  };
133
134  const call = fetch(provider.endpoint, {
135    method: "POST",
136    headers: {
137      authorization: `Bearer ${provider.apiKey}`,
138      "content-type": "application/json",
139    },
140    body: JSON.stringify(body),
141  });
142
143  let response: HttpResponseLike;
144  try {
145    const raced = await Promise.race([
146      call,
147      sleep(timeoutMs).then(() => TIMED_OUT),
148    ]);
149    if (raced === TIMED_OUT) {
150      return {
151        ok: false,
152        reason: `timed out after ${timeoutMs}ms`,
153        ms: since(),
154      };
155    }
156    response = raced as HttpResponseLike;
157  } catch (error) {
158    const detail = error instanceof Error ? error.message : String(error);
159    return { ok: false, reason: `request failed: ${detail}`, ms: since() };
160  }
161
162  if (!response) return { ok: false, reason: "no response", ms: since() };
163
164  if (!response.ok) {
165    return {
166      ok: false,
167      reason: `gateway said HTTP ${response.status}${gatewayNoteOf(response)}`,
168      ms: since(),
169    };
170  }
171
172  try {
173    const parsed = JSON.parse(response.text) as { answers?: unknown };
174    if (typeof parsed !== "object" || parsed === null || !parsed.answers) {
175      return { ok: false, reason: "response carried no answers", ms: since() };
176    }
177    return { ok: true, answers: parsed.answers, ms: since() };
178  } catch {
179    return { ok: false, reason: "response was not JSON", ms: since() };
180  }
181}
182
183/** The gateway's own error type, when it sent one, for the status line. */
184function gatewayNoteOf(response: HttpResponseLike): string {
185  try {
186    const body = JSON.parse(response.text) as { error?: { type?: string } };
187    const type = body?.error?.type;
188    return typeof type === "string" ? ` (${type})` : "";
189  } catch {
190    return "";
191  }
192}
193
hooks/label.ts 37 lines
1/**
2 * The footer label. `SessionMode` draws the strings it is handed, so this
3 * file's only job is to turn a decision into one of them.
4 */
5
6import type { Decision } from "./policy.ts";
7
8/** Below this, the pick is marked so a bad route is visible rather than silent. */
9export const LOW_CONFIDENCE = 0.5;
10
11/**
12 * The label for the footer, or null to add nothing.
13 *
14 * Null rather than a placeholder before the first turn: a footer that says
15 * nothing reads better than one that says the router has not run yet.
16 */
17export function labelOf(
18  decision: Decision | null,
19  enabled: boolean,
20): string | null {
21  if (!enabled) return "jev off";
22  if (!decision) return null;
23
24  const doubt = decision.confidence < LOW_CONFIDENCE ? "?" : "";
25  return `jev → ${decision.tier}·${decision.effort}${doubt}`;
26}
27
28/** The modes array `SessionMode` should draw, with our label on the end. */
29export function withLabel(
30  modes: readonly string[],
31  label: string | null,
32): readonly string[] {
33  if (label === null) return modes;
34  if (modes.includes(label)) return modes;
35  return [...modes, label];
36}
37
hooks/policy.ts 200 lines
1/**
2 * The routing policy: which tiers exist, what Jev is told each one is for,
3 * and how Jev's answers become a model and an effort level.
4 *
5 * Nothing here touches the engine or the network, so it runs under plain
6 * `node` in tests.
7 */
8
9export type Tier = "haiku" | "sonnet" | "opus" | "fable";
10
11export type Effort = "low" | "medium" | "high" | "xhigh" | "max";
12
13export type Decision = {
14  tier: Tier;
15  model: string;
16  effort: Effort;
17  /** Jev's confidence in the tier, 0 to 1. The gateway rounds to 2 places. */
18  confidence: number;
19  /**
20   * The tier Jev named, when stickiness kept the turn on the previous one
21   * instead. Absent on a turn that went where Jev pointed. Kept so the route
22   * line can say a hold happened; a hold nobody can see is indistinguishable
23   * from a router that is not running.
24   */
25  held?: Tier;
26};
27
28export const TIERS: readonly Tier[] = ["haiku", "sonnet", "opus", "fable"];
29
30export const EFFORTS: readonly Effort[] = [
31  "low",
32  "medium",
33  "high",
34  "xhigh",
35  "max",
36];
37
38/** Model ids as the engine names them. */
39export const MODEL_OF: Record<Tier, string> = {
40  haiku: "claude-haiku-4-5",
41  sonnet: "claude-sonnet-5",
42  opus: "claude-opus-5",
43  fable: "claude-fable-5-1",
44};
45
46/**
47 * What Jev is told each tier is for. This is the policy: edit these lines to
48 * change how the router behaves, and nothing else.
49 */
50export const TIER_CRITERIA: Record<Tier, string> = {
51  haiku:
52    "Trivial. A lookup, a rename, a yes or no question, reading one short file, " +
53    "restating something already on screen.",
54  sonnet:
55    "Straightforward and minor. A small edit whose shape is already obvious from " +
56    "the request, with no real decision to make.",
57  opus:
58    "Plain implementation carrying some complexity. Writing or changing real code, " +
59    "possibly across a few files, where the approach is known but the work is not " +
60    "mechanical.",
61  fable:
62    "High complexity needing higher-order reasoning. Planning, brainstorming, " +
63    "architecture, systematic debugging, weighing trade-offs, research. Anything " +
64    "where working out the approach is itself the hard part.",
65};
66
67/** Ordered low to high; the index Jev scores is the effort level. */
68export const EFFORT_CRITERIA: readonly string[] = [
69  "No thinking needed. The answer is immediate.",
70  "A little thinking. One or two steps.",
71  "Real thinking. Several steps, or a choice worth weighing.",
72  "Hard thinking. Many interacting parts, or a subtle failure to chase down.",
73  "As hard as it gets. Open-ended, ambiguous, or the cost of being wrong is high.",
74];
75
76/** Tiers dropped from the question entirely, lowercase, from the env var. */
77export function excludedTiers(raw: string | undefined): Set<Tier> {
78  const names = (raw ?? "")
79    .split(",")
80    .map((s) => s.trim().toLowerCase())
81    .filter(Boolean);
82  return new Set(
83    names.filter((n): n is Tier => (TIERS as string[]).includes(n)),
84  );
85}
86
87/** The tiers offered to Jev, in ladder order, never empty. */
88export function offeredTiers(excluded: Set<Tier>): Tier[] {
89  const kept = TIERS.filter((t) => !excluded.has(t));
90  return kept.length > 0 ? [...kept] : [...TIERS];
91}
92
93type ChoiceAnswer = { type: "choice"; choice?: unknown; confidence?: unknown };
94type ScoreAnswer = { type: "score"; score?: unknown; confidence?: unknown };
95
96function isRecord(v: unknown): v is Record<string, unknown> {
97  return typeof v === "object" && v !== null;
98}
99
100/** Jev's score across EFFORT_CRITERIA to the nearest effort level. */
101export function effortOf(score: unknown): Effort {
102  if (typeof score !== "number" || !Number.isFinite(score)) return "medium";
103  const i = Math.min(Math.max(Math.round(score), 0), EFFORTS.length - 1);
104  return EFFORTS[i] ?? "medium";
105}
106
107/**
108 * Turns the `answers` object of a Jev response into a decision.
109 *
110 * Returns null whenever the answer is missing, malformed, or names a tier
111 * that was not offered: the caller then leaves the turn alone.
112 */
113export function decisionOf(
114  answers: unknown,
115  offered: readonly Tier[] = TIERS,
116): Decision | null {
117  if (!isRecord(answers)) return null;
118
119  const tier = answers.tier as ChoiceAnswer | undefined;
120  if (!isRecord(tier) || tier.type !== "choice") return null;
121
122  const choice = tier.choice;
123  if (typeof choice !== "string") return null;
124  if (!offered.includes(choice as Tier)) return null;
125
126  const effort = answers.effort as ScoreAnswer | undefined;
127  const confidence =
128    typeof tier.confidence === "number" && Number.isFinite(tier.confidence)
129      ? tier.confidence
130      : 0;
131
132  return {
133    tier: choice as Tier,
134    model: MODEL_OF[choice as Tier],
135    effort: effortOf(isRecord(effort) ? effort.score : undefined),
136    confidence,
137  };
138}
139
140/**
141 * The confidence a switch must clear before the model moves, when stickiness
142 * is on.
143 *
144 * The prompt cache is per model: a session cached under one tier is cold for
145 * the next, so the turn that switches pays full input tokens. A router that
146 * flips on a 51% hunch can pick the cheaper model every time and still cost
147 * more. 0.75 is the starting point, not a measured optimum; retune it with
148 * `npm run try-prompts`.
149 */
150export const DEFAULT_STICKY_CONFIDENCE = 0.75;
151
152/** Whether stickiness is on. Off unless the env var says otherwise. */
153export function stickyOf(raw: string | undefined): boolean {
154  const flag = (raw ?? "").trim().toLowerCase();
155  return flag === "1" || flag === "true" || flag === "yes" || flag === "on";
156}
157
158/**
159 * The bar from the environment, or the default when it is unusable.
160 *
161 * A value above 1 is read as a percentage, since `JEV_ROUTER_STICKY_CONFIDENCE=80`
162 * is the likelier intent than a bar no turn can ever clear. 0 and 1 are both
163 * refused: one would hold every switch forever, the other would hold none,
164 * and each is better said by leaving the flag off.
165 */
166export function thresholdOf(raw: string | undefined): number {
167  const parsed = Number(raw);
168  if (!Number.isFinite(parsed) || parsed <= 0) return DEFAULT_STICKY_CONFIDENCE;
169  const ratio = parsed > 1 ? parsed / 100 : parsed;
170  if (ratio <= 0 || ratio >= 1) return DEFAULT_STICKY_CONFIDENCE;
171  return ratio;
172}
173
174/**
175 * Holds a shaky switch on the tier the last turn used.
176 *
177 * Only the model is held. The effort Jev asked for is applied either way,
178 * because effort does not change the model and so costs no cache: a held
179 * turn still gets to think harder or less hard than the one before it.
180 *
181 * `previous` is the tier the last routed turn ran on, or null on the first
182 * turn of a session, which has nothing to hold to.
183 */
184export function stickyDecision(
185  fresh: Decision,
186  previous: Decision | null,
187  threshold: number,
188): Decision {
189  if (previous === null) return fresh;
190  if (fresh.tier === previous.tier) return fresh;
191  if (fresh.confidence >= threshold) return fresh;
192  return {
193    tier: previous.tier,
194    model: previous.model,
195    effort: fresh.effort,
196    confidence: fresh.confidence,
197    held: fresh.tier,
198  };
199}
200
hooks/provider.ts 107 lines
1/**
2 * Provider resolution: which backend (TypeSafe direct or Vercel AI Gateway)
3 * should handle this request, based on available keys and user overrides.
4 *
5 * Precedence:
6 * 1. JEV_ROUTER_PROVIDER=typesafe|gateway forces one (and errors if its key is missing)
7 * 2. TypeSafe direct if TYPESAFE_API_KEY is set
8 * 3. Gateway if AI_GATEWAY_API_KEY is set
9 * 4. Error if neither is set
10 *
11 * TYPESAFE_BASE_URL overrides the TypeSafe endpoint base (defaults to https://api.typesafe.ai).
12 */
13
14export type ProviderResult =
15  | {
16      ok: true;
17      name: "typesafe" | "gateway";
18      endpoint: string;
19      model: string;
20      apiKey: string;
21    }
22  | {
23      ok: false;
24      reason: string;
25    };
26
27export type ProviderEnv = {
28  TYPESAFE_API_KEY: string | undefined;
29  AI_GATEWAY_API_KEY: string | undefined;
30  JEV_ROUTER_PROVIDER: string | undefined;
31  TYPESAFE_BASE_URL: string | undefined;
32};
33
34const TYPESAFE_BASE_DEFAULT = "https://api.typesafe.ai";
35const GATEWAY_BASE = "https://ai-gateway.vercel.sh";
36
37export function providerOf(env: ProviderEnv): ProviderResult {
38  const forced = (env.JEV_ROUTER_PROVIDER ?? "").toLowerCase().trim();
39
40  // Forced override takes absolute precedence.
41  if (forced === "typesafe") {
42    if (!env.TYPESAFE_API_KEY) {
43      return {
44        ok: false,
45        reason: "JEV_ROUTER_PROVIDER=typesafe but TYPESAFE_API_KEY is not set",
46      };
47    }
48    const base = (env.TYPESAFE_BASE_URL ?? TYPESAFE_BASE_DEFAULT).replace(
49      /\/$/,
50      "",
51    );
52    return {
53      ok: true,
54      name: "typesafe",
55      endpoint: `${base}/v1/systemone`,
56      model: "jev-latest",
57      apiKey: env.TYPESAFE_API_KEY,
58    };
59  }
60
61  if (forced === "gateway") {
62    if (!env.AI_GATEWAY_API_KEY) {
63      return {
64        ok: false,
65        reason: "JEV_ROUTER_PROVIDER=gateway but AI_GATEWAY_API_KEY is not set",
66      };
67    }
68    return {
69      ok: true,
70      name: "gateway",
71      endpoint: `${GATEWAY_BASE}/v1/evaluate`,
72      model: "typesafe-ai/jev",
73      apiKey: env.AI_GATEWAY_API_KEY,
74    };
75  }
76
77  // No forced override: use default precedence.
78  if (env.TYPESAFE_API_KEY) {
79    const base = (env.TYPESAFE_BASE_URL ?? TYPESAFE_BASE_DEFAULT).replace(
80      /\/$/,
81      "",
82    );
83    return {
84      ok: true,
85      name: "typesafe",
86      endpoint: `${base}/v1/systemone`,
87      model: "jev-latest",
88      apiKey: env.TYPESAFE_API_KEY,
89    };
90  }
91
92  if (env.AI_GATEWAY_API_KEY) {
93    return {
94      ok: true,
95      name: "gateway",
96      endpoint: `${GATEWAY_BASE}/v1/evaluate`,
97      model: "typesafe-ai/jev",
98      apiKey: env.AI_GATEWAY_API_KEY,
99    };
100  }
101
102  return {
103    ok: false,
104    reason: "no TYPESAFE_API_KEY or AI_GATEWAY_API_KEY",
105  };
106}
107
hooks/status.ts 440 lines
1/**
2 * What `/jev` prints. This is the router's only guaranteed-visible surface:
3 * a command's output row draws on every surface, where a footer label may
4 * not, so anything you need to be sure of belongs here.
5 */
6
7import type { JevResult } from "./jev.ts";
8import { LOW_CONFIDENCE } from "./label.ts";
9import {
10  decisionOf,
11  DEFAULT_STICKY_CONFIDENCE,
12  stickyDecision,
13  type Decision,
14  type Tier,
15} from "./policy.ts";
16import type { ProviderResult } from "./provider.ts";
17
18/**
19 * What the API said a turn cost, and which model it says answered. The
20 * shape of the engine's `TurnUsage`, spelled out here so this file stays
21 * free of engine types and runs under plain `node`.
22 */
23export type Usage = {
24  /** The model that answered, by the id the API reports. */
25  model: string;
26  input_tokens: number;
27  output_tokens: number;
28  cache_read_input_tokens: number;
29  cache_creation_input_tokens: number;
30};
31
32/**
33 * One turn's outcome, kept for the status report. `usage` arrives after the
34 * decision, from the `stop` chunk of each step, so it is filled in later and
35 * is absent for a turn still running or one whose response never came.
36 */
37export type Attempt = {
38  prompt: string;
39  ms: number;
40  usage?: Usage;
41  /**
42   * What started the turn, when it was not the person typing. Absent for a
43   * typed prompt. `notify`: the main loop woke because a background task
44   * finished, and the engine's `<task-notification>` was the turn's text.
45   * `agent`: a subagent's own loop, which no `turn.start` announces; its
46   * steps are seen but not routed. Without this, one prompt that spawned
47   * three reviewers read as one reply that changed model three times.
48   */
49  kind?: "notify" | "agent";
50  /** For `kind: 'agent'`: which subagent, as `$.agent.list()` describes it. */
51  agent?: AgentTag;
52} & ({ decision: Decision } | { skipped: string });
53
54/**
55 * Which subagent a turn ran in. `type` is the definition (`general-purpose`,
56 * `Explore`); `label` its row's description (`Review library-sync cluster`),
57 * or the id when the list has no row for it yet, in which case `type` is
58 * absent too.
59 */
60export type AgentTag = {
61  type?: string;
62  label: string;
63};
64
65/** The tag for a turn held on its previous tier: `held:haiku`. */
66export function heldMark(attempt: Attempt): string | null {
67  return "decision" in attempt && attempt.decision.held !== undefined
68    ? `held:${attempt.decision.held}`
69    : null;
70}
71
72/** The short tag for a turn nobody typed: `notify`, `agent:Explore`, `agent`. */
73export function kindMark(
74  attempt: Pick<Attempt, "kind" | "agent">,
75): string | null {
76  if (attempt.kind === "notify") return "notify";
77  if (attempt.kind === "agent")
78    return attempt.agent?.type ? `agent:${attempt.agent.type}` : "agent";
79  return null;
80}
81
82const NOTIFICATION = /^\s*<task-notification>/;
83const tagOf = (text: string, tag: string) =>
84  text.match(new RegExp(`<${tag}>([\\s\\S]*?)</${tag}>`))?.[1]?.trim();
85
86/**
87 * Reads the engine's task notification, when the turn's text is one: what
88 * the row should say instead of the XML envelope.
89 */
90export function notificationOf(text: string): string | null {
91  if (!NOTIFICATION.test(text)) return null;
92  return tagOf(text, "summary") ?? `task ${tagOf(text, "task-id") ?? "?"}`;
93}
94
95/**
96 * Folds one step's usage into its turn: counts sum, the model is the last
97 * step's, as the engine defines a turn's usage. Mutates, because the same
98 * object sits in the history and in the by-turn lookup.
99 */
100export function addUsage(attempt: Attempt, usage: Usage): void {
101  const prior = attempt.usage;
102  attempt.usage = {
103    model: usage.model,
104    input_tokens: (prior?.input_tokens ?? 0) + usage.input_tokens,
105    output_tokens: (prior?.output_tokens ?? 0) + usage.output_tokens,
106    cache_read_input_tokens:
107      (prior?.cache_read_input_tokens ?? 0) + usage.cache_read_input_tokens,
108    cache_creation_input_tokens:
109      (prior?.cache_creation_input_tokens ?? 0) +
110      usage.cache_creation_input_tokens,
111  };
112}
113
114/**
115 * How much of what the turn's requests carried was read from cache, 0 to 1.
116 * Everything carried is uncached input plus cache reads plus cache writes;
117 * this is the cost-relevant measure, since reads bill at a tenth.
118 */
119export function cacheRatio(usage: Usage): number {
120  const carried =
121    usage.input_tokens +
122    usage.cache_read_input_tokens +
123    usage.cache_creation_input_tokens;
124  return carried === 0 ? 0 : usage.cache_read_input_tokens / carried;
125}
126
127export type Status = {
128  enabled: boolean;
129  surface: string | null;
130  provider: ProviderResult;
131  timeoutMs: number;
132  /** The confidence a switch must clear, or null when stickiness is off. */
133  sticky: number | null;
134  offered: readonly Tier[];
135  excluded: readonly Tier[];
136  announce: boolean;
137  attempts: readonly Attempt[];
138};
139
140/**
141 * One turn's outcome from Jev's answer, so the three ways a turn can fail to
142 * route all land in one place and all get announced the same way.
143 */
144export function attemptOf(
145  text: string,
146  result: JevResult,
147  offered: readonly Tier[],
148  hold: { sticky: number | null; running: Decision | null } = {
149    sticky: null,
150    running: null,
151  },
152): Attempt {
153  const summary = notificationOf(text);
154  const head =
155    summary === null
156      ? { prompt: text }
157      : { prompt: summary, kind: "notify" as const };
158
159  if (!result.ok) return { ...head, ms: result.ms, skipped: result.reason };
160
161  const fresh = decisionOf(result.answers, offered);
162  const decision =
163    fresh && hold.sticky !== null
164      ? stickyDecision(fresh, hold.running, hold.sticky)
165      : fresh;
166  if (!decision) {
167    return {
168      ...head,
169      ms: result.ms,
170      skipped: "Jev answered but named no tier we offered",
171    };
172  }
173  return { ...head, ms: result.ms, decision };
174}
175
176/** The last few turns, newest first, so the report stays one screen. */
177export const HISTORY_LIMIT = 5;
178
179function shorten(text: string, width = 44): string {
180  const flat = text.replace(/\s+/g, " ").trim();
181  return flat.length > width ? `${flat.slice(0, width - 1)}…` : flat;
182}
183
184function attemptLine(attempt: Attempt): string {
185  const when = `${String(attempt.ms).padStart(4)}ms`;
186  const mark = kindMark(attempt);
187  const what = `${mark ? `[${mark}] ` : ""}${shorten(attempt.prompt)}`;
188  if ("skipped" in attempt) {
189    // A subagent's row names the agent, since "unrouted" is the whole story.
190    return attempt.kind === "agent"
191      ? `  ${when}  unrouted — ${what}`
192      : `  ${when}  unrouted — ${attempt.skipped}`;
193  }
194  const { tier, effort, confidence } = attempt.decision;
195  // A held turn is low-confidence by construction, so saying both is noise;
196  // the hold is the more useful of the two.
197  const held = heldMark(attempt);
198  const doubt = held ?? (confidence < LOW_CONFIDENCE ? "(low confidence)" : "");
199  return `  ${when}  ${tier}·${effort} ${confidence.toFixed(2)}${doubt ? ` ${doubt}` : ""}  ${what}`;
200}
201
202/** Thousands, rounded, for token counts: 130k, 2k, 0k. */
203function kOf(n: number): string {
204  return `${Math.round(n / 1000)}k`;
205}
206
207/**
208 * The line under a turn saying what the API reports actually answered, and
209 * what the requests carried. This is the intrinsic check: the route line is
210 * what we asked for; this is what we got.
211 *
212 * A dated id (`claude-opus-5-20260901`) still confirms `claude-opus-5`. A
213 * different model is marked `≠`, which is the one case worth looking at.
214 */
215function usageLine(attempt: Attempt): string | null {
216  const usage = attempt.usage;
217  if (!usage) return null;
218
219  let verdict = "";
220  if ("decision" in attempt) {
221    const asked = attempt.decision.model;
222    const matches =
223      usage.model === asked || usage.model.startsWith(`${asked}-`);
224    verdict = matches ? " ✓" : ` ≠ ${asked}`;
225  }
226
227  const carried =
228    usage.input_tokens +
229    usage.cache_read_input_tokens +
230    usage.cache_creation_input_tokens;
231  const pct = Math.round(cacheRatio(usage) * 100);
232  return (
233    `          answered ${usage.model}${verdict}  ` +
234    `cache ${pct}%  ${kOf(carried)} in  ${kOf(usage.output_tokens)} out`
235  );
236}
237
238/**
239 * The footer put at the end of a completed reply: what was asked for, and
240 * what the API says answered.
241 *
242 * It goes at the end because `usage` only exists once the response is whole —
243 * the stop chunk carries it. The route line at the top of the reply is the
244 * immediate signal; this is the settled one.
245 *
246 * Fenced, because markdown collapses leading whitespace and joins consecutive
247 * lines into one paragraph: unfenced, the rule and the two lines would render
248 * as a single run-on. A fence keeps the alignment and reads as data, not prose.
249 */
250export function usageFooter(attempt: Attempt): string | null {
251  const usage = attempt.usage;
252  if (!usage) return null;
253
254  const carried =
255    usage.input_tokens +
256    usage.cache_read_input_tokens +
257    usage.cache_creation_input_tokens;
258  const cost =
259    `cache ${Math.round(cacheRatio(usage) * 100)}% · ` +
260    `${kOf(carried)} in · ${kOf(usage.output_tokens)} out`;
261
262  let jev: string;
263  let api: string;
264
265  if ("decision" in attempt) {
266    const { tier, effort, confidence, model: asked } = attempt.decision;
267    const matches =
268      usage.model === asked || usage.model.startsWith(`${asked}-`);
269    const tags = [heldMark(attempt), kindMark(attempt)].filter(
270      (t) => t !== null,
271    );
272    jev =
273      `${tier}·${effort} · ${Math.round(confidence * 100)}%` +
274      `${tags.map((t) => ` · ${t}`).join("")} · ${attempt.ms}ms`;
275    api = `${usage.model}${matches ? " ✓" : ` ≠ ${asked}`} · ${cost}`;
276  } else {
277    jev = `unrouted — ${attempt.skipped}`;
278    api = `${usage.model} · ${cost}`;
279  }
280
281  const rows = [`jev  ${jev}`, `api  ${api}`];
282  // Count code points: the separators and check marks are multi-byte, and a
283  // rule measured in UTF-16 units would overshoot the text it sits above.
284  const width = Math.max(...rows.map((r) => [...r].length));
285  return ["```", "─".repeat(width), ...rows, "```"].join("\n");
286}
287
288/**
289 * The report, as plain lines. Written so the first three tell you whether
290 * the thing is on at all, which is the question that brings people here.
291 */
292export function statusReport(status: Status): string {
293  const lines: string[] = ["jev-router"];
294
295  lines.push(`  routing   ${status.enabled ? "on" : "off (/jev on)"}`);
296  lines.push(`  surface   ${status.surface ?? "unknown"}`);
297
298  if (status.provider.ok) {
299    const key =
300      status.provider.name === "typesafe"
301        ? "TYPESAFE_API_KEY"
302        : "AI_GATEWAY_API_KEY";
303    lines.push(`  provider  ${status.provider.name} · ${key} is set`);
304  } else {
305    lines.push(`  provider  NO KEYS — nothing will route`);
306  }
307
308  lines.push(`  budget    ${status.timeoutMs}ms`);
309  lines.push(
310    `  sticky    ${
311      status.sticky === null
312        ? "off (JEV_ROUTER_STICKY=1)"
313        : `on, switch needs ${Math.round(status.sticky * 100)}%`
314    }`,
315  );
316  lines.push(`  tiers     ${status.offered.join(", ")}`);
317  lines.push(
318    `  announce  ${status.announce ? "on, a line per turn" : "off (/jev loud)"}`,
319  );
320  if (status.excluded.length > 0) {
321    lines.push(`  excluded  ${status.excluded.join(", ")}`);
322  }
323
324  lines.push("");
325  if (status.attempts.length === 0) {
326    lines.push("  No turns yet. Send a prompt, then run /jev again.");
327    return lines.join("\n");
328  }
329
330  lines.push("  Recent turns, newest first:");
331  for (const attempt of status.attempts) {
332    lines.push(attemptLine(attempt));
333    const usage = usageLine(attempt);
334    if (usage !== null) lines.push(usage);
335  }
336
337  return lines.join("\n");
338}
339
340/** The reply to `/jev on`, `/jev off` and anything unrecognised. */
341export function toggleReply(enabled: boolean): string {
342  return enabled
343    ? "Jev routing on. The next turn picks its own model."
344    : "Jev routing off. Turns run on the session model.";
345}
346
347/**
348 * The line put at the top of each reply, as markdown.
349 *
350 * Markdown, because the line rides in the reply's own text and that is what
351 * the transcript renders. A blockquote sets it off from prose with a rail
352 * and dimmer text; the tier is inline code, which the theme colours. Neither
353 * a render hook nor `$.ui.log` drew anything in the desktop app, so this is
354 * the styling that is actually available.
355 */
356export function liveLine(attempt: Attempt): string {
357  if ("skipped" in attempt) {
358    return `> ⚠️ \`unrouted\` · ${attempt.skipped}`;
359  }
360
361  const { tier, effort, confidence } = attempt.decision;
362  const held = heldMark(attempt);
363  const doubt = held === null && confidence < LOW_CONFIDENCE ? "?" : "";
364  const pct = Math.round(confidence * 100);
365  const tags = [held, kindMark(attempt)].filter((t) => t !== null);
366  return (
367    `> ✳️ \`${tier}\` · ${effort} · ${pct}${"%"}${doubt}` +
368    `${tags.map((t) => ` · ${t}`).join("")} · ${attempt.ms}ms`
369  );
370}
371
372/**
373 * The rule drawn under the line, closing it off from the reply.
374 *
375 * It needs the blank line before it: `---` on the line after text is a setext
376 * heading underline, which would turn the route into a heading instead.
377 */
378export const REPLY_SEPARATOR = "\n\n---\n\n";
379
380/**
381 * What sits between the reply's last text and the footer. A blank line, so
382 * the fence opens a block of its own instead of joining the last paragraph.
383 */
384export const FOOTER_SEPARATOR = "\n\n";
385
386/** The reply to `/jev quiet` and `/jev loud`. */
387export function announceReply(announce: boolean): string {
388  return announce
389    ? "Jev will announce each route in the transcript."
390    : "Jev will route quietly. Run /jev to see what it has been doing.";
391}
392
393/**
394 * Reads `/jev sticky`, `/jev sticky off`, `/jev sticky 0.6` and says what the
395 * bar is now. `current` is the session's bar, or null when it is off.
396 *
397 * A bare `sticky` keeps a bar already set rather than resetting it to the
398 * default, so turning it off and on again does not silently lose a tuned
399 * value. A number that cannot be a confidence changes nothing and says so,
400 * rather than falling back to the default: an unnoticed 0.75 is worse than a
401 * refusal, since the point of the bar is knowing which one you are running.
402 */
403export function stickyCommand(
404  rest: string,
405  current: number | null,
406): { sticky: number | null; text: string } {
407  const arg = rest.trim().toLowerCase().replace(/%$/, "");
408
409  if (arg === "off") {
410    return {
411      sticky: null,
412      text: "Switching freely again. /jev sticky holds a shaky switch.",
413    };
414  }
415
416  if (arg === "" || arg === "on") {
417    const bar = current ?? DEFAULT_STICKY_CONFIDENCE;
418    return { sticky: bar, text: stuckAt(bar) };
419  }
420
421  const parsed = Number(arg);
422  const ratio = parsed > 1 ? parsed / 100 : parsed;
423  if (!Number.isFinite(parsed) || ratio <= 0 || ratio >= 1) {
424    return {
425      sticky: current,
426      text:
427        `"${rest.trim()}" is not a confidence. Give a number between 0 and 1 ` +
428        "(0.6), or a percentage (60).",
429    };
430  }
431  return { sticky: ratio, text: stuckAt(ratio) };
432}
433
434function stuckAt(bar: number): string {
435  return (
436    `Holding the tier until Jev is ${Math.round(bar * 100)}% sure of a switch. ` +
437    "/jev sticky off to switch freely."
438  );
439}
440