SLOPSHOPPER

jev-router

Asks TypeSafe Jev before each task which Claude model and effort fit best, then runs the turn on it

newbandguardcommandtoastprompt
★ 2v0.2.0MITupdated 2026-10-09dominicrico/jev-router
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · jev-router
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /jev ⎿ jev-router: ⎿ jev-router: ┌ Jev router ────────────────────────────┐ ⎿ jev-router: │ state ● on mode balanced │ ⎿ jev-router: │ pool haiku · sonnet · opus · fable │ ⎿ jev-router: │ sticky auto cache ○ cold │ ⎿ jev-router: │ cap high │ ◆ JEV ▏balanced▕ waiting for the first task ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
◆ JEV ▏balanced▕ waiting for the first task
README

<img src="assets/logo.jpg" width="200" alt="Jev Router, the smug switchman raccoon">

<h1 align="center">jev-router: automatic Claude model routing for Claude Code</h1>

<em>Switch between Opus, Sonnet and Haiku automatically. He pulls one lever; your task rides the right track.</em>

<img src="https://img.shields.io/github/v/release/dominicrico/jev-router?style=flat-square&color=111111&label=release&include_prereleases" alt="Release"> <img src="https://img.shields.io/badge/Claude%20Code-mod-111111?style=flat-square" alt="Claude Code mod"> <img src="https://img.shields.io/badge/models-haiku%20%C2%B7%20sonnet%20%C2%B7%20opus%20%C2%B7%20fable-111111?style=flat-square" alt="Models"> <img src="https://img.shields.io/badge/license-MIT-111111?style=flat-square" alt="MIT license">

<img src="assets/hero.jpg" width="880" alt="jev-router: right model, right effort, every task. Tasks routed to haiku, sonnet or opus, minus 97% cost on trivial tasks, minus 65% on standard, 253 ms per pick. And yet hard tasks cost more.">

<a href="https://dominicrico.github.io/jev-router/">Website</a> · <a href="#install-the-claude-code-plugin">Install</a> · <a href="#does-it-reduce-claude-code-costs">Benchmarks</a> · <a href="#faq">FAQ</a> · <a href="CHANGELOG.md">Changelog</a>


jev-router is a free, open-source Claude Code plugin that routes every prompt, step and subagent to the right Claude model (Haiku, Sonnet, Opus, Fable) and the right reasoning effort. It asks TypeSafe Jev which model a task needs, caps the effort so hard tasks do not burn tokens, protects your prompt cache from needless model switches, and shows the choice in a live status band above the prompt. Use it to cut Claude Code cost on easy work without giving up Opus when a task needs it.

You know the problem. You rename one variable and your most expensive model thinks about it for a minute. Or you start a gnarly refactor and a bargain model happily shrugs. jev-router puts a switchman in front of your session: before every task he reads it, asks Jev, pulls the lever, and the turn runs on the model and effort that fit.

Contents: Status band · Features · Does it reduce costs? · How it works · Install · Commands · Options · Privacy · FAQ

The status band

A live Claude Code status band above your prompt shows which model and effort were picked. Colour coded, with a spinner while a task runs.

<img src="assets/band.gif" width="880" alt="The band in all its states, one slide at a time: idle, running, haiku, sonnet, capped opus, !full unlock, cache kept, pinned, escalated, low confidence, cheap mode, paused, off.">

<img src="assets/band.png" width="880" alt="The band above the prompt: JEV, balanced mode, sonnet, medium effort, 88% confidence, warm cache of 42k tokens. Four colour-coded tiers below: haiku trivial, sonnet standard, opus hard, fable hardest.">

PartMeaning
balancedRouting mode: efficient, balanced or cheap
sonnet ▂▄__The model the task runs on, with its tier. Green haiku, blue sonnet, purple opus, gold fable
medium ▰▰▱▱▱Reasoning effort, low to max
conf ▮▮▮▮▯ 88%How sure Jev is. Green from 70%, amber from 50%, red below
cache 🔒 warm 42kWhether the prompt cache is still worth protecting, and how big it is

When the cache wins, the band says so: opus (kept 🔒 cache warm; wanted haiku).

Features: model routing, effort cap, cache guard, subagents

  • Picks the model and the effort, before every task, from haiku, sonnet, opus and fable. You can narrow the pool.
  • Three routing modes. efficient for the best result, balanced for quality and cost evenly, cheap for the cheapest model that can plausibly succeed.
  • Guards the prompt cache. Switching models throws the cache away. While it is warm and big enough to matter, he stays put unless Jev confidently asks for something stronger.
  • Caps the effort. Jev's effort pick is limited to high by default, because effort drives token use far more than the model does. Start a prompt with !full (or run /jev full) to lift the cap for that one prompt.
  • Escalates when a task struggles. After escalateAfter (default 3) failed tool calls in a row, the rest of the task moves up one model. The band shows ↑.
  • Pins a model for one prompt. Start a prompt with !opus, !sonnet, !haiku or !fable to skip Jev and use that model (at the effort cap). !cheap and !efficient set the routing mode for that prompt. Markers combine, like !opus !full.
  • Keeps a savings tally. /jev status shows an estimated saving against the model you would otherwise use (baselineModel, default opus), kept across sessions. It is a model-price estimate in relative units, blind to effort.
  • Never blocks you. After 3 failed Jev calls in a row it pauses for a minute instead of paying a timeout on every step. Optional fallback: heuristic guesses locally meanwhile. Jev slow or down? The session model keeps working.
  • Routes every step. Before each step after the first, Jev is asked again, so a task that turned out easier or harder moves to a fitting model (the cache guard still applies). Turn it off with routeSteps.
  • Routes subagents. Each subagent is routed when it is spawned, from its full task prompt, and the pick shows in the subagent list as a tag on its description (find importers · haiku/medium). Its steps then run on that model and effort. An explicit model on the Agent call wins. Turn it off with routeSubagents.
  • Keeps score. /jev status shows how often each model was used this session.

Does it reduce Claude Code costs?

Three benchmarks. First, 30 single-prompt tasks: 270 calls to Jev for the picks, then 147 real runs for token usage and cost, with the effort cap at its default (high) and uncapped. Second, 6 multi-step tasks with tools on a fixture repo, 72 real runs, checking per prompt against every-step-and-subagent routing. Third, 4 harder tasks graded by hidden tests, 100 real runs, where quality is measured.

<img src="assets/fit.png" width="880" alt="Routing fit: efficient 100%, balanced 87%, cheap 66% of picks fit the task. 253 ms added per task, 97 to 100% same pick on repeat.">

<img src="assets/cost.png" width="880" alt="Cost as a percentage of always opus with the effort capped at high. Trivial tasks 3%, standard 32 to 35%, hard 143%, 124% and 44% for efficient, balanced and cheap. Overall 99%, 87% and 36%. Uncapped hard tasks cost 342%, 307% and 61%.">

Cost vs always opus (no plugin)efficientbalancedcheap
Trivial tasks-97%-97%-97%
Standard tasks-65%-65%-68%
Hard tasks, cap high (default)+43%+24%-56%
Hard tasks, uncapped+242%+207%-39%
All 30, cap high (default)-1%-13%-64%
All 30, uncapped+123%+101%-53%

The honest read: the model switch saves a lot on easy and mid tasks, and the default cap of high brings every mode to or below always-opus overall. Hard tasks still cost 24% to 43% more in efficient and balanced, because Jev picks opus there and asks for deeper reasoning. !full lifts the cap for one prompt when you want that. The cap changes tokens spent, not answer quality, and that is not measured.

<img src="assets/cache.png" width="880" alt="Cache guard: following every pick in a warm session costs 4735 against 1500 for always opus. With the guard it is 1500. With cold cache and tasks 400 seconds apart it is 820.">

Multi-step tasks, with tools

<img src="assets/agentic.png" width="880" alt="Multi-step tasks with tools, cost as a percentage of always opus: always sonnet 52%, jev-router per prompt 52%, jev-router every step and subagents 47%. All strategies finished 18 of 18 tasks.">

6 tasks × 3 runs, real runs with toolstasks donecostvs always opusvs always sonnet
no plugin: always opus18/18$2.46
no plugin: always sonnet18/18$1.29-48%
jev-router, per prompt18/18$1.27-48%-2%
jev-router, every step + subagents18/18$1.15-53%-11%

Routing every step and subagent saved 11% on top of per-prompt routing. Most of that comes from one task: asked to use a subagent, it paired a sonnet main thread with a haiku subagent where per-prompt routing sometimes left the subagent on opus. On the other five tasks it matched per-prompt routing. Honest limits: these are easy tasks on a small repo, three runs each, and every strategy finished all of them, so this shows nothing was lost here, not that nothing is lost on hard work. Jev picked sonnet for nearly every one of these tasks, so against always-sonnet the gain is small.

Harder tasks, graded by hidden tests

<img src="assets/hard.png" width="880" alt="Harder tasks graded by hidden tests, cost as a percentage of always opus: always sonnet 39%, jev-router per prompt 86%, every step and subagents 67%, with !full 60%. Tasks done out of 20: opus 19, sonnet 19, per prompt 19, every step 18, !full 20.">

4 tasks × 5 runs, hidden teststasks donecostvs always opusvs always sonnet
no plugin: always opus19/20$3.10
no plugin: always sonnet19/20$1.19-61%
jev-router, per prompt19/20$2.65-14%+122%
jev-router, every step + subagents18/20$2.08-33%+74%
jev-router, !full20/20$1.87-40%+57%

The honest read: on these tasks plain sonnet did as well as opus, and was the cheapest. The tasks (a concurrency bug, an interval merger with open and closed bounds, a refactor that must keep an invariant, a flaky test with three causes) are harder than the first set, but not hard enough to need opus. jev-router lands between the two on cost because Jev sent about half of the spend to opus. Pass counts differ by at most one run, so none of the quality differences are significant at 5 runs. If your work looks like this, routing buys you less than just using sonnet; it earns its keep when tasks vary, or when your default is opus.

Without the cache guard, following every pick in a long warm session cost 3x more than just using opus. Inside one long warm session the guard keeps you on the first model, so the savings above show up mostly when the cache is cold or the context is small. Method, per-tier tables and all caveats: benchmarks/.

How it works

<img src="assets/how.png" width="880" alt="How it works in five steps: a task arrives, ask Jev which model and effort, check whether the prompt cache is warm, run the turn on the pick, show it above the prompt.">

Per prompt, the mod sends Jev the routing mode, the current model, the cache state, the last few messages and the task text, and asks two choice questions: which model, and how much effort. The answer is applied to every step of that task. After each step the mod notes which model the API cached and how many tokens, so the next task knows whether a switch is worth losing the cache.

Sent to api.typesafe.ai: the task text, recent conversation, and subagent descriptions. Each step and subagent adds one call (about 250 ms). Your API key travels in the request header and nowhere else.

Install the Claude Code plugin

Needs a Claude Code version with mods (hooks modules) and a TypeSafe API key.

/plugin marketplace add dominicrico/jev-router
/plugin install jev-router@jev-router

Update:

claude plugin marketplace update jev-router
claude plugin update jev-router@jev-router

API key

Any one of these, checked in this order:

  1. The plugin option apiKey (/config)
  2. /jev key <key>, stored by the mod across sessions (/jev key clear removes it)
  3. The TYPESAFE_API_KEY environment variable

That was it. He'd be proud. He won't say it.

Commands

CommandWhat it does
/jev or /jev statusSettings, cache state, last decision and per-model usage
`/jev efficient \balanced \cheap`Set the routing mode, from the next task
`/jev sticky off \auto \strict`Set cache stickiness
`/jev cap low \medium \high \xhigh \max \none`Set the effort cap
/jev fullLift the effort cap for the next prompt only
/jev stats clearReset the savings tally
/jev on / /jev offEnable routing / use the session model
/jev key <key>Store the TypeSafe API key
┌ Jev router ────────────────────────────┐
│ state    ● on    mode  balanced        │
│ pool     haiku · sonnet · opus · fable │
│ sticky   auto   cache ● warm           │
│ last     sonnet / medium  ▮▮▮▮▯ 0.88   │
│ usage    haiku   ░░░░░░░░░░   0  0%    │
│          sonnet  ██████████   1  100%  │
│          opus    ░░░░░░░░░░   0  0%    │
│          fable   ░░░░░░░░░░   0  0%    │
│ key      …fd1c  (TYPESAFE_API_KEY)     │
└────────────────────────────────────────┘

Options

OptionDefaultMeaning
modebalancedRouting mode
modelsall fourAliases Jev may choose from
effortCaphighHighest effort the router applies. none follows Jev. Per prompt: start with !full or run /jev full
pauseMs60000Stop asking Jev for this long after 3 failed calls in a row
escalateAfter3Failed tool calls in a row before the task moves up one model. 0 is off
baselineModelopusThe model the savings estimate compares against
sendHistorytrueSend recent messages to Jev. Off sends only the task text
redacttrueReplace API keys, tokens, private keys and KEY=value lines with [redacted] before sending
maxTaskChars8000How much of the prompt is sent
fallbacksessionWhen Jev is down: session keeps the session model, heuristic guesses locally
routeStepstrueAsk Jev again before every step after the first. Off: one call per prompt
routeSubagentstruePick a model and effort for each subagent from its description
stickinessautooff: always follow Jev. auto: while the cache is warm, only upgrade on high confidence. strict: never switch while warm
minConfidence0.7Confidence needed for a warm-cache upgrade
minContextTokens8000Below this a switch is free and stickiness is skipped
cacheTtlMs300000Time after the last request when the cache counts as cold
timeoutMs4000Keep the current model when Jev takes longer

Privacy

Sent to api.typesafe.ai: the prompt text (up to maxTaskChars), recent messages if sendHistory is on, and subagent descriptions. Before sending, redact replaces API keys, GitHub tokens, AWS keys, bearer tokens, private key blocks and NAME_KEY=value style lines with [redacted]. That is pattern matching, not a guarantee, so keep sendHistory off if your conversations carry things it would not catch. Your TypeSafe key is stored in the plugin store in plain text if you use /jev key, and the command line lands in the transcript; prefer the TYPESAFE_API_KEY environment variable.

FAQ

How do I switch between Opus, Sonnet and Haiku automatically in Claude Code? Install jev-router. It picks the model per prompt, again before each step, and once per subagent. Pin a model for one prompt with !opus, !sonnet or !haiku.

How do I reduce Claude Code token cost? Send easy tasks to cheaper models and cap the reasoning effort. jev-router does both; the benchmarks show where it saves (trivial tasks -97%, standard -65% against always Opus) and where it does not (on harder tasks plain Sonnet was as good and cheaper).

What is the Claude Code effort cap? Jev often asks for xhigh reasoning on hard tasks, which can use several times the tokens. jev-router limits it to high by default; start a prompt with !full to lift it once.

Does jev-router send my code to a third party? It sends the prompt, optionally the last messages, and subagent task text to api.typesafe.ai. Keys, tokens, private keys and password= style values are redacted first, and sendHistory can turn the history off. See Privacy.

Will it slow Claude Code down? Each Jev call adds about 250 ms. After three failed calls in a row it pauses for a minute and keeps your current model.

Does it need a config file? No. Set a key, pick a mode, done. Everything else has a default.

Why not just always use the best model? You can. /jev efficient and a pool of opus and fable. Your invoice will have opinions.

Why not just always use the cheap one? /jev cheap. The hard task will have opinions.

He kept my old model even though Jev wanted a cheaper one. Is it broken? No. Your cache was warm and a switch would have rewritten it. He did the maths. He was right.

What if Jev is down? Then you get your session model and one polite toast. He never blocks a turn.

Why a raccoon? Because he is a small creature that gets very serious about picking the right track.

Tests

claude plugin validate .
claude plugin test .

Benchmarks: see benchmarks/.

License

MIT.

<sub>Keywords: Claude Code plugin, Claude model router, Claude Code cost optimization, Opus Sonnet Haiku switching, Claude Code subagents, prompt cache, reasoning effort, LLM routing.</sub>

Source 2 files
hooks/register.tsx 689 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { Cache, Decision, Effort, EffortCap, Mode, Pending, Savings, Sticky } from '../types'
5
6export const JEV_URL = 'https://api.typesafe.ai/v1/systemone'
7const MODES: readonly Mode[] = ['efficient', 'balanced', 'cheap']
8const STICKY: readonly Sticky[] = ['off', 'auto', 'strict']
9const EFFORTS: readonly Effort[] = ['low', 'medium', 'high', 'xhigh', 'max']
10const CAPS: readonly EffortCap[] = [...EFFORTS, 'none']
11
12// weight: relative price per token, the multipliers the `about` lines state (fable has none, 10 is an assumption).
13const CLAUDE: Record<string, { id: string; about: string; weight: number }> = {
14  haiku: {
15    id: 'claude-haiku-5-5',
16    weight: 1,
17    about: 'Claude Haiku: fastest and cheapest (~1x). Simple edits, renames, lookups, short answers, boilerplate.',
18  },
19  sonnet: {
20    id: 'claude-sonnet-5-5',
21    weight: 3,
22    about: 'Claude Sonnet: strong general coding at mid cost (~3x). Features, refactors, tests, normal debugging.',
23  },
24  opus: {
25    id: 'claude-opus-5-5',
26    weight: 5,
27    about: 'Claude Opus: top reasoning at high cost (~5x). Hard debugging, architecture, large multi-file changes.',
28  },
29  fable: {
30    id: 'claude-fable-5-1',
31    weight: 10,
32    about: 'Claude Fable: most capable, highest cost. The hardest, longest, most ambiguous tasks.',
33  },
34}
35
36const MODE_RULE: Record<Mode, string> = {
37  efficient: 'Pick the model and effort that give the best result for this task; cost is secondary.',
38  balanced: 'Weigh result quality and cost evenly; use a stronger model only when the task clearly needs it.',
39  cheap: 'Pick the cheapest model and lowest effort that can plausibly succeed at this task.',
40}
41
42const EFFORT_ABOUT: Record<Effort, string> = {
43  low: 'Trivial or mechanical, little thinking needed',
44  medium: 'Ordinary task, some reasoning',
45  high: 'Non-trivial reasoning or multi-step work',
46  xhigh: 'Hard problem, careful deep reasoning',
47  max: 'Extremely hard, think as long as needed',
48}
49
50const S = { plugin: 'jev-router' } as const
51const modeAtom = atom({ ...S, key: 'mode' } as const, null)
52const enabledAtom = atom({ ...S, key: 'enabled' } as const, true)
53const lastAtom = atom({ ...S, key: 'last' } as const, null)
54const cacheAtom = atom({ ...S, key: 'cache' } as const, null)
55const stickyAtom = atom({ ...S, key: 'sticky' } as const, null)
56const capAtom = atom({ ...S, key: 'cap' } as const, null)
57const pendingAtom = atom({ ...S, key: 'pending' } as const, null)
58const frameAtom = atom({ ...S, key: 'frame' } as const, 0)
59const warnedAtom = atom({ ...S, key: 'warned' } as const, false)
60
61type JevChoice = { choice?: string; confidence?: number }
62type JevReply = { answers?: { model?: JevChoice; effort?: JevChoice } }
63
64export type Config = {
65  apiKey: string
66  timeoutMs: number
67  pool: string[]
68  defaultMode: Mode
69  defaultSticky: Sticky
70  defaultCap: EffortCap
71  pauseMs: number
72  routeSteps: boolean
73  routeSubagents: boolean
74  baseline: string
75  sendHistory: boolean
76  redact: boolean
77  maxTaskChars: number
78  fallback: 'session' | 'heuristic'
79  escalateAfter: number
80  minConfidence: number
81  minContextTokens: number
82  cacheTtlMs: number
83}
84type Picked = Omit<Decision, 'turnId'>
85
86// Estimated model-price units of one step: tokens as Anthropic bills them (cache reads 0.1x, cache writes 1.25x,
87// output 5x input) times the model's weight. Relative, not dollars, and blind to effort.
88type StepUsage = { input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number }
89export const stepUnits = (u: StepUsage, alias: string) =>
90  ((u.input_tokens + 1.25 * u.cache_creation_input_tokens + 0.1 * u.cache_read_input_tokens + 5 * u.output_tokens) * CLAUDE[alias]!.weight) / 1000
91
92export const savingsLine = (t: Savings) =>
93  t.actual === 0 ? 'no steps measured yet' : `est. ${Math.round((1 - t.actual / t.baseline) * 100)}% vs always ${t.name} over ${t.steps} steps (model price only; effort not compared)`
94
95let totals: Savings | null = null // loaded from the store once, then kept in memory
96
97// Cheapest to dearest (CLAUDE's key order); a switch up this ladder is an upgrade.
98const RANK = Object.keys(CLAUDE)
99const aliasOf = (model: string) => Object.keys(CLAUDE).find(a => CLAUDE[a]!.id === model)
100
101// Jev's effort drives token use more than its model pick does, so it is capped.
102// `capped` is what Jev wanted when the cap lowered it.
103export function capEffort(effort: Effort, cap: EffortCap): { effort: Effort; capped?: Effort } {
104  if (cap === 'none' || EFFORTS.indexOf(effort) <= EFFORTS.indexOf(cap)) return { effort }
105  return { effort: cap, capped: effort }
106}
107
108// Markers at the start of a prompt apply to that prompt only and are stripped before the model reads it:
109// !full lifts the effort cap, !<model> pins the model (no Jev call), !cheap and !efficient pick the routing mode.
110export function markers(text: string, pool: readonly string[]): { text: string; pending: Pending | null } {
111  const pending: Pending = {}
112  for (let m; (m = /^\s*!(full|cheap|efficient|[a-z]+)\b[ \t]*/i.exec(text)); ) {
113    const w = m[1]!.toLowerCase()
114    if (w === 'full') pending.full = true
115    else if (w === 'cheap' || w === 'efficient') pending.mode = w
116    else if (pool.includes(w)) pending.pin = w
117    else break
118    text = text.slice(m[0].length)
119  }
120
121  return { text, pending: Object.keys(pending).length ? pending : null }
122}
123
124export const isWarm = (cache: Cache | null, now: number, cfg: Config): cache is Cache =>
125  cache !== null && now - cache.at <= cfg.cacheTtlMs && cache.contextTokens >= cfg.minContextTokens
126
127// A switch discards the prompt cache. While it is warm and worth keeping, stay
128// on the cached model unless Jev confidently asks for a stronger one.
129export function decide(picked: Picked, cache: Cache | null, now: number, cfg: Config, sticky: Sticky): Picked {
130  if (sticky === 'off' || !isWarm(cache, now, cfg) || cache.model === picked.model) return picked
131  const from = aliasOf(cache.model)
132  if (!from) return picked
133  const upgrade = RANK.indexOf(picked.alias) > RANK.indexOf(from)
134  if (sticky === 'auto' && upgrade && picked.confidence >= cfg.minConfidence) return picked
135  return { ...picked, alias: from, model: cache.model, kept: picked.alias }
136}
137
138const bar = (n: number, on: string, off: string, len: number) =>
139  on.repeat(Math.round(n * len)) + off.repeat(len - Math.round(n * len))
140const gauge = (c: number) => bar(c, '▮', '▯', 5)
141const tier = (alias: string) => `${alias} ${'▂▄▆█'.slice(0, RANK.indexOf(alias) + 1).padEnd(4, '_')}`
142const effortBar = (e: Effort) => bar((EFFORTS.indexOf(e) + 1) / EFFORTS.length, '▰', '▱', 5)
143
144const TIER_COLOR: Record<string, string> = { haiku: '#4ade80', sonnet: '#60a5fa', opus: '#c084fc', fable: '#fbbf24' }
145const confColor = (c: number) => (c >= 0.7 ? '#4ade80' : c >= 0.5 ? '#fbbf24' : '#f87171')
146
147type Seg = { text: string; color?: string; bold?: boolean }
148
149const SPIN = '⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏'
150const GLOW = ['#c084fc', '#a78bfa', '#818cf8', '#60a5fa', '#38bdf8', '#60a5fa', '#818cf8', '#a78bfa']
151
152// frame is null while idle: static. While a task runs the mark spins and the label glows.
153export const segments = (mode: Mode, last: Decision | null, frame: number | null = null, cache: string | null = null, paused = false): Seg[] => {
154  const head: Seg[] = [
155    frame === null
156      ? { text: '◆ JEV ', color: '#c084fc', bold: true }
157      : { text: `${SPIN[frame % SPIN.length]} JEV `, color: GLOW[frame % GLOW.length], bold: true }, { text: `▏${mode}▕  `, color: '#94a3b8' },
158  ]
159  const warn: Seg[] = paused ? [{ text: '⚠ Jev paused  ', color: '#fbbf24', bold: true }] : []
160  if (!last) return [...head, ...warn, { text: 'waiting for the first task', color: '#64748b' }]
161  const model: Seg = last.kept
162    ? { text: `${last.alias} (kept 🔒 cache warm; wanted ${last.kept})`, color: '#fbbf24' }
163    : { text: `${tier(last.alias)}${last.pinned ? ' 📌' : ''}${last.escalated ? ' ↑' : ''}`, color: TIER_COLOR[last.alias], bold: true }
164  const sep: Seg = { text: '  │  ', color: '#475569' }
165  return [
166    ...head,
167    ...warn,
168    model,
169    sep,
170    { text: `${last.effort} ${effortBar(last.effort)}${last.capped ? ` ⤓${last.capped}` : ''}${last.unlocked ? ' 🔓' : ''}`, color: last.unlocked ? '#fbbf24' : '#38bdf8' },
171    sep,
172    { text: `conf ${gauge(last.confidence)} ${Math.round(last.confidence * 100)}%`, color: confColor(last.confidence) },
173    ...(cache ? [sep, { text: cache, color: cache.includes('warm') ? '#4ade80' : '#64748b' }] : []),
174  ]
175}
176
177// One row per model: steps that ran on it, with a share bar.
178const usageRows = (stats: Record<string, number>, pool: string[]): [string, string][] => {
179  const total = pool.reduce((n, m) => n + (stats[m] ?? 0), 0)
180  if (!total) return [['usage', 'no steps yet']]
181  return pool.map((m, i): [string, string] => [
182    i === 0 ? 'usage' : '',
183    `${m.padEnd(7)} ${bar((stats[m] ?? 0) / total, '█', '░', 10)} ${String(stats[m] ?? 0).padStart(3)}  ${Math.round(((stats[m] ?? 0) / total) * 100)}%`,
184  ])
185}
186
187function box(title: string, rows: [string, string][]) {
188  const body = rows.map(([k, v]) => `${k.padEnd(8)} ${v}`)
189  const w = Math.max(title.length + 3, ...body.map(l => l.length)) + 1
190  return [
191    `┌ ${title} ${'─'.repeat(w - title.length - 1)}┐`,
192    ...body.map(l => `│ ${l.padEnd(w)}│`),
193    `└${'─'.repeat(w + 1)}┘`,
194  ].join('\n')
195}
196
197async function cacheLabel($: EngineInterface, cfg: Config) {
198  const cache = await read($, cacheAtom)
199  if (!cache) return 'cache none yet'
200  const k = `${Math.round(cache.contextTokens / 1000)}k`
201  return isWarm(cache, await $.clock.now(), cfg) ? `cache 🔒 warm ${k}` : `cache cold ${k}`
202}
203
204async function currentMode($: EngineInterface, cfg: Config) {
205  return (await read($, modeAtom)) ?? cfg.defaultMode
206}
207
208async function buildState($: EngineInterface, text: string, mode: Mode, cache: Cache | null, now: number, cfg: Config) {
209  const current = await $.session.model()
210  const history = (await $.session.messages())
211    .slice(-8)
212    .map(m => `${m.role}: ${m.text.slice(0, 700)}`)
213    .join('\n')
214    .slice(-6000)
215
216  return composeState(text, mode, current, history, cache, now, cfg)
217}
218
219// Secrets that must not leave the machine inside a prompt or the conversation sent to Jev.
220const SECRETS: RegExp[] = [
221  /-----BEGIN [A-Z ]*PRIVATE KEY-----[\s\S]*?-----END [A-Z ]*PRIVATE KEY-----/g,
222  /\b(?:sk|pk|rk)-[A-Za-z0-9_-]{16,}/g,
223  /\bgh[pousr]_[A-Za-z0-9]{20,}/g,
224  /\b(?:AKIA|ASIA)[A-Z0-9]{16}\b/g,
225  /\bBearer\s+[A-Za-z0-9._~+/=-]{16,}/gi,
226]
227// NAME=value / "name": "value" for names that read like a secret, and user:password@ in URLs.
228const SECRET_NAME = /\b([\w.-]*(?:password|passwd|pwd|secret|token|api[_-]?key|(?:private|access|signing|encryption|deploy|ssh)[_-]?key)[\w.-]*)(["']?\s*[=:]\s*["']?)(?!(?:string|number|boolean|null|undefined|true|false|any|object)\b)[^\s"',;}]{3,}/gi
229const URL_CREDENTIALS = /\b([a-z][a-z0-9+.-]*:\/\/[^\s:/@]+):[^\s@/]+@/gi
230export const redact = (text: string) =>
231  SECRETS.reduce((t, re) => t.replace(re, '[redacted]'), text).replace(SECRET_NAME, '$1$2[redacted]').replace(URL_CREDENTIALS, '$1:[redacted]@')
232
233// Used only when Jev is down and `fallback` is 'heuristic'. ponytail: three buckets from length and a few words; real routing needs Jev.
234export function heuristic(text: string, pool: readonly string[]): Picked {
235  const hard = text.length > 1500 || /architect|design|migrat|root cause|concurren|race|deadlock|security|across .* (modules|services)/i.test(text)
236  const easy = text.length < 200 && /rename|typo|bump|format|lint|comment|where is|what does|show me/i.test(text)
237  const want = hard ? 'opus' : easy ? 'haiku' : 'sonnet'
238  const alias = RANK.slice(RANK.indexOf(want)).find(a => pool.includes(a)) ?? RANK.filter(a => pool.includes(a)).pop()!
239
240  return { alias, model: CLAUDE[alias]!.id, effort: hard ? 'high' : easy ? 'low' : 'medium', confidence: 0.3 }
241}
242
243// Pure, so the benchmark sends Jev exactly what the mod sends.
244export function composeState(text: string, mode: Mode, current: string, history: string, cache: Cache | null, now: number, cfg: Config) {
245  const clean = (t: string) => (cfg.redact ? redact(t) : t)
246
247  return [
248    `Routing mode: ${mode}. ${MODE_RULE[mode]}`,
249    `Current model: ${current}. Switching models discards the prompt cache, so prefer staying on it when the gain from switching is small.`,
250    cache
251      ? `Prompt cache: ${isWarm(cache, now, cfg) ? 'warm' : 'cold'}, ~${cache.contextTokens} tokens on ${cache.model}. Switching re-writes them.`
252      : '',
253    history && cfg.sendHistory ? `Recent conversation:\n${clean(history)}` : 'Recent conversation: (none, this is the first task)',
254    `New task from the user:\n${clean(text.slice(0, cfg.maxTaskChars))}`,
255  ]
256    .filter(Boolean)
257    .join('\n\n')
258}
259
260// Sensitive plugin options get no /config row, so the key also comes from
261// the plugin's store (`/jev key <key>`) or the TYPESAFE_API_KEY variable.
262async function resolveKey($: EngineInterface, cfg: Config): Promise<{ key: string; source: string } | null> {
263  if (cfg.apiKey) return { key: cfg.apiKey, source: 'plugin option' }
264  const stored = await $.store.get('apiKey')
265  if (typeof stored === 'string' && stored) return { key: stored, source: '/jev key' }
266  const env = await $.env.get('TYPESAFE_API_KEY')
267  if (env) return { key: env, source: 'TYPESAFE_API_KEY' }
268  return null
269}
270
271export function jevBody(pool: string[], state: string, mode: Mode) {
272  return {
273    model: 'jev-latest',
274    state,
275    questions: {
276      model: {
277        type: 'choice',
278        instructions: `Which Claude model should handle the new coding task? ${MODE_RULE[mode]}`,
279        criteria: Object.fromEntries(pool.map(m => [m, CLAUDE[m]!.about])),
280      },
281      effort: {
282        type: 'choice',
283        instructions: `How much reasoning effort does the new task need? ${MODE_RULE[mode]}`,
284        criteria: Object.fromEntries(EFFORTS.map(x => [x, EFFORT_ABOUT[x]])),
285      },
286    },
287  }
288}
289
290async function askJev($: EngineInterface, cfg: Config, key: string, state: string, mode: Mode) {
291  const body = jevBody(cfg.pool, state, mode)
292  const call = $.http.fetch(JEV_URL, {
293    method: 'POST',
294    headers: { 'content-type': 'application/json', authorization: `Bearer ${key}` },
295    body: JSON.stringify(body),
296  })
297  const timeout = $.clock.sleep(cfg.timeoutMs).then(() => {
298    throw new Error(`timed out after ${cfg.timeoutMs} ms`)
299  })
300  const res = await Promise.race([call, timeout])
301  if (!res.ok) throw new Error(`HTTP ${res.status}`)
302  const reply = JSON.parse(res.text) as JevReply
303  const alias = reply.answers?.model?.choice
304  const effort = reply.answers?.effort?.choice as Effort | undefined
305  if (!alias || !cfg.pool.includes(alias)) throw new Error(`unexpected model choice: ${alias}`)
306
307  return {
308    alias,
309    model: CLAUDE[alias]!.id,
310    effort: effort && EFFORTS.includes(effort) ? effort : 'medium',
311    confidence: reply.answers?.model?.confidence ?? 0,
312  }
313}
314
315// Circuit breaker: with a call per step, a slow Jev would cost timeoutMs on every step.
316// Three failures in a row pause Jev for pauseMs; the last pick keeps running meanwhile.
317let fails = 0
318let pausedUntil = 0
319export const isPaused = (now: number) => now < pausedUntil
320
321async function ask($: EngineInterface, cfg: Config, key: string, state: string, mode: Mode) {
322  const now = await $.clock.now()
323  if (isPaused(now)) throw new Error(`paused for ${Math.ceil((pausedUntil - now) / 1000)} s after repeated failures`)
324  try {
325    const picked = await askJev($, cfg, key, state, mode)
326    fails = 0
327
328    return picked
329  } catch (err) {
330    if (++fails >= 3) {
331      fails = 0
332      pausedUntil = now + cfg.pauseMs
333      $.ui.toast(`Jev failed 3 times in a row (${(err as Error).message}); paused for ${cfg.pauseMs / 1000} s, using the current model`)
334    }
335    throw err
336  }
337}
338
339let strikes = 0 // consecutive failed tool calls of the current task
340let skipNext = false // the next turn is a background notification, not a task of the user's
341let task = '' // the prompt the current task started with, for re-routing its later steps
342const agents = new Map<string, Promise<Picked | null>>()
343
344async function decideFor($: EngineInterface, cfg: Config, kind: string, key: string, text: string, state: string, mode: Mode, cache: Cache | null, now: number, cap: EffortCap) {
345  let asked: Picked
346  try {
347    asked = await ask($, cfg, key, state, mode)
348  } catch (err) {
349    if (cfg.fallback !== 'heuristic') throw err
350    asked = heuristic(text, cfg.pool)
351    $.ui.log(`jev ${kind}: Jev unavailable (${(err as Error).message}), heuristic pick ${asked.alias}`, { to: 'debug' })
352  }
353  const picked = decide(asked, cache, now, cfg, (await read($, stickyAtom)) ?? cfg.defaultSticky)
354  const d = { ...picked, ...capEffort(picked.effort, cap) }
355  $.ui.log(`jev ${kind}: ${d.alias}/${d.effort}${d.kept ? ` (kept, wanted ${d.kept})` : ''}${d.capped ? ` (capped from ${d.capped})` : ''} conf ${d.confidence.toFixed(2)} ${(await $.clock.now()) - now} ms`, { to: 'debug' })
356
357  return d
358}
359
360// Before every step after the first, ask again: the task may have turned easier or harder.
361// Failures keep the last pick quietly; the task's own first call already warned.
362async function restep($: EngineInterface, cfg: Config, e: { turnId: string; index: number }, last: Decision): Promise<Decision> {
363  const auth = await resolveKey($, cfg)
364  if (!auth || last.pinned) return last
365  try {
366    const mode = await currentMode($, cfg)
367    const now = await $.clock.now()
368    const cache = await read($, cacheAtom)
369    // strict never switches while the cache is warm, so asking could not change the pick
370    if (((await read($, stickyAtom)) ?? cfg.defaultSticky) === 'strict' && isWarm(cache, now, cfg)) return last
371    const text = `${task}\n\n[Routing step ${e.index + 1} of this task. The recent messages above show its progress.]`
372    const cap = last.unlocked ? 'none' : (await read($, capAtom)) ?? cfg.defaultCap
373    const d = await decideFor($, cfg, `step ${e.index + 1}`, auth.key, text, await buildState($, text, mode, cache, now, cfg), mode, cache, now, cap)
374    if (last.escalated && RANK.indexOf(d.alias) <= RANK.indexOf(last.alias)) return last // never de-escalate within a task
375    const next: Decision = { ...d, turnId: last.turnId, unlocked: last.unlocked }
376    await update($, lastAtom, () => next)
377
378    return next
379  } catch {
380    return last
381  }
382}
383
384// Routes one subagent from the text of its task. Called at spawn (the full prompt, so the pick is known before the
385// subagent exists and the subagent list can show it) and, for agents no spawn hook saw, from their description.
386async function routeSubagent($: EngineInterface, cfg: Config, kind: string, text: string, parentModel: string): Promise<Picked | null> {
387  const auth = await resolveKey($, cfg)
388  if (!auth) return null
389  try {
390    const now = await $.clock.now()
391    const mode = await currentMode($, cfg)
392    const cap = (await read($, lastAtom))?.unlocked ? 'none' : (await read($, capAtom)) ?? cfg.defaultCap
393
394    return await decideFor($, cfg, kind, auth.key, text, composeState(text, mode, parentModel, '', null, now, cfg), mode, null, now, cap)
395  } catch {
396    return null
397  }
398}
399
400async function routeAgent($: EngineInterface, cfg: Config, id: string, parentModel: string): Promise<Picked | null> {
401  const info = (await $.agent.list()).find(a => a.id === id)
402
403  return info?.description ? routeSubagent($, cfg, `agent ${info.type}: ${info.description}`, `Subagent task (${info.type}): ${info.description}`, parentModel) : null
404}
405
406// The running totals, loaded from the plugin store once (in session.start, so parallel steps never race on it).
407const emptyTally = (name: string): Savings => ({ actual: 0, baseline: 0, steps: 0, name, models: {} })
408async function tally($: EngineInterface, cfg: Config): Promise<Savings> {
409  const stored = (await $.store.get('savings')) as Savings | undefined
410  // a new baseline starts a new tally
411  return stored && stored.name === cfg.baseline ? { ...emptyTally(cfg.baseline), ...stored } : emptyTally(cfg.baseline)
412}
413
414// Add one step to the running totals (kept across sessions in the plugin store).
415async function record($: EngineInterface, cfg: Config, u: (StepUsage & { model: string }) | undefined) {
416  const alias = u && aliasOf(u.model)
417  if (!u || !alias) return
418  totals ??= await tally($, cfg)
419  totals.models[alias] = (totals.models[alias] ?? 0) + 1
420  totals.actual += stepUnits(u, alias)
421  totals.baseline += stepUnits(u, cfg.baseline)
422  totals.steps++
423  await $.store.set('savings', totals)
424}
425
426const usage = (cmd: string, list: readonly string[]) => ({ text: `Usage: /jev ${cmd} ${list.join(' | ')}` })
427const applied = (name: string, v: string) => ({ text: `Jev ${name}: ${v}. Applies from the next task.` })
428
429export const register: Register = (on, options) => {
430  const apiKey = String(options.apiKey ?? '')
431  const timeoutMs = Number(options.timeoutMs ?? 4000)
432  const allowed = (Array.isArray(options.models) ? options.models : Object.keys(CLAUDE))
433    .map(m => String(m).toLowerCase())
434    .filter(m => m in CLAUDE)
435  const defaultMode: Mode = MODES.includes(options.mode as Mode) ? (options.mode as Mode) : 'balanced'
436
437  const defaultSticky: Sticky = STICKY.includes(options.stickiness as Sticky) ? (options.stickiness as Sticky) : 'auto'
438  const defaultCap: EffortCap = CAPS.includes(options.effortCap as EffortCap) ? (options.effortCap as EffortCap) : 'high'
439  const flag = (v: unknown, d: boolean) => (v === undefined || v === '' ? d : String(v) !== 'false')
440  const num = (v: unknown, d: number) => (v !== undefined && v !== '' && Number.isFinite(Number(v)) ? Number(v) : d)
441
442  const cfg: Config = {
443    apiKey,
444    timeoutMs,
445    pool: allowed.length > 0 ? allowed : Object.keys(CLAUDE),
446    defaultMode,
447    defaultSticky,
448    defaultCap,
449    pauseMs: num(options.pauseMs, 60_000),
450    routeSteps: flag(options.routeSteps, true),
451    routeSubagents: flag(options.routeSubagents, true),
452    escalateAfter: num(options.escalateAfter, 3),
453    sendHistory: flag(options.sendHistory, true),
454    redact: flag(options.redact, true),
455    maxTaskChars: num(options.maxTaskChars, 8000),
456    fallback: options.fallback === 'heuristic' ? 'heuristic' : 'session',
457    baseline: String(options.baselineModel ?? 'opus') in CLAUDE ? String(options.baselineModel ?? 'opus') : 'opus',
458    minConfidence: num(options.minConfidence, 0.7),
459    minContextTokens: num(options.minContextTokens, 8000),
460    cacheTtlMs: num(options.cacheTtlMs, 300_000),
461  }
462
463  // The frame ticks only while a task runs, so an idle session redraws nothing.
464  let ticker: { cancel: () => void } | undefined
465  const stop = () => {
466    ticker?.cancel()
467    ticker = undefined
468  }
469
470  on('prompt.submit', async ($, e, next) => {
471    stop()
472    let ticks = 0 // self-cancel: a prompt that never reaches turn.complete must not tick forever
473    const t = $.clock.every(120, () => (++ticks > 5000 ? t.cancel() : update($, frameAtom, n => n + 1)))
474    ticker = t
475    skipNext = e.origin?.kind === 'task-notification'
476    const m = markers(e.text, cfg.pool)
477    if (!m.pending || m.text.trim() === '') return next(e)
478    await update($, pendingAtom, () => m.pending)
479
480    return next({ ...e, text: m.text })
481  })
482
483  on('turn.complete', async ($, e, next) => {
484    stop()
485    return next(e)
486  })
487
488  on('command.run', { command: 'jev' }, async ($, e, next) => {
489    stop() // a /jev command is not a task: no spinner
490    return next(e)
491  })
492
493  // Pick the subagent's model at spawn and show it in the subagent list.
494  on('agent.spawn', async ($, e, next) => {
495    if (!cfg.routeSubagents || !(await read($, enabledAtom)) || e.model) return next(e) // an explicit model on the Agent call wins
496    const pick = await routeSubagent($, cfg, `spawn ${e.subagentType}: ${e.description}`, `Subagent task (${e.subagentType}): ${e.description}\n\n${e.prompt}`, e.parentModel)
497    if (!pick) return next(e)
498    const res = await next({ ...e, model: pick.alias, description: `${e.description} · ${pick.alias}/${pick.effort}` })
499    if (res.agentId) agents.set(res.agentId, Promise.resolve(pick)) // its steps reuse the pick (and its effort)
500
501    return res
502  })
503
504  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
505    if (e.props.hasSurvey) return next(e)
506    const { Box, Text } = $.ui.resolve(e)
507    const segs: Seg[] = (await read($, enabledAtom))
508      ? segments(await currentMode($, cfg), await read($, lastAtom), e.props.isWorking ? await read($, frameAtom) : null, await cacheLabel($, cfg), isPaused(await $.clock.now()))
509      : [{ text: 'JEV ▏off▕', color: '#64748b' }]
510
511    return (
512      <Box>
513        <Text>
514          {segs.map(x => (
515            <Text color={x.color} bold={x.bold}>
516              {x.text}
517            </Text>
518          ))}
519        </Text>
520      </Box>
521    )
522  })
523
524  on('session.start', async ($, e, next) => {
525    agents.clear()
526    totals = await tally($, cfg)
527    await $.command.register({
528      name: 'jev',
529      description: 'Jev model routing (prompt markers: !full !haiku !sonnet !opus !fable !cheap !efficient): efficient | balanced | cheap | sticky <off|auto|strict> | cap <effort|none> | full | on | off | status | key <key>',
530    })
531
532    return next(e)
533  })
534
535  on('turn.start', async ($, e, next) => {
536    // A continuation, or a background notification, keeps the decision of the task it continues.
537    if (!(await read($, enabledAtom)) || e.text.trim() === '' || skipNext) return next(e)
538    task = e.text
539    strikes = 0
540    const pending = await read($, pendingAtom)
541    if (pending) await update($, pendingAtom, () => null) // one prompt only, even if nothing can be routed
542    const auth = await resolveKey($, cfg)
543    if (!auth && !pending?.pin) { // a pinned model needs no Jev call, so no key either
544      await update($, lastAtom, () => null) // never run this prompt on the previous task's pick
545      if (!(await read($, warnedAtom))) {
546        await update($, warnedAtom, () => true)
547        $.ui.toast('Jev: no TypeSafe API key. Run /jev key <key> or set TYPESAFE_API_KEY')
548      }
549      return next(e)
550    }
551
552    const unlocked = pending?.full === true
553    const mode = pending?.mode ?? (await currentMode($, cfg))
554    try {
555      const cache = await read($, cacheAtom)
556      const now = await $.clock.now()
557      const cap = unlocked ? 'none' : (await read($, capAtom)) ?? cfg.defaultCap
558      // A pinned model skips Jev; its effort is the cap (high by default).
559      const picked = pending?.pin
560        ? { alias: pending.pin, model: CLAUDE[pending.pin]!.id, effort: cap === 'none' ? 'high' : cap, confidence: 1, pinned: true }
561        : await decideFor($, cfg, 'prompt', auth!.key, e.text, await buildState($, e.text, mode, cache, now, cfg), mode, cache, now, cap)
562      const decision: Decision = { turnId: e.turnId, ...picked, unlocked }
563      await update($, lastAtom, () => decision)
564    } catch (err) {
565      await update($, lastAtom, () => null)
566      if (!(await read($, warnedAtom))) {
567        await update($, warnedAtom, () => true)
568        $.ui.toast(`Jev unavailable (${(err as Error).message}), using ${await $.session.model()}`)
569      }
570    }
571
572    return next(e)
573  })
574
575  // A task that keeps failing moves up one model for the rest of the task, whatever the cache guard says.
576  on('tool.call', async ($, e, next) => {
577    const res = await next(e)
578    if (e.agentId !== undefined || !cfg.escalateAfter) return res
579    strikes = 'isError' in res && res.isError ? strikes + 1 : 0
580    const last = await read($, lastAtom)
581    if (strikes >= cfg.escalateAfter && last && !last.escalated && !last.pinned) {
582      const to = RANK[Math.min(RANK.indexOf(last.alias) + 1, RANK.length - 1)]!
583      if (cfg.pool.includes(to) && to !== last.alias) {
584        await update($, lastAtom, () => ({ ...last, alias: to, model: CLAUDE[to]!.id, escalated: true }))
585        $.ui.log(`jev escalate: ${last.alias} -> ${to} after ${strikes} failed tool calls`, { to: 'debug' })
586      }
587    }
588
589    return res
590  })
591
592  on('turn.step', async function* ($, e, next) {
593    if (!(await read($, enabledAtom))) return yield* next(e)
594    if (e.agentId !== undefined) {
595      if (!cfg.routeSubagents) return yield* next(e)
596      const id = e.agentId
597      const pick = await (agents.get(id) ?? agents.set(id, routeAgent($, cfg, id, e.model)).get(id)!)
598
599      const res = yield* next(pick ? { ...e, model: pick.model, effort: pick.effort } : e)
600      await record($, cfg, res.usage)
601
602      return res
603    }
604    const prev = await read($, lastAtom)
605    const last = cfg.routeSteps && e.index > 0 && prev ? await restep($, cfg, e, prev) : prev
606
607    // A model that takes no effort setting is sent none, whatever is asked.
608    const res = yield* next(last ? { ...e, model: last.model, effort: last.effort } : e)
609
610    // Remember what the API actually cached, so the next turn can protect it.
611    const u = res.usage
612    if (u) {
613      const contextTokens = u.input_tokens + u.cache_read_input_tokens + u.cache_creation_input_tokens + u.output_tokens
614      const at = await $.clock.now()
615      await update($, cacheAtom, () => ({ model: u.model, at, contextTokens }))
616    }
617    await record($, cfg, res.usage)
618    return res
619  })
620
621  on('command.run', { command: 'jev' }, async ($, e) => {
622    const arg = e.args.trim().toLowerCase()
623
624    if (MODES.includes(arg as Mode)) {
625      await update($, modeAtom, () => arg as Mode)
626      await update($, enabledAtom, () => true)
627      return { text: `Jev routing mode: ${arg}. Applies from the next task.` }
628    }
629    if (arg.startsWith('sticky')) {
630      const v = arg.slice(6).trim()
631      if (!STICKY.includes(v as Sticky)) return usage('sticky', STICKY)
632      await update($, stickyAtom, () => v as Sticky)
633      return applied('cache stickiness', v)
634    }
635    if (arg === 'stats clear') {
636      totals = null
637      await $.store.delete('savings')
638      return { text: 'Jev savings tally cleared.' }
639    }
640    if (arg === 'full') {
641      await update($, pendingAtom, p => ({ ...p, full: true }))
642      return { text: 'Jev: the next prompt runs with the effort cap lifted.' }
643    }
644    if (arg.startsWith('cap')) {
645      const v = arg.slice(3).trim()
646      if (!CAPS.includes(v as EffortCap)) return usage('cap', CAPS)
647      await update($, capAtom, () => v as EffortCap)
648      return applied('effort cap', v)
649    }
650    if (arg === 'on' || arg === 'off') {
651      await update($, enabledAtom, () => arg === 'on')
652      if (arg === 'off') await update($, lastAtom, () => null)
653      return { text: `Jev routing ${arg === 'on' ? 'enabled' : 'disabled; the session model is used'}.` }
654    }
655    if (arg === 'key' || arg.startsWith('key ')) {
656      const key = e.args.trim().slice(3).trim()
657      if (key === '') return { text: 'Usage: /jev key <TypeSafe API key>   or   /jev key clear' }
658      if (key === 'clear') {
659        await $.store.delete('apiKey')
660        return { text: 'Stored TypeSafe API key removed.' }
661      }
662      await $.store.set('apiKey', key)
663      await update($, warnedAtom, () => false)
664      return { text: `TypeSafe API key saved (…${key.slice(-4)}). Applies from the next task.` }
665    }
666    if (arg === '' || arg === 'status') {
667      const auth = await resolveKey($, cfg)
668      const enabled = await read($, enabledAtom)
669      const last = await read($, lastAtom)
670      const cache = await read($, cacheAtom)
671      const warm = isWarm(cache, await $.clock.now(), cfg)
672      return {
673        text: '\n' + box('Jev router', [
674          ['state', `${enabled ? '● on ' : '○ off'}   mode  ${await currentMode($, cfg)}`],
675          ['pool', cfg.pool.join(' · ')],
676          ['sticky', `${(await read($, stickyAtom)) ?? cfg.defaultSticky}   cache ${warm ? '● warm' : '○ cold'}`],
677          ['cap', `${(await read($, capAtom)) ?? cfg.defaultCap}${(await read($, pendingAtom)) ? '   🔓 next prompt: overrides set' : ''}`],
678          ['last', last ? `${last.alias} / ${last.effort}  ${gauge(last.confidence)} ${last.confidence.toFixed(2)}${last.kept ? `  kept (wanted ${last.kept})` : ''}${last.capped ? `  capped from ${last.capped}` : ''}` : 'none yet'],
679          ...usageRows(totals?.models ?? {}, cfg.pool),
680          ['saved', savingsLine(totals ?? (await tally($, cfg)))],
681          ['jev', isPaused(await $.clock.now()) ? `⚠ paused ${Math.ceil((pausedUntil - (await $.clock.now())) / 1000)} s` : '● ok'],
682          ['key', auth ? `…${auth.key.slice(-4)}  (${auth.source})` : 'none: /jev key <key>'],
683        ]),
684      }
685    }
686    return { text: `Unknown option "${arg}". Use: /jev efficient | balanced | cheap | sticky <mode> | cap <effort|none> | full | on | off | status | key <key>` }
687  })
688}
689
types/index.d.ts 36 lines
1export type Mode = 'efficient' | 'balanced' | 'cheap'
2export type Sticky = 'off' | 'auto' | 'strict'
3export type Cache = { model: string; at: number; contextTokens: number }
4export type Effort = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
5export type EffortCap = Effort | 'none'
6export type Savings = { actual: number; baseline: number; steps: number; name: string; models: Record<string, number> }
7export type Pending = { full?: true; pin?: string; mode?: Mode }
8export type Decision = {
9  turnId: string
10  model: string
11  alias: string
12  effort: Effort
13  confidence: number
14  kept?: string
15  capped?: Effort
16  pinned?: boolean
17  escalated?: boolean
18  unlocked?: boolean
19}
20
21declare module 'claude-code' {
22  interface PluginState {
23    'jev-router': {
24      mode: Mode | null
25      enabled: boolean
26      last: Decision | null
27      cache: Cache | null
28      sticky: Sticky | null
29      cap: EffortCap | null
30      pending: Pending | null
31      frame: number
32      warned: boolean
33    }
34  }
35}
36