Asks TypeSafe Jev before each task which Claude model and effort fit best, then runs the turn on it

<img src="assets/logo.jpg" width="200" alt="Jev Router, the smug switchman raccoon">
<h1 align="center">jev-router: automatic Claude model routing for Claude Code</h1>
<em>Switch between Opus, Sonnet and Haiku automatically. He pulls one lever; your task rides the right track.</em>
<img src="https://img.shields.io/github/v/release/dominicrico/jev-router?style=flat-square&color=111111&label=release&include_prereleases" alt="Release"> <img src="https://img.shields.io/badge/Claude%20Code-mod-111111?style=flat-square" alt="Claude Code mod"> <img src="https://img.shields.io/badge/models-haiku%20%C2%B7%20sonnet%20%C2%B7%20opus%20%C2%B7%20fable-111111?style=flat-square" alt="Models"> <img src="https://img.shields.io/badge/license-MIT-111111?style=flat-square" alt="MIT license">
<img src="assets/hero.jpg" width="880" alt="jev-router: right model, right effort, every task. Tasks routed to haiku, sonnet or opus, minus 97% cost on trivial tasks, minus 65% on standard, 253 ms per pick. And yet hard tasks cost more.">
<a href="https://dominicrico.github.io/jev-router/">Website</a> · <a href="#install-the-claude-code-plugin">Install</a> · <a href="#does-it-reduce-claude-code-costs">Benchmarks</a> · <a href="#faq">FAQ</a> · <a href="CHANGELOG.md">Changelog</a>
jev-router is a free, open-source Claude Code plugin that routes every prompt, step and subagent to the right Claude model (Haiku, Sonnet, Opus, Fable) and the right reasoning effort. It asks TypeSafe Jev which model a task needs, caps the effort so hard tasks do not burn tokens, protects your prompt cache from needless model switches, and shows the choice in a live status band above the prompt. Use it to cut Claude Code cost on easy work without giving up Opus when a task needs it.
You know the problem. You rename one variable and your most expensive model thinks about it for a minute. Or you start a gnarly refactor and a bargain model happily shrugs. jev-router puts a switchman in front of your session: before every task he reads it, asks Jev, pulls the lever, and the turn runs on the model and effort that fit.
Contents: Status band · Features · Does it reduce costs? · How it works · Install · Commands · Options · Privacy · FAQ
A live Claude Code status band above your prompt shows which model and effort were picked. Colour coded, with a spinner while a task runs.
<img src="assets/band.gif" width="880" alt="The band in all its states, one slide at a time: idle, running, haiku, sonnet, capped opus, !full unlock, cache kept, pinned, escalated, low confidence, cheap mode, paused, off.">
<img src="assets/band.png" width="880" alt="The band above the prompt: JEV, balanced mode, sonnet, medium effort, 88% confidence, warm cache of 42k tokens. Four colour-coded tiers below: haiku trivial, sonnet standard, opus hard, fable hardest.">
| Part | Meaning |
|---|---|
balanced | Routing mode: efficient, balanced or cheap |
sonnet ▂▄__ | The model the task runs on, with its tier. Green haiku, blue sonnet, purple opus, gold fable |
medium ▰▰▱▱▱ | Reasoning effort, low to max |
conf ▮▮▮▮▯ 88% | How sure Jev is. Green from 70%, amber from 50%, red below |
cache 🔒 warm 42k | Whether the prompt cache is still worth protecting, and how big it is |
When the cache wins, the band says so: opus (kept 🔒 cache warm; wanted haiku).
efficient for the best result, balanced for quality and cost evenly, cheap for the cheapest model that can plausibly succeed.high by default, because effort drives token use far more than the model does. Start a prompt with !full (or run /jev full) to lift the cap for that one prompt.escalateAfter (default 3) failed tool calls in a row, the rest of the task moves up one model. The band shows ↑.!opus, !sonnet, !haiku or !fable to skip Jev and use that model (at the effort cap). !cheap and !efficient set the routing mode for that prompt. Markers combine, like !opus !full./jev status shows an estimated saving against the model you would otherwise use (baselineModel, default opus), kept across sessions. It is a model-price estimate in relative units, blind to effort.fallback: heuristic guesses locally meanwhile. Jev slow or down? The session model keeps working.routeSteps.find importers · haiku/medium). Its steps then run on that model and effort. An explicit model on the Agent call wins. Turn it off with routeSubagents./jev status shows how often each model was used this session.Three benchmarks. First, 30 single-prompt tasks: 270 calls to Jev for the picks, then 147 real runs for token usage and cost, with the effort cap at its default (high) and uncapped. Second, 6 multi-step tasks with tools on a fixture repo, 72 real runs, checking per prompt against every-step-and-subagent routing. Third, 4 harder tasks graded by hidden tests, 100 real runs, where quality is measured.
<img src="assets/fit.png" width="880" alt="Routing fit: efficient 100%, balanced 87%, cheap 66% of picks fit the task. 253 ms added per task, 97 to 100% same pick on repeat.">
<img src="assets/cost.png" width="880" alt="Cost as a percentage of always opus with the effort capped at high. Trivial tasks 3%, standard 32 to 35%, hard 143%, 124% and 44% for efficient, balanced and cheap. Overall 99%, 87% and 36%. Uncapped hard tasks cost 342%, 307% and 61%.">
| Cost vs always opus (no plugin) | efficient | balanced | cheap |
|---|---|---|---|
| Trivial tasks | -97% | -97% | -97% |
| Standard tasks | -65% | -65% | -68% |
| Hard tasks, cap high (default) | +43% | +24% | -56% |
| Hard tasks, uncapped | +242% | +207% | -39% |
| All 30, cap high (default) | -1% | -13% | -64% |
| All 30, uncapped | +123% | +101% | -53% |
The honest read: the model switch saves a lot on easy and mid tasks, and the default cap of high brings every mode to or below always-opus overall. Hard tasks still cost 24% to 43% more in efficient and balanced, because Jev picks opus there and asks for deeper reasoning. !full lifts the cap for one prompt when you want that. The cap changes tokens spent, not answer quality, and that is not measured.
<img src="assets/cache.png" width="880" alt="Cache guard: following every pick in a warm session costs 4735 against 1500 for always opus. With the guard it is 1500. With cold cache and tasks 400 seconds apart it is 820.">
<img src="assets/agentic.png" width="880" alt="Multi-step tasks with tools, cost as a percentage of always opus: always sonnet 52%, jev-router per prompt 52%, jev-router every step and subagents 47%. All strategies finished 18 of 18 tasks.">
| 6 tasks × 3 runs, real runs with tools | tasks done | cost | vs always opus | vs always sonnet |
|---|---|---|---|---|
| no plugin: always opus | 18/18 | $2.46 | ||
| no plugin: always sonnet | 18/18 | $1.29 | -48% | |
| jev-router, per prompt | 18/18 | $1.27 | -48% | -2% |
| jev-router, every step + subagents | 18/18 | $1.15 | -53% | -11% |
Routing every step and subagent saved 11% on top of per-prompt routing. Most of that comes from one task: asked to use a subagent, it paired a sonnet main thread with a haiku subagent where per-prompt routing sometimes left the subagent on opus. On the other five tasks it matched per-prompt routing. Honest limits: these are easy tasks on a small repo, three runs each, and every strategy finished all of them, so this shows nothing was lost here, not that nothing is lost on hard work. Jev picked sonnet for nearly every one of these tasks, so against always-sonnet the gain is small.
<img src="assets/hard.png" width="880" alt="Harder tasks graded by hidden tests, cost as a percentage of always opus: always sonnet 39%, jev-router per prompt 86%, every step and subagents 67%, with !full 60%. Tasks done out of 20: opus 19, sonnet 19, per prompt 19, every step 18, !full 20.">
| 4 tasks × 5 runs, hidden tests | tasks done | cost | vs always opus | vs always sonnet |
|---|---|---|---|---|
| no plugin: always opus | 19/20 | $3.10 | ||
| no plugin: always sonnet | 19/20 | $1.19 | -61% | |
| jev-router, per prompt | 19/20 | $2.65 | -14% | +122% |
| jev-router, every step + subagents | 18/20 | $2.08 | -33% | +74% |
jev-router, !full | 20/20 | $1.87 | -40% | +57% |
The honest read: on these tasks plain sonnet did as well as opus, and was the cheapest. The tasks (a concurrency bug, an interval merger with open and closed bounds, a refactor that must keep an invariant, a flaky test with three causes) are harder than the first set, but not hard enough to need opus. jev-router lands between the two on cost because Jev sent about half of the spend to opus. Pass counts differ by at most one run, so none of the quality differences are significant at 5 runs. If your work looks like this, routing buys you less than just using sonnet; it earns its keep when tasks vary, or when your default is opus.
Without the cache guard, following every pick in a long warm session cost 3x more than just using opus. Inside one long warm session the guard keeps you on the first model, so the savings above show up mostly when the cache is cold or the context is small. Method, per-tier tables and all caveats: benchmarks/.
<img src="assets/how.png" width="880" alt="How it works in five steps: a task arrives, ask Jev which model and effort, check whether the prompt cache is warm, run the turn on the pick, show it above the prompt.">
Per prompt, the mod sends Jev the routing mode, the current model, the cache state, the last few messages and the task text, and asks two choice questions: which model, and how much effort. The answer is applied to every step of that task. After each step the mod notes which model the API cached and how many tokens, so the next task knows whether a switch is worth losing the cache.
Sent to api.typesafe.ai: the task text, recent conversation, and subagent descriptions. Each step and subagent adds one call (about 250 ms). Your API key travels in the request header and nowhere else.
Needs a Claude Code version with mods (hooks modules) and a TypeSafe API key.
/plugin marketplace add dominicrico/jev-router
/plugin install jev-router@jev-router
Update:
claude plugin marketplace update jev-router
claude plugin update jev-router@jev-router
Any one of these, checked in this order:
apiKey (/config)/jev key <key>, stored by the mod across sessions (/jev key clear removes it)TYPESAFE_API_KEY environment variableThat was it. He'd be proud. He won't say it.
| Command | What it does | |||||
|---|---|---|---|---|---|---|
/jev or /jev status | Settings, cache state, last decision and per-model usage | |||||
| `/jev efficient \ | balanced \ | cheap` | Set the routing mode, from the next task | |||
| `/jev sticky off \ | auto \ | strict` | Set cache stickiness | |||
| `/jev cap low \ | medium \ | high \ | xhigh \ | max \ | none` | Set the effort cap |
/jev full | Lift the effort cap for the next prompt only | |||||
/jev stats clear | Reset the savings tally | |||||
/jev on / /jev off | Enable routing / use the session model | |||||
/jev key <key> | Store the TypeSafe API key |
┌ Jev router ────────────────────────────┐
│ state ● on mode balanced │
│ pool haiku · sonnet · opus · fable │
│ sticky auto cache ● warm │
│ last sonnet / medium ▮▮▮▮▯ 0.88 │
│ usage haiku ░░░░░░░░░░ 0 0% │
│ sonnet ██████████ 1 100% │
│ opus ░░░░░░░░░░ 0 0% │
│ fable ░░░░░░░░░░ 0 0% │
│ key …fd1c (TYPESAFE_API_KEY) │
└────────────────────────────────────────┘
| Option | Default | Meaning |
|---|---|---|
mode | balanced | Routing mode |
models | all four | Aliases Jev may choose from |
effortCap | high | Highest effort the router applies. none follows Jev. Per prompt: start with !full or run /jev full |
pauseMs | 60000 | Stop asking Jev for this long after 3 failed calls in a row |
escalateAfter | 3 | Failed tool calls in a row before the task moves up one model. 0 is off |
baselineModel | opus | The model the savings estimate compares against |
sendHistory | true | Send recent messages to Jev. Off sends only the task text |
redact | true | Replace API keys, tokens, private keys and KEY=value lines with [redacted] before sending |
maxTaskChars | 8000 | How much of the prompt is sent |
fallback | session | When Jev is down: session keeps the session model, heuristic guesses locally |
routeSteps | true | Ask Jev again before every step after the first. Off: one call per prompt |
routeSubagents | true | Pick a model and effort for each subagent from its description |
stickiness | auto | off: always follow Jev. auto: while the cache is warm, only upgrade on high confidence. strict: never switch while warm |
minConfidence | 0.7 | Confidence needed for a warm-cache upgrade |
minContextTokens | 8000 | Below this a switch is free and stickiness is skipped |
cacheTtlMs | 300000 | Time after the last request when the cache counts as cold |
timeoutMs | 4000 | Keep the current model when Jev takes longer |
Sent to api.typesafe.ai: the prompt text (up to maxTaskChars), recent messages if sendHistory is on, and subagent descriptions. Before sending, redact replaces API keys, GitHub tokens, AWS keys, bearer tokens, private key blocks and NAME_KEY=value style lines with [redacted]. That is pattern matching, not a guarantee, so keep sendHistory off if your conversations carry things it would not catch. Your TypeSafe key is stored in the plugin store in plain text if you use /jev key, and the command line lands in the transcript; prefer the TYPESAFE_API_KEY environment variable.
How do I switch between Opus, Sonnet and Haiku automatically in Claude Code? Install jev-router. It picks the model per prompt, again before each step, and once per subagent. Pin a model for one prompt with !opus, !sonnet or !haiku.
How do I reduce Claude Code token cost? Send easy tasks to cheaper models and cap the reasoning effort. jev-router does both; the benchmarks show where it saves (trivial tasks -97%, standard -65% against always Opus) and where it does not (on harder tasks plain Sonnet was as good and cheaper).
What is the Claude Code effort cap? Jev often asks for xhigh reasoning on hard tasks, which can use several times the tokens. jev-router limits it to high by default; start a prompt with !full to lift it once.
Does jev-router send my code to a third party? It sends the prompt, optionally the last messages, and subagent task text to api.typesafe.ai. Keys, tokens, private keys and password= style values are redacted first, and sendHistory can turn the history off. See Privacy.
Will it slow Claude Code down? Each Jev call adds about 250 ms. After three failed calls in a row it pauses for a minute and keeps your current model.
Does it need a config file? No. Set a key, pick a mode, done. Everything else has a default.
Why not just always use the best model? You can. /jev efficient and a pool of opus and fable. Your invoice will have opinions.
Why not just always use the cheap one? /jev cheap. The hard task will have opinions.
He kept my old model even though Jev wanted a cheaper one. Is it broken? No. Your cache was warm and a switch would have rewritten it. He did the maths. He was right.
What if Jev is down? Then you get your session model and one polite toast. He never blocks a turn.
Why a raccoon? Because he is a small creature that gets very serious about picking the right track.
claude plugin validate .
claude plugin test .
Benchmarks: see benchmarks/.
MIT.
<sub>Keywords: Claude Code plugin, Claude model router, Claude Code cost optimization, Opus Sonnet Haiku switching, Claude Code subagents, prompt cache, reasoning effort, LLM routing.</sub>
hooks/register.tsx 689 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { Cache, Decision, Effort, EffortCap, Mode, Pending, Savings, Sticky } from '../types'
5
6export const JEV_URL = 'https://api.typesafe.ai/v1/systemone'
7const MODES: readonly Mode[] = ['efficient', 'balanced', 'cheap']
8const STICKY: readonly Sticky[] = ['off', 'auto', 'strict']
9const EFFORTS: readonly Effort[] = ['low', 'medium', 'high', 'xhigh', 'max']
10const CAPS: readonly EffortCap[] = [...EFFORTS, 'none']
11
12// weight: relative price per token, the multipliers the `about` lines state (fable has none, 10 is an assumption).
13const CLAUDE: Record<string, { id: string; about: string; weight: number }> = {
14 haiku: {
15 id: 'claude-haiku-5-5',
16 weight: 1,
17 about: 'Claude Haiku: fastest and cheapest (~1x). Simple edits, renames, lookups, short answers, boilerplate.',
18 },
19 sonnet: {
20 id: 'claude-sonnet-5-5',
21 weight: 3,
22 about: 'Claude Sonnet: strong general coding at mid cost (~3x). Features, refactors, tests, normal debugging.',
23 },
24 opus: {
25 id: 'claude-opus-5-5',
26 weight: 5,
27 about: 'Claude Opus: top reasoning at high cost (~5x). Hard debugging, architecture, large multi-file changes.',
28 },
29 fable: {
30 id: 'claude-fable-5-1',
31 weight: 10,
32 about: 'Claude Fable: most capable, highest cost. The hardest, longest, most ambiguous tasks.',
33 },
34}
35
36const MODE_RULE: Record<Mode, string> = {
37 efficient: 'Pick the model and effort that give the best result for this task; cost is secondary.',
38 balanced: 'Weigh result quality and cost evenly; use a stronger model only when the task clearly needs it.',
39 cheap: 'Pick the cheapest model and lowest effort that can plausibly succeed at this task.',
40}
41
42const EFFORT_ABOUT: Record<Effort, string> = {
43 low: 'Trivial or mechanical, little thinking needed',
44 medium: 'Ordinary task, some reasoning',
45 high: 'Non-trivial reasoning or multi-step work',
46 xhigh: 'Hard problem, careful deep reasoning',
47 max: 'Extremely hard, think as long as needed',
48}
49
50const S = { plugin: 'jev-router' } as const
51const modeAtom = atom({ ...S, key: 'mode' } as const, null)
52const enabledAtom = atom({ ...S, key: 'enabled' } as const, true)
53const lastAtom = atom({ ...S, key: 'last' } as const, null)
54const cacheAtom = atom({ ...S, key: 'cache' } as const, null)
55const stickyAtom = atom({ ...S, key: 'sticky' } as const, null)
56const capAtom = atom({ ...S, key: 'cap' } as const, null)
57const pendingAtom = atom({ ...S, key: 'pending' } as const, null)
58const frameAtom = atom({ ...S, key: 'frame' } as const, 0)
59const warnedAtom = atom({ ...S, key: 'warned' } as const, false)
60
61type JevChoice = { choice?: string; confidence?: number }
62type JevReply = { answers?: { model?: JevChoice; effort?: JevChoice } }
63
64export type Config = {
65 apiKey: string
66 timeoutMs: number
67 pool: string[]
68 defaultMode: Mode
69 defaultSticky: Sticky
70 defaultCap: EffortCap
71 pauseMs: number
72 routeSteps: boolean
73 routeSubagents: boolean
74 baseline: string
75 sendHistory: boolean
76 redact: boolean
77 maxTaskChars: number
78 fallback: 'session' | 'heuristic'
79 escalateAfter: number
80 minConfidence: number
81 minContextTokens: number
82 cacheTtlMs: number
83}
84type Picked = Omit<Decision, 'turnId'>
85
86// Estimated model-price units of one step: tokens as Anthropic bills them (cache reads 0.1x, cache writes 1.25x,
87// output 5x input) times the model's weight. Relative, not dollars, and blind to effort.
88type StepUsage = { input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number }
89export const stepUnits = (u: StepUsage, alias: string) =>
90 ((u.input_tokens + 1.25 * u.cache_creation_input_tokens + 0.1 * u.cache_read_input_tokens + 5 * u.output_tokens) * CLAUDE[alias]!.weight) / 1000
91
92export const savingsLine = (t: Savings) =>
93 t.actual === 0 ? 'no steps measured yet' : `est. ${Math.round((1 - t.actual / t.baseline) * 100)}% vs always ${t.name} over ${t.steps} steps (model price only; effort not compared)`
94
95let totals: Savings | null = null // loaded from the store once, then kept in memory
96
97// Cheapest to dearest (CLAUDE's key order); a switch up this ladder is an upgrade.
98const RANK = Object.keys(CLAUDE)
99const aliasOf = (model: string) => Object.keys(CLAUDE).find(a => CLAUDE[a]!.id === model)
100
101// Jev's effort drives token use more than its model pick does, so it is capped.
102// `capped` is what Jev wanted when the cap lowered it.
103export function capEffort(effort: Effort, cap: EffortCap): { effort: Effort; capped?: Effort } {
104 if (cap === 'none' || EFFORTS.indexOf(effort) <= EFFORTS.indexOf(cap)) return { effort }
105 return { effort: cap, capped: effort }
106}
107
108// Markers at the start of a prompt apply to that prompt only and are stripped before the model reads it:
109// !full lifts the effort cap, !<model> pins the model (no Jev call), !cheap and !efficient pick the routing mode.
110export function markers(text: string, pool: readonly string[]): { text: string; pending: Pending | null } {
111 const pending: Pending = {}
112 for (let m; (m = /^\s*!(full|cheap|efficient|[a-z]+)\b[ \t]*/i.exec(text)); ) {
113 const w = m[1]!.toLowerCase()
114 if (w === 'full') pending.full = true
115 else if (w === 'cheap' || w === 'efficient') pending.mode = w
116 else if (pool.includes(w)) pending.pin = w
117 else break
118 text = text.slice(m[0].length)
119 }
120
121 return { text, pending: Object.keys(pending).length ? pending : null }
122}
123
124export const isWarm = (cache: Cache | null, now: number, cfg: Config): cache is Cache =>
125 cache !== null && now - cache.at <= cfg.cacheTtlMs && cache.contextTokens >= cfg.minContextTokens
126
127// A switch discards the prompt cache. While it is warm and worth keeping, stay
128// on the cached model unless Jev confidently asks for a stronger one.
129export function decide(picked: Picked, cache: Cache | null, now: number, cfg: Config, sticky: Sticky): Picked {
130 if (sticky === 'off' || !isWarm(cache, now, cfg) || cache.model === picked.model) return picked
131 const from = aliasOf(cache.model)
132 if (!from) return picked
133 const upgrade = RANK.indexOf(picked.alias) > RANK.indexOf(from)
134 if (sticky === 'auto' && upgrade && picked.confidence >= cfg.minConfidence) return picked
135 return { ...picked, alias: from, model: cache.model, kept: picked.alias }
136}
137
138const bar = (n: number, on: string, off: string, len: number) =>
139 on.repeat(Math.round(n * len)) + off.repeat(len - Math.round(n * len))
140const gauge = (c: number) => bar(c, '▮', '▯', 5)
141const tier = (alias: string) => `${alias} ${'▂▄▆█'.slice(0, RANK.indexOf(alias) + 1).padEnd(4, '_')}`
142const effortBar = (e: Effort) => bar((EFFORTS.indexOf(e) + 1) / EFFORTS.length, '▰', '▱', 5)
143
144const TIER_COLOR: Record<string, string> = { haiku: '#4ade80', sonnet: '#60a5fa', opus: '#c084fc', fable: '#fbbf24' }
145const confColor = (c: number) => (c >= 0.7 ? '#4ade80' : c >= 0.5 ? '#fbbf24' : '#f87171')
146
147type Seg = { text: string; color?: string; bold?: boolean }
148
149const SPIN = '⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏'
150const GLOW = ['#c084fc', '#a78bfa', '#818cf8', '#60a5fa', '#38bdf8', '#60a5fa', '#818cf8', '#a78bfa']
151
152// frame is null while idle: static. While a task runs the mark spins and the label glows.
153export const segments = (mode: Mode, last: Decision | null, frame: number | null = null, cache: string | null = null, paused = false): Seg[] => {
154 const head: Seg[] = [
155 frame === null
156 ? { text: '◆ JEV ', color: '#c084fc', bold: true }
157 : { text: `${SPIN[frame % SPIN.length]} JEV `, color: GLOW[frame % GLOW.length], bold: true }, { text: `▏${mode}▕ `, color: '#94a3b8' },
158 ]
159 const warn: Seg[] = paused ? [{ text: '⚠ Jev paused ', color: '#fbbf24', bold: true }] : []
160 if (!last) return [...head, ...warn, { text: 'waiting for the first task', color: '#64748b' }]
161 const model: Seg = last.kept
162 ? { text: `${last.alias} (kept 🔒 cache warm; wanted ${last.kept})`, color: '#fbbf24' }
163 : { text: `${tier(last.alias)}${last.pinned ? ' 📌' : ''}${last.escalated ? ' ↑' : ''}`, color: TIER_COLOR[last.alias], bold: true }
164 const sep: Seg = { text: ' │ ', color: '#475569' }
165 return [
166 ...head,
167 ...warn,
168 model,
169 sep,
170 { text: `${last.effort} ${effortBar(last.effort)}${last.capped ? ` ⤓${last.capped}` : ''}${last.unlocked ? ' 🔓' : ''}`, color: last.unlocked ? '#fbbf24' : '#38bdf8' },
171 sep,
172 { text: `conf ${gauge(last.confidence)} ${Math.round(last.confidence * 100)}%`, color: confColor(last.confidence) },
173 ...(cache ? [sep, { text: cache, color: cache.includes('warm') ? '#4ade80' : '#64748b' }] : []),
174 ]
175}
176
177// One row per model: steps that ran on it, with a share bar.
178const usageRows = (stats: Record<string, number>, pool: string[]): [string, string][] => {
179 const total = pool.reduce((n, m) => n + (stats[m] ?? 0), 0)
180 if (!total) return [['usage', 'no steps yet']]
181 return pool.map((m, i): [string, string] => [
182 i === 0 ? 'usage' : '',
183 `${m.padEnd(7)} ${bar((stats[m] ?? 0) / total, '█', '░', 10)} ${String(stats[m] ?? 0).padStart(3)} ${Math.round(((stats[m] ?? 0) / total) * 100)}%`,
184 ])
185}
186
187function box(title: string, rows: [string, string][]) {
188 const body = rows.map(([k, v]) => `${k.padEnd(8)} ${v}`)
189 const w = Math.max(title.length + 3, ...body.map(l => l.length)) + 1
190 return [
191 `┌ ${title} ${'─'.repeat(w - title.length - 1)}┐`,
192 ...body.map(l => `│ ${l.padEnd(w)}│`),
193 `└${'─'.repeat(w + 1)}┘`,
194 ].join('\n')
195}
196
197async function cacheLabel($: EngineInterface, cfg: Config) {
198 const cache = await read($, cacheAtom)
199 if (!cache) return 'cache none yet'
200 const k = `${Math.round(cache.contextTokens / 1000)}k`
201 return isWarm(cache, await $.clock.now(), cfg) ? `cache 🔒 warm ${k}` : `cache cold ${k}`
202}
203
204async function currentMode($: EngineInterface, cfg: Config) {
205 return (await read($, modeAtom)) ?? cfg.defaultMode
206}
207
208async function buildState($: EngineInterface, text: string, mode: Mode, cache: Cache | null, now: number, cfg: Config) {
209 const current = await $.session.model()
210 const history = (await $.session.messages())
211 .slice(-8)
212 .map(m => `${m.role}: ${m.text.slice(0, 700)}`)
213 .join('\n')
214 .slice(-6000)
215
216 return composeState(text, mode, current, history, cache, now, cfg)
217}
218
219// Secrets that must not leave the machine inside a prompt or the conversation sent to Jev.
220const SECRETS: RegExp[] = [
221 /-----BEGIN [A-Z ]*PRIVATE KEY-----[\s\S]*?-----END [A-Z ]*PRIVATE KEY-----/g,
222 /\b(?:sk|pk|rk)-[A-Za-z0-9_-]{16,}/g,
223 /\bgh[pousr]_[A-Za-z0-9]{20,}/g,
224 /\b(?:AKIA|ASIA)[A-Z0-9]{16}\b/g,
225 /\bBearer\s+[A-Za-z0-9._~+/=-]{16,}/gi,
226]
227// NAME=value / "name": "value" for names that read like a secret, and user:password@ in URLs.
228const SECRET_NAME = /\b([\w.-]*(?:password|passwd|pwd|secret|token|api[_-]?key|(?:private|access|signing|encryption|deploy|ssh)[_-]?key)[\w.-]*)(["']?\s*[=:]\s*["']?)(?!(?:string|number|boolean|null|undefined|true|false|any|object)\b)[^\s"',;}]{3,}/gi
229const URL_CREDENTIALS = /\b([a-z][a-z0-9+.-]*:\/\/[^\s:/@]+):[^\s@/]+@/gi
230export const redact = (text: string) =>
231 SECRETS.reduce((t, re) => t.replace(re, '[redacted]'), text).replace(SECRET_NAME, '$1$2[redacted]').replace(URL_CREDENTIALS, '$1:[redacted]@')
232
233// Used only when Jev is down and `fallback` is 'heuristic'. ponytail: three buckets from length and a few words; real routing needs Jev.
234export function heuristic(text: string, pool: readonly string[]): Picked {
235 const hard = text.length > 1500 || /architect|design|migrat|root cause|concurren|race|deadlock|security|across .* (modules|services)/i.test(text)
236 const easy = text.length < 200 && /rename|typo|bump|format|lint|comment|where is|what does|show me/i.test(text)
237 const want = hard ? 'opus' : easy ? 'haiku' : 'sonnet'
238 const alias = RANK.slice(RANK.indexOf(want)).find(a => pool.includes(a)) ?? RANK.filter(a => pool.includes(a)).pop()!
239
240 return { alias, model: CLAUDE[alias]!.id, effort: hard ? 'high' : easy ? 'low' : 'medium', confidence: 0.3 }
241}
242
243// Pure, so the benchmark sends Jev exactly what the mod sends.
244export function composeState(text: string, mode: Mode, current: string, history: string, cache: Cache | null, now: number, cfg: Config) {
245 const clean = (t: string) => (cfg.redact ? redact(t) : t)
246
247 return [
248 `Routing mode: ${mode}. ${MODE_RULE[mode]}`,
249 `Current model: ${current}. Switching models discards the prompt cache, so prefer staying on it when the gain from switching is small.`,
250 cache
251 ? `Prompt cache: ${isWarm(cache, now, cfg) ? 'warm' : 'cold'}, ~${cache.contextTokens} tokens on ${cache.model}. Switching re-writes them.`
252 : '',
253 history && cfg.sendHistory ? `Recent conversation:\n${clean(history)}` : 'Recent conversation: (none, this is the first task)',
254 `New task from the user:\n${clean(text.slice(0, cfg.maxTaskChars))}`,
255 ]
256 .filter(Boolean)
257 .join('\n\n')
258}
259
260// Sensitive plugin options get no /config row, so the key also comes from
261// the plugin's store (`/jev key <key>`) or the TYPESAFE_API_KEY variable.
262async function resolveKey($: EngineInterface, cfg: Config): Promise<{ key: string; source: string } | null> {
263 if (cfg.apiKey) return { key: cfg.apiKey, source: 'plugin option' }
264 const stored = await $.store.get('apiKey')
265 if (typeof stored === 'string' && stored) return { key: stored, source: '/jev key' }
266 const env = await $.env.get('TYPESAFE_API_KEY')
267 if (env) return { key: env, source: 'TYPESAFE_API_KEY' }
268 return null
269}
270
271export function jevBody(pool: string[], state: string, mode: Mode) {
272 return {
273 model: 'jev-latest',
274 state,
275 questions: {
276 model: {
277 type: 'choice',
278 instructions: `Which Claude model should handle the new coding task? ${MODE_RULE[mode]}`,
279 criteria: Object.fromEntries(pool.map(m => [m, CLAUDE[m]!.about])),
280 },
281 effort: {
282 type: 'choice',
283 instructions: `How much reasoning effort does the new task need? ${MODE_RULE[mode]}`,
284 criteria: Object.fromEntries(EFFORTS.map(x => [x, EFFORT_ABOUT[x]])),
285 },
286 },
287 }
288}
289
290async function askJev($: EngineInterface, cfg: Config, key: string, state: string, mode: Mode) {
291 const body = jevBody(cfg.pool, state, mode)
292 const call = $.http.fetch(JEV_URL, {
293 method: 'POST',
294 headers: { 'content-type': 'application/json', authorization: `Bearer ${key}` },
295 body: JSON.stringify(body),
296 })
297 const timeout = $.clock.sleep(cfg.timeoutMs).then(() => {
298 throw new Error(`timed out after ${cfg.timeoutMs} ms`)
299 })
300 const res = await Promise.race([call, timeout])
301 if (!res.ok) throw new Error(`HTTP ${res.status}`)
302 const reply = JSON.parse(res.text) as JevReply
303 const alias = reply.answers?.model?.choice
304 const effort = reply.answers?.effort?.choice as Effort | undefined
305 if (!alias || !cfg.pool.includes(alias)) throw new Error(`unexpected model choice: ${alias}`)
306
307 return {
308 alias,
309 model: CLAUDE[alias]!.id,
310 effort: effort && EFFORTS.includes(effort) ? effort : 'medium',
311 confidence: reply.answers?.model?.confidence ?? 0,
312 }
313}
314
315// Circuit breaker: with a call per step, a slow Jev would cost timeoutMs on every step.
316// Three failures in a row pause Jev for pauseMs; the last pick keeps running meanwhile.
317let fails = 0
318let pausedUntil = 0
319export const isPaused = (now: number) => now < pausedUntil
320
321async function ask($: EngineInterface, cfg: Config, key: string, state: string, mode: Mode) {
322 const now = await $.clock.now()
323 if (isPaused(now)) throw new Error(`paused for ${Math.ceil((pausedUntil - now) / 1000)} s after repeated failures`)
324 try {
325 const picked = await askJev($, cfg, key, state, mode)
326 fails = 0
327
328 return picked
329 } catch (err) {
330 if (++fails >= 3) {
331 fails = 0
332 pausedUntil = now + cfg.pauseMs
333 $.ui.toast(`Jev failed 3 times in a row (${(err as Error).message}); paused for ${cfg.pauseMs / 1000} s, using the current model`)
334 }
335 throw err
336 }
337}
338
339let strikes = 0 // consecutive failed tool calls of the current task
340let skipNext = false // the next turn is a background notification, not a task of the user's
341let task = '' // the prompt the current task started with, for re-routing its later steps
342const agents = new Map<string, Promise<Picked | null>>()
343
344async function decideFor($: EngineInterface, cfg: Config, kind: string, key: string, text: string, state: string, mode: Mode, cache: Cache | null, now: number, cap: EffortCap) {
345 let asked: Picked
346 try {
347 asked = await ask($, cfg, key, state, mode)
348 } catch (err) {
349 if (cfg.fallback !== 'heuristic') throw err
350 asked = heuristic(text, cfg.pool)
351 $.ui.log(`jev ${kind}: Jev unavailable (${(err as Error).message}), heuristic pick ${asked.alias}`, { to: 'debug' })
352 }
353 const picked = decide(asked, cache, now, cfg, (await read($, stickyAtom)) ?? cfg.defaultSticky)
354 const d = { ...picked, ...capEffort(picked.effort, cap) }
355 $.ui.log(`jev ${kind}: ${d.alias}/${d.effort}${d.kept ? ` (kept, wanted ${d.kept})` : ''}${d.capped ? ` (capped from ${d.capped})` : ''} conf ${d.confidence.toFixed(2)} ${(await $.clock.now()) - now} ms`, { to: 'debug' })
356
357 return d
358}
359
360// Before every step after the first, ask again: the task may have turned easier or harder.
361// Failures keep the last pick quietly; the task's own first call already warned.
362async function restep($: EngineInterface, cfg: Config, e: { turnId: string; index: number }, last: Decision): Promise<Decision> {
363 const auth = await resolveKey($, cfg)
364 if (!auth || last.pinned) return last
365 try {
366 const mode = await currentMode($, cfg)
367 const now = await $.clock.now()
368 const cache = await read($, cacheAtom)
369 // strict never switches while the cache is warm, so asking could not change the pick
370 if (((await read($, stickyAtom)) ?? cfg.defaultSticky) === 'strict' && isWarm(cache, now, cfg)) return last
371 const text = `${task}\n\n[Routing step ${e.index + 1} of this task. The recent messages above show its progress.]`
372 const cap = last.unlocked ? 'none' : (await read($, capAtom)) ?? cfg.defaultCap
373 const d = await decideFor($, cfg, `step ${e.index + 1}`, auth.key, text, await buildState($, text, mode, cache, now, cfg), mode, cache, now, cap)
374 if (last.escalated && RANK.indexOf(d.alias) <= RANK.indexOf(last.alias)) return last // never de-escalate within a task
375 const next: Decision = { ...d, turnId: last.turnId, unlocked: last.unlocked }
376 await update($, lastAtom, () => next)
377
378 return next
379 } catch {
380 return last
381 }
382}
383
384// Routes one subagent from the text of its task. Called at spawn (the full prompt, so the pick is known before the
385// subagent exists and the subagent list can show it) and, for agents no spawn hook saw, from their description.
386async function routeSubagent($: EngineInterface, cfg: Config, kind: string, text: string, parentModel: string): Promise<Picked | null> {
387 const auth = await resolveKey($, cfg)
388 if (!auth) return null
389 try {
390 const now = await $.clock.now()
391 const mode = await currentMode($, cfg)
392 const cap = (await read($, lastAtom))?.unlocked ? 'none' : (await read($, capAtom)) ?? cfg.defaultCap
393
394 return await decideFor($, cfg, kind, auth.key, text, composeState(text, mode, parentModel, '', null, now, cfg), mode, null, now, cap)
395 } catch {
396 return null
397 }
398}
399
400async function routeAgent($: EngineInterface, cfg: Config, id: string, parentModel: string): Promise<Picked | null> {
401 const info = (await $.agent.list()).find(a => a.id === id)
402
403 return info?.description ? routeSubagent($, cfg, `agent ${info.type}: ${info.description}`, `Subagent task (${info.type}): ${info.description}`, parentModel) : null
404}
405
406// The running totals, loaded from the plugin store once (in session.start, so parallel steps never race on it).
407const emptyTally = (name: string): Savings => ({ actual: 0, baseline: 0, steps: 0, name, models: {} })
408async function tally($: EngineInterface, cfg: Config): Promise<Savings> {
409 const stored = (await $.store.get('savings')) as Savings | undefined
410 // a new baseline starts a new tally
411 return stored && stored.name === cfg.baseline ? { ...emptyTally(cfg.baseline), ...stored } : emptyTally(cfg.baseline)
412}
413
414// Add one step to the running totals (kept across sessions in the plugin store).
415async function record($: EngineInterface, cfg: Config, u: (StepUsage & { model: string }) | undefined) {
416 const alias = u && aliasOf(u.model)
417 if (!u || !alias) return
418 totals ??= await tally($, cfg)
419 totals.models[alias] = (totals.models[alias] ?? 0) + 1
420 totals.actual += stepUnits(u, alias)
421 totals.baseline += stepUnits(u, cfg.baseline)
422 totals.steps++
423 await $.store.set('savings', totals)
424}
425
426const usage = (cmd: string, list: readonly string[]) => ({ text: `Usage: /jev ${cmd} ${list.join(' | ')}` })
427const applied = (name: string, v: string) => ({ text: `Jev ${name}: ${v}. Applies from the next task.` })
428
429export const register: Register = (on, options) => {
430 const apiKey = String(options.apiKey ?? '')
431 const timeoutMs = Number(options.timeoutMs ?? 4000)
432 const allowed = (Array.isArray(options.models) ? options.models : Object.keys(CLAUDE))
433 .map(m => String(m).toLowerCase())
434 .filter(m => m in CLAUDE)
435 const defaultMode: Mode = MODES.includes(options.mode as Mode) ? (options.mode as Mode) : 'balanced'
436
437 const defaultSticky: Sticky = STICKY.includes(options.stickiness as Sticky) ? (options.stickiness as Sticky) : 'auto'
438 const defaultCap: EffortCap = CAPS.includes(options.effortCap as EffortCap) ? (options.effortCap as EffortCap) : 'high'
439 const flag = (v: unknown, d: boolean) => (v === undefined || v === '' ? d : String(v) !== 'false')
440 const num = (v: unknown, d: number) => (v !== undefined && v !== '' && Number.isFinite(Number(v)) ? Number(v) : d)
441
442 const cfg: Config = {
443 apiKey,
444 timeoutMs,
445 pool: allowed.length > 0 ? allowed : Object.keys(CLAUDE),
446 defaultMode,
447 defaultSticky,
448 defaultCap,
449 pauseMs: num(options.pauseMs, 60_000),
450 routeSteps: flag(options.routeSteps, true),
451 routeSubagents: flag(options.routeSubagents, true),
452 escalateAfter: num(options.escalateAfter, 3),
453 sendHistory: flag(options.sendHistory, true),
454 redact: flag(options.redact, true),
455 maxTaskChars: num(options.maxTaskChars, 8000),
456 fallback: options.fallback === 'heuristic' ? 'heuristic' : 'session',
457 baseline: String(options.baselineModel ?? 'opus') in CLAUDE ? String(options.baselineModel ?? 'opus') : 'opus',
458 minConfidence: num(options.minConfidence, 0.7),
459 minContextTokens: num(options.minContextTokens, 8000),
460 cacheTtlMs: num(options.cacheTtlMs, 300_000),
461 }
462
463 // The frame ticks only while a task runs, so an idle session redraws nothing.
464 let ticker: { cancel: () => void } | undefined
465 const stop = () => {
466 ticker?.cancel()
467 ticker = undefined
468 }
469
470 on('prompt.submit', async ($, e, next) => {
471 stop()
472 let ticks = 0 // self-cancel: a prompt that never reaches turn.complete must not tick forever
473 const t = $.clock.every(120, () => (++ticks > 5000 ? t.cancel() : update($, frameAtom, n => n + 1)))
474 ticker = t
475 skipNext = e.origin?.kind === 'task-notification'
476 const m = markers(e.text, cfg.pool)
477 if (!m.pending || m.text.trim() === '') return next(e)
478 await update($, pendingAtom, () => m.pending)
479
480 return next({ ...e, text: m.text })
481 })
482
483 on('turn.complete', async ($, e, next) => {
484 stop()
485 return next(e)
486 })
487
488 on('command.run', { command: 'jev' }, async ($, e, next) => {
489 stop() // a /jev command is not a task: no spinner
490 return next(e)
491 })
492
493 // Pick the subagent's model at spawn and show it in the subagent list.
494 on('agent.spawn', async ($, e, next) => {
495 if (!cfg.routeSubagents || !(await read($, enabledAtom)) || e.model) return next(e) // an explicit model on the Agent call wins
496 const pick = await routeSubagent($, cfg, `spawn ${e.subagentType}: ${e.description}`, `Subagent task (${e.subagentType}): ${e.description}\n\n${e.prompt}`, e.parentModel)
497 if (!pick) return next(e)
498 const res = await next({ ...e, model: pick.alias, description: `${e.description} · ${pick.alias}/${pick.effort}` })
499 if (res.agentId) agents.set(res.agentId, Promise.resolve(pick)) // its steps reuse the pick (and its effort)
500
501 return res
502 })
503
504 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
505 if (e.props.hasSurvey) return next(e)
506 const { Box, Text } = $.ui.resolve(e)
507 const segs: Seg[] = (await read($, enabledAtom))
508 ? segments(await currentMode($, cfg), await read($, lastAtom), e.props.isWorking ? await read($, frameAtom) : null, await cacheLabel($, cfg), isPaused(await $.clock.now()))
509 : [{ text: 'JEV ▏off▕', color: '#64748b' }]
510
511 return (
512 <Box>
513 <Text>
514 {segs.map(x => (
515 <Text color={x.color} bold={x.bold}>
516 {x.text}
517 </Text>
518 ))}
519 </Text>
520 </Box>
521 )
522 })
523
524 on('session.start', async ($, e, next) => {
525 agents.clear()
526 totals = await tally($, cfg)
527 await $.command.register({
528 name: 'jev',
529 description: 'Jev model routing (prompt markers: !full !haiku !sonnet !opus !fable !cheap !efficient): efficient | balanced | cheap | sticky <off|auto|strict> | cap <effort|none> | full | on | off | status | key <key>',
530 })
531
532 return next(e)
533 })
534
535 on('turn.start', async ($, e, next) => {
536 // A continuation, or a background notification, keeps the decision of the task it continues.
537 if (!(await read($, enabledAtom)) || e.text.trim() === '' || skipNext) return next(e)
538 task = e.text
539 strikes = 0
540 const pending = await read($, pendingAtom)
541 if (pending) await update($, pendingAtom, () => null) // one prompt only, even if nothing can be routed
542 const auth = await resolveKey($, cfg)
543 if (!auth && !pending?.pin) { // a pinned model needs no Jev call, so no key either
544 await update($, lastAtom, () => null) // never run this prompt on the previous task's pick
545 if (!(await read($, warnedAtom))) {
546 await update($, warnedAtom, () => true)
547 $.ui.toast('Jev: no TypeSafe API key. Run /jev key <key> or set TYPESAFE_API_KEY')
548 }
549 return next(e)
550 }
551
552 const unlocked = pending?.full === true
553 const mode = pending?.mode ?? (await currentMode($, cfg))
554 try {
555 const cache = await read($, cacheAtom)
556 const now = await $.clock.now()
557 const cap = unlocked ? 'none' : (await read($, capAtom)) ?? cfg.defaultCap
558 // A pinned model skips Jev; its effort is the cap (high by default).
559 const picked = pending?.pin
560 ? { alias: pending.pin, model: CLAUDE[pending.pin]!.id, effort: cap === 'none' ? 'high' : cap, confidence: 1, pinned: true }
561 : await decideFor($, cfg, 'prompt', auth!.key, e.text, await buildState($, e.text, mode, cache, now, cfg), mode, cache, now, cap)
562 const decision: Decision = { turnId: e.turnId, ...picked, unlocked }
563 await update($, lastAtom, () => decision)
564 } catch (err) {
565 await update($, lastAtom, () => null)
566 if (!(await read($, warnedAtom))) {
567 await update($, warnedAtom, () => true)
568 $.ui.toast(`Jev unavailable (${(err as Error).message}), using ${await $.session.model()}`)
569 }
570 }
571
572 return next(e)
573 })
574
575 // A task that keeps failing moves up one model for the rest of the task, whatever the cache guard says.
576 on('tool.call', async ($, e, next) => {
577 const res = await next(e)
578 if (e.agentId !== undefined || !cfg.escalateAfter) return res
579 strikes = 'isError' in res && res.isError ? strikes + 1 : 0
580 const last = await read($, lastAtom)
581 if (strikes >= cfg.escalateAfter && last && !last.escalated && !last.pinned) {
582 const to = RANK[Math.min(RANK.indexOf(last.alias) + 1, RANK.length - 1)]!
583 if (cfg.pool.includes(to) && to !== last.alias) {
584 await update($, lastAtom, () => ({ ...last, alias: to, model: CLAUDE[to]!.id, escalated: true }))
585 $.ui.log(`jev escalate: ${last.alias} -> ${to} after ${strikes} failed tool calls`, { to: 'debug' })
586 }
587 }
588
589 return res
590 })
591
592 on('turn.step', async function* ($, e, next) {
593 if (!(await read($, enabledAtom))) return yield* next(e)
594 if (e.agentId !== undefined) {
595 if (!cfg.routeSubagents) return yield* next(e)
596 const id = e.agentId
597 const pick = await (agents.get(id) ?? agents.set(id, routeAgent($, cfg, id, e.model)).get(id)!)
598
599 const res = yield* next(pick ? { ...e, model: pick.model, effort: pick.effort } : e)
600 await record($, cfg, res.usage)
601
602 return res
603 }
604 const prev = await read($, lastAtom)
605 const last = cfg.routeSteps && e.index > 0 && prev ? await restep($, cfg, e, prev) : prev
606
607 // A model that takes no effort setting is sent none, whatever is asked.
608 const res = yield* next(last ? { ...e, model: last.model, effort: last.effort } : e)
609
610 // Remember what the API actually cached, so the next turn can protect it.
611 const u = res.usage
612 if (u) {
613 const contextTokens = u.input_tokens + u.cache_read_input_tokens + u.cache_creation_input_tokens + u.output_tokens
614 const at = await $.clock.now()
615 await update($, cacheAtom, () => ({ model: u.model, at, contextTokens }))
616 }
617 await record($, cfg, res.usage)
618 return res
619 })
620
621 on('command.run', { command: 'jev' }, async ($, e) => {
622 const arg = e.args.trim().toLowerCase()
623
624 if (MODES.includes(arg as Mode)) {
625 await update($, modeAtom, () => arg as Mode)
626 await update($, enabledAtom, () => true)
627 return { text: `Jev routing mode: ${arg}. Applies from the next task.` }
628 }
629 if (arg.startsWith('sticky')) {
630 const v = arg.slice(6).trim()
631 if (!STICKY.includes(v as Sticky)) return usage('sticky', STICKY)
632 await update($, stickyAtom, () => v as Sticky)
633 return applied('cache stickiness', v)
634 }
635 if (arg === 'stats clear') {
636 totals = null
637 await $.store.delete('savings')
638 return { text: 'Jev savings tally cleared.' }
639 }
640 if (arg === 'full') {
641 await update($, pendingAtom, p => ({ ...p, full: true }))
642 return { text: 'Jev: the next prompt runs with the effort cap lifted.' }
643 }
644 if (arg.startsWith('cap')) {
645 const v = arg.slice(3).trim()
646 if (!CAPS.includes(v as EffortCap)) return usage('cap', CAPS)
647 await update($, capAtom, () => v as EffortCap)
648 return applied('effort cap', v)
649 }
650 if (arg === 'on' || arg === 'off') {
651 await update($, enabledAtom, () => arg === 'on')
652 if (arg === 'off') await update($, lastAtom, () => null)
653 return { text: `Jev routing ${arg === 'on' ? 'enabled' : 'disabled; the session model is used'}.` }
654 }
655 if (arg === 'key' || arg.startsWith('key ')) {
656 const key = e.args.trim().slice(3).trim()
657 if (key === '') return { text: 'Usage: /jev key <TypeSafe API key> or /jev key clear' }
658 if (key === 'clear') {
659 await $.store.delete('apiKey')
660 return { text: 'Stored TypeSafe API key removed.' }
661 }
662 await $.store.set('apiKey', key)
663 await update($, warnedAtom, () => false)
664 return { text: `TypeSafe API key saved (…${key.slice(-4)}). Applies from the next task.` }
665 }
666 if (arg === '' || arg === 'status') {
667 const auth = await resolveKey($, cfg)
668 const enabled = await read($, enabledAtom)
669 const last = await read($, lastAtom)
670 const cache = await read($, cacheAtom)
671 const warm = isWarm(cache, await $.clock.now(), cfg)
672 return {
673 text: '\n' + box('Jev router', [
674 ['state', `${enabled ? '● on ' : '○ off'} mode ${await currentMode($, cfg)}`],
675 ['pool', cfg.pool.join(' · ')],
676 ['sticky', `${(await read($, stickyAtom)) ?? cfg.defaultSticky} cache ${warm ? '● warm' : '○ cold'}`],
677 ['cap', `${(await read($, capAtom)) ?? cfg.defaultCap}${(await read($, pendingAtom)) ? ' 🔓 next prompt: overrides set' : ''}`],
678 ['last', last ? `${last.alias} / ${last.effort} ${gauge(last.confidence)} ${last.confidence.toFixed(2)}${last.kept ? ` kept (wanted ${last.kept})` : ''}${last.capped ? ` capped from ${last.capped}` : ''}` : 'none yet'],
679 ...usageRows(totals?.models ?? {}, cfg.pool),
680 ['saved', savingsLine(totals ?? (await tally($, cfg)))],
681 ['jev', isPaused(await $.clock.now()) ? `⚠ paused ${Math.ceil((pausedUntil - (await $.clock.now())) / 1000)} s` : '● ok'],
682 ['key', auth ? `…${auth.key.slice(-4)} (${auth.source})` : 'none: /jev key <key>'],
683 ]),
684 }
685 }
686 return { text: `Unknown option "${arg}". Use: /jev efficient | balanced | cheap | sticky <mode> | cap <effort|none> | full | on | off | status | key <key>` }
687 })
688}
689types/index.d.ts 36 lines1export type Mode = 'efficient' | 'balanced' | 'cheap'
2export type Sticky = 'off' | 'auto' | 'strict'
3export type Cache = { model: string; at: number; contextTokens: number }
4export type Effort = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
5export type EffortCap = Effort | 'none'
6export type Savings = { actual: number; baseline: number; steps: number; name: string; models: Record<string, number> }
7export type Pending = { full?: true; pin?: string; mode?: Mode }
8export type Decision = {
9 turnId: string
10 model: string
11 alias: string
12 effort: Effort
13 confidence: number
14 kept?: string
15 capped?: Effort
16 pinned?: boolean
17 escalated?: boolean
18 unlocked?: boolean
19}
20
21declare module 'claude-code' {
22 interface PluginState {
23 'jev-router': {
24 mode: Mode | null
25 enabled: boolean
26 last: Decision | null
27 cache: Cache | null
28 sticky: Sticky | null
29 cap: EffortCap | null
30 pending: Pending | null
31 frame: number
32 warned: boolean
33 }
34 }
35}
36