TypeSafe's Jev decides how Claude Code works on each prompt: the main conversation's reasoning effort (low to xhigh, raised to max mid-turn when tool calls…

<img src="assets/banner.svg" alt="jev-pilot — let Jev steer Claude Code" width="100%">
<a href="https://github.com/Akramovic1/jev-pilot/actions/workflows/test.yml"><img alt="tests" src="https://img.shields.io/github/actions/workflow/status/Akramovic1/jev-pilot/test.yml?branch=main&style=flat-square&label=tests&labelColor=0b1020"></a> <a href="LICENSE"><img alt="License: MIT" src="https://img.shields.io/badge/license-MIT-d4ff4f?style=flat-square&labelColor=0b1020"></a> <a href="https://docs.claude.com/en/docs/claude-code"><img alt="Claude Code 2.1.278+" src="https://img.shields.io/badge/Claude%20Code-2.1.278%2B-7cf0c4?style=flat-square&labelColor=0b1020"></a> <a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev"><img alt="Powered by Jev" src="https://img.shields.io/badge/powered%20by-Jev%20(TypeSafe)-c9d2ea?style=flat-square&labelColor=0b1020"></a> <a href="https://openrouter.ai/~typesafe/jev-latest"><img alt="Jev on OpenRouter" src="https://img.shields.io/badge/runs%20on-OpenRouter-8d99b8?style=flat-square&labelColor=0b1020"></a>
<b>The right reasoning effort, subagent model and skill for every prompt, decided by a model built for decisions.</b>
<a href="#-install">Install</a> · <a href="#-how-it-works">How it works</a> · <a href="#%EF%B8%8F-meet-the-pilot">The pet</a> · <a href="#-the-crew-custom-models-and-other-agents">The crew</a> · <a href="#-see-if-its-paying-off">Report</a> · <a href="#%EF%B8%8F-configuration">Configuration</a> · <a href="#-acknowledgements">Acknowledgements</a>
jev-pilot is a Claude Code plugin. Before every turn, it asks Jev, TypeSafe's fast decision model, a few typed questions about your prompt and sets the turn up from the answers. You keep Opus for the conversation. Easy work runs at low effort and on cheaper subagents, and hard work gets the thinking it needs.
| Decision | When | |
|---|---|---|
| 🧠 | Reasoning effort, low → xhigh | at the start of each turn |
| 🚨 | Raise effort, up to max, when tool calls keep failing | mid-turn, at most once |
| 🤖 | Subagent model and effort: Haiku, Sonnet or Opus, low → xhigh | when a subagent starts |
| 🧭 | Strategy: do it directly, delegate, run in parallel, or plan a graph | at the start of each turn, as advice |
| 🧩 | The one skill the prompt needs, if any | at the start of each turn |
| 📊 | A record of every decision, with tuning suggestions | always, via /jev-pilot:report |
Jev never writes in your conversation. It talks through Claude the pilot, a small animated pet above the prompt that shows what Claude is doing and says what Jev decided.
<img src="assets/demo.svg" alt="An illustrative claude-jev session. A rename runs at low effort, and the pet's bubble says low, no skill, 99% sure. A failing-tests prompt starts at xhigh with the systematic-debugging skill: the pet reads, searches and runs the tests; after two failures the effort is raised to max; it writes the fix, the tests pass, and it jumps rope." width="860"> <sub>An illustrative session: the pet and its bubbles are drawn from the plugin's own code; the numbers are examples.</sub>
[!NOTE] jev-pilot runs on Claude Code's function hooks, which are early access: they need Claude Code 2.1.278 or newer and
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1. Theclaude-jevlauncher sets that for you.
curl -fsSL https://raw.githubusercontent.com/Akramovic1/jev-pilot/main/install.sh | bash
It asks for your OpenRouter key (create one here; Jev costs about $0.04 per million input tokens, with free output), installs jev-pilot as a regular Claude Code plugin, and adds the claude-jev command. Then start Claude Code with it:
claude-jev # takes the same arguments as claude: claude-jev -c, claude-jev -p "…"
claude plugin install jev-pilot@jev-pilot. The key is passed with --config, so Claude Code keeps it in its own credential store, not in plain settings.claude-jev into ~/.local/bin. It's plain claude with function hooks on, plus the local router for custom models.It changes nothing else. Re-running it updates jev-pilot and keeps your key. For scripted installs, set JEV_OPENROUTER_KEY=sk-or-… (or JEV_SKIP_KEY=1) to skip the prompt.
claude plugin marketplace add Akramovic1/jev-pilot
claude plugin install jev-pilot@jev-pilot --config openrouterApiKey=sk-or-v1-… --config timeoutMs=1500
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude
To skip typing the variable, put it in ~/.claude/settings.json and plain claude will do:
{ "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" } }
git clone https://github.com/Akramovic1/jev-pilot.git && cd jev-pilot
./install.sh
This loads the clone with --plugin-dir, so your edits take effect in the next session. Options go in ~/.claude/settings.json under pluginConfigs["jev-pilot"], and the installer adds your key there after a backup. claude-jev checks your clone's upstream once a day in the background and tells you when there's something new.
Start claude-jev and look above the prompt, at the right: the pilot appears with a bubble saying ready · openrouter. After your first prompt the bubble says what Jev decided, such as low · no skill · 99% sure.
If it says ready · no key, built-in, the key isn't being read. Run the installer again. To see every step Jev takes, turn on verboseLog (see Configuration).
claude-jev self-update # update jev-pilot, whichever way it was installed
curl -fsSL https://raw.githubusercontent.com/Akramovic1/jev-pilot/main/install.sh | bash -s -- --uninstall
[!IMPORTANT] If you ran
/jev-pilot:setup, run/jev-pilot:setup restorebefore uninstalling. Otherwise your skills stay hidden from Claude with nothing left to load them.
flowchart LR
P(["Your prompt<br/>+ recent messages"]) --> J{{"Jev<br/>≈0.5 s"}}
J -- "effort" --> T["Turn<br/>(Opus)"]
J -- "skill + SKILL.md" --> T
J -- "strategy advice" --> T
T -- "2 failed tool calls" --> R["Raise effort<br/>up to max"]
R --> T
T -- "spawns a subagent" --> J2{{"Jev"}}
J2 -- "haiku / sonnet / opus" --> A["Subagent"]
T --> L[("Decision<br/>ledger")]
L --> Rep["/jev-pilot:report"]
Effort. Jev picks one of five levels, each described by the kind of task it's for, not an amount. Every question to Jev is a choice like this, each option saying when to choose it:
| Level | Kind of task |
|---|---|
low | answered from what's known, or one mechanical step: a lookup, one command, a rename |
medium | an ordinary, well-specified change to a few files, or a direct question about code in view |
high | a change across several files, a described bug that must be traced, tests, a careful review |
xhigh | design across components, a bug with an unknown cause, a refactor with many dependents |
max | novel architecture, security or data integrity, a failure that resisted earlier attempts |
effortChanges: auto, per-turn or hold).high. If Jev's two most likely levels are within effortCloseMargin (0.15), the higher one wins, because under-thinking costs more than over-thinking.high, Jev has to be sure. xhigh needs Jev at least 60% sure the task is very hard (xhigh and max together). A near split between hard and very hard stays at high.high.xhigh at most (maxEffort). Only the mid-turn raise reaches max: after escalateAfterErrors (2) failed tool calls in a row, effort goes up at least one level, once per turn. Permission denials don't count as failures.What Jev reads.
Subagents. Following Anthropic's guidance for Sonnet 5.5 ("it fits best when the task has a clear spec and a way to check the result"): Haiku for read-only lookups where a mistake is cheap to spot (search, find a definition, read files, logs or test output, run a command and report). Sonnet for read-only work that needs some understanding (explain code, research across files, review a diff and report) and for code changes with a clear spec and a check: a fix whose cause is known, a feature to a written spec, tests for existing code, a scoped refactor. Opus for work that needs careful judgment or runs long: design, an open spec, an unknown cause, long multi-step builds, security, migrations, production or money. On 40 real subagent briefs from my own sessions, 16 builders and fixers with written findings and tests moved to Sonnet 5.5. The large milestone builds and the reviews stayed on Opus. A subagent that runs at low effort and may change code gets Anthropic's line for low effort appended to its brief: run a real check that exercises the change before reporting it done. A model named on the Agent call (you asked for one) is kept. Moving down a model needs Jev at least 60% sure. These are family names, so Claude Code uses its current release of each. No versions are hardcoded. It also gets an effort from the same decision, on the same rubric and bars as the main conversation, but Jev is asked how hard the brief is to carry out: a brief that already names the files, steps and tests has done the design, so builders and fixers usually get high, and xhigh is kept for briefs that ask for design or an unknown cause. The Agent tool has no effort setting, so jev-pilot sets it on each request the subagent makes.
Claude knows it's there. On the first prompt of each session (and after a compaction), Claude gets a short note listing what jev-pilot decides, so it leaves those decisions alone: it won't pin a subagent's model or effort, or make agent types just to fix one, unless you ask.
Strategy. The same request asks how to carry the work out:
direct: the usual case, and nothing is attached.delegate: one subagent on a cheaper model does the broad, mechanical part.parallel: fan out, then join. Independent pieces that share no files run as simultaneous background subagents; the results are integrated and tested once.graph: for large builds only, a small blueprint of plain subagents. Real nodes (a step you could do inline isn't one), waves that start together, one shared plan file, a separate read-only reviewer after each join, and bounds (at most 4 subagents at a time, 2 review rounds per wave). If it can't be explained in one breath, Claude works directly.Advice is attached only when Jev is confident (0.6, or 0.8 for graph) and it agrees with the tier. Claude may ignore it.
Skills. At most one per prompt. Jev reads every skill's description and the opening of its SKILL.md, next to a "none of these fits" option, and is asked to match the kind of work (debugging, planning, reviewing…), not a product the prompt happens to name. A skill for one platform (Vercel, Supabase, Firebase…) is picked only when the request, the conversation or the project uses that platform: jev-pilot reads what the project deploys with from its file names (cdk.json, vercel.json, Dockerfile…), so "deploy to production" in an AWS project doesn't get Vercel's deploy skill. A skill is picked when Jev is sure of it, or when the prompt needs a skill and it still fits. The winner's SKILL.md is added to the prompt.
Better code, not just cheaper. The same request also asks Jev what would make the work better. Jev can't judge code, since it never sees your repo, but it can judge the request. Each read acts only when Jev is sure:
| Jev reads the request as… | Claude gets | Bar |
|---|---|---|
| vague: "add caching", "make it better" | ask one short question, or state the assumption in one line, before coding | 85% |
| a bug: "it's off by one cent", "the test fails" | show the bug first with a failing test or a command, then fix it and show the same check passing | 80% |
| a costly area: money, auth, migrations, security | run the tests that cover it and add one for the changed case; if a reviewer (Codex or OpenCode) is working, get its review | 80% |
| your correction of the last turn: "it doesn't work", "not what I asked" | nothing: the last turn is marked in the ledger as corrected | 70% |
The questions were tuned on sample prompts. For example, a first wording rated "add a dark mode toggle" as vague as "add caching"; the final one separates them (0.19 against 0.86).
/claude-api prompt-audit flagged the first versions for rituals that make Opus 5.5 write more and repeat tool calls (a plan to sketch, a review after every wave, "show the bug, then show it fixed", the same test run asked for twice), and they were rewritten./jev-pilot:report now shows, for each starting effort, how often you corrected the turn, and /jev tune leans up when cheap starts keep getting corrected. That's how you find out whether low effort is really enough for your work./jev quality off switches all of this off. It costs no extra wait: the questions ride in the same request.UI design work gets the design pack. When Jev reads a request as UI design (a page, a screen, a component, a redesign, a mobile flow; measured: design requests 0.95 to 0.98, everything else 0.09 at most, a UI bug included), Claude is told to:
designSkills, by default design-taste-frontend and impeccable) for the direction and polish, and check the result against web-design-guidelines before calling it done. Only the ones you have installed are named;It rides in the same one request. /jev design off switches it off.
One request per prompt. The effort, model, strategy and skill questions all go to Jev together, in one request of about 0.5 s. Before 0.6 there were three requests one after another: effort and strategy, the skill ranking, then a re-check of the top skills, about 1.5 s in all. On 16 test prompts the single request picked the same skill 14 times, and the other two picks were better. A plain "continue" asks nothing: the work goes on as the last turn decided. (Catalogs over the API's 255-choice limit are ranked in parallel batches, so it's still one wait.) /jev-pilot:setup can hide your own skills from Claude's skill list entirely (it asks first; restore undoes it).
<img src="assets/pet.svg" alt="Claude the pilot above the Claude Code prompt: thinking with a thought cloud, reading a book, searching with a magnifying glass, running tests in a terminal, flying when the effort is raised, writing on paper, then jumping rope and waving while idle." width="860">
Claude the pilot, drawn as Claude Code's character, sits above the prompt at the right and shows what Claude is doing:
| Claude is… | The pilot | Its bubble |
|---|---|---|
| thinking | a thought cloud, ... filling in | ⠋ thinking · … |
| reading files or pages | an open book, the line being read lit up | ⠋ reading · … |
| searching (Grep, Glob, web) | a magnifying glass, sweeping | ⠋ searching · … |
| editing files or writing the answer | paper, a pencil writing lines | ⠋ writing · … |
| running commands | a terminal, output scrolling | ⠋ running · … |
| running subagents or other tools | flying: goggles down, jets on | ⠋ working · … |
| idle | hovering and blinking; every few seconds it jumps rope, waves or looks around | Jev's last decision |
The bubble says the turn's effort, the skill attached (or no skill), any strategy advice, and how sure Jev was of the effort: xhigh · /systematic-debugging · parallel · 88% sure. When Jev wanted a change but wasn't sure enough to make it, it says so: high kept · wanted low · 42% sure. A mid-turn raise shows as 2 fails → max ✈, and a subagent's model as Explore → haiku.
It draws only in the terminal (not in claude -p, the desktop app or mobile), and redraws only while something moves. /jev pet off hides it.
Everything is on by default except switching the main conversation's model. Type /jev to see the switches, and change them live:
/jev what is on
/jev skills off one switch: effort · raise · subagents · skills · strategy · quality · design · model · pet
/jev all off every switch (all on turns them back on)
/jev reset back to your settings' defaults
Switches are remembered across sessions. /jev skills off leaves skills exactly as Claude Code handles them.
N% sure at the end, for each turn.display to both or transcript, and each turn gets one line, such as jev · low (93% sure) · no skill · 1.3s.verboseLog to see each answer with its confidence, such as tier fast (0.99) · effort 0.0 → low (1.00) · risky 0.10 · strategy direct (1.00) and needs a skill 0.09.jev-pilot can bring more workers into a Claude Code session than Claude alone:
codex and opencode CLIs, with your own logins, as code reviewers. They run as themselves, not through OpenRouter.You choose how they're used with a mode. Within that mode, Jev decides task by task.
| Mode | What happens |
|---|---|
standard (default) | Claude only. Jev picks Haiku, Sonnet or Opus and the effort. |
budget | Subagent work that needs no judgment (searching, reading and reporting, boilerplate) can go to a custom model, when Jev is sure. Opus keeps the judgment. |
junior-lead | A junior on a custom model writes easy, well-specified code. Opus, as tech lead, reads its diff, runs the tests and sends it back once if something's wrong. You get a cheap implementation reviewed at Opus level. |
second-opinion | After a significant change, Claude asks Codex (or OpenCode) for a review before calling the work done, then fixes what's right and says why it disagrees with the rest. |
quality | Every subagent runs on Opus, plus the external review. |
Outside these modes you can still ask for a review at any time: "have Codex review this". Claude then spawns jev-pilot:codex-review.
Choose the reviewer's model. Say it in your request, "review this with Codex, Luna, high effort", or set a default that every session keeps:
/jev reviewer codex luna high Codex reviews on Luna, high effort
/jev reviewer codex effort xhigh just the effort
/jev reviewer opencode kimi-k3 an OpenCode model (checked against `opencode models`)
/jev reviewer codex default back to the CLI's own config
A Codex tier name (astra, sol, terra, luna) always means the newest model of that tier. jev-pilot reads Codex's model list at every session start, so when a newer Luna ships, reviews move to it with nothing to change. A full id such as gpt-5.6-luna pins that exact version. Efforts are checked against what the model takes. /jev status shows what each reviewer runs on now, e.g. luna (newest, now gpt-5.6-luna) · effort high.
Adding a model. Pick a name, find a model on OpenRouter's list of models that can call tools, and paste it after the name. The id, the page link or the model's name all work:
/jev flash deepseek/deepseek-v4.1-flash
/jev coder https://openrouter.ai/qwen/qwen3-coder
/jev cheap DeepSeek: DeepSeek V4.1 Flash
jev-pilot looks the model up in OpenRouter's live list, adds it and says what it is: coder is now qwen/qwen3-coder (Qwen: Qwen3 Coder 480B A35B · 262k context · $0.3 in · $1 out per million tokens). Then it checks the model answers.
-, starting with a letter. Words /jev already uses (status, mode, skills…) can't be names.~/.claude/jev-pilot/models.json, which every session reads, in any project and whichever way jev-pilot is installed. A session that's already open takes up a change at its next prompt./jev <name> <model> add a model, or replace the one under that name
/jev <name> that model, and the models you added before (to switch back)
/jev remove <name> delete it, from every session and from /model (also: /jev <name> off)
/jev status the mode, your models, and a health check of every worker
/jev mode junior-lead standard · budget · junior-lead · second-opinion · quality
/jev junior <name> which model the junior runs on (else the first you added)
/jev reviewer opencode which agent reviews: codex or opencode
Changes apply from the next turn. The
hooks/jev-pilot.ts 40 lines1/**
2 * jev-pilot — the plugin's one hooks module.
3 *
4 * Claude Code loads a single hooks module per plugin, so this entry registers
5 * both mods on the same `on` and the same options: the model router first,
6 * then the skill suggester. Each keeps its own handlers and state; on the
7 * events both hook (`prompt.submit`), the router's handler runs first and
8 * hands the prompt on to the suggester's.
9 *
10 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 and Claude Code >= 2.1.278.
11 */
12import type { Register } from 'claude-code'
13import { register as registerModelRouter } from './jev-model-router.ts'
14import { register as registerSkillSuggestion } from './jev-skill-suggestion.ts'
15import { register as registerPet } from './jev-pet.tsx'
16import { initFeatures } from './features.ts'
17import { initCrew } from './crew-state.ts'
18
19export const register: Register = (on, options) => {
20 // Every part's default comes from the options; /jev switches them live.
21 const flag = (key: string, fallback: boolean) => (typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback)
22 const display = typeof options.display === 'string' ? options.display : 'pet'
23 initFeatures({
24 effort: flag('routeMainEffort', true),
25 raise: typeof options.escalateAfterErrors === 'number' ? options.escalateAfterErrors > 0 : true,
26 subagents: flag('routeSubagentModel', true),
27 skills: flag('suggestSkills', true),
28 strategy: flag('suggestStrategy', true),
29 quality: flag('qualityAdvice', true),
30 design: flag('designPack', true),
31 model: flag('routeMainModel', false),
32 pet: display === 'pet' || display === 'both',
33 })
34 // The crew's defaults: mode, custom model slots, junior, reviewer.
35 initCrew(options)
36 registerModelRouter(on, options)
37 registerSkillSuggestion(on, options)
38 registerPet(on, options)
39}
40hooks/jev-model-router.ts 1177 lines1/**
2 * jev-model-router — Claude Mod (EARLY ACCESS)
3 *
4 * Picks the model each task runs on with TypeSafe's Jev, a System One
5 * decision model: unstructured state in, a typed choice with a probability
6 * distribution out.
7 *
8 * Jev is reached one of three ways, whichever key is configured: TypeSafe's
9 * own API (`typesafeApiKey`) or OpenRouter's Decisions API
10 * (`openrouterApiKey`), which report a calibrated confidence per answer, or
11 * the Vercel AI Gateway (`gatewayApiKey`), which does not. With none, the engine's own `$.model.classify` stands in, so the mod is
12 * useful without any account.
13 *
14 * Three things it can set, each on its own switch:
15 * agent.spawn — the model of each subagent (on by default)
16 * turn.step — the reasoning effort of the main loop (on by default)
17 * turn.step — the model of the main loop (off by default: switching
18 * models mid-session invalidates the prompt cache, which can
19 * cost more than the cheaper tier saves)
20 *
21 * Every one of them moves in both directions: a task the decision model reads
22 * as mechanical is routed down, one it reads as hard is routed up. The two
23 * mistakes do not cost the same, so they do not clear the same confidence bar
24 * (see `minUpgradeConfidence` / `minDowngradeConfidence` in policy.ts).
25 *
26 * The Agent tool has no effort parameter, so a subagent's effort is not ours
27 * to set; only its model is.
28 *
29 * The prompt is classified at `prompt.submit`, which runs before the turn
30 * starts, and the decision is applied at the turn's first request. The
31 * decision model reads the prompt with the last few messages before it
32 * (text and tool names only; see context.ts), so a follow-up such as "yes, do
33 * it" is read as the work it continues, not as a trivial message.
34 *
35 * The same request asks how the work should be carried out (`strategy`):
36 * directly, by one subagent, by parallel subagents, or as a small graph of
37 * subagents in waves. Anything but `direct`, answered confidently and
38 * consistent with the tier, is attached to the prompt as advice
39 * (`<execution_strategy>`); the main model decides whether it fits.
40 *
41 * Within a turn the effort holds, with one exception: when tool calls keep
42 * failing (`escalateAfterErrors` in a row, counted at `tool.call`), the
43 * effort goes up at least one rung, once per turn, as far as a fresh reading
44 * of the decision model says.
45 *
46 * Every failure path is fail-open: a classification that errors or runs past
47 * the latency budget leaves the request exactly as the engine built it.
48 *
49 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
50 * or "gatewayApiKey"). Never hardcode it in this file.
51 *
52 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (Claude Code >= 2.1.259). Typed
53 * against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
54 *
55 * Privacy: with a key set, the prompt text is sent to whichever backend the
56 * key belongs to.
57 */
58import type { HttpInit, HttpResponse, Register } from 'claude-code'
59import { NOT_A_TASK, platformsOf, recentContext, signalsOf, SKIPPED_DIRS } from './context.ts'
60import { clearSkillNotes, resetBriefing, takeBriefing, takeSkill, turnLine } from './summary.ts'
61import { moodOf, say, setBoost, subagentLabel, turnSpeech } from './pet-art.ts'
62import { feature } from './features.ts'
63import { crewNote, JUNIOR_AGENT, juniorSlot, REVIEWERS, reviewerAgent, slotAlias, slotsOffered } from './crew.ts'
64import { crew, reviewerHealthy, router, slotUsable } from './crew-state.ts'
65import { ensureCrew, newFallbacks, startSession, type CrewIo } from './crew-run.ts'
66import { isContinuation, offerPart, stillOffered, type RouterPart } from './jev-call.ts'
67import { recordCorrection, recordSubagent, recordSubagentUsage, recordTurn, resetStats, shortStats } from './session-stats.ts'
68import type { ContextMessage } from './context.ts'
69import { appendEntry, configKeysOf, entriesOf, LEDGER_KEY, markCorrected, reportPrompt, suggestions, summarize } from './ledger.ts'
70import { dueToPropose, effective, initTuning, setTuning, TUNING_KEY, tuningLoaded, tuningOf } from './tuning.ts'
71import type { LedgerEntry, TunableConfig } from './ledger.ts'
72import { CHECK_AT_LOW, effortClearsCache, QUALITY_BARS, qualityAdvice, spinOf, stepBackNote } from './model-router.policy.ts'
73import {
74 adviseStrategy,
75 EFFORT_ORDER,
76 effortLevel,
77 effortScoreOf,
78 engineMoved,
79 isFollowUp,
80 escalate,
81 DEFAULT_BASE_URL,
82 DEFAULT_MODEL,
83 describeDecision,
84 describeSetup,
85 describeStatus,
86 capabilityNote,
87 endpoint,
88 missOf,
89 pendingDecisions,
90 readDecision,
91 selectProvider,
92 requestBody,
93 requestHeaders,
94 modelIds,
95 requestModelId,
96 route,
97 TIER_ORDER,
98} from './model-router.policy.ts'
99import type { Miss, SlotChoice } from './model-router.policy.ts'
100import type { Decision, Effort, PolicyConfig, Provider, StrategyConfig, Tier } from './model-router.policy.ts'
101
102/** Where decisions are asked, when a backend is configured. */
103interface Backend {
104 provider: Provider
105 url: string
106 apiKey: string
107 modelId: string
108 timeoutMs: number
109}
110
111/**
112 * The engine calls the helpers below need. `$` is never handed to a helper:
113 * each hook builds this at its own call site, spelling every call on `$`
114 * there (see `io` in register).
115 */
116interface Io {
117 fetch: (url: string, init: HttpInit) => Promise<HttpResponse>
118 sleep: (ms: number) => Promise<void>
119 log: (text: string) => unknown
120 /** A line for the verbose log only: what is routine, not an error. */
121 detail: (text: string) => unknown
122 messages: () => Promise<readonly ContextMessage[]>
123}
124
125/**
126 * One request to the backend, read as a decision; null without a backend, or
127 * on timeout, error, a non-2xx or an unreadable answer, with why (`miss`).
128 * Every caller treats a null decision the same way: the request goes on as
129 * the engine built it. A timeout or a busy backend is routine (the pet says
130 * so); only an error is logged whatever the log level.
131 */
132async function askJev(
133 io: Io,
134 backend: Backend,
135 state: Record<string, unknown>,
136 withStrategy: boolean,
137 what: string,
138 subagent = false,
139 slots: readonly SlotChoice[] = [],
140 junior = false,
141 extra: Record<string, unknown> = {},
142 quality = false,
143): Promise<{ text: string | null; miss: Miss | null }> {
144 try {
145 const response = await Promise.race([
146 io.fetch(backend.url, {
147 method: 'POST',
148 headers: requestHeaders(backend.provider, backend.apiKey, backend.modelId),
149 body: requestBody(backend.provider, state, backend.modelId, withStrategy, subagent, slots, junior, extra, quality),
150 }),
151 io.sleep(backend.timeoutMs),
152 ])
153 if (response && response.ok) return { text: response.text, miss: null }
154 const miss = missOf(response ? response.status : null)
155 const tell = miss === 'error' ? io.log : io.detail
156 if (response) await tell(`[jev-model-router] ${backend.provider} responded ${response.status}: ${response.text.slice(0, 200)}`)
157 else await tell(`[jev-model-router] classification passed ${backend.timeoutMs}ms; leaving ${what} alone`)
158 return { text: null, miss }
159 } catch (error) {
160 await io.log(`[jev-model-router] classification failed: ${String(error)}`)
161 return { text: null, miss: 'error' }
162 }
163}
164
165async function classify(
166 io: Io,
167 backend: Backend | null,
168 state: Record<string, unknown>,
169 withStrategy: boolean,
170 what: string,
171 subagent = false,
172 slots: readonly SlotChoice[] = [],
173 junior = false,
174): Promise<{ decision: Decision | null; miss: Miss | null }> {
175 if (!backend) return { decision: null, miss: null }
176 const { text, miss } = await askJev(io, backend, state, withStrategy, what, subagent, slots, junior)
177 if (text === null) return { decision: null, miss }
178 const decision = readDecision(text, slots.map((slot) => slot.name))
179 return { decision, miss: decision ? null : 'error' }
180}
181
182/** The conversation so far, or none when not `wanted` or unreadable. */
183async function readMessages(io: Io, wanted: boolean): Promise<readonly ContextMessage[]> {
184 if (!wanted) return []
185 try {
186 return await io.messages()
187 } catch (error) {
188 await io.log(`[jev-model-router] could not read the conversation: ${String(error)}`)
189 return []
190 }
191}
192
193/** What prompt.submit knows of a turn's decision, waiting for the turn to start. */
194type Draft = Pick<
195 LedgerEntry,
196 'answered' | 'ms' | 'tier' | 'tierConfidence' | 'effortLevel' | 'effortConfidence' | 'strategy' | 'strategyConfidence' | 'advised'
197>
198
199export const register: Register = (on, options) => {
200 const text = (key: string, fallback: string) =>
201 typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
202 const number = (key: string, fallback: number) =>
203 typeof options[key] === 'number' ? (options[key] as number) : fallback
204 const flag = (key: string, fallback: boolean) =>
205 typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
206
207 // With several keys set, `auto` takes TypeSafe's own API, then OpenRouter,
208 // then the Gateway: the first two report the calibrated confidence the
209 // policy's threshold reads. `provider` forces one, "builtin" uses none.
210 const typesafeKey = text('typesafeApiKey', '')
211 const gatewayKey = text('gatewayApiKey', '')
212 const openrouterKey = text('openrouterApiKey', '')
213 const forced = text('provider', 'auto')
214 const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey, openrouterKey)
215
216 // Each backend keeps its own URL and model, so an override written for one
217 // can never be sent to the other when `auto` picks differently than expected.
218 const apiKey = !active ? '' : { typesafe: typesafeKey, gateway: gatewayKey, openrouter: openrouterKey }[active]
219 const modelId = !active ? '' : text(`${active}Model`, DEFAULT_MODEL[active])
220 const url = !active ? '' : endpoint(active, text(`${active}BaseUrl`, DEFAULT_BASE_URL[active]))
221
222 // A backend named in the options but missing its key degrades to the
223 // built-in classifier, which is silent; say so once, when a hook first runs.
224 let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
225
226 const timeoutMs = number('timeoutMs', 1500)
227 // Each part reads its switch live (features.ts): /jev turns it on or off.
228 const routeSubagentModel = () => feature('subagents')
229 const routeMainEffort = () => feature('effort')
230 const routeMainModel = () => feature('model')
231 const routeMainLoop = () => routeMainEffort() || routeMainModel()
232 const logDecisions = flag('logDecisions', true)
233 // Where jev-pilot talks: the pet at the bottom right (default), one line
234 // per turn in the transcript, both, or nowhere. verboseLog adds every step
235 // to the transcript whatever this says.
236 const display = text('display', 'pet')
237 const petOn = () => feature('pet')
238 const verbose = logDecisions && flag('verboseLog', false)
239 const lines = logDecisions && (verbose || display === 'transcript' || display === 'both')
240 const readyLine = () =>
241 verbose
242 ? `[jev-model-router] ${describeSetup(
243 active,
244 url,
245 { subagentModel: routeSubagentModel(), mainEffort: routeMainEffort(), mainModel: routeMainModel() },
246 forced === 'builtin',
247 )}`
248 : `jev-pilot · ready on ${active ?? `the built-in classifier${forced === 'builtin' ? '' : ' (no key set)'}`}`
249
250 // A reasoning level from the options; a value off the ladder is not guessed
251 // at and reads as the default.
252 const effortOption = (key: string, fallback: Effort): Effort => {
253 const value = text(key, fallback)
254 return (EFFORT_ORDER as readonly string[]).includes(value) ? (value as Effort) : fallback
255 }
256 // Two ceilings: where a turn may start, and how far trouble may raise it.
257 // Starting lower and raising only on evidence keeps `max` for the turns
258 // that show they need it.
259 const raisedCeiling: { maxEffort: Effort } = { maxEffort: effortOption('maxRaisedEffort', 'max') }
260 const policy: PolicyConfig = {
261 tiers: {
262 fast: text('fastModel', 'haiku'),
263 balanced: text('balancedModel', 'sonnet'),
264 deep: text('deepModel', 'opus'),
265 },
266 minUpgradeConfidence: number('minUpgradeConfidence', 0.3),
267 minHighConfidence: number('minHighConfidence', 0.5),
268 minDowngradeConfidence: number('minDowngradeConfidence', 0.6),
269 // Turns start at xhigh at most, and only when Jev is 60% sure a task is
270 // very hard. Anthropic warns that on Opus 5.5 xhigh thinks a lot more;
271 // replayed on 76 labelled prompts, every turn that started at xhigh
272 // needed it (4 of 4), and a high ceiling only started those four too
273 // low (exact 39 → 35). max is still reached only by the mid-turn raise.
274 maxEffort: effortOption('maxEffort', 'xhigh'),
275 // A near tie between two effort levels takes the higher one.
276 closeMargin: Math.max(0, number('effortCloseMargin', 0.15)),
277 }
278 let margin = policy.closeMargin ?? 0
279 // Each turn's decision and outcome, kept for /jev-pilot:report.
280 const recordDecisions = flag('recordDecisions', true)
281
282 // How much of the conversation the decision model reads beside a prompt.
283 const contextLimits = {
284 messages: Math.max(0, Math.round(number('contextMessages', 4))),
285 chars: Math.max(0, number('contextChars', 2000)),
286 }
287 const suggestStrategy = () => feature('strategy')
288 const strategyConfig: StrategyConfig = {
289 minConfidence: number('minStrategyConfidence', 0.6),
290 minGraphConfidence: number('minGraphConfidence', 0.8),
291 graphSkill: text('graphSkill', ''),
292 }
293 const escalateAfterErrors = Math.max(0, Math.round(number('escalateAfterErrors', 2)))
294 // A subagent's effort, set at its requests from the decision made when it
295 // started (the Agent tool itself takes none); under the subagents switch.
296 const subagentEffortOn = flag('routeSubagentEffort', true)
297 const routeSubagentEffort = () => routeSubagentModel() && subagentEffortOn
298 // Each subagent's decision, by the id core gives it when it starts; its
299 // effort, once its first request has settled it.
300 const subagents = new Map<string, { decision: Decision; label: string; model: string | null }>()
301 const subagentEffort = new Map<string, Effort | null>()
302 const MAX_SUBAGENTS = 64
303 /** What is switched on, for the note that tells the model. */
304 const capabilities = () => ({
305 effort: routeMainEffort(),
306 raise: routeMainEffort() && feature('raise') && escalateAfterErrors > 0,
307 subagents: routeSubagentModel(),
308 subagentEffort: routeSubagentEffort(),
309 skills: feature('skills'),
310 strategy: suggestStrategy(),
311 model: routeMainModel(),
312 quality: feature('quality'),
313 })
314
315 const backend: Backend | null = active ? { provider: active, url, apiKey, modelId, timeoutMs } : null
316 // What the ledger may tune (`/jev tune`): the settings' values, and how a
317 // learned change goes into force, here, at once.
318 initTuning(
319 {
320 timeoutMs,
321 minDowngradeConfidence: policy.minDowngradeConfidence,
322 effortCloseMargin: margin,
323 minHighConfidence: policy.minHighConfidence ?? 0.5,
324 },
325 (tuned: TunableConfig) => {
326 policy.minDowngradeConfidence = tuned.minDowngradeConfidence
327 policy.minHighConfidence = tuned.minHighConfidence
328 policy.closeMargin = tuned.effortCloseMargin
329 margin = tuned.effortCloseMargin
330 if (backend) backend.timeoutMs = tuned.timeoutMs
331 },
332 )
333
334 // The classification waiting for the turn that reads its prompt, and what
335 // the current turn settled on. Both are single slots: main-loop turns run
336 // one at a time, so nothing accumulates over a long session. `pending`
337 // reports no decision when two prompts are waiting at once, rather than
338 // routing a turn on a decision made for a different prompt.
339 // Each classified prompt waits with its decision and its ledger draft, so
340 // the turn that reads it knows which prompt it is working on.
341 const pending = pendingDecisions<{ decision: Decision | null; prompt: string; draft: Draft; miss: Miss | null }>()
342 // Said once, the first time a hook runs. A router that loaded and one that
343 // never loaded are otherwise told apart only by the absence of later lines,
344 // and absence is not evidence: the policy leaves most turns alone anyway.
345 let announced = false
346 let appliedTurnId: string | undefined
347 let applied: { model?: string; effort?: Effort } | null = null
348 // The model the engine named for the current turn's first request, before
349 // any rewrite: a later request naming another is the engine's fallback.
350 let turnEngineModel: string | null = null
351 // The prompt the current turn works on (null when its decision was
352 // withheld), for a re-reading mid-turn, and the turn whose effort was
353 // already raised: at most once each.
354 let turnPrompt: string | null = null
355 let escalatedTurnId: string | undefined
356 // The main loop's tool calls that failed in a row since its last success,
357 // counted as they finish (tool.call) and cleared when a turn starts.
358 let failedInARow = 0
359 // Said once: a custom model asked for with no router to serve it.
360 let warnedNoRouter = false
361 // Where changing the effort clears the cache (Bedrock, Google Cloud, a
362 // gateway): the effort chosen for the first turn is held for the session,
363 // and chosen again after a compaction (which rewrites the cache anyway).
364 const effortChanges = text('effortChanges', 'auto')
365 let holdEffort: boolean | null = effortChanges === 'hold' ? true : effortChanges === 'per-turn' ? false : null
366 let heldEffort: Effort | null | undefined
367 // What the project deploys with, found once per project folder.
368 let platformCache: { cwd: string; text: string } | null = null
369 // The turn id of the ledger entry this session finished last: the next
370 // prompt says whether it was right.
371 let lastEntryId: string | null = null
372 // A turn going in circles: each file's edits and each command's runs in
373 // the turn, and whether that was already said.
374 const edits = new Map<string, number>()
375 const runs = new Map<string, number>()
376 let spinning: string | null = null
377 let spunTurnId: string | undefined
378 // The last decision Jev made for a typed prompt: "continue" goes on with it.
379 let lastDecision: Decision | null = null
380 // Each family's current full id, learned from the requests the engine makes.
381 const ids = modelIds()
382 // The ledger: the turn in progress (its draft waits in `pending`).
383 let current: (LedgerEntry & { turnId: string }) | null = null
384
385 // Said as soon as the session opens, so a loaded jev-pilot is visible
386 // before the first prompt: a line in the transcript and one under the
387 // prompt. (The per-prompt lines follow once prompts arrive.)
388 on('session.start', async ($, e, next) => {
389 const result = await next(e)
390 // The crew: saved /jev changes, the router claude-jev started, the slots
391 // file it reads; then every worker's health, in the background.
392 const crewIo: CrewIo = {
393 fetch: (url, init) => $.http.fetch(url, init),
394 home: () => $.env.get('HOME'),
395 routerUrl: () => $.env.get('JEV_ROUTER_URL'),
396 write: (path, text) => $.fs.write(path, text),
397 read: async (path) => ((await $.fs.exists(path)) ? $.fs.read(path) : null),
398 run: (argv, timeoutMs) => $.process.run(argv, { timeoutMs }),
399 storeGet: (key) => $.store.get(key),
400 sleep: (ms) => $.clock.sleep(ms),
401 register: async (spec) => {
402 await $.agent.register(spec)
403 },
404 }
405 await startSession(crewIo, openrouterKey || null)
406 announced = true
407 if (lines) {
408 $.ui.log(readyLine())
409 $.ui.status(`jev · ready on ${active ?? 'the built-in classifier'}`)
410 }
411 if (petOn()) {
412 say(`ready · ${active ?? (forced === 'builtin' ? 'built-in' : 'no key, built-in')}`, 'ready')
413 $.ui.invalidate('ui.render')
414 }
415 return result
416 })
417
418 on('prompt.submit', async ($, e, next) => {
419 const io: Io = {
420 fetch: (url, init) => $.http.fetch(url, init),
421 sleep: (ms) => $.clock.sleep(ms),
422 log: (text) => $.ui.log(text),
423 detail: (text) => (verbose ? $.ui.log(text) : undefined),
424 messages: () => $.session.messages(),
425 }
426 // Before the routing guards: a module whose switches are all off has still
427 // loaded, and that is exactly when its silence is most misleading.
428 if (!announced) {
429 announced = true
430 if (lines) $.ui.log(readyLine())
431 }
432 const crewIo: CrewIo = {
433 fetch: (url, init) => $.http.fetch(url, init),
434 home: () => $.env.get('HOME'),
435 routerUrl: () => $.env.get('JEV_ROUTER_URL'),
436 write: (path, text) => $.fs.write(path, text),
437 read: async (path) => ((await $.fs.exists(path)) ? $.fs.read(path) : null),
438 run: (argv, timeoutMs) => $.process.run(argv, { timeoutMs }),
439 storeGet: (key) => $.store.get(key),
440 sleep: (ms) => $.clock.sleep(ms),
441 register: async (spec) => {
442 await $.agent.register(spec)
443 },
444 }
445 await ensureCrew(crewIo, openrouterKey || null)
446 // The tuning learned from the ledger, once per worker (a reload starts over).
447 if (!tuningLoaded()) setTuning(tuningOf(await $.store.get(TUNING_KEY).catch(() => undefined)))
448 // Only the person's own tasks are classified. A notification, a peer's
449 // message or a typed `/command` would otherwise take the pending slot and
450 // leave the next real prompt's turn without its decision.
451 const isTask = !!e.text.trim() && !/^\/\S/.test(e.text.trim()) && !(e.origin && NOT_A_TASK.has(e.origin.kind))
452 if (!isTask) return next(e)
453 // Once per session, and again after a compaction: what jev-pilot does,
454 // for the model, so it leaves those decisions to it.
455 const note = takeBriefing()
456 ? capabilityNote(
457 capabilities(),
458 crewNote(crew(), juniorSlot(crew(), router() !== null), REVIEWERS.filter(reviewerHealthy), slotsOffered(crew(), router() !== null).filter(slotUsable), routeSubagentModel()),
459 )
460 : null
461 const withNote = (input: typeof e, extra: string | null = null) => {
462 const blocks = [extra, note].filter((b): b is string => b !== null)
463 return blocks.length > 0 ? { ...input, context: [...(input.context ?? []), ...blocks] } : input
464 }
465 if (!routeMainLoop() && !suggestStrategy() && !feature('quality')) {
466 const passed = await next(withNote(e))
467 if (passed.drop && note) resetBriefing()
468 return passed
469 }
470 const planning = suggestStrategy()
471
472 if (!unusableReported) {
473 unusableReported = true
474 $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
475 }
476
477 const startedAt = await $.clock.now()
478 const messages = await readMessages(io, contextLimits.messages > 0)
479 const recent = recentContext(messages, e.text, contextLimits)
480 // What the answer settles, however it arrives: the strategy advice, the
481 // ledger's draft, and the decision waiting for the turn.
482 const finish = async (decision: Decision | null, miss: Miss | null, ms: number, reused = false): Promise<string | null> => {
483 if (decision && !reused) lastDecision = decision
484 // What the decision model actually answered, whatever the policy then
485 // does with it. This is the line that proves the classification ran.
486 if (verbose) {
487 const read = recent ? ` · read ${recent.split('\n').length} recent messages` : ''
488 $.ui.log(`[jev-model-router] jev: ${reused ? 'continuing, the last decision kept' : describeDecision(decision, ms, margin)}${read}`)
489 }
490 let block: string | null = null
491 let advisedStrategy = false
492 if (planning && !reused) {
493 const advice = adviseStrategy(decision, strategyConfig)
494 block = advice.block
495 advisedStrategy = block !== null
496 if (verbose) $.ui.log(`[jev-model-router] strategy: ${block ? 'advising ' : ''}${advice.reason}`)
497 }
498 // What makes the work better: ask first, test the bug first, check a costly change.
499 if (feature('quality') && !reused) {
500 const quality = qualityAdvice(decision)
501 if (quality) {
502 block = [block, quality].filter((b): b is string => b !== null).join('\n\n')
503 if (verbose) $.ui.log(`[jev-model-router] quality: ${quality.split('\n').filter((l) => l.startsWith('- ')).map((l) => l.slice(2, 40)).join(' · ')}`)
504 }
505 }
506 // Your reply said the last turn got it wrong: that turn is marked in the
507 // ledger, which is how jev-pilot learns where too little effort costs you.
508 if (!reused && typeof decision?.corrects === 'number' && lastEntryId !== null) {
509 const id = lastEntryId
510 lastEntryId = null
511 const corrected = decision.corrects >= QUALITY_BARS.corrects
512 if (corrected) recordCorrection()
513 try {
514 await $.store.set(LEDGER_KEY, markCorrected(await $.store.get(LEDGER_KEY), id, corrected))
515 } catch (error) {
516 if (verbose) $.ui.log(`[jev-model-router] could not mark the last turn: ${String(error)}`)
517 }
518 }
519 const draft: Draft = {
520 answered: decision !== null,
521 ms: active && !reused ? Math.round(ms) : null,
522 tier: decision?.tier ?? null,
523 tierConfidence: decision?.confidence ?? null,
524 effortLevel: decision ? effortScoreOf(decision, margin) : null,
525 effortConfidence: decision?.effortConfidence ?? null,
526 strategy: reused ? null : (decision?.strategy ?? null),
527 strategyConfidence: reused ? null : (decision?.strategyConfidence ?? null),
528 advised: advisedStrategy,
529 }
530 pending.put({ decision, prompt: e.text, draft, miss })
531 return block
532 }
533 const refused = (result: Awaited<ReturnType<typeof next>>) => {
534 // Refused further down: no turn will read this decision.
535 if (result.drop) {
536 pending.withdraw()
537 if (note) resetBriefing()
538 }
539 return result
540 }
541
542 // "continue": the work in progress goes on as the last turn decided.
543 if (active && lastDecision && isContinuation(e.text)) {
544 await finish(lastDecision, null, 0, true)
545 return refused(await next(withNote(e)))
546 }
547
548 if (active && backend) {
549 // One request for the prompt: the skill module, further down, adds its
550 // questions to these and sends them together (jev-call.ts).
551 const junior = (() => {
552 const slot = juniorSlot(crew(), router() !== null)
553 return !!slot && slotUsable(slot)
554 })()
555 // What the project deploys with, from its file names (once per folder):
556 // a platform's skill fits only a project on that platform.
557 let platforms = platformCache
558 try {
559 const cwd = await $.session.cwd()
560 if (platforms?.cwd !== cwd) {
561 const paths: string[] = []
562 const top = await $.fs.list(cwd)
563 for (const entry of top) paths.push(entry.kind === 'dir' ? `${entry.name}/` : entry.name)
564 const dirs = top.filter((entry) => entry.kind === 'dir' && !SKIPPED_DIRS.has(entry.name)).slice(0, 60)
565 for (const dir of dirs) {
566 for (const entry of await $.fs.list(`${cwd}/${dir.name}`).catch(() => [])) {
567 paths.push(`${dir.name}/${entry.kind === 'dir' ? `${entry.name}/` : entry.name}`)
568 }
569 }
570 platforms = { cwd, text: platformsOf(paths) }
571 platformCache = platforms
572 }
573 } catch {
574 platforms = null
575 }
576 const state = { prompt: e.text, recent_context: recent, signals: signalsOf(e.text, messages), ...(platforms ? { project_platforms: platforms.text } : {}) }
577 let settled = false
578 let settledBlock: string | null = null
579 const part: RouterPart = {
580 prompt: e.text,
581 ask: async (extra) => {
582 const asked = await askJev(io, backend, state, planning, 'the turn', false, [], junior, extra, feature('quality'))
583 return { ...asked, ms: (await $.clock.now()) - startedAt }
584 },
585 // Once: a second settle (never expected) gets the first one's block.
586 settle: async (answer) => {
587 if (settled) return settledBlock
588 settled = true
589 const decision = answer.text === null ? null : readDecision(answer.text)
590 settledBlock = await finish(decision, answer.text !== null && !decision ? 'error' : answer.miss, answer.ms)
591 return settledBlock
592 },
593 }
594 offerPart(part)
595 const result = await next(withNote(e))
596 // Nobody took it (the skill module never ran): asked now, before the
597 // turn starts; too late for advice, in time for the effort.
598 stillOffered(part)
599 if (!settled) await part.settle(await part.ask({}))
600 return refused(result)
601 }
602
603 // No backend: the engine's own small-model classifier answers the same
604 // question, without the confidence the policy's threshold reads, and
605 // without a strategy (it answers one label).
606 let decision: Decision | null = null
607 try {
608 const input = recent ? `Recent conversation:\n${recent}\n\nLatest request:\n${e.text}` : e.text
609 const label = await $.model.classify(input, TIER_ORDER)
610 if (label) {
611 decision = {
612 tier: label as Tier,
613 confidence: null,
614 risky: null,
615 effort: null,
616 effortConfidence: null,
617 }
618 }
619 } catch (error) {
620 $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
621 }
622 const block = await finish(decision, null, (await $.clock.now()) - startedAt)
623 // Attached on the way down: one block after the prompt as typed, read by
624 // the model and never shown to the person.
625 return refused(await next(withNote(e, block)))
626 })
627
628 on('turn.step', async function* ($, e, next) {
629 const io: Io = {
630 fetch: (url, init) => $.http.fetch(url, init),
631 sleep: (ms) => $.clock.sleep(ms),
632 log: (text) => $.ui.log(text),
633 detail: (text) => (verbose ? $.ui.log(text) : undefined),
634 messages: () => $.session.messages(),
635 }
636 // Every request names the id the engine resolved for it, a subagent's
637 // included: that is where the main loop's switch finds its ids.
638 ids.learn(e.model)
639 // A custom model where no router serves it (a plain `claude` session
640 // whose default was set to jev-…): Sonnet instead of a request that fails.
641 if (/^jev-/.test(e.model ?? '') && router() === null) {
642 const id = requestModelId(policy.tiers.balanced, ids)
643 if (id) {
644 if (!warnedNoRouter) {
645 warnedNoRouter = true
646 $.ui.log(`jev · ${e.model} needs claude-jev (its router isn't running here), so this session uses ${id}. /model default sets your default back.`)
647 }
648 return yield* next({ ...e, model: id })
649 }
650 }
651 // A subagent's request: the effort it was routed to when it started,
652 // settled at its first request (from the effort the engine built it with)
653 // and kept for the rest of its run.
654 if (e.agentId) {
655 const agentId = e.agentId
656 let effort = subagentEffort.get(agentId)
657 const known = subagents.get(agentId)
658 // A model without effort (the engine left it unset) is left that way.
659 if (effort === undefined && known && routeSubagentEffort() && e.effort !== undefined) {
660 const routing = route(known.decision, { model: e.model, effort: e.effort }, policy)
661 effort = routing.effort
662 subagentEffort.set(agentId, effort)
663 if (effort && lines) {
664 $.ui.log(verbose ? `[jev-model-router] ${known.label} → effort ${effort}: ${routing.reason}` : `jev · subagent ${known.label} → effort ${effort}`)
665 }
666 if (effort && petOn()) {
667 say(`${known.label} → ${[known.model, effort].filter(Boolean).join(' · ')}`, 'focused')
668 $.ui.invalidate('ui.render')
669 }
670 }
671 return yield* next(effort && routeSubagentEffort() ? { ...e, effort } : e)
672 }
673 if (!routeMainLoop()) return yield* next(e)
674
675 // Every request after the first reuses what the turn settled on, so
676 // neither the model nor the effort changes under its own tool loop —
677 // unless the loop is visibly struggling, and then only the effort, up.
678 if (e.index > 0 && e.turnId === appliedTurnId) {
679 if (routeMainEffort() && feature('raise') && escalateAfterErrors > 0 && escalatedTurnId !== e.turnId && !holdEffort) {
680 const failed = failedInARow
681 // Going in circles counts as struggling too, however the calls ended.
682 const circling = spinning
683 if (failed >= escalateAfterErrors || circling) {
684 spinning = null
685 escalatedTurnId = e.turnId
686 const effort = applied?.effort ?? e.effort
687 const turn = current
688 // The turn's own prompt: without one (its decision was withheld),
689 // there is nothing to re-read, and the raise is the one rung.
690 const prompt = turnPrompt
691 const messages = prompt === null ? [] : await readMessages(io, contextLimits.messages > 0)
692 const reread =
693 prompt === null
694 ? null
695 : (
696 await classify(
697 io,
698 backend,
699 {
700 prompt,
701 recent_context: recentContext(messages, prompt, contextLimits),
702 signals: signalsOf(prompt, messages),
703 trouble: circling ?? `${failed} tool calls in a row have failed while working on this request`,
704 },
705 false,
706 'the effort',
707 )
708 ).decision
709 const level = reread ? effortScoreOf(reread, margin) : null
710 const raised = escalate(effort, Math.max(failed, circling ? escalateAfterErrors : 0), escalateAfterErrors, level, raisedCeiling)
711 if (raised) {
712 applied = { ...(applied ?? {}), effort: raised }
713 if (turn && turn.turnId === e.turnId) turn.raisedTo = raised
714 if (lines) {
715 $.ui.log(
716 verbose
717 ? `[jev-model-router] main loop → effort ${raised}: ${circling ?? `${failed} tool calls failed in a row`}`
718 : `jev · ${circling ? 'going in circles' : `${failed} failed in a row`} → effort ${raised}`,
719 )
720 $.ui.status(`jev · struggling → ${raised}`)
721 }
722 if (petOn()) {
723 // At max, the pilot pulls its goggles down for the rest of the turn.
724 if (raised === 'max') setBoost(true)
725 say(`${circling ? 'circles' : `${failed} fails`} → ${raised} ✈`, raised === 'max' ? 'boost' : moodOf(raised))
726 $.ui.invalidate('ui.render')
727 }
728 } else if (verbose) {
729 $.ui.log(`[jev-model-router] ${failed} tool calls failed in a row; effort ${String(effort)} kept`)
730 }
731 }
732 }
733 // The engine moved the turn to another model (its overload fallback):
734 // that model stands for the rest of the turn; the routed effort stays.
735 if (applied?.model && engineMoved(turnEngineModel, e.model)) {
736 const { model: _dropped, ...rest } = applied
737 applied = Object.keys(rest).length > 0 ? rest : null
738 if (verbose) $.ui.log(`[jev-model-router] main loop: the engine moved the turn to ${e.model}; that model stands`)
739 if (lines) $.ui.status(`jev · engine moved to ${e.model}${applied?.effort ? ` · effort ${applied.effort}` : ''}`)
740 }
741 return yield* next(applied ? { ...e, ...applied } : e)
742 }
743
744 // A new turn: failures of the last one say nothing about this one.
745 failedInARow = 0
746 edits.clear()
747 runs.clear()
748 spinning = null
749 spunTurnId = undefined
750 const taken = pending.take()
751 const decision = taken?.decision ?? null
752 turnPrompt = taken?.prompt ?? null
753 const routing = route(decision, { model: e.model, effort: e.effort }, policy, { noLowering: taken ? isFollowUp(taken.prompt) : false })
754 const change: { model?: string; effort?: Effort } = {}
755 // The main loop's `model` is sent to the API as written, so an alias
756 // becomes the id the engine was seen using for it; a subagent's
757 // (agent.spawn) may stay an alias.
758 if (routeMainModel() && routing.model) {
759 const id = requestModelId(routing.model, ids)
760 if (id) change.model = id
761 else if (verbose) {
762 $.ui.log(`[jev-model-router] no ${routing.model} model seen yet this session; model left as ${e.model}`)
763 }
764 }
765 if (holdEffort === null) {
766 holdEffort = effortClearsCache({
767 bedrock: await $.env.get('CLAUDE_CODE_USE_BEDROCK'),
768 vertex: await $.env.get('CLAUDE_CODE_USE_VERTEX'),
769 upstream: (await $.env.get('JEV_ANTHROPIC_UPSTREAM')) ?? (await $.env.get('ANTHROPIC_BASE_URL')),
770 })
771 if (holdEffort && lines) $.ui.log('jev · effort held for this session: here, changing it would clear the cached conversation')
772 }
773 if (holdEffort && heldEffort !== undefined) {
774 // Held: every turn keeps the effort the first one got.
775 if (heldEffort && heldEffort !== e.effort) change.effort = heldEffort
776 } else {
777 if (routeMainEffort() && routing.effort) change.effort = routing.effort
778 if (holdEffort) heldEffort = change.effort ?? (typeof e.effort === 'string' ? (e.effort as Effort) : null)
779 }
780
781 appliedTurnId = e.turnId
782 turnEngineModel = e.model
783 applied = Object.keys(change).length > 0 ? change : null
784 if (recordDecisions) {
785 // A decision withheld (two prompts waiting) is recorded as unanswered:
786 // the turn ran on the engine's own settings.
787 const known = decision ? (taken?.draft ?? null) : null
788 const startEffort = change.effort ?? e.effort
789 current = {
790 turnId: e.turnId,
791 at: await $.clock.now(),
792 answered: known?.answered ?? false,
793 ms: known?.ms ?? null,
794 tier: known?.tier ?? null,
795 tierConfidence: known?.tierConfidence ?? null,
796 effortLevel: known?.effortLevel ?? null,
797 effortConfidence: known?.effortConfidence ?? null,
798 startedFrom: typeof e.effort === 'string' ? e.effort : null,
799 started: typeof startEffort === 'string' ? startEffort : null,
800 raisedTo: null,
801 toolCalls: 0,
802 failures: 0,
803 strategy: known?.strategy ?? null,
804 strategyConfidence: known?.strategyConfidence ?? null,
805 advised: known?.advised ?? false,
806 outcome: null,
807 durationMs: null,
808 outputTokens: null,
809 }
810 }
811 // A row in the transcript scrolls away; this line stays on screen.
812 if (lines) $.ui.status(describeStatus(decision, applied))
813 // The skill module's pick for this turn's prompt, read (and so released)
814 // whatever the log mode.
815 const skillNote = takeSkill(taken?.prompt ?? null)
816 const known = decision ? (taken?.draft ?? null) : null
817 const level = decision ? effortScoreOf(decision, margin) : null
818 // A turn nobody typed (a subagent's or a background task's notification)
819 // was never put to the decision model: the bubble keeps what it said.
820 if (petOn() && taken) {
821 say(
822 turnSpeech({
823 answered: decision !== null,
824 miss: taken?.miss ?? null,
825 applied: change.effort ?? null,
826 current: typeof e.effort === 'string' ? e.effort : null,
827 wanted: level === null ? null : effortLevel(level),
828 confidence: decision?.effortConfidence ?? null,
829 skill: skillNote ? skillNote.skill : undefined,
830 advised: known?.advised ? (known.strategy ?? null) : null,
831 }).text,
832 decision ? moodOf(change.effort ?? (typeof e.effort === 'string' ? e.effort : null)) : 'alert',
833 )
834 $.ui.invalidate('ui.render')
835 }
836 if (lines && !verbose) {
837 // One line for the whole decision: effort, skill, advice, time.
838 $.ui.log(
839 turnLine({
840 answered: decision !== null,
841 confidence: decision?.effortConfidence ?? null,
842 applied: change.effort ?? null,
843 current: typeof e.effort === 'string' ? e.effort : null,
844 wanted: level === null ? null : effortLevel(level),
845 jevMs: known?.ms ?? null,
846 skill: skillNote,
847 advised: known?.advised ? (known.strategy ?? null) : null,
848 }),
849 )
850 }
851
852 if (!applied) {
853 // A turn left alone is the common case, and it used to be silent, which
854 // made a working mod look like one that never loaded. Say what happened.
855 if (verbose) {
856 const suppressed = routing.model && !routeMainModel() ? ' (main-loop model routing off)' : ''
857 $.ui.log(`[jev-model-router] main loop: ${routing.reason}${suppressed}`)
858 }
859 return yield* next(e)
860 }
861 if (verbose) {
862 const what = [change.model, change.effort && `effort ${change.effort}`]
863 .filter(Boolean)
864 .join(', ')
865 $.ui.log(`[jev-model-router] main loop → ${what}: ${routing.reason}`)
866 }
867 return yield* next({ ...e, ...change })
868 })
869
870 // Observation only: the main loop's tool calls, counted as they finish, so
871 // a struggling turn is seen without re-reading the transcript every step,
872 // and so the ledger knows how much each turn did.
873 // A refusal (a denied permission) is the person's choice, not the task
874 // going wrong: it neither counts nor clears the run.
875 on('tool.call', async ($, e, next) => {
876 const result = await next(e)
877 if (!e.agentId && !result.deny) {
878 failedInARow = result.isError ? failedInARow + 1 : 0
879 if (current) {
880 current.toolCalls++
881 current.failures = Math.max(current.failures, failedInARow)
882 }
883 // Going in circles: the same file edited again and again, or the same
884 // command run again and failing. Said once a turn, to the model only.
885 const input = e as unknown as { tool?: string; file_path?: unknown; command?: unknown }
886 const circle = spinOf(input, !!result.isError, edits, runs)
887 if (circle && feature('quality') && current && spunTurnId !== current.turnId) {
888 spunTurnId = current.turnId
889 spinning = circle
890 if (verbose) $.ui.log(`[jev-model-router] going in circles: ${circle}`)
891 if (petOn()) {
892 say('going in circles · stepping back', 'alert')
893 $.ui.invalidate('ui.render')
894 }
895 const now = applied?.effort ?? current?.started ?? null
896 const atTop = now === 'xhigh' || now === 'max'
897 return { ...result, context: [...(result.context ?? []), stepBackNote(circle, atTop)] } as typeof result
898 }
899 }
900 return result
901 })
902
903 // The end of a main-loop turn closes its ledger entry: how it ended, how
904 // long it took, what it produced. Written after the engine's own handling.
905 on('turn.complete', async ($, e, next) => {
906 const result = await next(e)
907 // A subagent that finished: its routing is done with.
908 if (e.agentId) {
909 subagents.delete(e.agentId)
910 subagentEffort.delete(e.agentId)
911 recordSubagentUsage(e.agentId, e.usage as Record<string, unknown> | null | undefined)
912 }
913 // A custom model that failed during the turn: said, since Claude's answer hides it.
914 if (!e.agentId && crew().slots.length > 0 && router()) {
915 const failed = await newFallbacks({
916 fetch: (url, init) => $.http.fetch(url, init),
917 sleep: (ms) => $.clock.sleep(ms),
918 })
919 for (const line of failed) $.ui.log(`jev · ${line}. /jev status shows the details`)
920 if (failed.length > 0 && petOn()) {
921 say('custom model failed · Claude answered', 'alert')
922 $.ui.invalidate('ui.render')
923 }
924 }
925 if (!e.agentId && current && current.turnId === e.turnId) {
926 const { turnId, ...entry } = current
927 current = null
928 // Every tenth turn, the bubble says what the session came to.
929 const turns = recordTurn(entry.startedFrom, entry.started)
930 if (turns % 10 === 0 && petOn()) {
931 say(shortStats(), 'calm')
932 $.ui.invalidate('ui.render')
933 }
934 const finished: LedgerEntry = {
935 ...entry,
936 id: turnId,
937 outcome: e.reason,
938 durationMs: e.durationMs,
939 outputTokens: e.usage?.output_tokens ?? null,
940 }
941 try {
942 const entries = appendEntry(await $.store.get(LEDGER_KEY), finished)
943 await $.store.set(LEDGER_KEY, entries)
944 lastEntryId = turnId
945 // Every 20 turns: what the ledger now suggests, in one line.
946 const found = dueToPropose(entries)
947 if (found.length > 0) {
948 const first = found[0] as (typeof found)[number]
949 if (lines) $.ui.log(`jev · learned from your last turns: ${first.why}. /jev tune shows the change, /jev tune apply takes it`)
950 if (petOn()) {
951 say(`tune? ${first.option} ${first.from}→${first.to} · /jev tune`, 'ready')
952 $.ui.invalidate('ui.render')
953 }
954 }
955 } catch (error) {
956 $.ui.log(`[jev-model-router] could not record the turn: ${String(error)}`)
957 }
958 }
959 return result
960 })
961
962 // `/clear` or a resume starts another session in this worker: nothing
963 // waiting or in progress carries over. The learned model ids stay: they
964 // are the engine's own and still valid. (Under a match-all matcher: the
965 // skill module hooks session.end too, and one unmatched hook per plugin.)
966 on('session.end', { sessionId: /(?:)/ }, async ($, e, next) => {
967 pending.clear()
968 resetBriefing()
969 resetStats()
970 lastDecision = null
971 lastEntryId = null
972 heldEffort = undefined
973 spunTurnId = undefined
974 subagents.clear()
975 subagentEffort.clear()
976 clearSkillNotes()
977 current = null
978 appliedTurnId = undefined
979 applied = null
980 escalatedTurnId = undefined
981 turnPrompt = null
982 failedInARow = 0
983 return next(e)
984 })
985
986 // A compaction rewrites the cached conversation anyway: a held effort may
987 // be chosen again. (Under a matcher: the skill module hooks it too.)
988 on('session.compact', { trigger: /(?:)/ }, async ($, e, next) => {
989 heldEffort = undefined
990 return next(e)
991 })
992
993 // `/jev-pilot:report`: the ledger summarised, with suggested changes;
994 // `/jev-pilot:report reset` clears it. The command's markdown is a
995 // placeholder: the prompt the model reads is written here.
996 on('skill.prompt', { skill: 'jev-pilot:report' }, async ($, e, next) => {
997 try {
998 if (/\breset\b/i.test(e.text)) {
999 await $.store.delete(LEDGER_KEY)
1000 return next({ ...e, text: 'Tell the user the jev-pilot decision ledger was cleared. Change nothing else.' })
1001 }
1002 const entries = entriesOf(await $.store.get(LEDGER_KEY))
1003 const home = (await $.env.get('HOME')) ?? '~'
1004 const settingsPath = `${home}/.claude/settings.json`
1005 // The key jev-pilot's options live under depends on how it was
1006 // installed (marketplace, --mod, --plugin-dir): read which exist.
1007 let keys: string[] = []
1008 try {
1009 if (await $.fs.exists(settingsPath)) keys = configKeysOf(await $.fs.read(settingsPath))
1010 } catch {
1011 keys = []
1012 }
1013 const text = reportPrompt(summarize(entries, effective()), settingsPath, suggestions(entries, effective()).length > 0, keys)
1014 return next({ ...e, text })
1015 } catch (error) {
1016 return next({ ...e, text: `Tell the user the jev-pilot ledger could not be read: ${String(error)}. Change nothing.` })
1017 }
1018 })
1019
1020 on('agent.spawn', async ($, e, next) => {
1021 const io: Io = {
1022 fetch: (url, init) => $.http.fetch(url, init),
1023 sleep: (ms) => $.clock.sleep(ms),
1024 log: (text) => $.ui.log(text),
1025 detail: (text) => (verbose ? $.ui.log(text) : undefined),
1026 messages: () => $.session.messages(),
1027 }
1028 // Before the routing guards: a module whose switches are all off has still
1029 // loaded, and that is exactly when its silence is most misleading.
1030 if (!announced) {
1031 announced = true
1032 if (lines) $.ui.log(readyLine())
1033 }
1034
1035 // A fork inherits its parent's model; `model` is ignored for it.
1036 if (!routeSubagentModel() || e.fork) return next(e)
1037
1038 if (!unusableReported) {
1039 unusableReported = true
1040 $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
1041 }
1042
1043 const crewIo: CrewIo = {
1044 fetch: (url, init) => $.http.fetch(url, init),
1045 home: () => $.env.get('HOME'),
1046 routerUrl: () => $.env.get('JEV_ROUTER_URL'),
1047 write: (path, text) => $.fs.write(path, text),
1048 read: async (path) => ((await $.fs.exists(path)) ? $.fs.read(path) : null),
1049 run: (argv, timeoutMs) => $.process.run(argv, { timeoutMs }),
1050 storeGet: (key) => $.store.get(key),
1051 sleep: (ms) => $.clock.sleep(ms),
1052 register: async (spec) => {
1053 await $.agent.register(spec)
1054 },
1055 }
1056 await ensureCrew(crewIo, openrouterKey || null)
1057
1058 // A reviewer only relays to its CLI: the smallest model does.
1059 if (REVIEWERS.some((r) => e.subagentType === reviewerAgent(r))) {
1060 if (petOn()) {
1061 say(`${subagentLabel(e.description, 'review')} → ${e.subagentType.replace(/^jev-pilot:|-review$/g, '')}`, 'focused')
1062 $.ui.invalidate('ui.render')
1063 }
1064 const reviewed = await next({ ...e, model: policy.tiers.fast })
1065 recordSubagent(reviewed.agentId, policy.tiers.fast, e.parentModel ?? e.model)
1066 return reviewed
1067 }
1068
1069 // The junior runs on its slot; with the slot down, on Sonnet instead.
1070 if (e.subagentType === JUNIOR_AGENT) {
1071 const slot = juniorSlot(crew(), router() !== null)
1072 const model = slot && slotUsable(slot) ? slotAlias(slot.name) : policy.tiers.balanced
1073 if (petOn()) {
1074 say(`${subagentLabel(e.description, 'junior')} → junior on ${model}`, 'focused')
1075 $.ui.invalidate('ui.render')
1076 }
1077 if (lines) $.ui.log(`jev · junior on ${model}`)
1078 const spawned = await next({ ...e, model })
1079 recordSubagent(spawned.agentId, model, e.parentModel ?? e.model)
1080 return spawned
1081 }
1082
1083 // A model the caller named is kept: asked for by the user, or chosen on
1084 // purpose by the main model. Jev still sets the subagent's effort.
1085 const named = e.model ?? null
1086
1087 // Quality mode: Opus for every subagent, no cheaper models.
1088 if (crew().mode === 'quality' && !named) {
1089 if (verbose) $.ui.log(`[jev-model-router] ${e.subagentType}: quality mode, ${policy.tiers.deep}`)
1090 return next({ ...e, model: policy.tiers.deep })
1091 }
1092
1093 // The custom models Jev may choose here: this mode's, the router up, each
1094 // one answering its last check.
1095 const offered = slotsOffered(crew(), router() !== null).filter(slotUsable)
1096 const startedAt = await $.clock.now()
1097 let decision: Decision | null = null
1098 if (active) {
1099 // A subagent's brief is self-contained by design: no conversation added.
1100 decision = (
1101 await classify(
1102 io,
1103 backend,
1104 { prompt: e.prompt, description: e.description, agentType: e.subagentType },
1105 false,
1106 'the subagent',
1107 true,
1108 offered,
1109 )
1110 ).decision
1111 } else {
1112 try {
1113 const label = await $.model.classify(e.prompt, TIER_ORDER)
1114 if (label) {
1115 decision = {
1116 tier: label as Tier,
1117 confidence: null,
1118 risky: null,
1119 effort: null,
1120 effortConfidence: null,
1121 }
1122 }
1123 } catch (error) {
1124 $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
1125 }
1126 }
1127
1128 if (logDecisions) {
1129 const ms = (await $.clock.now()) - startedAt
1130 if (verbose) $.ui.log(`[jev-model-router] jev (${e.subagentType}): ${describeDecision(decision, ms)}`)
1131 }
1132
1133 // The subagent's own model wins when the caller named one; otherwise it
1134 // would inherit the parent's, so that is what a change is measured from.
1135 // The Agent tool takes no effort: that is set at the subagent's requests
1136 // (turn.step), from this same decision, kept by the id it starts with.
1137 const current = e.model ?? e.parentModel
1138 let { model, reason } = route(decision, { model: current }, policy)
1139 // A custom model: moving down to it takes the same confidence as any
1140 // cheaper model; the router sends `jev-<slot>` to it.
1141 const slot = decision?.slot ? offered.find((choice) => choice.name === decision.slot) : undefined
1142 if (named) {
1143 model = null
1144 reason = `the caller named ${named}`
1145 } else if (slot) {
1146 const sure = (decision?.confidence ?? 0) >= policy.minDowngradeConfidence
1147 model = sure ? slotAlias(slot.name) : null
1148 reason = sure ? `${slot.name}: ${slot.model} (confidence ${decision?.confidence?.toFixed(2)})` : `${slot.name} wanted, confidence ${decision?.confidence?.toFixed(2)} < ${policy.minDowngradeConfidence}`
1149 }
1150 // Named by its task in the bubble and the log, not its generic type.
1151 const label = subagentLabel(e.description, e.subagentType)
1152 if (!model) {
1153 if (verbose) $.ui.log(`[jev-model-router] ${label} (${e.subagentType}): model kept (${reason})`)
1154 } else {
1155 if (lines) {
1156 $.ui.log(verbose ? `[jev-model-router] ${label} (${e.subagentType}) → ${model}: ${reason}` : `jev · subagent ${label} → ${model}`)
1157 }
1158 if (petOn()) {
1159 say(`${label} → ${model}`, 'focused')
1160 $.ui.invalidate('ui.render')
1161 }
1162 }
1163 // A brief that will run at low effort and may change code gets Anthropic's
1164 // line for low effort: at `low` the check that exercises a change can be skipped.
1165 const low =
1166 routeSubagentEffort() && decision !== null && decision.tier !== 'fast' && effortLevel(effortScoreOf(decision, margin) ?? 1) === 'low'
1167 const brief = low && !e.prompt.includes(CHECK_AT_LOW) ? { prompt: `${e.prompt}\n\n${CHECK_AT_LOW}` } : {}
1168 const result = await next({ ...e, ...(model ? { model } : {}), ...brief })
1169 recordSubagent(result.agentId, model ?? current ?? null, e.parentModel ?? current)
1170 if (decision && result.agentId && routeSubagentEffort()) {
1171 subagents.set(result.agentId, { decision, label, model: model ?? null })
1172 while (subagents.size > MAX_SUBAGENTS) subagents.delete(subagents.keys().next().value as string)
1173 }
1174 return result
1175 })
1176}
1177hooks/jev-skill-suggestion.ts 708 lines1/**
2 * jev-skill-suggestion — Claude Mod (EARLY ACCESS)
3 *
4 * Takes the skill listing out of the context window and has TypeSafe's Jev,
5 * a System One decision model, suggest at most one skill per prompt, going
6 * by the skills' descriptions. The skills stay installed and loadable; what
7 * goes away is the listing the engine sends the model every session, one
8 * line per skill, whether the prompt has anything to do with any of them.
9 *
10 * The decision follows TypeSafe's "Skill suggestion" cookbook: two requests
11 * per prompt, one to rank every skill and ask whether the prompt needs a
12 * skill at all, one to re-read the top few with their full text and let each
13 * be rejected on its own. Either may come back empty-handed.
14 *
15 * Three hooks:
16 * prompt.attachment — the engine's `skill_listing` attachment is answered
17 * with `{ text: null }` (left out) or trimmed to the
18 * names in `alwaysListed`. Its names are remembered:
19 * they are the engine's word on which skills the model
20 * may invoke.
21 * prompt.submit — the two requests run, and the winner (if any) is
22 * attached to the prompt as a `<skill_relevance>`
23 * block: with the skill's own SKILL.md inside it
24 * (`inject: "content"`, the default), so the skill
25 * loads even when `skillOverrides` hides it from the
26 * model, or with its name for the Skill tool
27 * (`inject: "suggest"`).
28 * skill.prompt — observation: whether the model took the suggestion,
29 * or loaded a skill on its own. Also writes the prompt
30 * of the plugin's own `/jev-pilot:setup`,
31 * which hides every skill from the engine's listing
32 * (user-invocable-only) once the person has seen the
33 * list and said yes; the model makes the edit with its
34 * own tools, so it shows and asks like any other.
35 *
36 * Jev is reached one of three ways, whichever key is configured: TypeSafe's
37 * own API (`typesafeApiKey`) or OpenRouter's Decisions API
38 * (`openrouterApiKey`), which report a calibrated confidence, or the Vercel
39 * AI Gateway (`gatewayApiKey`), which does not. With none, the
40 * engine's own `$.model.classify` stands in with a single request and no
41 * gate, so the mod is useful without any account.
42 *
43 * The candidates come from `$.command.list()`, not from the listing: the
44 * listing is only rendered at the turn's first request, after `prompt.submit`
45 * has run, so the first prompt of a session would otherwise have nothing to
46 * choose from. The listing, once seen, narrows the candidates to what the
47 * engine itself would have shown.
48 *
49 * The second request reads the opening of each shortlisted skill's SKILL.md,
50 * found on disk by how Claude Code lays skills out (project and user
51 * `.claude/skills` and `.claude/commands`, a plugin's install path from
52 * `~/.claude/plugins/installed_plugins.json`). A body that cannot be found
53 * leaves that skill with its one-line description; nothing fails over it.
54 *
55 * Only the main conversation is handled. A subagent's own listing is left as
56 * the engine renders it: its prompt is a tool call's argument, not a
57 * `prompt.submit`, so nothing here could suggest for it.
58 *
59 * Every failure path is fail-open: a request that errors or runs past the
60 * latency budget lets the prompt through with no suggestion, and the listing
61 * hook always answers the same way, so the model's prompt cache holds.
62 *
63 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
64 * or "gatewayApiKey"). Never hardcode it in this file.
65 *
66 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 and Claude Code >= 2.1.278: the
67 * `prompt.attachment` event is that release's. Typed against Anthropic's
68 * declarations: https://github.com/anthropics/claude-code/tree/main/mods
69 *
70 * Privacy: with a key set, the prompt text, every candidate skill's name and
71 * description, and the opening of each shortlisted skill's SKILL.md are sent
72 * to whichever backend the key belongs to.
73 */
74import type { Register } from 'claude-code'
75import { missOf } from './model-router.policy.ts'
76import { NOT_A_TASK, recentContext } from './context.ts'
77import { noteSkill, resetBriefing } from './summary.ts'
78import { feature } from './features.ts'
79import {
80 DEFAULT_BASE_URL,
81 DEFAULT_MODEL,
82 NONE,
83 SETUP_COMMAND,
84 builtinWide,
85 catalog,
86 classifyText,
87 commandLike,
88 decide,
89 pickSkill,
90 designQuestion,
91 readDesign,
92 DESIGN_BAR,
93 packSkills,
94 designBlock,
95 describeRerank,
96 describeSetup,
97 describeStatus,
98 describeStillListed,
99 describeWide,
100 detailOf,
101 displayIds,
102 canonical,
103 endpoint,
104 injectionBlock,
105 installPathsOf,
106 parseListing,
107 parseNames,
108 passesGate,
109 pluginFileCandidates,
110 readRerank,
111 readSkillSettings,
112 modelInvocable,
113 readWide,
114 rerankQuestions,
115 requestBody,
116 requestHeaders,
117 selectProvider,
118 setupAborted,
119 setupInstructions,
120 setupPlan,
121 shortlistOf,
122 validBackup,
123 skillFileCandidates,
124 suggestionBlock,
125 syncedFileCandidates,
126 trimListing,
127 wideQuestions,
128 batchesOf,
129 MAX_CHOICES,
130 mergeWide,
131} from './skill-suggestion.policy.ts'
132import type { Candidate, PolicyConfig, Provider, Rerank, Skill, Suggestion, Wide } from './skill-suggestion.policy.ts'
133import { isContinuation, takePart } from './jev-call.ts'
134
135export const register: Register = (on, options) => {
136 const text = (key: string, fallback: string) =>
137 typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
138 const number = (key: string, fallback: number) =>
139 typeof options[key] === 'number' ? (options[key] as number) : fallback
140 const flag = (key: string, fallback: boolean) =>
141 typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
142
143 // With several keys set, `auto` takes TypeSafe's own API, then OpenRouter,
144 // then the Gateway: the first two report a calibrated confidence.
145 // `provider` forces one, "builtin" uses none.
146 const typesafeKey = text('typesafeApiKey', '')
147 const gatewayKey = text('gatewayApiKey', '')
148 const openrouterKey = text('openrouterApiKey', '')
149 const forced = text('provider', 'auto')
150 const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey, openrouterKey)
151
152 // Each backend keeps its own URL and model, so an override written for one
153 // can never be sent to the other when `auto` picks differently than expected.
154 const apiKey = !active ? '' : { typesafe: typesafeKey, gateway: gatewayKey, openrouter: openrouterKey }[active]
155 const modelId = !active ? '' : text(`${active}Model`, DEFAULT_MODEL[active])
156 const url = !active ? '' : endpoint(active, text(`${active}BaseUrl`, DEFAULT_BASE_URL[active]))
157
158 // A backend named in the options but missing its key degrades to the
159 // built-in classifier, which is silent; say so once, when a hook first runs.
160 let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
161
162 const hideListing = flag('hideListing', true)
163 // "content": the mod reads the chosen skill's SKILL.md and attaches it, so
164 // the skill loads even when the engine lists it as user-invocable-only or
165 // off. "suggest": the cookbook's block alone, and the model loads the skill
166 // with the Skill tool, which honours the engine's skillOverrides.
167 const injectContent = text('inject', 'content') !== 'suggest'
168 const alwaysListed = parseNames(text('alwaysListed', ''))
169 const neverSuggested = parseNames(text('neverSuggested', ''))
170 // Off (the default): the one request decides. On: a second request re-reads the shortlist.
171 const rerankEnabled = flag('rerank', false)
172 const excerptChars = number('excerptChars', 300)
173 const timeoutMs = number('timeoutMs', 1500)
174 const logDecisions = flag('logDecisions', true)
175 // One line per turn (written by the router) by default; every step here with verboseLog.
176 const verbose = logDecisions && flag('verboseLog', false)
177 // The setup tip goes to the transcript only where jev-pilot talks there.
178 const display = text('display', 'pet')
179 const lines = logDecisions && (verbose || display === 'transcript' || display === 'both')
180 // Shared with the model router: how much of the conversation Jev reads.
181 const contextLimits = {
182 messages: Math.max(0, Math.round(number('contextMessages', 4))),
183 chars: Math.max(0, number('contextChars', 2000)),
184 }
185 const policy: PolicyConfig = {
186 // The rerank is one Choice too, under the same 255-option limit.
187 shortlist: Math.min(MAX_CHOICES, Math.max(1, Math.round(number('shortlist', 3)))),
188 gateThreshold: number('gateThreshold', 0.3),
189 fitsThreshold: number('fitsThreshold', 0.3),
190 }
191
192 // The names every skill_listing attachment carried so far. Once non-empty,
193 // only these are offered to the decision model: the listing is the engine's
194 // word on which skills the model is allowed to invoke, and `$.command.list()`
195 // also names commands the model may not.
196 const listed = new Set<string>()
197 // The skill suggested for the current prompt, so a skill.prompt that loads
198 // it can be told apart from one the model reached for on its own.
199 let suggested: string | null = null
200 // Each skill's SKILL.md as first found, with where, or null when nowhere:
201 // read once per session, since the second request wants it on every prompt
202 // it is on, and the injection wants it whole.
203 const files = new Map<string, { path: string; markdown: string } | null>()
204 // The skills whose instructions were already attached this session: a
205 // second time, the block only names the skill again.
206 const injected = new Set<string>()
207 // Said once, the first time a hook runs. A mod that loaded and one that
208 // never loaded are otherwise told apart only by the absence of later lines,
209 // and absence is not evidence: with the listing gone, silence is the norm.
210 let announced = false
211 // The setup hint, once per session.
212 let hintedSetup = false
213 // The design pack's skills, and the tool names (for the screen libraries), read once.
214 const designSkillsOption = text('designSkills', 'design-taste-frontend, impeccable, web-design-guidelines')
215 let toolNames: string[] | null = null
216
217 // A skill whose frontmatter `name:` has spaces ("PocketBase API Rules") is
218 // reported by `$.command.list()` under that name, but the engine lists,
219 // runs and overrides it by its directory name (`pb-api-rules`). The map
220 // from one to the other is read from disk once per session, in whichever
221 // hook first needs it.
222 let displayToId: Map<string, string> | null = null
223
224 on('prompt.attachment', { type: 'skill_listing' }, async ($, e, next) => {
225 if (!announced) {
226 announced = true
227 if (verbose) $.ui.log(`[jev-skill-suggestion] ${describeSetup(active, url, hideListing, forced === 'builtin')}`)
228 }
229
230 // A subagent's listing is not ours: nothing here suggests for a subagent,
231 // so hiding it would leave the subagent with no skills, and adding its
232 // names would narrow the main conversation's roster to the subagent's.
233 if (e.agentId) return next(e)
234 // Skills switched off (/jev skills off): the listing reaches the model as
235 // the engine built it, and nothing is picked.
236 if (!feature('skills')) return next(e)
237
238 const skills = parseListing(e.text)
239 for (const skill of skills) listed.add(skill.name)
240
241 // With the mod loading skills itself, a listing that still names any is
242 // context the setup command would have saved: say so once.
243 if (injectContent && !hintedSetup && skills.length > 0) {
244 hintedSetup = true
245 if (lines) {
246 $.ui.log(
247 verbose
248 ? `[jev-skill-suggestion] ${describeStillListed(skills.length)}`
249 : `jev · tip: /jev-pilot:setup takes the ${skills.length} listed skills out of context`,
250 )
251 }
252 }
253
254 if (!hideListing) return next(e)
255
256 const kept = trimListing(e.text, alwaysListed)
257 if (verbose) {
258 const keptNames = kept
259 ? parseListing(kept)
260 .map((skill) => skill.name)
261 .join(', ')
262 : 'none'
263 $.ui.log(
264 `[jev-skill-suggestion] withheld the skill listing (${skills.length} skills, ${e.text.length} characters); kept listed: ${keptNames}`,
265 )
266 }
267 // Answered without `next`: the engine's text never reaches the model.
268 return { text: kept }
269 })
270
271 // Every prompt, under a matcher: jev-pilot.ts registers the model
272 // router first, and a plugin may hook an event once without a matcher. The
273 // two nest, router outermost, so this runs inside it on the same prompt.
274 on('prompt.submit', { text: /(?:)/ }, async ($, e, next) => {
275 if (!announced) {
276 announced = true
277 if (verbose) $.ui.log(`[jev-skill-suggestion] ${describeSetup(active, url, hideListing, forced === 'builtin')}`)
278 }
279 suggested = null
280 // The router's questions for this prompt, to send with ours in one
281 // request (jev-call.ts). Whatever happens here, they are sent and settled.
282 const part = takePart(e.text)
283 /** Passes the prompt on, the router's part asked alone when no ranking carried it. */
284 const alone = async (input: typeof e) => {
285 if (!part) return next(input)
286 const block = await part.settle(await part.ask({}))
287 return next(block ? { ...input, context: [...(input.context ?? []), block] } : input)
288 }
289 if (!feature('skills')) return alone(e)
290
291 // Notifications and peer messages are not tasks; a typed `/name` already
292 // names its skill. Neither gets a suggestion. Nor does "continue": the
293 // work in progress already has what it loaded.
294 if (!e.text.trim() || /^\/\S/.test(e.text.trim())) return alone(e)
295 if (e.origin && NOT_A_TASK.has(e.origin.kind)) return alone(e)
296 if (isContinuation(e.text)) return alone(e)
297
298 // The conversation before the prompt, so a follow-up is read as the work
299 // it continues. Read only when a backend will receive it, and not when
300 // the router's part already carries it.
301 let recent = ''
302 if (active && !part && contextLimits.messages > 0) {
303 try {
304 recent = recentContext(await $.session.messages(), e.text, contextLimits)
305 } catch (error) {
306 $.ui.log(`[jev-skill-suggestion] could not read the conversation: ${String(error)}`)
307 }
308 }
309
310 /** One request to the active backend, or null on timeout, error or a non-2xx. */
311 const ask = async (
312 prompt: string,
313 questions: Record<string, unknown>,
314 what: string,
315 ): Promise<string | null> => {
316 if (!active) return null
317 try {
318 const response = await Promise.race([
319 $.http.fetch(url, {
320 method: 'POST',
321 headers: requestHeaders(active, apiKey, modelId),
322 body: requestBody(active, prompt, questions, modelId, recent),
323 }),
324 $.clock.sleep(timeoutMs),
325 ])
326 if (response && response.ok) return response.text
327 // A timeout or a busy backend is routine (the pet says so): verbose
328 // only. Any other status is an error, and its body says why (a limit,
329 // a bad field): worth the one line.
330 const miss = missOf(response ? response.status : null)
331 const tell = (text: string) => (miss === 'error' || verbose ? $.ui.log(text) : undefined)
332 if (response) {
333 tell(`[jev-skill-suggestion] ${active} responded ${response.status} to the ${what}: ${response.text.slice(0, 200)}`)
334 } else tell(`[jev-skill-suggestion] ${what} passed ${timeoutMs}ms; no suggestion`)
335 } catch (error) {
336 $.ui.log(`[jev-skill-suggestion] ${what} failed: ${String(error)}`)
337 }
338 return null
339 }
340
341 /** A skill's file, found on disk by Claude Code's layout, or null. */
342 const fileOf = async (
343 skill: Skill,
344 plugin: string | undefined,
345 ): Promise<{ path: string; markdown: string } | null> => {
346 const cached = files.get(skill.name)
347 if (cached !== undefined) return cached
348 let found: { path: string; markdown: string } | null = null
349 try {
350 const home = (await $.env.get('HOME')) ?? ''
351 const relative = skillFileCandidates(skill.name, plugin)
352 // The engine reads the project's `.claude/` (the working directory
353 // only, not its ancestors) and the user's.
354 const candidates = [...relative, ...(home ? relative.map((file) => `${home}/${file}`) : [])]
355 if (plugin && home) {
356 const installed = `${home}/.claude/plugins/installed_plugins.json`
357 if (await $.fs.exists(installed)) {
358 for (const path of installPathsOf(await $.fs.read(installed), plugin)) {
359 candidates.push(...pluginFileCandidates(path, skill.name, plugin))
360 }
361 }
362 }
363 // A claude.ai-synced skill sits under an account directory only
364 // `$.fs.list` can name.
365 if (home) {
366 const synced = `${home}/.claude/skills/synced`
367 if (await $.fs.exists(synced)) {
368 const accounts = (await $.fs.list(synced)).filter((entry) => entry.kind === 'dir').map((entry) => entry.name)
369 candidates.push(...syncedFileCandidates(home, accounts, skill.name))
370 }
371 }
372 for (const file of candidates) {
373 if (await $.fs.exists(file)) {
374 found = { path: file, markdown: await $.fs.read(file) }
375 break
376 }
377 }
378 } catch (error) {
379 $.ui.log(`[jev-skill-suggestion] could not read /${skill.name}: ${String(error)}`)
380 }
381 files.set(skill.name, found)
382 return found
383 }
384 /** The opening of a skill's body, or null when its file is nowhere. */
385 const bodyOf = async (skill: Skill, plugin: string | undefined): Promise<string | null> =>
386 (await fileOf(skill, plugin))?.markdown ?? null
387
388 if (!unusableReported) {
389 unusableReported = true
390 $.ui.log(`[jev-skill-suggestion] provider "${forced}" has no key set; using the built-in classifier`)
391 }
392
393 let commands: Awaited<ReturnType<typeof $.command.list>>
394 try {
395 commands = await $.command.list()
396 if (!displayToId && commands.some((command) => !commandLike(command.name))) {
397 const found: { dir: string; markdown: string }[] = []
398 for (const root of [await $.session.cwd(), (await $.env.get('HOME')) ?? '']) {
399 const dir = root && `${root}/.claude/skills`
400 if (!dir || !(await $.fs.exists(dir))) continue
401 for (const entry of await $.fs.list(dir)) {
402 const file = `${dir}/${entry.name}/SKILL.md`
403 if (entry.kind === 'dir' && (await $.fs.exists(file))) found.push({ dir: entry.name, markdown: await $.fs.read(file) })
404 }
405 }
406 displayToId = displayIds(found)
407 }
408 commands = canonical(commands, displayToId ?? new Map())
409 } catch (error) {
410 $.ui.log(`[jev-skill-suggestion] could not list the skills: ${String(error)}`)
411 return alone(e)
412 }
413 // Loading the skill itself, the mod is not bound to what the engine would
414 // list: a skill hidden with skillOverrides is still a candidate.
415 const listedSkills = catalog(commands, injectContent ? new Set() : listed, neverSuggested)
416 if (listedSkills.length === 0) {
417 if (verbose) $.ui.log('[jev-skill-suggestion] no candidate skills; nothing to suggest')
418 return alone(e)
419 }
420 const pluginOf = new Map(commands.map((command) => [command.name, command.plugin]))
421 // Each skill as Jev reads it, in the one request: its description and the
422 // opening of its SKILL.md (read once per session), without the skills
423 // whose own frontmatter says the model may not invoke them.
424 const barred: string[] = []
425 const skills: Skill[] = []
426 for (const skill of listedSkills) {
427 const body = active && !rerankEnabled ? await bodyOf(skill, pluginOf.get(skill.name)) : null
428 if (body !== null && !modelInvocable(body)) {
429 barred.push(skill.name)
430 continue
431 }
432 skills.push(active && !rerankEnabled ? { ...skill, description: detailOf(skill, body, excerptChars) } : skill)
433 }
434 if (skills.length === 0) return alone(e)
435
436 // Request 1: rank everything, and ask whether the prompt wants a skill at
437 // all. The first batch carries the router's questions too: one request.
438 const startedAt = await $.clock.now()
439 let wide: Wide | null = null
440 let parts: (Wide | null)[] = []
441 let routerBlock: string | null = null
442 let routerSettled = false
443 let designRead: number | null = null
444 // The shortlist grows with the batches, so every batch's leaders reach
445 // the rerank (their scores do not compare across batches).
446 let picked: PolicyConfig = policy
447 if (active) {
448 // One Choice takes at most MAX_CHOICES options: a larger catalog is
449 // ranked in batches, side by side, and the gate asked once.
450 const batches = batchesOf(skills, MAX_CHOICES - 1)
451 picked = { ...policy, shortlist: Math.min(MAX_CHOICES, policy.shortlist * batches.length) }
452 const withNone = !rerankEnabled
453 const answers = await Promise.all(
454 batches.map(async (batch, index) => {
455 const questions = wideQuestions(active, batch, index === 0, withNone)
456 // UI design work? Asked in the same request, for the design pack.
457 if (index === 0 && feature('design')) questions.ui_design = designQuestion(active)
458 const what = batches.length > 1 ? `ranking ${index + 1}/${batches.length}` : 'ranking'
459 if (index > 0 || !part) {
460 const text = await ask(e.text, questions, what)
461 if (index === 0) designRead = readDesign(text)
462 return text
463 }
464 const answer = await part.ask(questions)
465 routerSettled = true
466 routerBlock = await part.settle(answer)
467 designRead = readDesign(answer.text)
468 return answer.text
469 }),
470 )
471 parts = answers.map((answer) => (answer ? readWide(answer) : null))
472 wide = mergeWide(parts)
473 } else {
474 // No backend: the engine's own small-model classifier answers the same
475 // question, with the descriptions folded into the text it reads. One
476 // label, no gate, no rerank.
477 try {
478 const label = await $.model.classify(classifyText(e.text, skills), [
479 NONE,
480 ...skills.map((skill) => skill.name),
481 ])
482 wide = builtinWide(label)
483 parts = [wide]
484 } catch (error) {
485 $.ui.log(`[jev-skill-suggestion] built-in classifier failed: ${String(error)}`)
486 }
487 }
488 if (part && !routerSettled) routerBlock = await part.settle(await part.ask({}))
489 // What the decision model actually answered, whatever the policy then
490 // does with it. This is the line that proves the ranking ran.
491 if (verbose) {
492 const ms = (await $.clock.now()) - startedAt
493 $.ui.log(`[jev-skill-suggestion] jev: ${describeWide(wide, skills.length, ms)}`)
494 }
495
496 let decision: Suggestion
497 if (active && rerankEnabled) {
498 // Request 2 (the `rerank` option): re-read the shortlist with each
499 // skill's full text, and let every candidate be rejected on its own.
500 let rerank: Rerank | null = null
501 let rerankAttempted = false
502 if (wide && passesGate(wide, policy)) {
503 const candidates: Candidate[] = []
504 for (const skill of shortlistOf(wide, skills, picked.shortlist)) {
505 const body = await bodyOf(skill, pluginOf.get(skill.name))
506 if (!modelInvocable(body)) {
507 barred.push(skill.name)
508 continue
509 }
510 candidates.push({ ...skill, detail: detailOf(skill, body, excerptChars) })
511 }
512 if (candidates.length > 0) {
513 const rerankStartedAt = await $.clock.now()
514 rerankAttempted = true
515 const answer = await ask(e.text, rerankQuestions(active, candidates), 'rerank')
516 if (answer) rerank = readRerank(answer)
517 if (verbose) {
518 const ms = (await $.clock.now()) - rerankStartedAt
519 const read = candidates.filter((candidate) => files.get(candidate.name)).length
520 $.ui.log(`[jev-skill-suggestion] jev: ${describeRerank(rerank, ms)} · ${read}/${candidates.length} bodies read`)
521 }
522 }
523 }
524 const offered = barred.length > 0 ? skills.filter((skill) => !barred.includes(skill.name)) : skills
525 decision = decide(wide, rerank, offered, picked, rerankAttempted)
526 } else if (active) {
527 decision = pickSkill(parts, skills, policy)
528 } else {
529 decision = decide(wide, null, skills, picked, false)
530 }
531 let pick = decision.name ? (skills.find((skill) => skill.name === decision.name) ?? null) : null
532 // The winner's own frontmatter has the last word, whichever path picked it.
533 if (pick && !barred.includes(pick.name) && !modelInvocable(await bodyOf(pick, pluginOf.get(pick.name)))) {
534 barred.push(pick.name)
535 decision = { name: null, reason: `/${pick.name} has disable-model-invocation` }
536 pick = null
537 }
538 if (verbose && barred.length > 0) {
539 $.ui.log(
540 `[jev-skill-suggestion] not model-invocable, left out: ${barred.map((name) => `/${name}`).join(', ')}`,
541 )
542 }
543 // The router writes the turn's one line (and its status) from this note.
544 noteSkill(e.text, { skill: pick?.name ?? null, ms: active ? Math.round((await $.clock.now()) - startedAt) : null })
545 if (verbose) $.ui.status(describeStatus(pick?.name ?? null))
546 if (verbose) {
547 $.ui.log(
548 pick
549 ? `[jev-skill-suggestion] suggesting /${pick.name}: ${decision.reason}`
550 : `[jev-skill-suggestion] no suggestion: ${decision.reason}`,
551 )
552 }
553
554 suggested = pick?.name ?? null
555 let block: string | null
556 if (injectContent && pick) {
557 const file = await fileOf(pick, pluginOf.get(pick.name))
558 const projectDir = await $.session.cwd()
559 block = injectionBlock(pick, file?.markdown ?? null, file?.path ?? null, projectDir, injected.has(pick.name))
560 if (verbose) {
561 $.ui.log(
562 file
563 ? injected.has(pick.name)
564 ? `[jev-skill-suggestion] /${pick.name} already injected this session; named again`
565 : `[jev-skill-suggestion] injected /${pick.name} from ${file.path} (${file.markdown.length} characters)`
566 : `[jev-skill-suggestion] no file found for /${pick.name}; suggested by name only`,
567 )
568 }
569 if (file) injected.add(pick.name)
570 } else {
571 block = suggestionBlock(pick, hideListing)
572 }
573 // Attached on the way down, the router's advice first: blocks after the
574 // prompt as typed, read by the model and never shown to the person.
575 // UI design work: the design pack, with the design skills installed and
576 // the screen libraries connected.
577 let design: string | null = null
578 // (Set inside the request's callbacks above, which the compiler can't follow.)
579 const designScore = designRead as number | null
580 if (feature('design') && designScore !== null && designScore >= DESIGN_BAR) {
581 const skills = packSkills([...parseNames(designSkillsOption)], listedSkills.map((skill) => skill.name))
582 if (!toolNames) toolNames = (await $.tool.list().catch(() => [])).map((tool) => tool.name)
583 design = designBlock(skills, {
584 mobbin: toolNames.some((name) => name.startsWith('mcp__mobbin__')),
585 inspo: toolNames.some((name) => name.startsWith('mcp__inspo__')),
586 })
587 if (verbose) $.ui.log(`[jev-skill-suggestion] design work (${designScore.toFixed(2)}): ${skills.map((name) => `/${name}`).join(', ') || 'no design skills installed'}`)
588 }
589 const blocks = [routerBlock, block, design].filter((b): b is string => b !== null)
590 if (blocks.length === 0) return next(e)
591 return next({ ...e, context: [...(e.context ?? []), ...blocks] })
592 })
593
594 // An injected skill lives in the conversation, not the process: `/clear`
595 // or a resume starts another under the same worker, and a compaction may
596 // summarize the block away. Either way the next pick goes in whole again.
597 on('session.end', async ($, e, next) => {
598 injected.clear()
599 suggested = null
600 // Everything learned about this session's skills goes with it: the next
601 // one may have another roster, and its SKILL.md files may have changed.
602 listed.clear()
603 files.clear()
604 displayToId = null
605 announced = false
606 hintedSetup = false
607 return next(e)
608 })
609 on('session.compact', async ($, e, next) => {
610 if (!e.agentId) {
611 injected.clear()
612 // The compacted conversation no longer holds the router's note to the
613 // model about what jev-pilot does: it is given again on the next prompt.
614 resetBriefing()
615 }
616 return next(e)
617 })
618
619 on('skill.prompt', { skill: 'jev-pilot:setup' }, async ($, e, next) => {
620 // The plugin's own setup command: its markdown is a placeholder, and the
621 // prompt the model reads is written here, from the roster as the engine
622 // has it and the user settings as they are. The model does the editing
623 // with its own tools, so the change shows as a diff and asks permission.
624 const mode = /\brestore\b/i.test(e.text) ? 'restore' : 'apply'
625 // No plan from a partial roster or unreadable settings: the edit would
626 // hide too little, and the backup would save the wrong values.
627 let commands: Awaited<ReturnType<typeof $.command.list>> = []
628 try {
629 commands = await $.command.list()
630 if (!displayToId && commands.some((command) => !commandLike(command.name))) {
631 const found: { dir: string; markdown: string }[] = []
632 for (const root of [await $.session.cwd(), (await $.env.get('HOME')) ?? '']) {
633 const dir = root && `${root}/.claude/skills`
634 if (!dir || !(await $.fs.exists(dir))) continue
635 for (const entry of await $.fs.list(dir)) {
636 const file = `${dir}/${entry.name}/SKILL.md`
637 if (entry.kind === 'dir' && (await $.fs.exists(file))) found.push({ dir: entry.name, markdown: await $.fs.read(file) })
638 }
639 }
640 displayToId = displayIds(found)
641 }
642 commands = canonical(commands, displayToId ?? new Map())
643 } catch (error) {
644 $.ui.log(`[jev-skill-suggestion] setup: could not list the skills: ${String(error)}`)
645 return next({ ...e, text: setupAborted(`the skills could not be listed (${String(error)})`) })
646 }
647 const home = (await $.env.get('HOME')) ?? '~'
648 const settingsPath = `${home}/.claude/settings.json`
649 const backupPath = `${home}/.claude/jev-pilot.skill-overrides.backup.json`
650 let json: string | null = null
651 try {
652 if (await $.fs.exists(settingsPath)) json = await $.fs.read(settingsPath)
653 } catch (error) {
654 $.ui.log(`[jev-skill-suggestion] setup: could not read ${settingsPath}: ${String(error)}`)
655 return next({ ...e, text: setupAborted(`${settingsPath} exists but could not be read (${String(error)})`) })
656 }
657 const settings = readSkillSettings(json)
658 // Never plan edits on a file that could not be parsed: the backup would
659 // miss what it holds, and the edit could destroy it.
660 if (settings.invalid) {
661 $.ui.log(`[jev-skill-suggestion] setup: ${settingsPath} is not valid JSON`)
662 return next({ ...e, text: setupAborted(`${settingsPath} is not a valid JSON object; ask the user to fix it first`) })
663 }
664 const plan = setupPlan(commands, settings, new Set([SETUP_COMMAND]))
665 // An earlier run's backup is reused only if it is one: a file that is
666 // not this mod's, or is corrupt, is nothing restore could apply, so no
667 // setup is built on top of it.
668 let backupExists = false
669 try {
670 backupExists = await $.fs.exists(backupPath)
671 if (backupExists && !validBackup(await $.fs.read(backupPath))) {
672 $.ui.log(`[jev-skill-suggestion] setup: ${backupPath} is not a valid backup`)
673 return next({
674 ...e,
675 text: setupAborted(
676 `${backupPath} exists but is not a backup this mod wrote (expected {"skillOverrides": {...}, "disableBundledSkills": true|false|null}); ask the user to inspect it and move it away, or fix it, before running the setup again`,
677 ),
678 })
679 }
680 } catch (error) {
681 $.ui.log(`[jev-skill-suggestion] setup: could not read ${backupPath}: ${String(error)}`)
682 return next({ ...e, text: setupAborted(`${backupPath} could not be read (${String(error)})`) })
683 }
684 if (logDecisions) {
685 $.ui.log(
686 `[jev-skill-suggestion] setup (${mode}): ${plan.hide.length} to hide, ${plan.alreadyHidden.length} already hidden, ${plan.locked.length} locked by a plugin`,
687 )
688 }
689 return next({ ...e, text: setupInstructions(mode, plan, settings, settingsPath, backupPath, backupExists) })
690 })
691
692 on('skill.prompt', async ($, e, next) => {
693 // Observation only: whether the model took the suggestion, or reached for
694 // a skill it was never told about, is the one measure of this mod's worth.
695 // The plugin's own commands (setup, report) are not skills to measure.
696 if (verbose && !e.skill.startsWith('jev-pilot:')) {
697 const how =
698 suggested === e.skill
699 ? 'as suggested'
700 : suggested
701 ? `suggested was /${suggested}`
702 : 'nothing was suggested'
703 $.ui.log(`[jev-skill-suggestion] skill /${e.skill} loaded (${how})`)
704 }
705 return next(e)
706 })
707}
708hooks/jev-pet.tsx 408 lines1/**
2 * jev-pet — Jev as a companion above the prompt, at the right: Claude the
3 * pilot (Claude Code's character in pilot gear), with a speech bubble saying what Jev just
4 * decided, so the decisions stay out of the conversation. It also owns
5 * `/jev`, the switches for every part of jev-pilot.
6 *
7 * While a turn runs the pilot shows what Claude is doing, and the bubble says
8 * it beside a spinner: thinking (a thought cloud), reading (a book),
9 * searching (a magnifier), writing (paper and a pencil), running a command
10 * (a terminal), anything else (a subagent) flying, goggles down. With
11 * subagents still working in the background after a turn, it cruises,
12 * goggles down, until they finish; at max effort the goggles stay down for
13 * the turn. Idle, it hovers, its scarf's end dipping every few seconds; it
14 * blinks, and every few seconds plays for a moment: jumps rope, waves, looks
15 * around. Resting, it redraws only when something moves.
16 *
17 * The router and the skill module set what it says (pet-art.ts `say`) and ask
18 * for a redraw; this module draws it, in the `AbovePrompt` band on the
19 * terminal. Nothing draws in `claude -p`, the desktop app or mobile.
20 */
21import type { Register, Timer } from 'claude-code'
22import {
23 describeFeatures,
24 feature,
25 FEATURE_INFO,
26 featureOverrides,
27 loadFeatureOverrides,
28 parseJevCommand,
29 setFeature,
30} from './features.ts'
31import { applyCrewCommand, describeChoice, describeCrew, describeSlot, MAX_MODELS, MODELS_PAGE, parseCrewCommand, resolveCodexChoice, resolveModel, resolveOpencodeChoice, slotAlias, type CrewCommand, type ReviewerResolved } from './crew.ts'
32import { resetBriefing } from './summary.ts'
33import { describeStats } from './session-stats.ts'
34import { entriesOf, LEDGER_KEY } from './ledger.ts'
35import { applied, currentTuning, describeTuning, proposals, setTuning, TUNING_KEY, tuningLoaded, tuningOf } from './tuning.ts'
36import { codexModels, crew, crewOverrides, describeHealth, router, setCrewOverrides } from './crew-state.ts'
37import { CREW_KEY, checkCrew, checkSlot, ensureCrew, modelCatalog, opencodeModels, publishSlots, refreshCodexModels, registerCrew, type CrewIo } from './crew-run.ts'
38import {
39 type Act,
40 ACT_LABEL,
41 actOfTool,
42 CANVAS_W,
43 currentSpeech,
44 isBoosted,
45 MOOD_COLOR,
46 PLAY_FRAMES,
47 PLAYS,
48 sceneRows,
49 setBoost,
50 WORK_ACTS,
51} from './pet-art.ts'
52
53const FEATURES_KEY = 'features'
54const FLY_MS = 200
55// Background subagents still working, no turn running: the pilot cruises,
56// goggles down, at a calmer rate; running agents are checked this often.
57const CRUISE_MS = 450
58const AGENTS_EVERY_MS = 1500
59const BLINK_EVERY_MS = 4600
60const BLINK_MS = 170
61const PLAY_EVERY_MS = 9000
62const PLAY_STEP_MS = 180
63/** How many loops each idle play runs for: about three seconds each. */
64const PLAY_LOOPS = { rope: 4, wave: 4, look: 2 } as const
65const SPINNER = '⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏'
66
67export const register: Register = (on, options) => {
68 let act: Act = 'rest'
69 let frame = 0
70 let blink = false
71 let workingTurn: string | null = null
72 let flying: Timer | null = null
73 let blinker: Timer | null = null
74 let player: Timer | null = null
75 let playing: Timer | null = null
76 let plays = 0
77 let cruising: Timer | null = null
78 let watcher: Timer | null = null
79 // The main loop's tool calls in flight, each with the act it shows.
80 const running = new Map<string, Act>()
81 let calls = 0
82
83 const stopPlay = () => {
84 playing?.cancel()
85 playing = null
86 }
87 const stopCruise = () => {
88 cruising?.cancel()
89 cruising = null
90 }
91
92 // Session setup: the saved switches, the /jev command, the blink. (Under a
93 // match-all matcher: the router hooks session.start too, one unmatched
94 // registration per plugin.)
95 on('session.start', { cwd: /(?:)/ }, async ($, e, next) => {
96 const result = await next(e)
97 loadFeatureOverrides(await $.store.get(FEATURES_KEY).catch(() => undefined))
98 await $.command
99 .register({
100 name: 'jev',
101 description: 'jev-pilot switches: /jev shows them, /jev <effort|raise|subagents|skills|strategy|model|pet> on|off',
102 argumentHint: '[<feature> on|off | all on|off | reset]',
103 immediate: true,
104 })
105 .catch((error) => $.ui.log(`[jev-pilot] /jev not registered: ${String(error)}`))
106 blinker?.cancel()
107 blinker = $.clock.every(BLINK_EVERY_MS, () => {
108 if (!feature('pet') || workingTurn || playing || cruising) return
109 blink = true
110 $.ui.invalidate('ui.render')
111 $.clock.after(BLINK_MS, () => {
112 blink = false
113 $.ui.invalidate('ui.render')
114 })
115 })
116 // Idle play: every few seconds one of the plays, in turn, for a moment.
117 player?.cancel()
118 player = $.clock.every(PLAY_EVERY_MS, () => {
119 if (!feature('pet') || workingTurn || playing || cruising) return
120 const play = PLAYS[plays++ % PLAYS.length] as (typeof PLAYS)[number]
121 const steps = PLAY_FRAMES[play] * PLAY_LOOPS[play]
122 act = play
123 frame = 0
124 playing = $.clock.every(PLAY_STEP_MS, () => {
125 frame++
126 if (frame >= steps) {
127 stopPlay()
128 act = 'rest'
129 frame = 0
130 }
131 $.ui.invalidate('ui.render')
132 })
133 $.ui.invalidate('ui.render')
134 })
135 // Subagents working in the background, with no turn running: the pilot
136 // cruises until they are all done. Asked of the engine every few seconds,
137 // so one that was stopped or failed never leaves it flying.
138 watcher?.cancel()
139 watcher = $.clock.every(AGENTS_EVERY_MS, async () => {
140 if (!feature('pet')) return
141 let busy = false
142 try {
143 busy = (await $.agent.list()).some((agent) => agent.status === 'running')
144 } catch {
145 busy = false
146 }
147 if (busy && !workingTurn && !cruising) {
148 stopPlay()
149 act = 'fly'
150 frame = 0
151 cruising = $.clock.every(CRUISE_MS, () => {
152 frame++
153 $.ui.invalidate('ui.render')
154 })
155 } else if (!busy && cruising) {
156 stopCruise()
157 if (!workingTurn) {
158 act = 'rest'
159 frame = 0
160 $.ui.invalidate('ui.render')
161 }
162 }
163 })
164 return result
165 })
166
167 // A new session in this worker (/clear, resume): the old timers stop.
168 on('session.end', { reason: /(?:)/ }, async ($, e, next) => {
169 flying?.cancel()
170 flying = null
171 stopPlay()
172 stopCruise()
173 setBoost(false)
174 act = 'rest'
175 workingTurn = null
176 running.clear()
177 return next(e)
178 })
179
180 on('command.run', { command: 'jev' }, async ($, e) => {
181 // `/jev tune [apply|reset]`: the changes the ledger suggests.
182 const tune = /^\s*tune(?:\s+(apply|reset))?\s*$/i.exec(e.args)
183 if (tune) {
184 if (!tuningLoaded()) setTuning(tuningOf(await $.store.get(TUNING_KEY).catch(() => undefined)))
185 const entries = entriesOf(await $.store.get(LEDGER_KEY).catch(() => undefined))
186 const action = tune[1]?.toLowerCase()
187 if (action === 'reset') {
188 setTuning({ values: {}, since: Date.now() })
189 await $.store.delete(TUNING_KEY).catch(() => undefined)
190 return { text: `tuning cleared; back to your settings.\n\n${describeTuning(entries)}` }
191 }
192 if (action === 'apply') {
193 const found = proposals(entries)
194 if (found.length === 0) return { text: `nothing to apply.\n\n${describeTuning(entries)}` }
195 const next = applied(currentTuning(), found, Date.now())
196 setTuning(next)
197 await $.store.set(TUNING_KEY, next).catch((error) => $.ui.log(`[jev-pilot] tuning not saved: ${String(error)}`))
198 return {
199 text: `applied ${found.map((f) => `${f.option} ${f.from} → ${f.to}`).join(', ')}. The next suggestion waits for 20 new turns.\n\n${describeTuning(entries)}`,
200 }
201 }
202 return { text: describeTuning(entries) }
203 }
204 // The crew: `/jev status`, `/jev mode ...`, `/jev alpha <model>|off`,
205 // `/jev junior <slot>`, `/jev reviewer <codex|opencode>`.
206 const crewCommand = /^\s*status\s*$/i.test(e.args) ? ({ kind: 'show' } as const) : parseCrewCommand(e.args)
207 if (crewCommand) {
208 const crewIo: CrewIo = {
209 fetch: (url, init) => $.http.fetch(url, init),
210 home: () => $.env.get('HOME'),
211 routerUrl: () => $.env.get('JEV_ROUTER_URL'),
212 write: (path, text) => $.fs.write(path, text),
213 read: async (path) => ((await $.fs.exists(path)) ? $.fs.read(path) : null),
214 run: (argv, timeoutMs) => $.process.run(argv, { timeoutMs }),
215 storeGet: (key) => $.store.get(key),
216 sleep: (ms) => $.clock.sleep(ms),
217 register: async (spec) => {
218 await $.agent.register(spec)
219 },
220 }
221 const openrouterKey = typeof options.openrouterApiKey === 'string' && options.openrouterApiKey ? options.openrouterApiKey : null
222 await ensureCrew(crewIo, openrouterKey)
223 if (crewCommand.kind === 'unknown') return { text: `jev-pilot: unknown "${crewCommand.text}".\n${describeCrew(crew(), router() !== null, codexModels())}` }
224 if (crewCommand.kind === 'show') {
225 await checkCrew(crewIo, openrouterKey)
226 await registerCrew(crewIo)
227 return { text: `${describeCrew(crew(), router() !== null, codexModels())}\n${describeHealth()}\n\n${describeStats()}` }
228 }
229 if (crewCommand.kind === 'slot') return { text: describeSlot(crew(), crewCommand.slot, crewOverrides().recent ?? []) }
230 // What was pasted, checked against OpenRouter's list before it's set.
231 let change: CrewCommand = crewCommand
232 let about: string | null = null
233 // A reviewer's model and effort, checked against the CLI's own list before it's set.
234 if (crewCommand.kind === 'reviewer-paste') {
235 const current = crew().reviewerChoices[crewCommand.reviewer] ?? {}
236 let resolved: ReviewerResolved
237 if (crewCommand.reviewer === 'codex') {
238 if (codexModels().length === 0) await refreshCodexModels(crewIo)
239 resolved = resolveCodexChoice(crewCommand.input, crewCommand.effort, codexModels(), current)
240 } else {
241 resolved = resolveOpencodeChoice(crewCommand.input, crewCommand.effort, crewCommand.input ? await opencodeModels(crewIo) : [], current)
242 }
243 if (!resolved.ok) {
244 const list = resolved.suggestions.length > 0 ? `\nAvailable:\n${resolved.suggestions.map((line) => ` ${line}`).join('\n')}` : ''
245 return { text: `The ${crewCommand.reviewer} reviewer is not changed: ${resolved.why}.${list}` }
246 }
247 change = { kind: 'reviewer-choice', reviewer: crewCommand.reviewer, choice: resolved.choice }
248 about = resolved.about || null
249 }
250 if (crewCommand.kind === 'paste') {
251 const resolved = resolveModel(crewCommand.input, await modelCatalog(crewIo))
252 if (!resolved.ok) {
253 const tries = resolved.suggestions.length > 0 ? `\nDid you mean:\n${resolved.suggestions.map((id) => ` /jev ${crewCommand.slot} ${id}`).join('\n')}` : ''
254 return { text: `${crewCommand.slot} not changed: ${resolved.why}.${tries}\nFind one at ${MODELS_PAGE} and paste its id, link or name.` }
255 }
256 // A new name past the limit: each custom model is an option Jev weighs.
257 if (!crew().slots.some((slot) => slot.name === crewCommand.slot) && crew().slots.length >= MAX_MODELS) {
258 return { text: `${crewCommand.slot} not added: ${MAX_MODELS} custom models is the most at once. Remove one first: /jev remove <name>` }
259 }
260 change = { kind: 'model', slot: crewCommand.slot, model: resolved.id, about: resolved.about }
261 about = resolved.about
262 }
263 if (change.kind === 'model' && !change.model && !crew().slots.some((slot) => slot.name === change.slot)) {
264 return { text: `No custom model is called ${change.slot}.\n\n${describeCrew(crew(), router() !== null, codexModels())}` }
265 }
266 setCrewOverrides(applyCrewCommand(crewOverrides(), change))
267 await $.store.set(CREW_KEY, crewOverrides()).catch((error) => $.ui.log(`[jev-pilot] crew not saved: ${String(error)}`))
268 await publishSlots(crewIo).catch((error) => $.ui.log(`[jev-pilot] custom models not published: ${String(error)}`))
269 if (change.kind === 'model' && change.model) await checkSlot(crewIo, openrouterKey, change.slot, change.model)
270 // A new junior, say: its agent type from the next turn, and the model
271 // told about the change with its next prompt.
272 await registerCrew(crewIo)
273 resetBriefing()
274 const headline =
275 change.kind === 'model'
276 ? change.model
277 ? `${change.slot} is now ${change.model}${about ? ` (${about})` : ''}. Saved for every session.\n` +
278 `It's in /model from your next claude-jev session (press s there to use it for that session only). To use it as the main model right now: /model ${slotAlias(change.slot)}. That also makes it your default for new sessions, and /model default undoes that.\n\n`
279 : `${change.slot} is removed, from every session and from /model.\n\n`
280 : change.kind === 'reviewer-choice'
281 ? `${change.reviewer} reviews now run on ${describeChoice(change.choice, codexModels())}${about ? ` (${about})` : ''}. Saved for every session. To use another model for one review, just say so ("review it with codex, luna, high").\n\n`
282 : ''
283 // A mode that hands work to custom models, with none set: say how to set one.
284 const unset =
285 change.kind === 'mode' && (change.mode === 'budget' || change.mode === 'junior-lead') && crew().slots.length === 0
286 ? `\n\nNo custom model is set yet, so this mode has nothing to hand work to. Add one from ${MODELS_PAGE}:\n /jev <name you choose> <model>`
287 : ''
288 return { text: `${headline}${describeCrew(crew(), router() !== null, codexModels())}\n${describeHealth()}${unset}` }
289 }
290 const command = parseJevCommand(e.args)
291 if (command.kind === 'unknown') {
292 return { text: `jev-pilot: unknown "${command.text}".\n${describeFeatures()}` }
293 }
294 if (command.kind === 'reset') loadFeatureOverrides({})
295 if (command.kind === 'set') for (const name of command.features) setFeature(name, command.on)
296 if (command.kind !== 'status') {
297 await $.store.set(FEATURES_KEY, featureOverrides()).catch((error) => $.ui.log(`[jev-pilot] switches not saved: ${String(error)}`))
298 $.ui.invalidate('ui.render')
299 }
300 if (command.kind === 'set' && command.features.length === 1) {
301 const name = command.features[0] as keyof typeof FEATURE_INFO
302 return { text: `jev-pilot: ${name} ${command.on ? 'on' : 'off'} (${FEATURE_INFO[name]}) · /jev ${name} ${command.on ? 'off' : 'on'} undoes it` }
303 }
304 return { text: `${describeFeatures()}\n\n${describeStats()}` }
305 })
306
307 // While a turn runs, the pilot shows what Claude is doing: thinking first.
308 on('turn.start', async ($, e, next) => {
309 const result = await next(e)
310 workingTurn = e.turnId
311 flying?.cancel()
312 stopPlay()
313 stopCruise()
314 setBoost(false)
315 running.clear()
316 act = 'think'
317 frame = 0
318 flying = $.clock.every(FLY_MS, () => {
319 if (!feature('pet')) return
320 frame++
321 $.ui.invalidate('ui.render')
322 })
323 return result
324 })
325
326 // The model's response as it streams: thinking shows as thinking, the
327 // answer's text as writing. Observed only: every chunk passes on as it came.
328 on('turn.step', { turnId: /(?:)/ }, async function* ($, e, next) {
329 if (e.agentId || e.turnId !== workingTurn) return yield* next(e)
330 if (running.size === 0) act = 'think'
331 for await (const chunk of next(e)) {
332 if (running.size === 0) {
333 if (chunk.kind === 'thinking') act = 'think'
334 else if (chunk.kind === 'text') act = 'write'
335 }
336 yield chunk
337 }
338 })
339
340 // A main-loop tool call shows what the tool does: reading, searching,
341 // editing, running a command; anything else (a subagent) flies. When it is
342 // done, the next call still running shows, or thinking.
343 on('tool.call', { tool: /(?:)/ }, async ($, e, next) => {
344 if (e.agentId || !workingTurn) return next(e)
345 const id = e.tool_use_id ?? `call-${++calls}`
346 act = actOfTool(e.tool)
347 running.set(id, act)
348 try {
349 return await next(e)
350 } finally {
351 running.delete(id)
352 if (workingTurn) act = [...running.values()].at(-1) ?? 'think'
353 }
354 })
355
356 on('turn.complete', { turnId: /(?:)/ }, async ($, e, next) => {
357 const result = await next(e)
358 if (e.turnId === workingTurn) {
359 running.clear()
360 workingTurn = null
361 flying?.cancel()
362 flying = null
363 setBoost(false)
364 act = 'rest'
365 frame = 0
366 $.ui.invalidate('ui.render')
367 }
368 return result
369 })
370
371 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
372 if (!feature('pet') || e.props.hasSurvey || e.surface !== 'terminal') return next(e)
373 const { Box, Text } = $.ui.resolve(e)
374 // The band's own width: the transcript column's while a Pane is docked.
375 const columns = e.props.bodyColumns || (e.viewport?.columns ?? 80)
376 const speech = currentSpeech()
377 const color = MOOD_COLOR[speech.mood]
378 // Goggles down when flying, and at full power (effort raised to max).
379 const rows = sceneRows(act, frame, blink, frame, act === 'fly' || isBoosted())
380 const working = (WORK_ACTS as readonly Act[]).includes(act)
381 const text = working ? `${SPINNER[frame % SPINNER.length]} ${ACT_LABEL[act as (typeof WORK_ACTS)[number]]} · ${speech.text}` : speech.text
382 return (
383 <Box flexDirection="column">
384 <Box key="jev:pet" flexDirection="row" justifyContent="flex-end" alignItems="center" columnGap={1} width={columns} paddingRight={4}>
385 {/* The bubble gives way: a long line is cut, the pilot is never squeezed. */}
386 <Box key="jev:bubble" borderStyle="round" borderColor={color} paddingX={1} flexShrink={1}>
387 <Text key="jev:say" color={color} wrap="truncate-end">
388 {text}
389 </Text>
390 </Box>
391 <Box key="jev:sprite" flexDirection="column" flexShrink={0} width={CANVAS_W} minWidth={CANVAS_W}>
392 {rows.map((row, y) => (
393 <Text key={`jev:row${y}`} wrap="truncate-end">
394 {row.map((cell, x) => (
395 <Text key={`jev:${y}:${x}`} color={cell.fg} backgroundColor={cell.bg}>
396 {cell.ch}
397 </Text>
398 ))}
399 </Text>
400 ))}
401 </Box>
402 </Box>
403 {await next(e)}
404 </Box>
405 )
406 })
407}
408hooks/features.ts 109 lines1/**
2 * jev-pilot — what is switched on. Every part can be turned off on its own,
3 * from the settings (defaults) or live with `/jev <feature> on|off`, which is
4 * remembered in the plugin's store across sessions.
5 *
6 * Shared: the modules run in the plugin's one worker, so they all read these.
7 */
8
9export type Feature = 'effort' | 'raise' | 'subagents' | 'skills' | 'strategy' | 'quality' | 'design' | 'model' | 'pet'
10
11export const FEATURES: readonly Feature[] = ['effort', 'raise', 'subagents', 'skills', 'strategy', 'quality', 'design', 'model', 'pet']
12
13/** What each switch does, for `/jev`. */
14export const FEATURE_INFO: Record<Feature, string> = {
15 effort: 'sets the reasoning effort of each turn',
16 raise: 'raises the effort when tool calls keep failing',
17 subagents: 'picks each subagent’s model and effort',
18 skills: 'picks the one skill a prompt needs (off: the full skill list stays)',
19 strategy: 'advises splitting big work across subagents',
20 quality: 'asks before guessing, tests bugs first, checks costly changes, and steps back when a turn goes in circles',
21 design: 'gives UI design work your design skills, a real product\u2019s DESIGN.md or Mobbin screens for direction, and the web guidelines to check against',
22 model: 'switches the main conversation’s model (resets the prompt cache)',
23 pet: 'shows Claude the pilot above the prompt',
24}
25
26let defaults: Record<Feature, boolean> = {
27 effort: true,
28 raise: true,
29 subagents: true,
30 skills: true,
31 strategy: true,
32 quality: true,
33 design: true,
34 model: false,
35 pet: true,
36}
37let overrides: Partial<Record<Feature, boolean>> = {}
38
39/** The defaults, from the plugin's options (once, when the plugin registers). */
40export function initFeatures(from: Record<Feature, boolean>): void {
41 defaults = { ...from }
42 overrides = {}
43}
44
45export function feature(name: Feature): boolean {
46 return overrides[name] ?? defaults[name]
47}
48
49export function setFeature(name: Feature, on: boolean): void {
50 if (on === defaults[name]) delete overrides[name]
51 else overrides[name] = on
52}
53
54/** What `/jev` changed, to keep in the store. */
55export function featureOverrides(): Partial<Record<Feature, boolean>> {
56 return { ...overrides }
57}
58
59/** Overrides read back from the store; anything that is not one is ignored. */
60export function loadFeatureOverrides(stored: unknown): void {
61 overrides = {}
62 if (!stored || typeof stored !== 'object' || Array.isArray(stored)) return
63 for (const [name, on] of Object.entries(stored as Record<string, unknown>)) {
64 if ((FEATURES as readonly string[]).includes(name) && typeof on === 'boolean') {
65 setFeature(name as Feature, on)
66 }
67 }
68}
69
70export type JevCommand =
71 | { kind: 'status' }
72 | { kind: 'set'; features: Feature[]; on: boolean }
73 | { kind: 'reset' }
74 | { kind: 'unknown'; text: string }
75
76/**
77 * `/jev` arguments:
78 * (none) what is on
79 * <feature> on|off one switch
80 * all on|off every switch
81 * on | off the pet (a shortcut)
82 * reset back to the settings' defaults
83 */
84export function parseJevCommand(args: string): JevCommand {
85 const words = args.trim().toLowerCase().split(/\s+/).filter(Boolean)
86 if (words.length === 0) return { kind: 'status' }
87 if (words.length === 1 && words[0] === 'reset') return { kind: 'reset' }
88 const onOff = (word: string | undefined) => (word === 'on' ? true : word === 'off' ? false : null)
89 if (words.length === 1) {
90 const on = onOff(words[0])
91 return on === null ? { kind: 'unknown', text: words[0] as string } : { kind: 'set', features: ['pet'], on }
92 }
93 const on = onOff(words[1])
94 if (on === null || words.length > 2) return { kind: 'unknown', text: args.trim() }
95 if (words[0] === 'all') return { kind: 'set', features: [...FEATURES], on }
96 if ((FEATURES as readonly string[]).includes(words[0] as string)) return { kind: 'set', features: [words[0] as Feature], on }
97 return { kind: 'unknown', text: words[0] as string }
98}
99
100/** `/jev`'s answer: every switch with its state. */
101export function describeFeatures(): string {
102 const width = Math.max(...FEATURES.map((name) => name.length))
103 return [
104 'jev-pilot switches (/jev <name> on|off, /jev all on|off, /jev reset):',
105 ...FEATURES.map((name) => ` ${feature(name) ? 'on ' : 'off'} ${name.padEnd(width)} ${FEATURE_INFO[name]}`),
106 'also: /jev status (the crew and its health) · /jev mode <name> · /jev tune (changes learned from your turns)',
107 ].join('\n')
108}
109hooks/crew-state.ts 124 lines1/**
2 * jev-pilot — the crew as this session has it: the options, what `/jev`
3 * changed, whether the router is up, and each worker's last health check.
4 * Shared, like features.ts: the router module routes with it, the pet module's
5 * `/jev` changes and shows it.
6 */
7import { crewOf, type CodexModel, type Crew, type CrewOverrides, type Reviewer, type Slot } from './crew.ts'
8
9export type Health = { ok: boolean; detail: string; at: number }
10
11let options: Record<string, unknown> = {}
12let overrides: CrewOverrides = {}
13let routerUrl: string | null = null
14let started = false
15const health = new Map<string, Health>()
16// Codex's model list as it was last read (`codex debug models`): what a tier such as `luna` means now.
17let codexList: CodexModel[] = []
18
19export function codexModels(): CodexModel[] {
20 return codexList
21}
22
23export function setCodexModels(list: CodexModel[]): void {
24 codexList = [...list]
25}
26
27export function initCrew(from: Record<string, unknown>): void {
28 options = { ...from }
29 overrides = {}
30 routerUrl = null
31 started = false
32 health.clear()
33}
34
35/** Whether this worker has set the crew up: after a reload it hasn't, until a hook does. */
36export function crewStarted(): boolean {
37 return started
38}
39
40export function markCrewStarted(): void {
41 started = true
42}
43
44export function crew(): Crew {
45 return crewOf(options, overrides)
46}
47
48export function crewOverrides(): CrewOverrides {
49 return { ...overrides }
50}
51
52export function setCrewOverrides(next: CrewOverrides): void {
53 overrides = { ...next }
54}
55
56/** The router's address when `claude-jev` started one and it answered; null otherwise. */
57export function router(): string | null {
58 return routerUrl
59}
60
61export function setRouter(url: string | null): void {
62 routerUrl = url
63}
64
65/** Health keys: `router`, `slot:alpha`, `agent:codex`, ... */
66export function setHealth(key: string, value: Health): void {
67 health.set(key, value)
68}
69
70export function healthOf(key: string): Health | null {
71 return health.get(key) ?? null
72}
73
74/** A slot is usable once its last check passed (unchecked counts as not yet). */
75export function slotHealthy(slot: Slot): boolean {
76 const h = health.get(`slot:${slot.name}`)
77 return !!h && h.ok && h.detail.startsWith(slot.model)
78}
79
80/**
81 * A slot that may be used: answering its last check, or not checked yet (a
82 * headless run's first prompt arrives before the check). An unchecked slot
83 * is safe to try: if it fails, the router gives the request to Claude.
84 * Only a slot that failed its check is left out.
85 */
86export function slotUsable(slot: Slot): boolean {
87 const h = health.get(`slot:${slot.name}`)
88 return !h || !h.detail.startsWith(slot.model) || h.ok
89}
90
91export function reviewerHealthy(reviewer: Reviewer): boolean {
92 return health.get(`agent:${reviewer}`)?.ok === true
93}
94
95/** The health lines for `/jev status`. */
96export function describeHealth(): string {
97 const c = crew()
98 const mark = (h: Health | null) => (h === null ? '· not checked' : h.ok ? `✓ ${h.detail}` : `✗ ${h.detail}`)
99 const lines = ['health:']
100 lines.push(` router ${routerUrl ? mark(health.get('router') ?? null) : '✗ not running (start Claude Code with claude-jev)'}`)
101 for (const slot of c.slots) lines.push(` ${slot.name.padEnd(10)} ${mark(health.get(`slot:${slot.name}`) ?? null)}`)
102 for (const agent of ['codex', 'opencode']) lines.push(` ${agent.padEnd(10)} ${mark(health.get(`agent:${agent}`) ?? null)}`)
103 return lines.join('\n')
104}
105
106// ---- the checks' verdicts, pure: the hooks run them and hand the output here ----
107
108/** Codex is working when `codex login status` exits 0 and says it's logged in. */
109export function codexVerdict(exitCode: number, output: string): { ok: boolean; detail: string } {
110 if (exitCode !== 0) return { ok: false, detail: `codex login status failed: ${output.trim().slice(0, 80) || `exit ${exitCode}`}` }
111 const line = output.trim().split('\n')[0] ?? ''
112 return /^logged in\b/i.test(line) ? { ok: true, detail: line } : { ok: false, detail: `not logged in: ${line.slice(0, 80)}` }
113}
114
115/** OpenCode is working when it runs and has at least one provider logged in. */
116export function opencodeVerdict(versionExit: number, version: string, authOutput: string): { ok: boolean; detail: string } {
117 if (versionExit !== 0) return { ok: false, detail: 'opencode not found or not starting' }
118 const plain = authOutput.replace(/\x1b\[[0-9;]*m/g, '')
119 const providers = [...plain.matchAll(/●\s+([^\n]+?)\s+(?:oauth|api|wellknown)\b/g)].map((m) => (m[1] as string).trim())
120 return providers.length > 0
121 ? { ok: true, detail: `opencode ${version.trim()} · ${providers.slice(0, 3).join(', ')}${providers.length > 3 ? '…' : ''}` }
122 : { ok: false, detail: `opencode ${version.trim()} has no provider logged in (opencode auth login)` }
123}
124hooks/context.ts 203 lines1/**
2 * jev-pilot — what the decision model is shown of the conversation, and
3 * which prompts are tasks at all.
4 *
5 * Pure, like the policy modules: the hooks read `$.session.messages()` and
6 * hand the list here. Only message text and tool names travel, never a tool's
7 * input or output: those hold file contents and command output, which is more
8 * than a classifier needs and more than should leave the machine.
9 */
10
11/** Prompt origins that are not a task of the person's: nothing to plan or suggest for. */
12export const NOT_A_TASK: ReadonlySet<string> = new Set([
13 'task-notification',
14 'peer',
15 'peer-send-message',
16 'projects-relay',
17 'observer',
18 'observer-activity',
19 'scheduled-trigger',
20 'slack-ping',
21])
22
23/** The part of a transcript message this module reads (`SessionMessage`). */
24export interface ContextMessage {
25 role: 'user' | 'assistant'
26 text: string
27 toolUses?: readonly { tool: string; isError?: true }[]
28 toolResults?: readonly { isError: boolean }[]
29}
30
31export interface ContextLimits {
32 /** How many messages before the prompt to include; 0 sends none. */
33 messages: number
34 /** The most characters all of them may take together. */
35 chars: number
36}
37
38/** A message as one line: who, what they said, and which tools ran. */
39function lineOf(message: ContextMessage, cap: number, tail = 0): string | null {
40 const text = message.text.replace(/\s+/g, ' ').trim()
41 const tools = (message.toolUses ?? []).map((use) => (use.isError ? `${use.tool} (failed)` : use.tool))
42 if (!text && tools.length === 0) return null
43 // With a tail, a long message keeps its beginning and its end: where a
44 // reply asks its question ("Shall I start?") is usually the end.
45 const said =
46 text.length <= cap ? text : tail > 0 && tail < cap ? `${text.slice(0, cap - tail)} … ${text.slice(-tail)}` : `${text.slice(0, cap)}…`
47 const ran = tools.length > 0 ? ` [tools: ${tools.join(', ')}]` : ''
48 return `${message.role}: ${said}${ran}`
49}
50
51/**
52 * The conversation just before `prompt`, newest last, as the text the
53 * decision model reads beside it; '' when there is none or `messages` is 0.
54 *
55 * The prompt itself is dropped when the transcript already holds it, so it is
56 * never counted twice. Messages that carry only tool results (no text) are
57 * skipped: their outcome already shows on the tool call as "(failed)". The
58 * newest messages win the character budget; older ones are dropped, not
59 * squeezed. The newest is always sent, cut to the budget if it must be.
60 */
61export function recentContext(
62 messages: readonly ContextMessage[],
63 prompt: string,
64 limits: ContextLimits,
65): string {
66 if (limits.messages <= 0 || limits.chars <= 0) return ''
67 let list = messages
68 const last = list.at(-1)
69 if (last && last.role === 'user' && last.text.trim() === prompt.trim()) list = list.slice(0, -1)
70
71 // The newest assistant message is what a short reply answers ("yes", "1",
72 // "fix all and continue"), and its proposal is usually at its end: it gets
73 // about half the budget and keeps its beginning and its end; the others
74 // share the rest.
75 const cap = Math.max(200, Math.floor(limits.chars / limits.messages))
76 let newestAssistant = -1
77 for (let index = list.length - 1; index >= 0; index--) {
78 if ((list[index] as ContextMessage).role === 'assistant' && (list[index] as ContextMessage).text.trim()) {
79 newestAssistant = index
80 break
81 }
82 }
83 const bigCap = Math.min(limits.chars, Math.max(cap, Math.floor(limits.chars * 0.55)))
84 const otherCap = newestAssistant >= 0 && limits.messages > 1 ? Math.max(200, Math.floor((limits.chars - bigCap) / (limits.messages - 1))) : cap
85 const lines: string[] = []
86 let used = 0
87 for (let index = list.length - 1; index >= 0 && lines.length < limits.messages; index--) {
88 const message = list[index] as ContextMessage
89 const line = index === newestAssistant ? lineOf(message, bigCap, Math.floor(bigCap * 0.72)) : lineOf(message, otherCap)
90 if (!line) continue
91 if (used + line.length > limits.chars) {
92 // A budget smaller than one message still carries the newest one.
93 if (lines.length === 0) lines.unshift(`${line.slice(0, Math.max(0, limits.chars - 1))}…`)
94 break
95 }
96 lines.unshift(line)
97 used += line.length + 1
98 }
99 return lines.join('\n')
100}
101
102/**
103 * Plain facts about a request, sent beside it so the decision model does not
104 * have to infer them from prose: how long it is, how many files it names,
105 * whether it carries code or an error, whether it is phrased as a question,
106 * and what the recent turns did with their tools. Counts and flags only;
107 * nothing here quotes the conversation.
108 */
109export interface Signals {
110 prompt_chars: number
111 files_mentioned: number
112 has_code_or_error: boolean
113 is_question: boolean
114 /** Tool use over the last `window` messages, by kind. */
115 recent_tools: { edits: number; commands: number; reads: number; subagents: number; failed: number }
116}
117
118const TOOL_KINDS: Record<string, keyof Omit<Signals['recent_tools'], 'failed'>> = {
119 Edit: 'edits',
120 MultiEdit: 'edits',
121 Write: 'edits',
122 NotebookEdit: 'edits',
123 Bash: 'commands',
124 PowerShell: 'commands',
125 Read: 'reads',
126 Grep: 'reads',
127 Glob: 'reads',
128 LS: 'reads',
129 WebFetch: 'reads',
130 WebSearch: 'reads',
131 Agent: 'subagents',
132 Task: 'subagents',
133}
134
135const PATH = /(?:^|[\s`'"(])((?:[\w.-]+\/)+[\w.-]+|[\w-]+\.(?:tsx?|jsx?|mjs|py|go|rs|java|kt|swift|rb|php|cs|c|cc|cpp|h|hpp|sql|json|ya?ml|toml|md|css|scss|html|sh|lock))(?=$|[\s`'"),:;])/g
136const CODE_OR_ERROR = /```|Traceback|Exception|\berror\b|\bfailed\b|stack ?trace|\bat \S+:\d+/i
137const QUESTION = /^(what|why|how|when|where|which|who|can|could|does|do|did|is|are|should|would|will)\b/i
138
139export function signalsOf(prompt: string, messages: readonly ContextMessage[], window = 10): Signals {
140 const text = prompt.trim()
141 const files = new Set<string>()
142 for (const match of text.matchAll(PATH)) files.add(match[1] as string)
143 const tools = { edits: 0, commands: 0, reads: 0, subagents: 0, failed: 0 }
144 for (const message of messages.slice(-window)) {
145 for (const use of message.toolUses ?? []) {
146 const kind = TOOL_KINDS[use.tool]
147 if (kind) tools[kind]++
148 if (use.isError) tools.failed++
149 }
150 }
151 return {
152 prompt_chars: text.length,
153 files_mentioned: files.size,
154 has_code_or_error: CODE_OR_ERROR.test(text),
155 is_question: text.endsWith('?') || QUESTION.test(text),
156 recent_tools: tools,
157 }
158}
159
160// ---- what the project deploys with ------------------------------------------------------
161
162/**
163 * The files that show which platform a project deploys to or builds on, by
164 * what they're called. Read from the project's top folder and one level down
165 * (infra/cdk.json, apps/api/Dockerfile), never their contents.
166 */
167const PLATFORM_MARKERS: [RegExp, string][] = [
168 [/(^|\/)vercel\.json$|(^|\/)\.vercel\/$/, 'Vercel'],
169 [/(^|\/)cdk\.json$/, 'AWS CDK'],
170 [/(^|\/)serverless\.(yml|yaml|ts|js)$/, 'AWS Serverless'],
171 [/(^|\/)template\.ya?ml$|(^|\/)samconfig\.toml$/, 'AWS SAM'],
172 [/(^|\/)buildspec\.ya?ml$/, 'AWS CodeBuild'],
173 [/(^|\/)amplify\.ya?ml$|(^|\/)amplify\/$/, 'AWS Amplify'],
174 [/(^|\/)netlify\.toml$/, 'Netlify'],
175 [/(^|\/)firebase\.json$/, 'Firebase'],
176 [/(^|\/)supabase\/config\.toml$/, 'Supabase'],
177 [/(^|\/)fly\.toml$/, 'Fly.io'],
178 [/(^|\/)wrangler\.(toml|json|jsonc)$/, 'Cloudflare Workers'],
179 [/(^|\/)app\.ya?ml$|(^|\/)cloudbuild\.ya?ml$/, 'Google Cloud'],
180 [/(^|\/)render\.ya?ml$/, 'Render'],
181 [/(^|\/)railway\.(json|toml)$/, 'Railway'],
182 [/(^|\/)(docker-)?compose\.ya?ml$/, 'Docker Compose'],
183 [/(^|\/)Dockerfile$/, 'Docker'],
184 [/(^|\/)\.github\/workflows\/$/, 'GitHub Actions'],
185 [/(^|\/)bitbucket-pipelines\.yml$/, 'Bitbucket Pipelines'],
186]
187
188/** Folders never looked into for platform files: dependencies and build output. */
189export const SKIPPED_DIRS = new Set(['node_modules', '.git', 'dist', 'build', '.next', 'out', 'cdk.out', 'vendor', 'target', '.venv', 'venv', '__pycache__', 'coverage'])
190
191/**
192 * What a project deploys to or builds on, from its file names (folders end
193 * in "/"): "AWS CDK, Docker", or "none found". Jev reads it so a platform's
194 * skill (Vercel's, say) fits only a project that uses that platform.
195 */
196export function platformsOf(paths: readonly string[]): string {
197 const found: string[] = []
198 for (const [marker, name] of PLATFORM_MARKERS) {
199 if (!found.includes(name) && paths.some((path) => marker.test(path))) found.push(name)
200 }
201 return found.length > 0 ? found.join(', ') : 'none found'
202}
203hooks/summary.ts 105 lines1/**
2 * jev-pilot — the one line each turn gets in the transcript, and the note the
3 * skill module leaves for it.
4 *
5 * Both modules run in the plugin's one worker, so this module is shared: the
6 * skill module notes its pick per prompt at `prompt.submit`, and the router
7 * reads it when the turn starts and writes a single line for both, instead
8 * of each module logging every step (that detail is `verboseLog`).
9 */
10
11import { sure } from './pet-art.ts'
12
13/** What the skill module decided for one prompt. */
14export interface SkillNote {
15 /** The skill attached, or null for none. */
16 skill: string | null
17 /** How long its requests took, ms; null when none was made. */
18 ms: number | null
19}
20
21const notes = new Map<string, SkillNote>()
22const MAX_NOTES = 32
23
24/** Records the skill module's pick for `prompt`, bounded to the last few. */
25export function noteSkill(prompt: string, note: SkillNote): void {
26 notes.delete(prompt)
27 notes.set(prompt, note)
28 while (notes.size > MAX_NOTES) notes.delete(notes.keys().next().value as string)
29}
30
31/** The pick noted for `prompt`, removed as it is read; null when none. */
32export function takeSkill(prompt: string | null): SkillNote | null {
33 if (prompt === null) return null
34 const note = notes.get(prompt) ?? null
35 notes.delete(prompt)
36 return note
37}
38
39/** A new session: no pick carries over. */
40export function clearSkillNotes(): void {
41 notes.clear()
42}
43
44/** What the turn line says. */
45export interface TurnFacts {
46 /** Whether the decision model answered for this turn. */
47 answered: boolean
48 /** How sure the decision model was of the effort it picked. */
49 confidence: number | null
50 /** The effort this turn was set to, or null when left as built. */
51 applied: string | null
52 /** The effort the engine built the turn with. */
53 current: string | null
54 /** The level the decision model's answer pointed to. */
55 wanted: string | null
56 /** The router's request time, ms. */
57 jevMs: number | null
58 skill: SkillNote | null
59 /** The strategy advised to the model, if any. */
60 advised: string | null
61}
62
63/**
64 * One line for a turn, the effort with how sure the decision model was of it:
65 * jev · low (93% sure) · no skill · 1.3s
66 * jev · xhigh (88% sure) · skill /systematic-debugging · parallel advised · 1.4s
67 * jev · high kept (wanted low, 42% sure) · no skill · 1.2s
68 */
69export function turnLine(facts: TurnFacts): string {
70 const parts = ['jev']
71 if (!facts.answered) {
72 parts.push('no answer, turn left as built')
73 } else {
74 const read = facts.confidence === null ? '' : sure(facts.confidence)
75 if (facts.applied) {
76 parts.push(`${facts.applied}${read ? ` (${read})` : ''}`)
77 } else if (facts.wanted && facts.current && facts.wanted !== facts.current) {
78 parts.push(`${facts.current} kept (wanted ${facts.wanted}${read ? `, ${read}` : ''})`)
79 } else {
80 parts.push(`${facts.current ?? 'default effort'}${read ? ` (${read})` : ''}`)
81 }
82 }
83 if (facts.skill) parts.push(facts.skill.skill ? `skill /${facts.skill.skill}` : 'no skill')
84 if (facts.advised) parts.push(`${facts.advised} advised`)
85 const ms = (facts.jevMs ?? 0) + (facts.skill?.ms ?? 0)
86 if (ms > 0) parts.push(`${(ms / 1000).toFixed(1)}s`)
87 return parts.join(' · ')
88}
89
90// ---- the note to the model: once per session, and again after a compaction ----
91
92let briefed = false
93
94/** Whether the model still needs jev-pilot's capability note; marks it given. */
95export function takeBriefing(): boolean {
96 if (briefed) return false
97 briefed = true
98 return true
99}
100
101/** A new session or a compacted conversation: the note is due again. */
102export function resetBriefing(): void {
103 briefed = false
104}
105hooks/pet-art.ts 409 lines1/**
2 * jev-pilot — the pet: Claude Code's character as a pilot, drawn in
3 * half-block characters at the bottom right, and what it says.
4 *
5 * Pure: the sprite as rows of cells, the speech texts, and the one shared
6 * speech state the hooks update and the pet's render hook reads (the modules
7 * run in the plugin's one worker, so this state is shared between them).
8 */
9
10/** How the pet feels about the last thing it did: sets the bubble's color. */
11export type Mood = 'ready' | 'calm' | 'focused' | 'boost' | 'alert'
12
13export const MOOD_COLOR: Record<Mood, string> = {
14 ready: '#8b949e',
15 calm: '#7cf0c4',
16 focused: '#c9d2ea',
17 boost: '#d4ff4f',
18 alert: '#ff8f8f',
19}
20
21export interface Speech {
22 text: string
23 mood: Mood
24}
25
26let speech: Speech = { text: 'ready', mood: 'ready' }
27
28export function say(text: string, mood: Mood): void {
29 speech = { text, mood }
30}
31
32export function currentSpeech(): Speech {
33 return speech
34}
35
36// Full power: the effort was raised to max this turn. The pilot wears its
37// goggles down until the turn ends.
38let boosted = false
39
40export function setBoost(on: boolean): void {
41 boosted = on
42}
43
44export function isBoosted(): boolean {
45 return boosted
46}
47
48/**
49 * A subagent's name in the bubble: its task's short description (the Agent
50 * tool's `description`, e.g. "Fix S2a Codex findings"), cut to fit; its type
51 * (`general-purpose`, `Explore`) when it has none.
52 */
53export function subagentLabel(description: string | undefined, type: string, max = 32): string {
54 const text = (description ?? '').replace(/\s+/g, ' ').trim()
55 if (!text) return type
56 return text.length <= max ? text : `${text.slice(0, max - 1).trimEnd()}…`
57}
58
59/** The mood an effort level reads as. */
60export function moodOf(effort: string | null): Mood {
61 if (effort === 'low' || effort === 'medium') return 'calm'
62 if (effort === 'high') return 'focused'
63 if (effort === 'xhigh' || effort === 'max') return 'boost'
64 return 'ready'
65}
66
67/** How sure the decision model was of the effort, as people read it: `89% sure`. */
68export function sure(confidence: number): string {
69 return `${Math.round(confidence * 100)}% sure`
70}
71
72/**
73 * What the pet says when a turn starts:
74 * low · no skill · 89% sure
75 * jev busy · left as is (the backend overloaded; the turn runs as set)
76 * xhigh · /systematic-debugging · parallel · 88% sure
77 * high kept · wanted low · 42% sure
78 */
79export function turnSpeech(facts: {
80 answered: boolean
81 /** Why the decision model gave no answer, when it gave none. */
82 miss?: 'timeout' | 'busy' | 'error' | null
83 applied: string | null
84 current: string | null
85 wanted: string | null
86 confidence: number | null
87 skill: string | null | undefined
88 advised: string | null
89}): Speech {
90 if (!facts.answered) {
91 const why = facts.miss === 'busy' ? 'jev busy' : facts.miss === 'error' ? 'jev error' : 'no answer in time'
92 return { text: `${why} · left as is`, mood: 'alert' }
93 }
94 const parts: string[] = []
95 const effort = facts.applied ?? facts.current
96 if (facts.applied) parts.push(facts.applied)
97 else if (facts.wanted && facts.current && facts.wanted !== facts.current) parts.push(`${facts.current} kept`, `wanted ${facts.wanted}`)
98 else parts.push(facts.current ?? 'effort as set')
99 if (facts.skill !== undefined) parts.push(facts.skill ? `/${facts.skill}` : 'no skill')
100 if (facts.advised) parts.push(facts.advised)
101 if (facts.confidence !== null) parts.push(sure(facts.confidence))
102 return { text: parts.join(' · '), mood: moodOf(effort) }
103}
104
105// ---- the sprite -------------------------------------------------------------
106
107/** One character of the sprite: a glyph and its colors. */
108export interface Cell {
109 ch: string
110 fg?: string
111 bg?: string
112}
113
114const CORAL = '#D97757'
115export const PALETTE: Record<string, string> = {
116 C: CORAL, // the pilot's body
117 E: '#0b1020', // eyes
118 L: '#d4ff4f', // goggle lenses
119 W: '#ffffff', // the lenses' glint
120 G: '#2b3a67', // goggle strap
121 S: '#7cf0c4', // scarf
122 F: '#ffd166', // flame
123 f: '#ff7a59', // flame, outer
124 R: '#e6edf3', // jump rope, thought dots, cursor
125 B: '#5b8def', // book cover
126 P: '#f5f0e1', // book pages
127 M: '#56607d', // magnifier rim
128 H: '#a06a3f', // magnifier handle, pencil wood
129 r: '#ff5f57', // terminal: close
130 y: '#febc2e', // terminal: minimise
131 g: '#28c840', // terminal: zoom
132 l: '#9fe7ff', // magnifier lens
133 K: '#8b949e', // laptop
134 k: '#0b1020', // laptop screen
135}
136
137/**
138 * The pilot: Claude Code's character (the same pilot as the README banner and
139 * the demo video) at 3/4 of the banner's 16x12, every part kept: goggles
140 * pushed up on the forehead (pulled down over the eyes to fly), their lenses
141 * glinting, the head, two-pixel eyes, both rows of arms, a teal scarf, long
142 * legs, and the jet flames under them.
143 */
144const CLAWD = [
145 '.GWLGGGGWLG.',
146 '.CCCCCCCCCC.',
147 '.CCECCCCECC.',
148 'CCCECCCCECCC',
149 'CCCCCCCCCCCC',
150 '.SSSSSSSSSS.',
151 '..C.C..C.C..',
152 '..C.C..C.C..',
153]
154const FLAMES = ['..F.f..F.f..', '..f.F..f.F..']
155
156/**
157 * The canvas every scene draws on: fixed, so the band never changes size.
158 * The pilot's own part is the left BODY_W columns; to its right, what it holds.
159 */
160export const BODY_W = 14
161export const CANVAS_W = 24
162export const CANVAS_H = 10
163
164/**
165 * What the pet is doing. While a turn runs, what Claude is doing:
166 * think thinking: eyes up, thought dots rising
167 * read reading files or pages: a book, its pages turning
168 * search searching: a magnifier sweeping, the eyes following it
169 * write editing files: typing code on a laptop
170 * run running commands: a terminal prompt, the cursor blinking
171 * fly anything else (subagents, other tools): flying, goggles down
172 * Idle:
173 * rest hovering, the scarf flapping, blinking now and then
174 * rope play: jumping rope
175 * wave play: waving
176 * look play: looking around
177 */
178export type Act = 'think' | 'read' | 'search' | 'write' | 'run' | 'fly' | 'rest' | 'rope' | 'wave' | 'look'
179
180/** The acts a turn shows, by what Claude is doing. */
181export const WORK_ACTS = ['think', 'read', 'search', 'write', 'run', 'fly'] as const
182
183/** What the bubble says Claude is doing, for each working act. */
184export const ACT_LABEL: Record<(typeof WORK_ACTS)[number], string> = {
185 think: 'thinking',
186 read: 'reading',
187 search: 'searching',
188 write: 'writing',
189 run: 'running',
190 fly: 'working',
191}
192
193/** The act a tool call shows: what the tool does, by its name. */
194export function actOfTool(tool: string): Act {
195 if (/^(Read|NotebookRead|WebFetch|ReadMcpResource)/.test(tool)) return 'read'
196 if (/^(Grep|Glob|LS|WebSearch|ToolSearch)$/.test(tool)) return 'search'
197 if (/^(Edit|MultiEdit|Write|NotebookEdit)$/.test(tool)) return 'write'
198 if (/^(Bash|PowerShell|Monitor)$/.test(tool)) return 'run'
199 return 'fly'
200}
201
202/** The idle plays, in the order they take turns. */
203export const PLAYS = ['rope', 'wave', 'look'] as const
204export type Play = (typeof PLAYS)[number]
205
206/** How many frames one loop of each idle play has. */
207export const PLAY_FRAMES: Record<Play, number> = { rope: 4, wave: 2, look: 4 }
208
209type Grid = string[][]
210
211function blank(): Grid {
212 return Array.from({ length: CANVAS_H }, () => Array.from({ length: CANVAS_W }, () => '.'))
213}
214
215function paste(grid: Grid, rows: readonly string[], ox: number, oy: number): void {
216 rows.forEach((row, y) => {
217 for (let x = 0; x < row.length; x++) {
218 const ch = row[x] as string
219 if (ch !== '.') dot(grid, ox + x, oy + y, ch)
220 }
221 })
222}
223
224function dot(grid: Grid, x: number, y: number, ch: string): void {
225 if (y >= 0 && y < CANVAS_H && x >= 0 && x < CANVAS_W) (grid[y] as string[])[x] = ch
226}
227
228function setAt(row: string, x: number, ch: string): string {
229 return row.slice(0, x) + ch + row.slice(x + 1)
230}
231
232/** The pilot for one frame of one act, before it is placed. */
233function clawdFor(act: Act, frame: number, blink: boolean, goggles: boolean): string[] {
234 let rows = [...CLAWD]
235 if (act === 'look' && !blink) {
236 // The eyes glance left, back, right, back.
237 const shift = [-1, 0, 1, 0][frame % 4] as number
238 rows = rows.map((row, y) => {
239 if (y !== 2 && y !== 3) return row
240 return setAt(setAt(row.replace(/E/g, 'C'), 3 + shift, 'E'), 8 + shift, 'E')
241 })
242 }
243 if (act === 'think') rows[3] = (rows[3] as string).replace(/E/g, 'C') // eyes up
244 if (act === 'read' || act === 'search' || act === 'write' || act === 'run') {
245 // The eyes turn to what the pilot holds at its side (down, for the page).
246 rows = rows.map((row, y) => (y === 2 || y === 3 ? setAt(setAt(row.replace(/E/g, 'C'), 4, 'E'), 9, 'E') : row))
247 if (act !== 'search') rows[2] = (rows[2] as string).replace(/E/g, 'C')
248 }
249 if (blink) rows = rows.map((row) => row.replace(/E/g, 'C'))
250 if (act === 'wave' && frame % 2 === 0) {
251 // The right arm up beside the head.
252 for (const y of [3, 4]) rows[y] = setAt(rows[y] as string, 11, '.')
253 for (const y of [1, 2]) rows[y] = setAt(rows[y] as string, 11, 'C')
254 }
255 if (goggles) {
256 // Goggles down over the eyes, the strap round the head: flying, or at
257 // full power.
258 rows[0] = '.CCCCCCCCCC.'
259 rows[1] = '.CCCCCCCCCC.'
260 rows[2] = '.GWLGGGGWLG.'
261 rows[3] = 'CCLLCCCCLLCC'
262 }
263 if (act === 'fly') {
264 // The flames flicker long and short, their colors steady.
265 rows.push(...(frame % 2 === 1 ? [FLAMES[0] as string] : FLAMES))
266 } else if (act !== 'rope') {
267 // Hovering: the flames on, steady. (Jumping rope, the feet do the work.)
268 rows.push(FLAMES[0] as string)
269 }
270 return rows
271}
272
273/**
274 * The pixel canvas for one frame of one act. `wind` is no longer used (the
275 * scarf's end that flapped in it is gone) and is kept only so callers need
276 * not change. `goggles` pulls them down over the eyes: always when flying.
277 */
278export function scenePixels(act: Act, frame = 0, blink = false, wind = frame, goggles = act === 'fly'): string[] {
279 const grid = blank()
280 const clawd = clawdFor(act, frame, blink, goggles)
281 const X = 1
282 if (act === 'fly') {
283 // Flying: a bob, down a pixel on the short flame.
284 paste(grid, clawd, X, frame % 2)
285 } else if (act === 'rope') {
286 // Four beats: the rope overhead, coming down in front, under the feet
287 // (the pilot up in the air), coming round behind.
288 const beat = frame % 4
289 const lift = beat === 2 ? 2 : beat === 1 ? 1 : 0
290 const top = CANVAS_H - clawd.length - lift
291 paste(grid, clawd, X, top)
292 const hands = top + 3
293 const right = BODY_W - 1
294 if (beat === 0) {
295 for (let x = 1; x < right; x++) dot(grid, x, 0, 'R')
296 for (let y = 1; y < hands; y++) dot(grid, 0, y, 'R'), dot(grid, right, y, 'R')
297 } else if (beat === 2) {
298 for (let x = 1; x < right; x++) dot(grid, x, CANVAS_H - 1, 'R')
299 for (let y = hands + 1; y < CANVAS_H - 1; y++) dot(grid, 0, y, 'R'), dot(grid, right, y, 'R')
300 } else if (beat === 1) {
301 for (let y = hands; y < CANVAS_H; y++) dot(grid, 0, y, 'R'), dot(grid, right, y, 'R')
302 } else {
303 for (let y = 0; y <= hands; y++) dot(grid, 0, y, 'R'), dot(grid, right, y, 'R')
304 }
305 } else {
306 paste(grid, clawd, X, CANVAS_H - clawd.length)
307 drawProp(grid, act, frame)
308 }
309 return grid.map((row) => row.join(''))
310}
311
312/** Pixel art for what the pilot holds, each drawn to read as the thing at a glance. */
313const CLOUD = ['.RR.RRR..', 'RRRRRRRRR', 'RRRRRRRRR', '.RRRRRRR.']
314const BOOK = [
315 '.PPP.PPP.',
316 'BkkPMkkPB',
317 'BPPPMPPPB',
318 'BkkPMkPPB',
319 'BBBBBBBBB',
320]
321const LENS = ['.MMMM.', 'MRRllM', 'MRlllM', 'MllllM', 'MllllM', '.MMMM.']
322const PAPER = ['PPPPM.', 'PPPPPP', 'PPPPPP', 'PPPPPP', 'PPPPPP', 'PPPPPP', 'PPPPPP', 'PPPPPP']
323const TERMINAL = [
324 'KKKKKKKKK',
325 'KrKyKgKKK',
326 'KkkkkkkkK',
327 'KkkkkkkkK',
328 'KkkkkkkkK',
329 'KkkkkkkkK',
330 'KkkkkkkkK',
331 'KKKKKKKKK',
332]
333
334/**
335 * What the pilot holds at its side while working, right of its body, by the
336 * right hand (resting, the arm ends at column 12, rows 4 and 5).
337 */
338function drawProp(grid: Grid, act: Act, frame: number): void {
339 const art = (x: number, y: number, rows: readonly string[]) => paste(grid, rows, x, y)
340 const line = (x: number, y: number, pixels: string) => paste(grid, [pixels], x, y)
341 const X = BODY_W // the first column right of the pilot
342 if (act === 'think') {
343 // Thought bubbles rising from the head into a cloud, "..." filling in.
344 art(X, 0, CLOUD)
345 dot(grid, 12, 3, 'R')
346 dot(grid, X - 1, 2, 'R')
347 const dots = Math.floor(frame / 2) % 4
348 for (let i = 0; i < dots; i++) dot(grid, X + 2 + i * 2, 2, 'k')
349 } else if (act === 'read') {
350 // An open book: two pages of text, the fold between them; the line
351 // being read lights up, down the left page, then the right.
352 art(X - 1, 3, BOOK)
353 const lines: [number, number, number][] = [[X, 4, 2], [X, 6, 2], [X + 4, 4, 2], [X + 4, 6, 1]]
354 const [lx, ly, len] = lines[frame % lines.length] as [number, number, number]
355 line(lx, ly, 'S'.repeat(len))
356 } else if (act === 'search') {
357 // A magnifying glass held out by its handle, from the hand to the rim,
358 // sweeping in and out, the lens glinting.
359 const dx = [0, 1, 2, 1][frame % 4] as number
360 art(X + 2 + dx, 1, LENS)
361 for (let x = X - 1; x < X + 2 + dx; x++) dot(grid, x, 5, 'H')
362 if (frame % 4 === 1) dot(grid, X + 5 + dx, 2, 'R')
363 } else if (act === 'write') {
364 // A sheet of paper, its corner turned; lines of writing appear as a
365 // pencil moves along them.
366 art(X + 1, 1, PAPER)
367 const done = frame % 12
368 for (let i = 0; i < done; i++) dot(grid, X + 2 + (i % 4), 3 + Math.floor(i / 4) * 2, 'k')
369 const px = X + 2 + (done % 4)
370 const py = 3 + Math.floor(done / 4) * 2
371 // The pencil: point, wood, yellow body, eraser, leaning up and right.
372 dot(grid, px, py, 'G')
373 dot(grid, px + 1, py - 1, 'H')
374 dot(grid, px + 2, py - 2, 'F')
375 dot(grid, px + 3, py - 3, 'f')
376 } else if (act === 'run') {
377 // A terminal window: its three buttons, output scrolling, a > prompt
378 // and a blinking cursor.
379 art(X, 1, TERMINAL)
380 const outputs = ['LLL.LLL', 'LL.LLL.', 'LLLLL.L', 'L.LLLL.', 'LLL.LL.']
381 for (let i = 0; i < 2; i++) line(X + 1, 3 + i, outputs[(frame + i) % outputs.length] as string)
382 dot(grid, X + 1, 5, 'S')
383 dot(grid, X + 2, 6, 'S')
384 dot(grid, X + 1, 7, 'S')
385 if (frame % 2 === 0) line(X + 4, 7, 'RR')
386 }
387}
388
389/** The canvas as terminal lines of cells, two pixel rows per line. */
390export function sceneRows(act: Act, frame = 0, blink = false, wind = frame, goggles = act === 'fly'): Cell[][] {
391 const pixels = scenePixels(act, frame, blink, wind, goggles)
392 const rows: Cell[][] = []
393 for (let y = 0; y < pixels.length; y += 2) {
394 const top = pixels[y] as string
395 const bottom = pixels[y + 1] ?? ''
396 const row: Cell[] = []
397 for (let x = 0; x < top.length; x++) {
398 const up = PALETTE[top[x] as string]
399 const down = PALETTE[bottom[x] ?? '.']
400 if (up && down) row.push({ ch: '▀', fg: up, bg: down })
401 else if (up) row.push({ ch: '▀', fg: up })
402 else if (down) row.push({ ch: '▄', fg: down })
403 else row.push({ ch: ' ' })
404 }
405 rows.push(row)
406 }
407 return rows
408}
409hooks/crew.ts 790 lines1/**
2 * jev-pilot — the crew: which workers a session can use and how the user wants
3 * them used. Pure: the hooks read options and the store, and act on this.
4 *
5 * Workers:
6 * Claude haiku, sonnet, opus (the plan's own models)
7 * model slots alpha, beta, gamma: any OpenRouter model, reached through the
8 * local router (router/jev-router.mjs) as `jev-<slot>`
9 * agents Codex and OpenCode, their own CLIs, as reviewers
10 *
11 * Modes (the user's choice, `/jev mode <name>`); Jev decides within one:
12 * standard Claude only (the default)
13 * budget model slots take subagent work that needs no judgment
14 * junior-lead a junior on a slot writes easy, well-specified code;
15 * the main conversation reviews it as the tech lead
16 * second-opinion an external agent reviews significant changes
17 * quality Opus for every subagent, and an external review
18 */
19
20import { FEATURES } from './features.ts'
21
22export type Mode = 'standard' | 'budget' | 'junior-lead' | 'second-opinion' | 'quality'
23export const MODES: readonly Mode[] = ['standard', 'budget', 'junior-lead', 'second-opinion', 'quality']
24
25export const MODE_INFO: Record<Mode, string> = {
26 standard: 'Claude only: Jev picks Haiku, Sonnet or Opus and the effort for each task',
27 budget: 'custom models take subagent work that needs no judgment; Opus keeps the judgment',
28 'junior-lead': 'a junior on a custom model writes easy, well-specified code; Opus reviews it as tech lead',
29 'second-opinion': 'an external agent (Codex or OpenCode) reviews significant changes before they are done',
30 quality: 'Opus for every subagent, and an external review of significant changes',
31}
32
33export type Reviewer = 'codex' | 'opencode'
34export const REVIEWERS: readonly Reviewer[] = ['codex', 'opencode']
35
36export interface Slot {
37 /** Its name, as you type it after /jev: `flash` in `/jev flash deepseek/...`. */
38 name: string
39 /** The OpenRouter model id, e.g. deepseek/deepseek-v4.1-flash. */
40 model: string
41 /** When to choose it: what the decision model reads. */
42 when: string
43 /** What OpenRouter says it is, when known: "DeepSeek: DeepSeek V4.1 Flash · 1M context · $0.14 in · …". */
44 about?: string
45}
46
47/**
48 * A custom model's name, as you type it after `/jev`: lowercase letters,
49 * digits and "-", starting with a letter. Claude Code sees it as `jev-<name>`.
50 */
51export const SLOT_NAME = /^[a-z][a-z0-9-]{0,23}$/
52
53/** Words `/jev` already means something by: never a custom model's name. */
54export const RESERVED_NAMES: ReadonlySet<string> = new Set([
55 ...FEATURES,
56 'all', 'reset', 'status', 'mode', 'models', 'crew', 'junior', 'reviewer', 'tune',
57 'remove', 'delete', 'on', 'off', 'help', 'list', 'default', 'set', 'when',
58])
59
60/** A name `/jev <name> <model>` may use. */
61export function validName(name: string): boolean {
62 return SLOT_NAME.test(name) && !RESERVED_NAMES.has(name)
63}
64
65/** At most this many custom models at once: each is an option in Jev's question. */
66export const MAX_MODELS = 8
67
68/** How many entries a record (the store, models.json) is read for, removed ones included. */
69export const RECORD_LIMIT = 64
70
71/** The names the `alphaModel`/`betaModel`/`gammaModel` settings fill (before names were yours to choose). */
72export const LEGACY_NAMES = ['alpha', 'beta', 'gamma'] as const
73
74/**
75 * Read-only bulk work only: code is written on the main model (Anthropic's
76 * own guidance for Opus 5.5, and our junior-mode runs agreed: a cheaper
77 * model writing code cost more once the lead's review was counted).
78 */
79export const DEFAULT_SLOT_WHEN =
80 'Choose for read-only bulk work where cost matters more than precision: searching or reading across many files and reporting what is there, summarizing logs or test output, listing. Never for writing or changing code.'
81
82/** The model name Claude Code uses for a slot; the router maps it to the slot's model. */
83export function slotAlias(name: string): string {
84 return `jev-${name}`
85}
86
87/** What `/jev` has changed. */
88export interface CrewOverrides {
89 mode?: Mode
90 /** Each custom model by its name; '' once removed (so a setting can't bring it back). */
91 models?: Record<string, string>
92 junior?: string
93 reviewer?: Reviewer
94 /** The models set before, newest first, so switching back is one command. */
95 recent?: string[]
96 /** What OpenRouter said each model is, by id. */
97 about?: Record<string, string>
98 /** The model and effort each reviewer runs with, when not its CLI's own default. */
99 reviewerChoices?: Partial<Record<Reviewer, ReviewerChoice>>
100}
101
102/** A reviewer's model and effort; either may be left to the CLI's own config. */
103export interface ReviewerChoice {
104 model?: string
105 effort?: string
106}
107
108/** How many models `/jev <name>` remembers. */
109export const RECENT_MODELS = 5
110
111export interface Crew {
112 mode: Mode
113 /** The custom models set, in the order they were added. */
114 slots: Slot[]
115 /** The junior's model name ('' when there is none). */
116 junior: string
117 reviewer: Reviewer
118 reviewerChoices: Partial<Record<Reviewer, ReviewerChoice>>
119}
120
121/** The crew from the options and what `/jev` changed. */
122export function crewOf(options: Record<string, unknown>, overrides: CrewOverrides = {}): Crew {
123 const text = (key: string, fallback: string) => (typeof options[key] === 'string' ? (options[key] as string).trim() : fallback)
124 const models: Record<string, string> = { ...(overrides.models ?? {}) }
125 for (const name of LEGACY_NAMES) if (!(name in models)) models[name] = text(`${name}Model`, '')
126 const slots: Slot[] = []
127 for (const [name, model] of Object.entries(models)) {
128 // A record written elsewhere with more than the limit: the first ones count.
129 if (slots.length >= MAX_MODELS) break
130 if (!model || !validName(name)) continue
131 const about = overrides.about?.[model]
132 slots.push({ name, model, when: text(`${name}When`, '') || DEFAULT_SLOT_WHEN, ...(about ? { about } : {}) })
133 }
134 const modeOption = text('mode', 'standard') as Mode
135 const reviewerOption = text('reviewer', 'codex') as Reviewer
136 const juniorWanted = overrides.junior ?? text('junior', '')
137 return {
138 mode: overrides.mode ?? (MODES.includes(modeOption) ? modeOption : 'standard'),
139 slots,
140 // The one named, while it's set; else the first model.
141 junior: slots.some((slot) => slot.name === juniorWanted) ? juniorWanted : (slots[0]?.name ?? ''),
142 reviewer: overrides.reviewer ?? (REVIEWERS.includes(reviewerOption) ? reviewerOption : 'codex'),
143 reviewerChoices: { ...(overrides.reviewerChoices ?? {}) },
144 }
145}
146
147/** A reviewer model id as the CLIs name them (checked against their lists when set). */
148const REVIEWER_MODEL_ID = /^[\w.:~-]+(?:\/[\w.:~-]+)?$/
149/** A reasoning effort word (the CLIs' own names: low, medium, high, xhigh, max, ultra…). */
150const EFFORT_WORD = /^[a-z]{2,12}$/
151
152/** A reviewer choice read back from a record; anything that isn't one is dropped. */
153function reviewerChoicesOf(raw: unknown): Partial<Record<Reviewer, ReviewerChoice>> | undefined {
154 if (!raw || typeof raw !== 'object' || Array.isArray(raw)) return undefined
155 const out: Partial<Record<Reviewer, ReviewerChoice>> = {}
156 for (const reviewer of REVIEWERS) {
157 const entry = (raw as Record<string, unknown>)[reviewer]
158 if (!entry || typeof entry !== 'object') continue
159 const { model, effort } = entry as { model?: unknown; effort?: unknown }
160 const choice: ReviewerChoice = {}
161 if (typeof model === 'string' && model.length <= 100 && REVIEWER_MODEL_ID.test(model)) choice.model = model
162 if (typeof effort === 'string' && EFFORT_WORD.test(effort)) choice.effort = effort
163 if (choice.model || choice.effort) out[reviewer] = choice
164 }
165 return out
166}
167
168/** Overrides read back from the store or models.json; anything that isn't one is dropped. */
169export function overridesOf(stored: unknown): CrewOverrides {
170 if (!stored || typeof stored !== 'object' || Array.isArray(stored)) return {}
171 const raw = stored as Record<string, unknown>
172 const out: CrewOverrides = {}
173 if (typeof raw.mode === 'string' && MODES.includes(raw.mode as Mode)) out.mode = raw.mode as Mode
174 if (typeof raw.junior === 'string' && validName(raw.junior)) out.junior = raw.junior
175 if (typeof raw.reviewer === 'string' && REVIEWERS.includes(raw.reviewer as Reviewer)) out.reviewer = raw.reviewer as Reviewer
176 if (Array.isArray(raw.recent)) {
177 out.recent = raw.recent.filter((id): id is string => typeof id === 'string' && MODEL_ID.test(id)).slice(0, RECENT_MODELS)
178 }
179 if (raw.models && typeof raw.models === 'object' && !Array.isArray(raw.models)) {
180 const models: Record<string, string> = {}
181 for (const [name, value] of Object.entries(raw.models as Record<string, unknown>).slice(0, RECORD_LIMIT)) {
182 if (validName(name) && typeof value === 'string' && (value.trim() === '' || MODEL_ID.test(value.trim()))) models[name] = value.trim()
183 }
184 out.models = models
185 }
186 const choices = reviewerChoicesOf(raw.reviewerChoices)
187 if (choices) out.reviewerChoices = choices
188 if (raw.about && typeof raw.about === 'object' && !Array.isArray(raw.about)) {
189 const about: Record<string, string> = {}
190 for (const [id, value] of Object.entries(raw.about as Record<string, unknown>).slice(0, RECORD_LIMIT)) {
191 if (MODEL_ID.test(id) && typeof value === 'string') about[id] = value.slice(0, 200)
192 }
193 out.about = about
194 }
195 return out
196}
197
198/**
199 * ~/.claude/jev-pilot/models.json: the router's table, and the record of
200 * what you set: every model by its name (a removed one as ""), what
201 * OpenRouter said it is, and the models set before. Every session, in any
202 * project and whichever way jev-pilot is installed, reads the same choice back.
203 */
204export function routerTable(crew: Crew, overrides: Pick<CrewOverrides, 'models' | 'recent' | 'reviewerChoices'> = {}): string {
205 const slots: Record<string, { model: string; about?: string }> = {}
206 for (const [name, model] of Object.entries(overrides.models ?? {})) if (model === '' && validName(name)) slots[name] = { model: '' }
207 for (const slot of crew.slots) slots[slot.name] = { model: slot.model, ...(slot.about ? { about: slot.about } : {}) }
208 const reviewers = overrides.reviewerChoices ?? {}
209 return JSON.stringify({ slots, recent: (overrides.recent ?? []).slice(0, RECENT_MODELS), ...(Object.keys(reviewers).length > 0 ? { reviewers } : {}) }, null, 2)
210}
211
212/**
213 * The models recorded in models.json, as overrides; null when the file
214 * isn't one (missing, or not jev-pilot's).
215 */
216export function recordedModels(fileText: string | null): Pick<CrewOverrides, 'models' | 'recent' | 'about' | 'reviewerChoices'> | null {
217 if (!fileText) return null
218 let parsed: unknown
219 try {
220 parsed = JSON.parse(fileText)
221 } catch {
222 return null
223 }
224 const slots = (parsed as { slots?: unknown })?.slots
225 if (!slots || typeof slots !== 'object' || Array.isArray(slots)) return null
226 const models: Record<string, string> = {}
227 const about: Record<string, string> = {}
228 for (const [name, entry] of Object.entries(slots as Record<string, { model?: unknown; about?: unknown } | undefined>).slice(0, RECORD_LIMIT)) {
229 const model = entry?.model
230 if (!validName(name) || typeof model !== 'string' || !(model === '' || MODEL_ID.test(model))) continue
231 models[name] = model
232 if (model && typeof entry?.about === 'string') about[model] = entry.about.slice(0, 200)
233 }
234 const recent = overridesOf({ recent: (parsed as { recent?: unknown }).recent }).recent
235 const reviewerChoices = reviewerChoicesOf((parsed as { reviewers?: unknown }).reviewers) ?? {}
236 return { models, ...(recent ? { recent } : {}), ...(Object.keys(about).length > 0 ? { about } : {}), reviewerChoices }
237}
238
239/**
240 * The rows jev-pilot adds to Claude Code's `/model` list, one per custom
241 * model (the `modelPicker` setting, in a file `claude-jev` passes with
242 * `--settings`, so they're there only where the router is). `behavesAs`
243 * lets Claude Code, which doesn't know the model, treat it like Sonnet on
244 * its side (prompt, capabilities, effort); the requests still go to the
245 * model itself.
246 */
247export function pickerSettings(crew: Crew): string {
248 const options = crew.slots.map((slot) => {
249 const [name, ...details] = (slot.about ?? '').split(' · ')
250 return {
251 model: slotAlias(slot.name),
252 label: `${slot.name}${name ? ` · ${name.replace(/^[^:]+:\s*/, '')}` : ''}`,
253 description: [`${slot.model} on OpenRouter`, ...details.filter((d) => !/context$/.test(d))].join(' · ') + ' · claude-jev only',
254 behavesAs: 'sonnet',
255 }
256 })
257 return JSON.stringify({ modelPicker: { options } }, null, 2)
258}
259
260/**
261 * The slots Jev may choose for a subagent. Only with the router running, and
262 * only in the modes that hand work to custom models.
263 */
264export function slotsOffered(crew: Crew, routerOn: boolean): Slot[] {
265 if (!routerOn) return []
266 return crew.mode === 'budget' || crew.mode === 'junior-lead' ? crew.slots : []
267}
268
269/** The junior's slot, when the junior can work: junior-lead mode, the router up, the slot set. */
270export function juniorSlot(crew: Crew, routerOn: boolean): Slot | null {
271 if (!routerOn || crew.mode !== 'junior-lead') return null
272 return crew.slots.find((slot) => slot.name === crew.junior) ?? null
273}
274
275/** Whether significant changes get an external review. */
276export function reviews(crew: Crew): boolean {
277 return crew.mode === 'second-opinion' || crew.mode === 'quality'
278}
279
280export type CrewCommand =
281 | { kind: 'show' }
282 | { kind: 'mode'; mode: Mode }
283 /** `model` is the resolved id ('' removes the model); `about` what OpenRouter says it is. */
284 | { kind: 'model'; slot: string; model: string; about?: string }
285 /** `/jev <name> <what you pasted>`: resolved against OpenRouter's list before it's set. */
286 | { kind: 'paste'; slot: string; input: string }
287 /** `/jev <name>`: that model, and the ones set before. */
288 | { kind: 'slot'; slot: string }
289 | { kind: 'junior'; slot: string }
290 | { kind: 'reviewer'; reviewer: Reviewer }
291 /** `/jev reviewer codex luna high`: resolved against the CLI's own model list before it's set. */
292 | { kind: 'reviewer-paste'; reviewer: Reviewer; input: string; effort?: string }
293 /** A reviewer's choice as set (after resolving); an empty choice goes back to the CLI's config. */
294 | { kind: 'reviewer-choice'; reviewer: Reviewer; choice: ReviewerChoice }
295 | { kind: 'unknown'; text: string }
296
297/**
298 * `/jev` arguments about the crew, or null when they're about something else
299 * (a switch such as `/jev skills off`):
300 * models the crew: mode, custom models, junior, reviewer
301 * mode <name> standard · budget · junior-lead · second-opinion · quality
302 * <name> <openrouter model> add a custom model under a name you choose, pasted as its
303 * id, page link or name; the same name again replaces it
304 * <name> that model, and the ones set before
305 * remove <name> delete it (also: <name> off)
306 * junior <name> which custom model the junior runs on
307 * reviewer <codex|opencode> which external agent reviews
308 */
309export function parseCrewCommand(args: string): CrewCommand | null {
310 const words = args.trim().split(/\s+/).filter(Boolean)
311 const head = words[0]?.toLowerCase()
312 if (!head) return null
313 if (head === 'models' || head === 'crew') return words.length === 1 ? { kind: 'show' } : { kind: 'unknown', text: args.trim() }
314 if (head === 'mode') {
315 const mode = words[1]?.toLowerCase() as Mode | undefined
316 return mode && MODES.includes(mode) && words.length === 2 ? { kind: 'mode', mode } : { kind: 'unknown', text: args.trim() }
317 }
318 if (head === 'remove' || head === 'delete') {
319 const name = words[1]?.toLowerCase()
320 return name && validName(name) && words.length === 2 ? { kind: 'model', slot: name, model: '' } : { kind: 'unknown', text: args.trim() }
321 }
322 if (head === 'junior') {
323 const name = words[1]?.toLowerCase()
324 return name && validName(name) && words.length === 2 ? { kind: 'junior', slot: name } : { kind: 'unknown', text: args.trim() }
325 }
326 if (head === 'reviewer') {
327 const reviewer = words[1]?.toLowerCase() as Reviewer | undefined
328 if (!reviewer || !REVIEWERS.includes(reviewer)) return { kind: 'unknown', text: args.trim() }
329 if (words.length === 2) return { kind: 'reviewer', reviewer }
330 const third = (words[2] as string).toLowerCase()
331 // reviewer codex default back to the CLI's own model and effort
332 // reviewer codex effort high the effort alone
333 // reviewer codex luna [high] a model, and an effort
334 if (third === 'default' && words.length === 3) return { kind: 'reviewer-choice', reviewer, choice: {} }
335 if (third === 'effort' && words.length === 4 && EFFORT_WORD.test((words[3] as string).toLowerCase())) {
336 return { kind: 'reviewer-paste', reviewer, input: '', effort: (words[3] as string).toLowerCase() }
337 }
338 if (words.length === 3 || (words.length === 4 && EFFORT_WORD.test((words[3] as string).toLowerCase()))) {
339 return { kind: 'reviewer-paste', reviewer, input: words[2] as string, ...(words[3] ? { effort: (words[3] as string).toLowerCase() } : {}) }
340 }
341 return { kind: 'unknown', text: args.trim() }
342 }
343 if (validName(head)) {
344 const input = args.trim().slice(head.length).trim()
345 if (!input) return { kind: 'slot', slot: head }
346 if (/^(off|remove|delete)$/i.test(input)) return { kind: 'model', slot: head, model: '' }
347 return { kind: 'paste', slot: head, input }
348 }
349 return null
350}
351
352/** Applies a command to the overrides. */
353export function applyCrewCommand(overrides: CrewOverrides, command: CrewCommand): CrewOverrides {
354 if (command.kind === 'mode') return { ...overrides, mode: command.mode }
355 if (command.kind === 'model') {
356 const recent = command.model ? [command.model, ...(overrides.recent ?? []).filter((id) => id !== command.model)].slice(0, RECENT_MODELS) : overrides.recent
357 const about = command.model && command.about ? { ...(overrides.about ?? {}), [command.model]: command.about } : overrides.about
358 const next: CrewOverrides = { ...overrides, models: { ...(overrides.models ?? {}), [command.slot]: command.model } }
359 if (recent) next.recent = recent
360 if (about) next.about = about
361 // A removed model can't stay the junior.
362 if (!command.model && next.junior === command.slot) delete next.junior
363 return next
364 }
365 if (command.kind === 'junior') return { ...overrides, junior: command.slot }
366 if (command.kind === 'reviewer') return { ...overrides, reviewer: command.reviewer }
367 if (command.kind === 'reviewer-choice') {
368 const choices = { ...(overrides.reviewerChoices ?? {}) }
369 if (command.choice.model || command.choice.effort) choices[command.reviewer] = { ...command.choice }
370 else delete choices[command.reviewer]
371 return { ...overrides, reviewerChoices: choices }
372 }
373 return overrides
374}
375
376/** `/jev models`: the crew as it stands. */
377export function describeCrew(crew: Crew, routerOn: boolean, codexList: readonly CodexModel[] = []): string {
378 const lines = [
379 `jev-pilot mode: ${crew.mode} (${MODE_INFO[crew.mode]})`,
380 ` /jev mode <${MODES.join('|')}>`,
381 `custom models (${routerOn ? 'router running' : 'router not running: start Claude Code with claude-jev'}):`,
382 ]
383 const width = Math.max(6, ...crew.slots.map((slot) => slot.name.length))
384 for (const slot of crew.slots) lines.push(` ${slot.name.padEnd(width)} ${slot.model}${slot.name === crew.junior ? ' (the junior)' : ''}`)
385 if (crew.slots.length === 0) lines.push(' none yet')
386 lines.push(` add: /jev <name> <model from openrouter.ai/models> · remove: /jev remove <name> · junior: /jev junior <name>`)
387 const said = (r: Reviewer) => describeChoice(crew.reviewerChoices[r], r === 'codex' ? codexList : [])
388 lines.push(`reviewer: ${crew.reviewer} /jev reviewer <codex|opencode>`)
389 lines.push(` codex ${said('codex')} /jev reviewer codex <model> [effort] · default`)
390 lines.push(` opencode ${said('opencode')} /jev reviewer opencode <provider/model> [effort] · default`)
391 return lines.join('\n')
392}
393
394// ---- the junior: a coder on a custom model, reviewed by the lead ----------------
395
396/** An agent type jev-pilot registers, as `$.agent.register` takes it. */
397export interface AgentSpec {
398 name: string
399 description: string
400 prompt: string
401 tools: string[]
402 model: string
403 maxTurns: number
404}
405
406export const JUNIOR_AGENT = 'jev-pilot:junior'
407
408/** The junior's agent definition, registered as `jev-pilot:junior`. */
409export function juniorSpec(model: string): AgentSpec {
410 return {
411 name: 'junior',
412 description:
413 'A junior developer on a cheaper model. Give it one easy, well-specified coding change: the files to change, the exact behavior, and the command that proves it (usually the tests). It implements and reports; review its diff yourself before calling the work done.',
414 prompt: [
415 'You are the junior developer on a small team. Your lead gives you one well-specified coding task.',
416 'Do exactly that: change only what the brief names, follow the existing code style, and run the command the brief gives to prove it (usually the tests).',
417 "Don't refactor, rename or add anything that wasn't asked for.",
418 "If the brief is unclear, or the task turns out bigger than it says, stop and say so instead of guessing.",
419 'When you are done, reply with: what you changed (file by file, one line each), the command you ran and its result, and anything you were unsure about.',
420 ].join('\n'),
421 tools: ['Read', 'Edit', 'Write', 'Grep', 'Glob', 'Bash'],
422 model,
423 maxTurns: 40,
424 }
425}
426
427// ---- the reviewers: Codex and OpenCode, their own CLI agents --------------------
428
429export function reviewerAgent(reviewer: Reviewer): string {
430 return `jev-pilot:${reviewer}-review`
431}
432
433const REVIEWER_NAME: Record<Reviewer, string> = { codex: 'Codex', opencode: 'OpenCode' }
434
435/** What the external agent is asked, ahead of the lead's brief. */
436export const REVIEW_INSTRUCTIONS = [
437 "Review a code change in this repository, read-only: don't edit anything.",
438 'Read the changed files and `git diff` (or `git diff HEAD~1` when the change is already committed).',
439 'Check that it does what the brief says, and look for wrong behavior, input that is not checked, broken edge cases and missing tests.',
440 'Report each finding as P1 (wrong behavior or lost data), P2 (a likely bug or a missing check) or P3 (minor), with file:line and a one-line fix.',
441 'End with one verdict: PASS, PASS-WITH-FOLLOWUP or NEEDS FIXES. Keep it tight.',
442].join('\n')
443
444/** A model Codex offers, as `codex debug models` lists it. */
445export interface CodexModel {
446 slug: string
447 name: string
448 description: string
449 efforts: string[]
450}
451
452/** Codex's model list from `codex debug models` (the ones it shows; hidden ones left out). */
453export function codexCatalog(json: string): CodexModel[] {
454 let parsed: unknown
455 try {
456 parsed = JSON.parse(json)
457 } catch {
458 return []
459 }
460 const list = Array.isArray(parsed) ? parsed : (parsed as { models?: unknown })?.models
461 if (!Array.isArray(list)) return []
462 return list
463 .filter((m): m is Record<string, unknown> => !!m && typeof m === 'object' && typeof (m as { slug?: unknown }).slug === 'string')
464 .filter((m) => m.visibility === undefined || m.visibility === 'list')
465 .map((m) => ({
466 slug: m.slug as string,
467 name: typeof m.display_name === 'string' ? m.display_name : (m.slug as string),
468 description: typeof m.description === 'string' ? m.description : '',
469 efforts: Array.isArray(m.supported_reasoning_levels)
470 ? (m.supported_reasoning_levels as { effort?: unknown }[]).map((level) => level?.effort).filter((e): e is string => typeof e === 'string')
471 : [],
472 }))
473}
474
475/** A model's short name, its tier, as people say it: `gpt-5.6-luna` → `luna`, `gpt-6-astra` → `astra`. */
476export function shortName(slug: string): string {
477 return slug.split('-').pop() ?? slug
478}
479
480/** The version numbers in a model id, to order a tier's models: `gpt-5.6-luna` → [5, 6]. */
481function versionOf(slug: string): number[] {
482 return (slug.match(/\d+/g) ?? []).map(Number)
483}
484
485function newer(a: number[], b: number[]): number {
486 for (let i = 0; i < Math.max(a.length, b.length); i++) {
487 const diff = (a[i] ?? 0) - (b[i] ?? 0)
488 if (diff !== 0) return diff
489 }
490 return 0
491}
492
493/** Codex's tiers (astra, sol, terra, luna…), each with its newest model now. */
494export function codexTiers(catalog: readonly CodexModel[]): Map<string, CodexModel> {
495 const tiers = new Map<string, CodexModel>()
496 for (const model of catalog) {
497 const tier = shortName(model.slug).toLowerCase()
498 if (!/^[a-z]+$/.test(tier)) continue
499 const held = tiers.get(tier)
500 if (!held || newer(versionOf(model.slug), versionOf(held.slug)) > 0) tiers.set(tier, model)
501 }
502 return tiers
503}
504
505/**
506 * The model a choice runs on now: a tier (`luna`) is its newest model in
507 * Codex's list at this moment, so a new generation is taken up by itself;
508 * an id (`gpt-5.6-luna`) stays that model. Unknown to the list: as given.
509 */
510export function resolvedChoice(reviewer: Reviewer, choice: ReviewerChoice = {}, catalog: readonly CodexModel[] = []): ReviewerChoice {
511 if (reviewer !== 'codex' || !choice.model || choice.model.includes('-')) return choice
512 const latest = codexTiers(catalog).get(choice.model.toLowerCase())
513 return latest ? { ...choice, model: latest.slug } : choice
514}
515
516export type ReviewerResolved = { ok: true; choice: ReviewerChoice; about: string } | { ok: false; why: string; suggestions: string[] }
517
518/**
519 * `/jev reviewer codex <model> [effort]`, checked against Codex's own list:
520 * the id (`gpt-5.6-luna`), its name (`GPT-5.6-Luna`) or its short name
521 * (`luna`); the effort must be one that model takes.
522 */
523export function resolveCodexChoice(input: string, effort: string | undefined, catalog: readonly CodexModel[], current: ReviewerChoice = {}): ReviewerResolved {
524 const wanted = input.trim().toLowerCase()
525 const tiers = codexTiers(catalog)
526 // A tier (`luna`) is kept as the tier: its newest model is used each time.
527 // An id or a full name (`gpt-5.6-luna`) pins that exact model.
528 let kept: string | undefined = current.model
529 let model: CodexModel | undefined
530 if (wanted) {
531 const tier = tiers.get(wanted)
532 const pinned = catalog.find((m) => m.slug.toLowerCase() === wanted) ?? catalog.find((m) => m.name.toLowerCase() === wanted)
533 model = tier ?? pinned
534 kept = tier ? wanted : pinned?.slug
535 if (!model || !kept) {
536 return {
537 ok: false,
538 why: catalog.length > 0 ? `Codex has no model "${input}"` : "Codex's model list couldn't be read (codex debug models)",
539 suggestions: [...tiers].map(([name, m]) => `${name} (now ${m.slug}): ${m.description}`),
540 }
541 }
542 } else if (kept) {
543 model = tiers.get(kept.toLowerCase()) ?? catalog.find((m) => m.slug === kept)
544 }
545 const efforts = model?.efforts ?? []
546 if (effort && efforts.length > 0 && !efforts.includes(effort)) {
547 return { ok: false, why: `${model?.slug ?? 'that model'} doesn't take effort "${effort}"`, suggestions: efforts }
548 }
549 const keptEffort = effort ?? (wanted ? undefined : current.effort)
550 const choice: ReviewerChoice = { ...(kept ? { model: kept } : {}), ...(keptEffort ? { effort: keptEffort } : {}) }
551 const about = model ? (wanted && tiers.get(wanted) ? `the newest ${wanted}, now ${model.slug}: ${model.description}` : `${model.name}: ${model.description}`) : ''
552 return { ok: true, choice, about }
553}
554
555/**
556 * `/jev reviewer opencode <model> [effort]`, checked against `opencode
557 * models` (provider/model): the full id, or a model name only one provider has.
558 */
559export function resolveOpencodeChoice(input: string, effort: string | undefined, models: readonly string[], current: ReviewerChoice = {}): ReviewerResolved {
560 const wanted = input.trim().toLowerCase()
561 let id: string | undefined = current.model
562 if (wanted) {
563 const exact = models.find((m) => m.toLowerCase() === wanted)
564 const byName = models.filter((m) => m.toLowerCase().endsWith(`/${wanted}`))
565 id = exact ?? (byName.length === 1 ? byName[0] : undefined)
566 if (!id) {
567 const close = byName.length > 1 ? byName : models.filter((m) => m.toLowerCase().includes(wanted))
568 return { ok: false, why: byName.length > 1 ? `more than one provider has "${input}"` : `OpenCode has no model "${input}"`, suggestions: close.slice(0, 5) }
569 }
570 }
571 return { ok: true, choice: { ...(id ? { model: id } : {}), ...((effort ?? (wanted ? undefined : current.effort)) ? { effort: effort ?? current.effort } : {}) }, about: '' }
572}
573
574/** "luna (newest, now gpt-5.6-luna) · effort high", or the CLI's own when nothing is chosen. */
575export function describeChoice(choice: ReviewerChoice | undefined, catalog: readonly CodexModel[] = []): string {
576 const model = choice?.model
577 const now = model && !model.includes('-') ? codexTiers(catalog).get(model.toLowerCase())?.slug : undefined
578 const parts = [model ? (now ? `${model} (newest, now ${now})` : model) : undefined, choice?.effort ? `effort ${choice.effort}` : undefined].filter(Boolean)
579 return parts.length > 0 ? parts.join(' · ') : "the CLI's own model and effort"
580}
581
582const shellQuote = (value: string) => `'${value.replace(/'/g, `'\\''`)}'`
583
584/** The model and effort flags for a reviewer's CLI, as `set --` arguments. */
585export function reviewerArgs(reviewer: Reviewer, choice: ReviewerChoice = {}): string {
586 const args: string[] = []
587 if (choice.model) args.push('-m', shellQuote(choice.model))
588 if (choice.effort) args.push(...(reviewer === 'codex' ? ['-c', shellQuote(`model_reasoning_effort="${choice.effort}"`)] : ['--variant', shellQuote(choice.effort)]))
589 return `set --${args.length > 0 ? ` ${args.join(' ')}` : ''}`
590}
591
592/** The command that runs the review, reading the brief from `$brief` and the flags from `set --`. */
593function reviewCommand(reviewer: Reviewer): string {
594 return reviewer === 'codex'
595 ? 'codex exec "$@" -s read-only --skip-git-repo-check -C "$PWD" -o "$brief.out" - < "$brief" > /dev/null 2> "$brief.err"; echo "exit $?"; cat "$brief.out" 2>/dev/null || tail -20 "$brief.err"'
596 : 'opencode run "$@" --agent plan --dir "$PWD" "$(cat "$brief")" < /dev/null 2> "$brief.err"; echo "exit $?"; [ -s "$brief.err" ] && tail -5 "$brief.err"'
597}
598
599/** How the reviewer is told which models it may switch to, when the brief asks for one. */
600function modelMenu(reviewer: Reviewer, codexModels: readonly CodexModel[]): string[] {
601 if (reviewer === 'codex') {
602 if (codexModels.length === 0) return ['If the brief asks for a Codex model or effort, pass it with -m <model id> and -c model_reasoning_effort="<effort>" in the set -- line.']
603 return [
604 'If the brief asks for a particular Codex model or effort for this review, change the set -- line to it (and only then). A model named by its tier means its newest model:',
605 ...[...codexTiers(codexModels)].map(([tier, m]) => ` ${tier} = -m '${m.slug}' (${m.description}; efforts: ${m.efforts.join(', ')})`),
606 ` an exact id the brief gives (e.g. ${codexModels[0]?.slug ?? 'gpt-…'}) = -m '<that id>'`,
607 ` effort: -c 'model_reasoning_effort="<effort>"'`,
608 "If it names a model or effort that isn't listed, don't guess: reply that it isn't available, with the list.",
609 ]
610 }
611 return [
612 "If the brief asks for a particular OpenCode model or effort, change the set -- line to -m '<provider/model>' and --variant '<effort>'. If you're unsure of the id, run `opencode models | grep -i <name>` first and use the exact line it prints; if it prints none or several, reply with them instead of guessing.",
613 ]
614}
615
616/**
617 * A reviewer's agent definition, registered as `jev-pilot:<reviewer>-review`:
618 * a small Claude model that hands the brief to the external CLI and brings
619 * its findings back, so the long review stays out of the main conversation.
620 */
621export function reviewerSpec(reviewer: Reviewer, model: string, choice: ReviewerChoice = {}, codexModels: readonly CodexModel[] = []): AgentSpec {
622 const name = REVIEWER_NAME[reviewer]
623 return {
624 name: `${reviewer}-review`,
625 description: `Gets a code review from ${name}, an external coding agent (its own CLI, not Claude), on ${describeChoice(choice)}. Brief it with what changed and why, the files, and what to check; to use another ${name} model or effort for this review, say which in the brief. It returns ${name}'s findings (P1/P2/P3) and verdict. Takes a few minutes: run it in the background when there is other work.`,
626 prompt: [
627 `You hand a code review to ${name}, an external coding agent, and bring back what it finds. You don't review the code yourself.`,
628 'Run one Bash command (timeout 600000), with the brief you were given pasted between the JEV_BRIEF lines exactly as given:',
629 '',
630 reviewerArgs(reviewer, choice),
631 'brief=$(mktemp /tmp/jev-review-XXXXXX); cat > "$brief" <<\'JEV_BRIEF\'',
632 REVIEW_INSTRUCTIONS,
633 '',
634 'The change:',
635 '<the brief>',
636 'JEV_BRIEF',
637 reviewCommand(reviewer),
638 '',
639 ...modelMenu(reviewer, codexModels),
640 '',
641 `Then reply with ${name}'s findings and verdict as it gave them, without adding your own, and say which model and effort it ran on.`,
642 `If the command fails or times out, reply with the exit code and the error lines instead, and say the review didn't run. Retry at most once, and only on a timeout.`,
643 ].join('\n'),
644 tools: ['Bash'],
645 model,
646 maxTurns: 6,
647 }
648}
649
650/**
651 * The crew's lines for the note to the main model: the mode the user chose,
652 * the junior, and the reviewers that are working.
653 */
654export function crewNote(crew: Crew, junior: Slot | null, working: Reviewer[], offered: readonly Slot[] = [], workflows = false): string[] {
655 const lines: string[] = []
656 if (crew.mode !== 'standard') lines.push(`The user chose the ${crew.mode} mode: ${MODE_INFO[crew.mode]}.`)
657 // Workflow agents never pass the Agent tool, so jev-pilot can't route them:
658 // the script sets each one's model (agent(prompt, { model })).
659 if (workflows) {
660 lines.push(
661 "Agents a Workflow script starts (agent()) don't pass through jev-pilot, so choose each one's model in the script with opts.model: 'haiku' for searching, reading and reporting, 'sonnet' for ordinary well-specified work, and leave it out for work that needs judgment.",
662 )
663 if (offered.length > 0) {
664 lines.push(
665 `In this mode, the user wants bulk work on their custom model: for a workflow agent searching, reading and reporting, or doing other work with nothing to judge, use ${offered.map((slot) => `opts.model: '${slotAlias(slot.name)}' (${slot.model})`).join(' or ')} rather than 'haiku'.`,
666 )
667 }
668 }
669 if (junior) {
670 lines.push(
671 `${JUNIOR_AGENT} is a junior developer on ${junior.model}. Give it easy, well-specified coding changes (the files, the exact behavior, the command that proves it); then read its diff. Rerun the tests only if its report doesn't show them passing or the diff goes beyond the brief. If it falls short, send your findings back once or fix small things yourself.`,
672 )
673 }
674 if (working.length > 0) {
675 lines.push(
676 `External reviewers (their own CLI agents, not Claude): ${working.map((r) => `${reviewerAgent(r)} (${REVIEWER_NAME[r]}, on ${describeChoice(crew.reviewerChoices[r])})`).join(', ')}. Use one when the user asks for a review by it. When the user names a model or effort for the review (for Codex: astra, sol, terra, luna…), put it in the brief as "Model: <name>, effort: <level>".`,
677 )
678 }
679 if (reviews(crew)) {
680 const chosen = working.includes(crew.reviewer) ? crew.reviewer : (working[0] ?? null)
681 lines.push(
682 chosen
683 ? `Before calling done a change that adds a feature or touches security, money or stored data, spawn ${reviewerAgent(chosen)} with a brief: what changed and why, the files, what to check. Fix the findings you agree with; mention briefly any you set aside.`
684 : `No external reviewer is working (${REVIEWER_NAME[crew.reviewer]} failed its check; /jev status shows why), so there is no external review: say so once when one would have been due.`,
685 )
686 }
687 return lines
688}
689
690// ---- setting a slot: what you paste from OpenRouter, checked against its list -----
691
692/** An OpenRouter model id: provider/model, optionally with a :variant or a ~ alias. */
693export const MODEL_ID = /^~?[\w.-]+\/[\w.:-]+$/
694
695/** Where to find a model to paste: OpenRouter's list, filtered to models that can call tools. */
696export const MODELS_PAGE = 'https://openrouter.ai/models?supported_parameters=tools'
697
698/** A model as OpenRouter's list (GET /api/v1/models) describes it. */
699export interface OpenRouterModel {
700 id: string
701 name?: string
702 canonical_slug?: string
703 context_length?: number
704 /** When OpenRouter added it, seconds since the epoch: the newest are suggested first. */
705 created?: number
706 pricing?: { prompt?: string; completion?: string }
707 supported_parameters?: string[]
708}
709
710export type Resolved =
711 | { ok: true; id: string; about: string }
712 | { ok: false; why: string; suggestions: string[] }
713
714const clean = (text: string) => text.trim().replace(/^[`'"<]+|[`'">]+$/g, '').trim()
715
716/** Per million tokens, as OpenRouter shows it: "$0.14 in · $0.42 out". */
717function price(model: OpenRouterModel): string {
718 const perMillion = (value: string | undefined) => {
719 const n = Number(value)
720 return Number.isFinite(n) ? `$${(n * 1e6).toFixed(n * 1e6 < 1 ? 3 : 2).replace(/0+$/, '').replace(/\.$/, '')}` : '?'
721 }
722 return `${perMillion(model.pricing?.prompt)} in · ${perMillion(model.pricing?.completion)} out per million tokens`
723}
724
725/** "DeepSeek: DeepSeek V4.1 Flash · 1M context · $0.14 in · $0.42 out per million tokens" */
726export function aboutModel(model: OpenRouterModel): string {
727 const context = model.context_length
728 ? ` · ${model.context_length >= 1e6 ? `${Math.round(model.context_length / 1e5) / 10}M` : `${Math.round(model.context_length / 1000)}k`} context`
729 : ''
730 return `${model.name ?? model.id}${context} · ${price(model)}`
731}
732
733/**
734 * What was pasted, as a model OpenRouter serves and a subagent can use.
735 * Accepted: the id (`deepseek/deepseek-v4.1-flash`), its page link
736 * (`https://openrouter.ai/deepseek/deepseek-v4.1-flash`), its dated slug, or
737 * its name as the list shows it (`DeepSeek: DeepSeek V4.1 Flash`, or without
738 * the `DeepSeek: ` prefix). Refused: anything not in the list (with the
739 * closest ids to try), and models that can't call tools: a subagent works
740 * through tools, so one without them could do nothing.
741 *
742 * With no list (OpenRouter unreachable), an id-shaped paste is taken as is;
743 * the slot's own check (a 1-token request) then says whether it answers.
744 */
745export function resolveModel(pasted: string, catalog: readonly OpenRouterModel[] | null): Resolved {
746 let text = clean(pasted)
747 const link = /^(?:https?:\/\/)?(?:www\.)?openrouter\.ai\/(?:models\/)?([^?#\s]+)/i.exec(text)
748 if (link) text = (link[1] as string).split('/').slice(0, 2).join('/')
749 if (!text) return { ok: false, why: 'nothing pasted', suggestions: [] }
750 if (!catalog) {
751 return MODEL_ID.test(text)
752 ? { ok: true, id: text, about: `${text} (OpenRouter's list couldn't be read to check it)` }
753 : { ok: false, why: `"${text}" isn't a model id (provider/model), and OpenRouter's list couldn't be read to look it up`, suggestions: [] }
754 }
755 const lower = text.toLowerCase()
756 const bare = (name: string | undefined) => (name ?? '').toLowerCase().replace(/^[^:]+:\s*/, '')
757 const found =
758 catalog.find((m) => m.id.toLowerCase() === lower) ??
759 catalog.find((m) => (m.canonical_slug ?? '').toLowerCase() === lower) ??
760 catalog.find((m) => (m.name ?? '').toLowerCase() === lower) ??
761 catalog.find((m) => bare(m.name) === lower.replace(/^[^:]+:\s*/, ''))
762 if (!found) {
763 const words = lower.split(/[^a-z0-9.]+/).filter((w) => w.length > 1)
764 const score = (m: OpenRouterModel) => words.filter((w) => `${m.id} ${m.name ?? ''}`.toLowerCase().includes(w)).length
765 const suggestions = catalog
766 .filter((m) => (m.supported_parameters ?? []).includes('tools') && score(m) > 0)
767 .sort((a, b) => score(b) - score(a) || (b.created ?? 0) - (a.created ?? 0) || a.id.length - b.id.length)
768 .slice(0, 3)
769 .map((m) => m.id)
770 return { ok: false, why: `OpenRouter has no model "${text}"`, suggestions }
771 }
772 if (!(found.supported_parameters ?? []).includes('tools')) {
773 return { ok: false, why: `${found.id} can't call tools on OpenRouter, so it can't work as a subagent`, suggestions: [] }
774 }
775 return { ok: true, id: found.id, about: aboutModel(found) }
776}
777
778/** `/jev alpha`: the slot, its model, and the models set before (to switch back). */
779export function describeSlot(crew: Crew, name: string, recent: readonly string[] = []): string {
780 const slot = crew.slots.find((s) => s.name === name)
781 const lines = [
782 slot ? `${name}: ${slot.model}${slot.about ? ` (${slot.about})` : ''}${name === crew.junior ? ', the junior' : ''}` : `${name}: no model by that name yet`,
783 ` ${slot ? 'replace it' : 'add it'}: /jev ${name} <model>, pasting its id, page link or name from ${MODELS_PAGE}`,
784 ]
785 if (slot) lines.push(` use it as the main model: /model ${slotAlias(name)} (in claude-jev) · delete it: /jev remove ${name}`)
786 const others = recent.filter((id) => id !== slot?.model)
787 if (others.length > 0) lines.push(' set before:', ...others.map((id) => ` /jev ${name} ${id}`))
788 return lines.join('\n')
789}
790hooks/crew-run.ts 286 lines1/**
2 * jev-pilot — starting the crew for a session and checking every worker.
3 *
4 * The hooks hand in the engine calls this needs (`CrewIo`); `$` never leaves
5 * the hook. Checks are cheap: a 1-token call per custom model on OpenRouter,
6 * `codex login status`, `opencode --version` and `opencode auth list`.
7 */
8import { aboutModel, codexCatalog, juniorSlot, juniorSpec, overridesOf, pickerSettings, resolvedChoice, recordedModels, REVIEWERS, reviewerSpec, routerTable, slotAlias, type AgentSpec, type OpenRouterModel } from './crew.ts'
9import { codexModels, codexVerdict, crew, crewOverrides, crewStarted, setCodexModels, markCrewStarted, opencodeVerdict, reviewerHealthy, router, setCrewOverrides, setHealth, setRouter } from './crew-state.ts'
10
11export interface CrewIo {
12 fetch: (url: string, init: { method?: string; headers?: Record<string, string>; body?: string }) => Promise<{ ok: boolean; status: number; text: string }>
13 /** $HOME and $JEV_ROUTER_URL: the only variables this reads. */
14 home: () => Promise<string | undefined>
15 routerUrl: () => Promise<string | undefined>
16 write: (path: string, text: string) => Promise<void>
17 /** A file's text, or null when it isn't there. */
18 read: (path: string) => Promise<string | null>
19 run: (argv: string[], timeoutMs: number) => Promise<{ exitCode: number; stdout: string; stderr: string }>
20 storeGet: (key: string) => Promise<unknown>
21 sleep: (ms: number) => Promise<void>
22 /** Registers an agent type, `jev-pilot:<name>`. */
23 register: (spec: AgentSpec) => Promise<void>
24}
25
26export const CREW_KEY = 'crew'
27
28/** The router's address as shown to you: without the secret in its path. */
29export function shownUrl(url: string): string {
30 return url.replace(/\/[0-9a-f]{16,}$/i, '')
31}
32
33/** Where the router reads the slots from. */
34export async function modelsFile(io: CrewIo): Promise<string> {
35 const home = (await io.home()) ?? '~'
36 return `${home}/.claude/jev-pilot/models.json`
37}
38
39/** The `/model` rows `claude-jev` passes to Claude Code with `--settings`. */
40export async function pickerFile(io: CrewIo): Promise<string> {
41 const home = (await io.home()) ?? '~'
42 return `${home}/.claude/jev-pilot/picker.json`
43}
44
45/**
46 * Writes the router's table and the record of what's set (models.json), and
47 * the `/model` rows (picker.json), from the crew as it stands now.
48 */
49export async function publishSlots(io: CrewIo): Promise<void> {
50 await io.write(await modelsFile(io), routerTable(crew(), crewOverrides()))
51 await io.write(await pickerFile(io), pickerSettings(crew()))
52}
53
54async function within<T>(io: Pick<CrewIo, 'sleep'>, ms: number, work: Promise<T>): Promise<T | null> {
55 return Promise.race([work, io.sleep(ms).then(() => null)])
56}
57
58/** A session's crew: the saved `/jev` changes, the router, the slots file. */
59export async function startCrew(io: CrewIo): Promise<void> {
60 markCrewStarted()
61 // The mode, junior and reviewer from the store; the models from
62 // models.json, which every install and project shares, and which holds
63 // the latest change made anywhere.
64 const stored = overridesOf(await io.storeGet(CREW_KEY).catch(() => undefined))
65 const recorded = recordedModels(await io.read(await modelsFile(io)).catch(() => null))
66 setCrewOverrides(recorded ? { ...stored, ...recorded } : stored)
67 const url = (await io.routerUrl())?.replace(/\/$/, '') || null
68 setRouter(null)
69 if (url) {
70 const answer = await within(io, 1500, io.fetch(`${url}/jev-router/health`, { method: 'GET' }).catch(() => null))
71 const ok = !!answer && answer.ok
72 setHealth('router', { ok, detail: ok ? shownUrl(url) : `no answer at ${shownUrl(url)}`, at: Date.now() })
73 if (ok) setRouter(url)
74 }
75 await publishSlots(io).catch(() => undefined)
76}
77
78/** Checks one custom model: a 1-token request on OpenRouter. */
79export async function checkSlot(io: CrewIo, key: string | null, name: string, model: string): Promise<void> {
80 if (!key) {
81 setHealth(`slot:${name}`, { ok: false, detail: `${model} · no OpenRouter key`, at: Date.now() })
82 return
83 }
84 const answer = await within(
85 io,
86 8000,
87 io
88 .fetch('https://openrouter.ai/api/v1/messages', {
89 method: 'POST',
90 headers: { 'content-type': 'application/json', authorization: `Bearer ${key}`, 'anthropic-version': '2023-06-01' },
91 body: JSON.stringify({ model, max_tokens: 1, messages: [{ role: 'user', content: 'ok' }] }),
92 })
93 .catch(() => null),
94 )
95 const ok = !!answer && answer.ok
96 const why = answer ? `HTTP ${answer.status}: ${answer.text.slice(0, 80)}` : 'no answer in 8s'
97 setHealth(`slot:${name}`, { ok, detail: ok ? `${model} · answering` : `${model} · ${why}`, at: Date.now() })
98}
99
100/** Checks the external agents: installed, and logged in. */
101export async function checkAgents(io: CrewIo): Promise<void> {
102 try {
103 const codex = await io.run(['codex', 'login', 'status'], 15_000)
104 const verdict = codexVerdict(codex.exitCode, `${codex.stdout}\n${codex.stderr}`)
105 setHealth('agent:codex', { ...verdict, at: Date.now() })
106 // Its models now: what a tier such as `luna` resolves to this session.
107 if (verdict.ok) await refreshCodexModels(io)
108 } catch {
109 setHealth('agent:codex', { ok: false, detail: 'codex not installed', at: Date.now() })
110 }
111 try {
112 const version = await io.run(['opencode', '--version'], 15_000)
113 const auth = version.exitCode === 0 ? await io.run(['opencode', 'auth', 'list'], 15_000) : { stdout: '', stderr: '', exitCode: 1 }
114 setHealth('agent:opencode', { ...opencodeVerdict(version.exitCode, version.stdout, `${auth.stdout}\n${auth.stderr}`), at: Date.now() })
115 } catch {
116 setHealth('agent:opencode', { ok: false, detail: 'opencode not installed', at: Date.now() })
117 }
118}
119
120/** Reads Codex's model list (`codex debug models`); kept as it was when it can't be read. */
121export async function refreshCodexModels(io: CrewIo): Promise<void> {
122 const listed = await io.run(['codex', 'debug', 'models'], 15_000).catch(() => null)
123 const models = listed && listed.exitCode === 0 ? codexCatalog(listed.stdout) : []
124 if (models.length > 0) setCodexModels(models)
125}
126
127/** OpenCode's models (`opencode models`), one provider/model id per line; [] when it can't be read. */
128export async function opencodeModels(io: CrewIo): Promise<string[]> {
129 const listed = await io.run(['opencode', 'models'], 30_000).catch(() => null)
130 if (!listed || listed.exitCode !== 0) return []
131 return listed.stdout
132 .split('\n')
133 .map((line) => line.replace(/\x1b\[[0-9;]*m/g, '').trim())
134 .filter((line) => /^[\w.:~-]+\/[\w.:~\/-]+$/.test(line))
135}
136
137/** The router still answering, and any custom model it had to hand to Claude. */
138export async function checkRouter(io: CrewIo): Promise<void> {
139 const url = router()
140 if (!url) return
141 const answer = await within(io, 1500, io.fetch(`${url}/jev-router/health`, { method: 'GET' }).catch(() => null))
142 if (!answer || !answer.ok) {
143 setHealth('router', { ok: false, detail: `no answer at ${shownUrl(url)}`, at: Date.now() })
144 return
145 }
146 setHealth('router', { ok: true, detail: `${shownUrl(url)}${fallbackNote(answer.text)}`, at: Date.now() })
147}
148
149/** "· alpha fell back to claude-sonnet-5 2× (OpenRouter 503)", from the router's health answer. */
150export function fallbackNote(healthText: string): string {
151 try {
152 const fallbacks = (JSON.parse(healthText) as { fallbacks?: Record<string, { count?: number; to?: string; why?: string }> }).fallbacks ?? {}
153 const notes = Object.entries(fallbacks)
154 .filter(([, f]) => typeof f?.count === 'number' && f.count > 0)
155 .map(([slot, f]) => `${slot} fell back to ${f.to ?? 'Claude'} ${f.count}× (last: ${String(f.why ?? '').slice(0, 60)})`)
156 return notes.length > 0 ? ` · ${notes.join('; ')}` : ''
157 } catch {
158 return ''
159 }
160}
161
162/** Every check, in parallel; then what OpenRouter says each model is, where that's missing. */
163export async function checkCrew(io: CrewIo, key: string | null): Promise<void> {
164 await Promise.all([...crew().slots.map((slot) => checkSlot(io, key, slot.name, slot.model)), checkAgents(io), checkRouter(io)])
165 await fillAbout(io).catch(() => undefined)
166}
167
168/**
169 * A model set without its description (before descriptions were kept, or
170 * while OpenRouter's list was out of reach): looked up once it's there, for
171 * `/model` and `/jev status`.
172 */
173export async function fillAbout(io: CrewIo): Promise<void> {
174 const missing = crew().slots.filter((slot) => !slot.about)
175 if (missing.length === 0) return
176 const catalog = await modelCatalog(io)
177 if (!catalog) return
178 const about = { ...(crewOverrides().about ?? {}) }
179 for (const slot of missing) {
180 const found = catalog.find((model) => model.id === slot.model)
181 if (found) about[slot.model] = aboutModel(found)
182 }
183 if (Object.keys(about).length === Object.keys(crewOverrides().about ?? {}).length) return
184 setCrewOverrides({ ...crewOverrides(), about })
185 await publishSlots(io)
186}
187
188/** The small Claude model a reviewer runs on: it only relays. */
189export const REVIEWER_MODEL = 'haiku'
190
191/**
192 * The crew's agent types: the junior in junior-lead mode (with the router
193 * up), and each reviewer whose CLI passed its check.
194 */
195export async function registerCrew(io: CrewIo): Promise<void> {
196 const junior = juniorSlot(crew(), router() !== null)
197 const specs = [
198 ...(junior ? [juniorSpec(slotAlias(junior.name))] : []),
199 // Each reviewer on its chosen model, a tier resolved to its newest now.
200 ...REVIEWERS.filter(reviewerHealthy).map((r) => reviewerSpec(r, REVIEWER_MODEL, resolvedChoice(r, crew().reviewerChoices[r], codexModels()), codexModels())),
201 ]
202 for (const spec of specs) await io.register(spec).catch(() => undefined)
203}
204
205/** A session's crew, start to finish: set up, agents registered, then checked and registered again. */
206export async function startSession(io: CrewIo, key: string | null): Promise<void> {
207 await startCrew(io).catch(() => undefined)
208 await registerCrew(io)
209 void checkCrew(io, key)
210 .then(() => registerCrew(io))
211 .catch(() => undefined)
212}
213
214/**
215 * The crew, set up if this worker hasn't yet: a reload mid-session (a plugin
216 * update) starts a fresh worker without a session start, and routing must
217 * not think the router is gone. Checks run in the background.
218 */
219export async function ensureCrew(io: CrewIo, key: string | null): Promise<void> {
220 if (!crewStarted()) await startSession(io, key)
221 else await refreshModels(io, key)
222}
223
224/**
225 * The models set in another session since this one started (models.json
226 * changed): taken up here too, and a newly set model checked. One small
227 * file read per prompt.
228 */
229export async function refreshModels(io: CrewIo, key: string | null): Promise<void> {
230 const recorded = recordedModels(await io.read(await modelsFile(io)).catch(() => null))
231 if (!recorded) return
232 const current = crewOverrides()
233 const before = crew().slots
234 const same = (a: unknown, b: unknown) => JSON.stringify(a ?? null) === JSON.stringify(b ?? null)
235 if (same(recorded.models, current.models) && same(recorded.recent, current.recent) && same(recorded.about, current.about) && same(recorded.reviewerChoices, current.reviewerChoices)) return
236 setCrewOverrides({ ...current, ...recorded })
237 for (const slot of crew().slots) {
238 if (!before.some((b) => b.name === slot.name && b.model === slot.model)) void checkSlot(io, key, slot.name, slot.model).catch(() => undefined)
239 }
240 await registerCrew(io)
241}
242
243let catalogCache: { at: number; models: OpenRouterModel[] } | null = null
244
245/**
246 * OpenRouter's model list, to check what `/jev <slot>` was given against
247 * (public, no key). Kept 10 minutes; null when it can't be read in 6 s.
248 */
249export async function modelCatalog(io: CrewIo): Promise<OpenRouterModel[] | null> {
250 if (catalogCache && Date.now() - catalogCache.at < 600_000) return catalogCache.models
251 const answer = await within(io, 6000, io.fetch('https://openrouter.ai/api/v1/models', { method: 'GET' }).catch(() => null))
252 if (!answer || !answer.ok) return null
253 try {
254 const models = (JSON.parse(answer.text) as { data?: unknown }).data
255 if (!Array.isArray(models)) return null
256 catalogCache = { at: Date.now(), models: models.filter((m): m is OpenRouterModel => !!m && typeof (m as OpenRouterModel).id === 'string') }
257 return catalogCache.models
258 } catch {
259 return null
260 }
261}
262
263let fallbacksSeen: number | null = null
264
265/**
266 * New fallbacks since the last look: each custom model that failed and was
267 * answered by Claude instead, with why. Empty on the first look (it only
268 * counts from there) and when the router isn't answering.
269 */
270export async function newFallbacks(io: Pick<CrewIo, 'fetch' | 'sleep'>): Promise<string[]> {
271 const url = router()
272 if (!url) return []
273 const answer = await within(io, 800, io.fetch(`${url}/jev-router/health`, { method: 'GET' }).catch(() => null))
274 if (!answer || !answer.ok) return []
275 try {
276 const fallbacks = (JSON.parse(answer.text) as { fallbacks?: Record<string, { count?: number; to?: string; why?: string }> }).fallbacks ?? {}
277 const total = Object.values(fallbacks).reduce((sum, f) => sum + (typeof f?.count === 'number' ? f.count : 0), 0)
278 const before = fallbacksSeen
279 fallbacksSeen = total
280 if (before === null || total <= before) return []
281 return Object.entries(fallbacks).map(([name, f]) => `${name} failed (${String(f.why ?? '').slice(0, 90)}), so ${f.to ?? 'Claude'} answered instead`)
282 } catch {
283 return []
284 }
285}
286hooks/jev-call.ts 92 lines1/**
2 * jev-pilot — one request to Jev per prompt.
3 *
4 * Two modules ask Jev about the same prompt: the router (effort, model,
5 * strategy) and the skill picker (which skill, if any). They used to ask one
6 * after the other, three requests in a row (the router's, the skill ranking,
7 * and a re-read of the shortlist): about 1.5 s before each turn. Now the
8 * router hands its questions down, and the skill module, which runs inside
9 * it on the same prompt, sends them with its own in one request. One request
10 * with both sets of questions takes as long as the slower of the two alone,
11 * and answers them the same (measured: same effort, tier and strategy on
12 * every prompt tried).
13 *
14 * The hand-off: the router's prompt.submit offers its part, then passes the
15 * prompt on; the skill module's takes it, asks (`ask`, with its own
16 * questions added), and gives the answer back (`settle`), which returns the
17 * router's block for the prompt (strategy advice). A part nobody took is
18 * asked for by the router itself once the prompt comes back up.
19 */
20import type { Miss } from './model-router.policy.ts'
21
22export interface Answer {
23 /** The response body, or null with why (`miss`). */
24 text: string | null
25 miss: Miss | null
26 ms: number
27}
28
29export interface RouterPart {
30 /** The prompt this part is for: a part is only ever taken for its own prompt. */
31 prompt: string
32 /**
33 * One request with the router's questions and `extra` (the skill
34 * module's), in the router's state (the prompt, recent context, signals).
35 */
36 ask: (extra: Record<string, unknown>) => Promise<Answer>
37 /** Hands the answer to the router; returns its block for the prompt, if any. */
38 settle: (answer: Answer) => Promise<string | null>
39}
40
41/**
42 * The parts waiting, by prompt text, oldest first. The engine gives a
43 * prompt no id of its own, so its text is the key; two prompts with the same
44 * text (a quick double submit) can only swap parts built for identical text,
45 * and a take for one prompt never touches another's part.
46 */
47const offered = new Map<string, RouterPart[]>()
48const MAX_WAITING = 16
49
50/** The router's part for this prompt, waiting for the skill module. */
51export function offerPart(part: RouterPart): void {
52 const queue = offered.get(part.prompt) ?? []
53 queue.push(part)
54 offered.set(part.prompt, queue)
55 // Bounded: parts nobody takes (a module that never ran) don't pile up.
56 while ([...offered.values()].reduce((n, q) => n + q.length, 0) > MAX_WAITING) {
57 const [oldest] = offered.keys()
58 if (oldest === undefined) break
59 const q = offered.get(oldest) as RouterPart[]
60 q.shift()
61 if (q.length === 0) offered.delete(oldest)
62 }
63}
64
65/** Takes the oldest part waiting for this prompt, once: null when there is none for it. */
66export function takePart(prompt: string): RouterPart | null {
67 const queue = offered.get(prompt)
68 const part = queue?.shift() ?? null
69 if (queue && queue.length === 0) offered.delete(prompt)
70 return part
71}
72
73/** Whether this part is still waiting (nobody took it); it's withdrawn either way. */
74export function stillOffered(part: RouterPart): boolean {
75 const queue = offered.get(part.prompt)
76 const at = queue ? queue.indexOf(part) : -1
77 if (!queue || at < 0) return false
78 queue.splice(at, 1)
79 if (queue.length === 0) offered.delete(part.prompt)
80 return true
81}
82
83/**
84 * A prompt that only says to go on with the work in progress ("continue",
85 * "keep going"): the turn before already decided how to do that work, so
86 * Jev isn't asked again. An approval ("yes", "go ahead", "do it") is not one:
87 * it often starts the very work the turn before only proposed.
88 */
89export function isContinuation(prompt: string): boolean {
90 return /^(?:please\s+)?(?:continue|go on|keep going|carry on|resume)(?:\s+please)?[.!]*$/i.test(prompt.trim())
91}
92