SLOPSHOPPER

jev-pilot

TypeSafe's Jev decides how Claude Code works on each prompt: the main conversation's reasoning effort (low to xhigh, raised to max mid-turn when tool calls…

newbandguardcommandstatusprompt
★ 9v0.12.1MITupdated 2026-09-30Akramovic1/jev-pilot
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · jev-pilot
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /jev ⎿ jev-pilot: jev-pilot switches (/jev <name> on|off, /jev all on|off, /jev reset): ⎿ jev-pilot: on effort sets the reasoning effort of each turn ⎿ jev-pilot: on raise raises the effort when tool calls keep failing ⎿ jev-pilot: on subagents picks each subagent’s model and effort ⎿ jev-pilot: on skills picks the one skill a prompt needs (off: the full skill list stays) ⎿ jev-pilot: on strategy advises splitting big work across subagents ╭──────────────────────────────────────────────────────────────────────────────────────────────╮ ▄▄▄▄▄▄▄▄▄▄ │ ready · no key, built-in │ ▀▀▀▀▀▀▀▀▀▀ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ ▀▀▀▀▀▀▀▀▀▀▀▀ ▀▀▀▀▀▀▀▀▀▀ ▀ ▀ ▀ ▀ ⟨Claude Code's own drawing⟩ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
╭──────────────────────────────────────────────────────────────────────────────────────────────╮ ▄ │ ready · no key, built-in │ ▀ ╰──────────────────────────────────────────────────────────────────────────────────────────────╯ ▀▀ ▀ ⟨Claude Code's own drawing⟩
README

<img src="assets/banner.svg" alt="jev-pilot — let Jev steer Claude Code" width="100%">

<a href="https://github.com/Akramovic1/jev-pilot/actions/workflows/test.yml"><img alt="tests" src="https://img.shields.io/github/actions/workflow/status/Akramovic1/jev-pilot/test.yml?branch=main&style=flat-square&label=tests&labelColor=0b1020"></a> <a href="LICENSE"><img alt="License: MIT" src="https://img.shields.io/badge/license-MIT-d4ff4f?style=flat-square&labelColor=0b1020"></a> <a href="https://docs.claude.com/en/docs/claude-code"><img alt="Claude Code 2.1.278+" src="https://img.shields.io/badge/Claude%20Code-2.1.278%2B-7cf0c4?style=flat-square&labelColor=0b1020"></a> <a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev"><img alt="Powered by Jev" src="https://img.shields.io/badge/powered%20by-Jev%20(TypeSafe)-c9d2ea?style=flat-square&labelColor=0b1020"></a> <a href="https://openrouter.ai/~typesafe/jev-latest"><img alt="Jev on OpenRouter" src="https://img.shields.io/badge/runs%20on-OpenRouter-8d99b8?style=flat-square&labelColor=0b1020"></a>

<b>The right reasoning effort, subagent model and skill for every prompt, decided by a model built for decisions.</b>

<a href="#-install">Install</a> · <a href="#-how-it-works">How it works</a> · <a href="#%EF%B8%8F-meet-the-pilot">The pet</a> · <a href="#-the-crew-custom-models-and-other-agents">The crew</a> · <a href="#-see-if-its-paying-off">Report</a> · <a href="#%EF%B8%8F-configuration">Configuration</a> · <a href="#-acknowledgements">Acknowledgements</a>


jev-pilot is a Claude Code plugin. Before every turn, it asks Jev, TypeSafe's fast decision model, a few typed questions about your prompt and sets the turn up from the answers. You keep Opus for the conversation. Easy work runs at low effort and on cheaper subagents, and hard work gets the thinking it needs.

DecisionWhen
🧠Reasoning effort, low → xhighat the start of each turn
🚨Raise effort, up to max, when tool calls keep failingmid-turn, at most once
🤖Subagent model and effort: Haiku, Sonnet or Opus, low → xhighwhen a subagent starts
🧭Strategy: do it directly, delegate, run in parallel, or plan a graphat the start of each turn, as advice
🧩The one skill the prompt needs, if anyat the start of each turn
📊A record of every decision, with tuning suggestionsalways, via /jev-pilot:report

Jev never writes in your conversation. It talks through Claude the pilot, a small animated pet above the prompt that shows what Claude is doing and says what Jev decided.

<img src="assets/demo.svg" alt="An illustrative claude-jev session. A rename runs at low effort, and the pet's bubble says low, no skill, 99% sure. A failing-tests prompt starts at xhigh with the systematic-debugging skill: the pet reads, searches and runs the tests; after two failures the effort is raised to max; it writes the fix, the tests pass, and it jumps rope." width="860"> <sub>An illustrative session: the pet and its bubbles are drawn from the plugin's own code; the numbers are examples.</sub>

[!NOTE] jev-pilot runs on Claude Code's function hooks, which are early access: they need Claude Code 2.1.278 or newer and CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1. The claude-jev launcher sets that for you.

🚀 Install

curl -fsSL https://raw.githubusercontent.com/Akramovic1/jev-pilot/main/install.sh | bash

It asks for your OpenRouter key (create one here; Jev costs about $0.04 per million input tokens, with free output), installs jev-pilot as a regular Claude Code plugin, and adds the claude-jev command. Then start Claude Code with it:

claude-jev          # takes the same arguments as claude: claude-jev -c, claude-jev -p "…"
  1. Checks that Claude Code is installed and new enough.
  2. Adds this repo as a Claude Code plugin marketplace and runs claude plugin install jev-pilot@jev-pilot. The key is passed with --config, so Claude Code keeps it in its own credential store, not in plain settings.
  3. Links claude-jev into ~/.local/bin. It's plain claude with function hooks on, plus the local router for custom models.

It changes nothing else. Re-running it updates jev-pilot and keeps your key. For scripted installs, set JEV_OPENROUTER_KEY=sk-or-… (or JEV_SKIP_KEY=1) to skip the prompt.

claude plugin marketplace add Akramovic1/jev-pilot
claude plugin install jev-pilot@jev-pilot --config openrouterApiKey=sk-or-v1-… --config timeoutMs=1500
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude

To skip typing the variable, put it in ~/.claude/settings.json and plain claude will do:

{ "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" } }
git clone https://github.com/Akramovic1/jev-pilot.git && cd jev-pilot
./install.sh

This loads the clone with --plugin-dir, so your edits take effect in the next session. Options go in ~/.claude/settings.json under pluginConfigs["jev-pilot"], and the installer adds your key there after a backup. claude-jev checks your clone's upstream once a day in the background and tells you when there's something new.

Check it's working

Start claude-jev and look above the prompt, at the right: the pilot appears with a bubble saying ready · openrouter. After your first prompt the bubble says what Jev decided, such as low · no skill · 99% sure.

If it says ready · no key, built-in, the key isn't being read. Run the installer again. To see every step Jev takes, turn on verboseLog (see Configuration).

Update and uninstall

claude-jev self-update                                            # update jev-pilot, whichever way it was installed
curl -fsSL https://raw.githubusercontent.com/Akramovic1/jev-pilot/main/install.sh | bash -s -- --uninstall

[!IMPORTANT] If you ran /jev-pilot:setup, run /jev-pilot:setup restore before uninstalling. Otherwise your skills stay hidden from Claude with nothing left to load them.

🧠 How it works

flowchart LR
    P(["Your prompt<br/>+ recent messages"]) --> J{{"Jev<br/>≈0.5 s"}}
    J -- "effort" --> T["Turn<br/>(Opus)"]
    J -- "skill + SKILL.md" --> T
    J -- "strategy advice" --> T
    T -- "2 failed tool calls" --> R["Raise effort<br/>up to max"]
    R --> T
    T -- "spawns a subagent" --> J2{{"Jev"}}
    J2 -- "haiku / sonnet / opus" --> A["Subagent"]
    T --> L[("Decision<br/>ledger")]
    L --> Rep["/jev-pilot:report"]

Effort. Jev picks one of five levels, each described by the kind of task it's for, not an amount. Every question to Jev is a choice like this, each option saying when to choose it:

LevelKind of task
lowanswered from what's known, or one mechanical step: a lookup, one command, a rename
mediuman ordinary, well-specified change to a few files, or a direct question about code in view
higha change across several files, a described bug that must be traced, tests, a careful review
xhighdesign across components, a bug with an unknown cause, a refactor with many dependents
maxnovel architecture, security or data integrity, a failure that resisted earlier attempts
  • Effort changes keep the cache, where they can. With an API key or a Claude subscription, changing the effort mid-session keeps the cached conversation on Opus 5.5, so a new effort every turn costs nothing extra. On Amazon Bedrock, Google Cloud or a gateway in front of the API, a change clears the cache, and the next request pays the cache-write price on the whole conversation. There jev-pilot holds the first turn's effort for the session and chooses again only after a compaction (effortChanges: auto, per-turn or hold).
  • Close calls lean up, as far as high. If Jev's two most likely levels are within effortCloseMargin (0.15), the higher one wins, because under-thinking costs more than over-thinking.
  • Above high, Jev has to be sure. xhigh needs Jev at least 60% sure the task is very hard (xhigh and max together). A near split between hard and very hard stays at high.
  • Raising and lowering have different bars. Raising effort needs confidence of 0.3 or more; lowering it needs 0.6.
  • Risky work gets real thought. If carrying the task out would itself deploy, move money or destroy data, effort goes to at least high.
  • Turns start at xhigh at most (maxEffort). Only the mid-turn raise reaches max: after escalateAfterErrors (2) failed tool calls in a row, effort goes up at least one level, once per turn. Permission denials don't count as failures.

What Jev reads.

  • Your prompt.
  • The last 4 messages, capped at 2000 characters: text and tool names only, never tool input or output. So "yes, do it" is judged as the work it agrees to.
  • Plain facts about the request: its length, how many files it names, whether it contains code or an error, whether it's a question, and what recent turns did with their tools.

Subagents. Following Anthropic's guidance for Sonnet 5.5 ("it fits best when the task has a clear spec and a way to check the result"): Haiku for read-only lookups where a mistake is cheap to spot (search, find a definition, read files, logs or test output, run a command and report). Sonnet for read-only work that needs some understanding (explain code, research across files, review a diff and report) and for code changes with a clear spec and a check: a fix whose cause is known, a feature to a written spec, tests for existing code, a scoped refactor. Opus for work that needs careful judgment or runs long: design, an open spec, an unknown cause, long multi-step builds, security, migrations, production or money. On 40 real subagent briefs from my own sessions, 16 builders and fixers with written findings and tests moved to Sonnet 5.5. The large milestone builds and the reviews stayed on Opus. A subagent that runs at low effort and may change code gets Anthropic's line for low effort appended to its brief: run a real check that exercises the change before reporting it done. A model named on the Agent call (you asked for one) is kept. Moving down a model needs Jev at least 60% sure. These are family names, so Claude Code uses its current release of each. No versions are hardcoded. It also gets an effort from the same decision, on the same rubric and bars as the main conversation, but Jev is asked how hard the brief is to carry out: a brief that already names the files, steps and tests has done the design, so builders and fixers usually get high, and xhigh is kept for briefs that ask for design or an unknown cause. The Agent tool has no effort setting, so jev-pilot sets it on each request the subagent makes.

Claude knows it's there. On the first prompt of each session (and after a compaction), Claude gets a short note listing what jev-pilot decides, so it leaves those decisions alone: it won't pin a subagent's model or effort, or make agent types just to fix one, unless you ask.

Strategy. The same request asks how to carry the work out:

  • direct: the usual case, and nothing is attached.
  • delegate: one subagent on a cheaper model does the broad, mechanical part.
  • parallel: fan out, then join. Independent pieces that share no files run as simultaneous background subagents; the results are integrated and tested once.
  • graph: for large builds only, a small blueprint of plain subagents. Real nodes (a step you could do inline isn't one), waves that start together, one shared plan file, a separate read-only reviewer after each join, and bounds (at most 4 subagents at a time, 2 review rounds per wave). If it can't be explained in one breath, Claude works directly.

Advice is attached only when Jev is confident (0.6, or 0.8 for graph) and it agrees with the tier. Claude may ignore it.

Skills. At most one per prompt. Jev reads every skill's description and the opening of its SKILL.md, next to a "none of these fits" option, and is asked to match the kind of work (debugging, planning, reviewing…), not a product the prompt happens to name. A skill for one platform (Vercel, Supabase, Firebase…) is picked only when the request, the conversation or the project uses that platform: jev-pilot reads what the project deploys with from its file names (cdk.json, vercel.json, Dockerfile…), so "deploy to production" in an AWS project doesn't get Vercel's deploy skill. A skill is picked when Jev is sure of it, or when the prompt needs a skill and it still fits. The winner's SKILL.md is added to the prompt.

Better code, not just cheaper. The same request also asks Jev what would make the work better. Jev can't judge code, since it never sees your repo, but it can judge the request. Each read acts only when Jev is sure:

Jev reads the request as…Claude getsBar
vague: "add caching", "make it better"ask one short question, or state the assumption in one line, before coding85%
a bug: "it's off by one cent", "the test fails"show the bug first with a failing test or a command, then fix it and show the same check passing80%
a costly area: money, auth, migrations, securityrun the tests that cover it and add one for the changed case; if a reviewer (Codex or OpenCode) is working, get its review80%
your correction of the last turn: "it doesn't work", "not what I asked"nothing: the last turn is marked in the ledger as corrected70%

The questions were tuned on sample prompts. For example, a first wording rated "add a dark mode toggle" as vague as "add caching"; the final one separates them (0.19 against 0.86).

  • Advice, not procedures. The blocks are short and ask for each check in one place only. Anthropic's /claude-api prompt-audit flagged the first versions for rituals that make Opus 5.5 write more and repeat tool calls (a plan to sketch, a review after every wave, "show the bug, then show it fixed", the same test run asked for twice), and they were rewritten.
  • The corrections are the quality signal. Until now jev-pilot only saw tool failures. /jev-pilot:report now shows, for each starting effort, how often you corrected the turn, and /jev tune leans up when cheap starts keep getting corrected. That's how you find out whether low effort is really enough for your work.
  • A turn going in circles. When the same file is edited 4 times in a turn, or the same command fails a third time, Claude gets a note after that tool call (you don't see it): step back, read the error in full, say what's causing it, and try something else. The effort goes up a level, once. Before, only tool calls failing back to back raised it, and the edit, test, edit, test loop never does that.
  • /jev quality off switches all of this off. It costs no extra wait: the questions ride in the same request.

UI design work gets the design pack. When Jev reads a request as UI design (a page, a screen, a component, a redesign, a mobile flow; measured: design requests 0.95 to 0.98, everything else 0.09 at most, a UI bug included), Claude is told to:

  • load your design skills (designSkills, by default design-taste-frontend and impeccable) for the direction and polish, and check the result against web-design-guidelines before calling it done. Only the ones you have installed are named;
  • follow the project's own design system first; otherwise take a direction from a real product's DESIGN.md in awesome-design-md (Linear, Stripe, Vercel, Notion, Apple, Airbnb…), or from real app screens with the Mobbin MCP (and Inspo), when those are connected.

It rides in the same one request. /jev design off switches it off.

One request per prompt. The effort, model, strategy and skill questions all go to Jev together, in one request of about 0.5 s. Before 0.6 there were three requests one after another: effort and strategy, the skill ranking, then a re-check of the top skills, about 1.5 s in all. On 16 test prompts the single request picked the same skill 14 times, and the other two picks were better. A plain "continue" asks nothing: the work goes on as the last turn decided. (Catalogs over the API's 255-choice limit are ranked in parallel batches, so it's still one wait.) /jev-pilot:setup can hide your own skills from Claude's skill list entirely (it asks first; restore undoes it).

🛩️ Meet the pilot

<img src="assets/pet.svg" alt="Claude the pilot above the Claude Code prompt: thinking with a thought cloud, reading a book, searching with a magnifying glass, running tests in a terminal, flying when the effort is raised, writing on paper, then jumping rope and waving while idle." width="860">

Claude the pilot, drawn as Claude Code's character, sits above the prompt at the right and shows what Claude is doing:

Claude is…The pilotIts bubble
thinkinga thought cloud, ... filling in⠋ thinking · …
reading files or pagesan open book, the line being read lit up⠋ reading · …
searching (Grep, Glob, web)a magnifying glass, sweeping⠋ searching · …
editing files or writing the answerpaper, a pencil writing lines⠋ writing · …
running commandsa terminal, output scrolling⠋ running · …
running subagents or other toolsflying: goggles down, jets on⠋ working · …
idlehovering and blinking; every few seconds it jumps rope, waves or looks aroundJev's last decision

The bubble says the turn's effort, the skill attached (or no skill), any strategy advice, and how sure Jev was of the effort: xhigh · /systematic-debugging · parallel · 88% sure. When Jev wanted a change but wasn't sure enough to make it, it says so: high kept · wanted low · 42% sure. A mid-turn raise shows as 2 fails → max ✈, and a subagent's model as Explore → haiku.

It draws only in the terminal (not in claude -p, the desktop app or mobile), and redraws only while something moves. /jev pet off hides it.

Switch any part on or off

Everything is on by default except switching the main conversation's model. Type /jev to see the switches, and change them live:

/jev                    what is on
/jev skills off         one switch: effort · raise · subagents · skills · strategy · quality · design · model · pet
/jev all off            every switch (all on turns them back on)
/jev reset              back to your settings' defaults

Switches are remembered across sessions. /jev skills off leaves skills exactly as Claude Code handles them.

How sure is Jev?

  • The bubble: the N% sure at the end, for each turn.
  • A line per turn in the conversation: set display to both or transcript, and each turn gets one line, such as jev · low (93% sure) · no skill · 1.3s.
  • Every raw score: turn on verboseLog to see each answer with its confidence, such as tier fast (0.99) · effort 0.0 → low (1.00) · risky 0.10 · strategy direct (1.00) and needs a skill 0.09.

👥 The crew: custom models and other agents

jev-pilot can bring more workers into a Claude Code session than Claude alone:

  • Custom models: any OpenRouter model you add, under a name you choose (up to 8). None is set until you add one.
  • Codex and OpenCode: your own codex and opencode CLIs, with your own logins, as code reviewers. They run as themselves, not through OpenRouter.

You choose how they're used with a mode. Within that mode, Jev decides task by task.

ModeWhat happens
standard (default)Claude only. Jev picks Haiku, Sonnet or Opus and the effort.
budgetSubagent work that needs no judgment (searching, reading and reporting, boilerplate) can go to a custom model, when Jev is sure. Opus keeps the judgment.
junior-leadA junior on a custom model writes easy, well-specified code. Opus, as tech lead, reads its diff, runs the tests and sends it back once if something's wrong. You get a cheap implementation reviewed at Opus level.
second-opinionAfter a significant change, Claude asks Codex (or OpenCode) for a review before calling the work done, then fixes what's right and says why it disagrees with the rest.
qualityEvery subagent runs on Opus, plus the external review.

Outside these modes you can still ask for a review at any time: "have Codex review this". Claude then spawns jev-pilot:codex-review.

Choose the reviewer's model. Say it in your request, "review this with Codex, Luna, high effort", or set a default that every session keeps:

/jev reviewer codex luna high          Codex reviews on Luna, high effort
/jev reviewer codex effort xhigh       just the effort
/jev reviewer opencode kimi-k3         an OpenCode model (checked against `opencode models`)
/jev reviewer codex default            back to the CLI's own config

A Codex tier name (astra, sol, terra, luna) always means the newest model of that tier. jev-pilot reads Codex's model list at every session start, so when a newer Luna ships, reviews move to it with nothing to change. A full id such as gpt-5.6-luna pins that exact version. Efforts are checked against what the model takes. /jev status shows what each reviewer runs on now, e.g. luna (newest, now gpt-5.6-luna) · effort high.

Adding a model. Pick a name, find a model on OpenRouter's list of models that can call tools, and paste it after the name. The id, the page link or the model's name all work:

/jev flash deepseek/deepseek-v4.1-flash
/jev coder https://openrouter.ai/qwen/qwen3-coder
/jev cheap DeepSeek: DeepSeek V4.1 Flash

jev-pilot looks the model up in OpenRouter's live list, adds it and says what it is: coder is now qwen/qwen3-coder (Qwen: Qwen3 Coder 480B A35B · 262k context · $0.3 in · $1 out per million tokens). Then it checks the model answers.

  • Refused: a model OpenRouter doesn't have gets the three closest ones to try, newest first. So does one that can't call tools, since Claude Code works through tools.
  • Names: lowercase letters, digits and -, starting with a letter. Words /jev already uses (status, mode, skills…) can't be names.
  • Kept for every session: the models are recorded in ~/.claude/jev-pilot/models.json, which every session reads, in any project and whichever way jev-pilot is installed. A session that's already open takes up a change at its next prompt.
/jev <name> <model>                  add a model, or replace the one under that name
/jev <name>                          that model, and the models you added before (to switch back)
/jev remove <name>                   delete it, from every session and from /model (also: /jev <name> off)
/jev status                          the mode, your models, and a health check of every worker
/jev mode junior-lead                standard · budget · junior-lead · second-opinion · quality
/jev junior <name>                   which model the junior runs on (else the first you added)
/jev reviewer opencode               which agent reviews: codex or opencode

Changes apply from the next turn. The

Source 17 files
hooks/jev-pilot.ts 40 lines
1/**
2 * jev-pilot — the plugin's one hooks module.
3 *
4 * Claude Code loads a single hooks module per plugin, so this entry registers
5 * both mods on the same `on` and the same options: the model router first,
6 * then the skill suggester. Each keeps its own handlers and state; on the
7 * events both hook (`prompt.submit`), the router's handler runs first and
8 * hands the prompt on to the suggester's.
9 *
10 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 and Claude Code >= 2.1.278.
11 */
12import type { Register } from 'claude-code'
13import { register as registerModelRouter } from './jev-model-router.ts'
14import { register as registerSkillSuggestion } from './jev-skill-suggestion.ts'
15import { register as registerPet } from './jev-pet.tsx'
16import { initFeatures } from './features.ts'
17import { initCrew } from './crew-state.ts'
18
19export const register: Register = (on, options) => {
20  // Every part's default comes from the options; /jev switches them live.
21  const flag = (key: string, fallback: boolean) => (typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback)
22  const display = typeof options.display === 'string' ? options.display : 'pet'
23  initFeatures({
24    effort: flag('routeMainEffort', true),
25    raise: typeof options.escalateAfterErrors === 'number' ? options.escalateAfterErrors > 0 : true,
26    subagents: flag('routeSubagentModel', true),
27    skills: flag('suggestSkills', true),
28    strategy: flag('suggestStrategy', true),
29    quality: flag('qualityAdvice', true),
30    design: flag('designPack', true),
31    model: flag('routeMainModel', false),
32    pet: display === 'pet' || display === 'both',
33  })
34  // The crew's defaults: mode, custom model slots, junior, reviewer.
35  initCrew(options)
36  registerModelRouter(on, options)
37  registerSkillSuggestion(on, options)
38  registerPet(on, options)
39}
40
hooks/jev-model-router.ts 1177 lines
1/**
2 * jev-model-router — Claude Mod (EARLY ACCESS)
3 *
4 * Picks the model each task runs on with TypeSafe's Jev, a System One
5 * decision model: unstructured state in, a typed choice with a probability
6 * distribution out.
7 *
8 * Jev is reached one of three ways, whichever key is configured: TypeSafe's
9 * own API (`typesafeApiKey`) or OpenRouter's Decisions API
10 * (`openrouterApiKey`), which report a calibrated confidence per answer, or
11 * the Vercel AI Gateway (`gatewayApiKey`), which does not. With none, the engine's own `$.model.classify` stands in, so the mod is
12 * useful without any account.
13 *
14 * Three things it can set, each on its own switch:
15 *   agent.spawn  — the model of each subagent (on by default)
16 *   turn.step    — the reasoning effort of the main loop (on by default)
17 *   turn.step    — the model of the main loop (off by default: switching
18 *                  models mid-session invalidates the prompt cache, which can
19 *                  cost more than the cheaper tier saves)
20 *
21 * Every one of them moves in both directions: a task the decision model reads
22 * as mechanical is routed down, one it reads as hard is routed up. The two
23 * mistakes do not cost the same, so they do not clear the same confidence bar
24 * (see `minUpgradeConfidence` / `minDowngradeConfidence` in policy.ts).
25 *
26 * The Agent tool has no effort parameter, so a subagent's effort is not ours
27 * to set; only its model is.
28 *
29 * The prompt is classified at `prompt.submit`, which runs before the turn
30 * starts, and the decision is applied at the turn's first request. The
31 * decision model reads the prompt with the last few messages before it
32 * (text and tool names only; see context.ts), so a follow-up such as "yes, do
33 * it" is read as the work it continues, not as a trivial message.
34 *
35 * The same request asks how the work should be carried out (`strategy`):
36 * directly, by one subagent, by parallel subagents, or as a small graph of
37 * subagents in waves. Anything but `direct`, answered confidently and
38 * consistent with the tier, is attached to the prompt as advice
39 * (`<execution_strategy>`); the main model decides whether it fits.
40 *
41 * Within a turn the effort holds, with one exception: when tool calls keep
42 * failing (`escalateAfterErrors` in a row, counted at `tool.call`), the
43 * effort goes up at least one rung, once per turn, as far as a fresh reading
44 * of the decision model says.
45 *
46 * Every failure path is fail-open: a classification that errors or runs past
47 * the latency budget leaves the request exactly as the engine built it.
48 *
49 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
50 * or "gatewayApiKey"). Never hardcode it in this file.
51 *
52 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (Claude Code >= 2.1.259). Typed
53 * against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods
54 *
55 * Privacy: with a key set, the prompt text is sent to whichever backend the
56 * key belongs to.
57 */
58import type { HttpInit, HttpResponse, Register } from 'claude-code'
59import { NOT_A_TASK, platformsOf, recentContext, signalsOf, SKIPPED_DIRS } from './context.ts'
60import { clearSkillNotes, resetBriefing, takeBriefing, takeSkill, turnLine } from './summary.ts'
61import { moodOf, say, setBoost, subagentLabel, turnSpeech } from './pet-art.ts'
62import { feature } from './features.ts'
63import { crewNote, JUNIOR_AGENT, juniorSlot, REVIEWERS, reviewerAgent, slotAlias, slotsOffered } from './crew.ts'
64import { crew, reviewerHealthy, router, slotUsable } from './crew-state.ts'
65import { ensureCrew, newFallbacks, startSession, type CrewIo } from './crew-run.ts'
66import { isContinuation, offerPart, stillOffered, type RouterPart } from './jev-call.ts'
67import { recordCorrection, recordSubagent, recordSubagentUsage, recordTurn, resetStats, shortStats } from './session-stats.ts'
68import type { ContextMessage } from './context.ts'
69import { appendEntry, configKeysOf, entriesOf, LEDGER_KEY, markCorrected, reportPrompt, suggestions, summarize } from './ledger.ts'
70import { dueToPropose, effective, initTuning, setTuning, TUNING_KEY, tuningLoaded, tuningOf } from './tuning.ts'
71import type { LedgerEntry, TunableConfig } from './ledger.ts'
72import { CHECK_AT_LOW, effortClearsCache, QUALITY_BARS, qualityAdvice, spinOf, stepBackNote } from './model-router.policy.ts'
73import {
74  adviseStrategy,
75  EFFORT_ORDER,
76  effortLevel,
77  effortScoreOf,
78  engineMoved,
79  isFollowUp,
80  escalate,
81  DEFAULT_BASE_URL,
82  DEFAULT_MODEL,
83  describeDecision,
84  describeSetup,
85  describeStatus,
86  capabilityNote,
87  endpoint,
88  missOf,
89  pendingDecisions,
90  readDecision,
91  selectProvider,
92  requestBody,
93  requestHeaders,
94  modelIds,
95  requestModelId,
96  route,
97  TIER_ORDER,
98} from './model-router.policy.ts'
99import type { Miss, SlotChoice } from './model-router.policy.ts'
100import type { Decision, Effort, PolicyConfig, Provider, StrategyConfig, Tier } from './model-router.policy.ts'
101
102/** Where decisions are asked, when a backend is configured. */
103interface Backend {
104  provider: Provider
105  url: string
106  apiKey: string
107  modelId: string
108  timeoutMs: number
109}
110
111/**
112 * The engine calls the helpers below need. `$` is never handed to a helper:
113 * each hook builds this at its own call site, spelling every call on `$`
114 * there (see `io` in register).
115 */
116interface Io {
117  fetch: (url: string, init: HttpInit) => Promise<HttpResponse>
118  sleep: (ms: number) => Promise<void>
119  log: (text: string) => unknown
120  /** A line for the verbose log only: what is routine, not an error. */
121  detail: (text: string) => unknown
122  messages: () => Promise<readonly ContextMessage[]>
123}
124
125/**
126 * One request to the backend, read as a decision; null without a backend, or
127 * on timeout, error, a non-2xx or an unreadable answer, with why (`miss`).
128 * Every caller treats a null decision the same way: the request goes on as
129 * the engine built it. A timeout or a busy backend is routine (the pet says
130 * so); only an error is logged whatever the log level.
131 */
132async function askJev(
133  io: Io,
134  backend: Backend,
135  state: Record<string, unknown>,
136  withStrategy: boolean,
137  what: string,
138  subagent = false,
139  slots: readonly SlotChoice[] = [],
140  junior = false,
141  extra: Record<string, unknown> = {},
142  quality = false,
143): Promise<{ text: string | null; miss: Miss | null }> {
144  try {
145    const response = await Promise.race([
146      io.fetch(backend.url, {
147        method: 'POST',
148        headers: requestHeaders(backend.provider, backend.apiKey, backend.modelId),
149        body: requestBody(backend.provider, state, backend.modelId, withStrategy, subagent, slots, junior, extra, quality),
150      }),
151      io.sleep(backend.timeoutMs),
152    ])
153    if (response && response.ok) return { text: response.text, miss: null }
154    const miss = missOf(response ? response.status : null)
155    const tell = miss === 'error' ? io.log : io.detail
156    if (response) await tell(`[jev-model-router] ${backend.provider} responded ${response.status}: ${response.text.slice(0, 200)}`)
157    else await tell(`[jev-model-router] classification passed ${backend.timeoutMs}ms; leaving ${what} alone`)
158    return { text: null, miss }
159  } catch (error) {
160    await io.log(`[jev-model-router] classification failed: ${String(error)}`)
161    return { text: null, miss: 'error' }
162  }
163}
164
165async function classify(
166  io: Io,
167  backend: Backend | null,
168  state: Record<string, unknown>,
169  withStrategy: boolean,
170  what: string,
171  subagent = false,
172  slots: readonly SlotChoice[] = [],
173  junior = false,
174): Promise<{ decision: Decision | null; miss: Miss | null }> {
175  if (!backend) return { decision: null, miss: null }
176  const { text, miss } = await askJev(io, backend, state, withStrategy, what, subagent, slots, junior)
177  if (text === null) return { decision: null, miss }
178  const decision = readDecision(text, slots.map((slot) => slot.name))
179  return { decision, miss: decision ? null : 'error' }
180}
181
182/** The conversation so far, or none when not `wanted` or unreadable. */
183async function readMessages(io: Io, wanted: boolean): Promise<readonly ContextMessage[]> {
184  if (!wanted) return []
185  try {
186    return await io.messages()
187  } catch (error) {
188    await io.log(`[jev-model-router] could not read the conversation: ${String(error)}`)
189    return []
190  }
191}
192
193/** What prompt.submit knows of a turn's decision, waiting for the turn to start. */
194type Draft = Pick<
195  LedgerEntry,
196  'answered' | 'ms' | 'tier' | 'tierConfidence' | 'effortLevel' | 'effortConfidence' | 'strategy' | 'strategyConfidence' | 'advised'
197>
198
199export const register: Register = (on, options) => {
200  const text = (key: string, fallback: string) =>
201    typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
202  const number = (key: string, fallback: number) =>
203    typeof options[key] === 'number' ? (options[key] as number) : fallback
204  const flag = (key: string, fallback: boolean) =>
205    typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
206
207  // With several keys set, `auto` takes TypeSafe's own API, then OpenRouter,
208  // then the Gateway: the first two report the calibrated confidence the
209  // policy's threshold reads. `provider` forces one, "builtin" uses none.
210  const typesafeKey = text('typesafeApiKey', '')
211  const gatewayKey = text('gatewayApiKey', '')
212  const openrouterKey = text('openrouterApiKey', '')
213  const forced = text('provider', 'auto')
214  const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey, openrouterKey)
215
216  // Each backend keeps its own URL and model, so an override written for one
217  // can never be sent to the other when `auto` picks differently than expected.
218  const apiKey = !active ? '' : { typesafe: typesafeKey, gateway: gatewayKey, openrouter: openrouterKey }[active]
219  const modelId = !active ? '' : text(`${active}Model`, DEFAULT_MODEL[active])
220  const url = !active ? '' : endpoint(active, text(`${active}BaseUrl`, DEFAULT_BASE_URL[active]))
221
222  // A backend named in the options but missing its key degrades to the
223  // built-in classifier, which is silent; say so once, when a hook first runs.
224  let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
225
226  const timeoutMs = number('timeoutMs', 1500)
227  // Each part reads its switch live (features.ts): /jev turns it on or off.
228  const routeSubagentModel = () => feature('subagents')
229  const routeMainEffort = () => feature('effort')
230  const routeMainModel = () => feature('model')
231  const routeMainLoop = () => routeMainEffort() || routeMainModel()
232  const logDecisions = flag('logDecisions', true)
233  // Where jev-pilot talks: the pet at the bottom right (default), one line
234  // per turn in the transcript, both, or nowhere. verboseLog adds every step
235  // to the transcript whatever this says.
236  const display = text('display', 'pet')
237  const petOn = () => feature('pet')
238  const verbose = logDecisions && flag('verboseLog', false)
239  const lines = logDecisions && (verbose || display === 'transcript' || display === 'both')
240  const readyLine = () =>
241    verbose
242      ? `[jev-model-router] ${describeSetup(
243          active,
244          url,
245          { subagentModel: routeSubagentModel(), mainEffort: routeMainEffort(), mainModel: routeMainModel() },
246          forced === 'builtin',
247        )}`
248      : `jev-pilot · ready on ${active ?? `the built-in classifier${forced === 'builtin' ? '' : ' (no key set)'}`}`
249
250  // A reasoning level from the options; a value off the ladder is not guessed
251  // at and reads as the default.
252  const effortOption = (key: string, fallback: Effort): Effort => {
253    const value = text(key, fallback)
254    return (EFFORT_ORDER as readonly string[]).includes(value) ? (value as Effort) : fallback
255  }
256  // Two ceilings: where a turn may start, and how far trouble may raise it.
257  // Starting lower and raising only on evidence keeps `max` for the turns
258  // that show they need it.
259  const raisedCeiling: { maxEffort: Effort } = { maxEffort: effortOption('maxRaisedEffort', 'max') }
260  const policy: PolicyConfig = {
261    tiers: {
262      fast: text('fastModel', 'haiku'),
263      balanced: text('balancedModel', 'sonnet'),
264      deep: text('deepModel', 'opus'),
265    },
266    minUpgradeConfidence: number('minUpgradeConfidence', 0.3),
267    minHighConfidence: number('minHighConfidence', 0.5),
268    minDowngradeConfidence: number('minDowngradeConfidence', 0.6),
269    // Turns start at xhigh at most, and only when Jev is 60% sure a task is
270    // very hard. Anthropic warns that on Opus 5.5 xhigh thinks a lot more;
271    // replayed on 76 labelled prompts, every turn that started at xhigh
272    // needed it (4 of 4), and a high ceiling only started those four too
273    // low (exact 39 → 35). max is still reached only by the mid-turn raise.
274    maxEffort: effortOption('maxEffort', 'xhigh'),
275    // A near tie between two effort levels takes the higher one.
276    closeMargin: Math.max(0, number('effortCloseMargin', 0.15)),
277  }
278  let margin = policy.closeMargin ?? 0
279  // Each turn's decision and outcome, kept for /jev-pilot:report.
280  const recordDecisions = flag('recordDecisions', true)
281
282  // How much of the conversation the decision model reads beside a prompt.
283  const contextLimits = {
284    messages: Math.max(0, Math.round(number('contextMessages', 4))),
285    chars: Math.max(0, number('contextChars', 2000)),
286  }
287  const suggestStrategy = () => feature('strategy')
288  const strategyConfig: StrategyConfig = {
289    minConfidence: number('minStrategyConfidence', 0.6),
290    minGraphConfidence: number('minGraphConfidence', 0.8),
291    graphSkill: text('graphSkill', ''),
292  }
293  const escalateAfterErrors = Math.max(0, Math.round(number('escalateAfterErrors', 2)))
294  // A subagent's effort, set at its requests from the decision made when it
295  // started (the Agent tool itself takes none); under the subagents switch.
296  const subagentEffortOn = flag('routeSubagentEffort', true)
297  const routeSubagentEffort = () => routeSubagentModel() && subagentEffortOn
298  // Each subagent's decision, by the id core gives it when it starts; its
299  // effort, once its first request has settled it.
300  const subagents = new Map<string, { decision: Decision; label: string; model: string | null }>()
301  const subagentEffort = new Map<string, Effort | null>()
302  const MAX_SUBAGENTS = 64
303  /** What is switched on, for the note that tells the model. */
304  const capabilities = () => ({
305    effort: routeMainEffort(),
306    raise: routeMainEffort() && feature('raise') && escalateAfterErrors > 0,
307    subagents: routeSubagentModel(),
308    subagentEffort: routeSubagentEffort(),
309    skills: feature('skills'),
310    strategy: suggestStrategy(),
311    model: routeMainModel(),
312    quality: feature('quality'),
313  })
314
315  const backend: Backend | null = active ? { provider: active, url, apiKey, modelId, timeoutMs } : null
316  // What the ledger may tune (`/jev tune`): the settings' values, and how a
317  // learned change goes into force, here, at once.
318  initTuning(
319    {
320      timeoutMs,
321      minDowngradeConfidence: policy.minDowngradeConfidence,
322      effortCloseMargin: margin,
323      minHighConfidence: policy.minHighConfidence ?? 0.5,
324    },
325    (tuned: TunableConfig) => {
326      policy.minDowngradeConfidence = tuned.minDowngradeConfidence
327      policy.minHighConfidence = tuned.minHighConfidence
328      policy.closeMargin = tuned.effortCloseMargin
329      margin = tuned.effortCloseMargin
330      if (backend) backend.timeoutMs = tuned.timeoutMs
331    },
332  )
333
334  // The classification waiting for the turn that reads its prompt, and what
335  // the current turn settled on. Both are single slots: main-loop turns run
336  // one at a time, so nothing accumulates over a long session. `pending`
337  // reports no decision when two prompts are waiting at once, rather than
338  // routing a turn on a decision made for a different prompt.
339  // Each classified prompt waits with its decision and its ledger draft, so
340  // the turn that reads it knows which prompt it is working on.
341  const pending = pendingDecisions<{ decision: Decision | null; prompt: string; draft: Draft; miss: Miss | null }>()
342  // Said once, the first time a hook runs. A router that loaded and one that
343  // never loaded are otherwise told apart only by the absence of later lines,
344  // and absence is not evidence: the policy leaves most turns alone anyway.
345  let announced = false
346  let appliedTurnId: string | undefined
347  let applied: { model?: string; effort?: Effort } | null = null
348  // The model the engine named for the current turn's first request, before
349  // any rewrite: a later request naming another is the engine's fallback.
350  let turnEngineModel: string | null = null
351  // The prompt the current turn works on (null when its decision was
352  // withheld), for a re-reading mid-turn, and the turn whose effort was
353  // already raised: at most once each.
354  let turnPrompt: string | null = null
355  let escalatedTurnId: string | undefined
356  // The main loop's tool calls that failed in a row since its last success,
357  // counted as they finish (tool.call) and cleared when a turn starts.
358  let failedInARow = 0
359  // Said once: a custom model asked for with no router to serve it.
360  let warnedNoRouter = false
361  // Where changing the effort clears the cache (Bedrock, Google Cloud, a
362  // gateway): the effort chosen for the first turn is held for the session,
363  // and chosen again after a compaction (which rewrites the cache anyway).
364  const effortChanges = text('effortChanges', 'auto')
365  let holdEffort: boolean | null = effortChanges === 'hold' ? true : effortChanges === 'per-turn' ? false : null
366  let heldEffort: Effort | null | undefined
367  // What the project deploys with, found once per project folder.
368  let platformCache: { cwd: string; text: string } | null = null
369  // The turn id of the ledger entry this session finished last: the next
370  // prompt says whether it was right.
371  let lastEntryId: string | null = null
372  // A turn going in circles: each file's edits and each command's runs in
373  // the turn, and whether that was already said.
374  const edits = new Map<string, number>()
375  const runs = new Map<string, number>()
376  let spinning: string | null = null
377  let spunTurnId: string | undefined
378  // The last decision Jev made for a typed prompt: "continue" goes on with it.
379  let lastDecision: Decision | null = null
380  // Each family's current full id, learned from the requests the engine makes.
381  const ids = modelIds()
382  // The ledger: the turn in progress (its draft waits in `pending`).
383  let current: (LedgerEntry & { turnId: string }) | null = null
384
385  // Said as soon as the session opens, so a loaded jev-pilot is visible
386  // before the first prompt: a line in the transcript and one under the
387  // prompt. (The per-prompt lines follow once prompts arrive.)
388  on('session.start', async ($, e, next) => {
389    const result = await next(e)
390    // The crew: saved /jev changes, the router claude-jev started, the slots
391    // file it reads; then every worker's health, in the background.
392    const crewIo: CrewIo = {
393      fetch: (url, init) => $.http.fetch(url, init),
394      home: () => $.env.get('HOME'),
395      routerUrl: () => $.env.get('JEV_ROUTER_URL'),
396      write: (path, text) => $.fs.write(path, text),
397      read: async (path) => ((await $.fs.exists(path)) ? $.fs.read(path) : null),
398      run: (argv, timeoutMs) => $.process.run(argv, { timeoutMs }),
399      storeGet: (key) => $.store.get(key),
400      sleep: (ms) => $.clock.sleep(ms),
401      register: async (spec) => {
402        await $.agent.register(spec)
403      },
404    }
405    await startSession(crewIo, openrouterKey || null)
406    announced = true
407    if (lines) {
408      $.ui.log(readyLine())
409      $.ui.status(`jev · ready on ${active ?? 'the built-in classifier'}`)
410    }
411    if (petOn()) {
412      say(`ready · ${active ?? (forced === 'builtin' ? 'built-in' : 'no key, built-in')}`, 'ready')
413      $.ui.invalidate('ui.render')
414    }
415    return result
416  })
417
418  on('prompt.submit', async ($, e, next) => {
419    const io: Io = {
420      fetch: (url, init) => $.http.fetch(url, init),
421      sleep: (ms) => $.clock.sleep(ms),
422      log: (text) => $.ui.log(text),
423      detail: (text) => (verbose ? $.ui.log(text) : undefined),
424      messages: () => $.session.messages(),
425    }
426    // Before the routing guards: a module whose switches are all off has still
427    // loaded, and that is exactly when its silence is most misleading.
428    if (!announced) {
429      announced = true
430      if (lines) $.ui.log(readyLine())
431    }
432    const crewIo: CrewIo = {
433      fetch: (url, init) => $.http.fetch(url, init),
434      home: () => $.env.get('HOME'),
435      routerUrl: () => $.env.get('JEV_ROUTER_URL'),
436      write: (path, text) => $.fs.write(path, text),
437      read: async (path) => ((await $.fs.exists(path)) ? $.fs.read(path) : null),
438      run: (argv, timeoutMs) => $.process.run(argv, { timeoutMs }),
439      storeGet: (key) => $.store.get(key),
440      sleep: (ms) => $.clock.sleep(ms),
441      register: async (spec) => {
442        await $.agent.register(spec)
443      },
444    }
445    await ensureCrew(crewIo, openrouterKey || null)
446    // The tuning learned from the ledger, once per worker (a reload starts over).
447    if (!tuningLoaded()) setTuning(tuningOf(await $.store.get(TUNING_KEY).catch(() => undefined)))
448    // Only the person's own tasks are classified. A notification, a peer's
449    // message or a typed `/command` would otherwise take the pending slot and
450    // leave the next real prompt's turn without its decision.
451    const isTask = !!e.text.trim() && !/^\/\S/.test(e.text.trim()) && !(e.origin && NOT_A_TASK.has(e.origin.kind))
452    if (!isTask) return next(e)
453    // Once per session, and again after a compaction: what jev-pilot does,
454    // for the model, so it leaves those decisions to it.
455    const note = takeBriefing()
456      ? capabilityNote(
457          capabilities(),
458          crewNote(crew(), juniorSlot(crew(), router() !== null), REVIEWERS.filter(reviewerHealthy), slotsOffered(crew(), router() !== null).filter(slotUsable), routeSubagentModel()),
459        )
460      : null
461    const withNote = (input: typeof e, extra: string | null = null) => {
462      const blocks = [extra, note].filter((b): b is string => b !== null)
463      return blocks.length > 0 ? { ...input, context: [...(input.context ?? []), ...blocks] } : input
464    }
465    if (!routeMainLoop() && !suggestStrategy() && !feature('quality')) {
466      const passed = await next(withNote(e))
467      if (passed.drop && note) resetBriefing()
468      return passed
469    }
470    const planning = suggestStrategy()
471
472    if (!unusableReported) {
473      unusableReported = true
474      $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
475    }
476
477    const startedAt = await $.clock.now()
478    const messages = await readMessages(io, contextLimits.messages > 0)
479    const recent = recentContext(messages, e.text, contextLimits)
480    // What the answer settles, however it arrives: the strategy advice, the
481    // ledger's draft, and the decision waiting for the turn.
482    const finish = async (decision: Decision | null, miss: Miss | null, ms: number, reused = false): Promise<string | null> => {
483      if (decision && !reused) lastDecision = decision
484      // What the decision model actually answered, whatever the policy then
485      // does with it. This is the line that proves the classification ran.
486      if (verbose) {
487        const read = recent ? ` · read ${recent.split('\n').length} recent messages` : ''
488        $.ui.log(`[jev-model-router] jev: ${reused ? 'continuing, the last decision kept' : describeDecision(decision, ms, margin)}${read}`)
489      }
490      let block: string | null = null
491      let advisedStrategy = false
492      if (planning && !reused) {
493        const advice = adviseStrategy(decision, strategyConfig)
494        block = advice.block
495        advisedStrategy = block !== null
496        if (verbose) $.ui.log(`[jev-model-router] strategy: ${block ? 'advising ' : ''}${advice.reason}`)
497      }
498      // What makes the work better: ask first, test the bug first, check a costly change.
499      if (feature('quality') && !reused) {
500        const quality = qualityAdvice(decision)
501        if (quality) {
502          block = [block, quality].filter((b): b is string => b !== null).join('\n\n')
503          if (verbose) $.ui.log(`[jev-model-router] quality: ${quality.split('\n').filter((l) => l.startsWith('- ')).map((l) => l.slice(2, 40)).join(' · ')}`)
504        }
505      }
506      // Your reply said the last turn got it wrong: that turn is marked in the
507      // ledger, which is how jev-pilot learns where too little effort costs you.
508      if (!reused && typeof decision?.corrects === 'number' && lastEntryId !== null) {
509        const id = lastEntryId
510        lastEntryId = null
511        const corrected = decision.corrects >= QUALITY_BARS.corrects
512        if (corrected) recordCorrection()
513        try {
514          await $.store.set(LEDGER_KEY, markCorrected(await $.store.get(LEDGER_KEY), id, corrected))
515        } catch (error) {
516          if (verbose) $.ui.log(`[jev-model-router] could not mark the last turn: ${String(error)}`)
517        }
518      }
519      const draft: Draft = {
520        answered: decision !== null,
521        ms: active && !reused ? Math.round(ms) : null,
522        tier: decision?.tier ?? null,
523        tierConfidence: decision?.confidence ?? null,
524        effortLevel: decision ? effortScoreOf(decision, margin) : null,
525        effortConfidence: decision?.effortConfidence ?? null,
526        strategy: reused ? null : (decision?.strategy ?? null),
527        strategyConfidence: reused ? null : (decision?.strategyConfidence ?? null),
528        advised: advisedStrategy,
529      }
530      pending.put({ decision, prompt: e.text, draft, miss })
531      return block
532    }
533    const refused = (result: Awaited<ReturnType<typeof next>>) => {
534      // Refused further down: no turn will read this decision.
535      if (result.drop) {
536        pending.withdraw()
537        if (note) resetBriefing()
538      }
539      return result
540    }
541
542    // "continue": the work in progress goes on as the last turn decided.
543    if (active && lastDecision && isContinuation(e.text)) {
544      await finish(lastDecision, null, 0, true)
545      return refused(await next(withNote(e)))
546    }
547
548    if (active && backend) {
549      // One request for the prompt: the skill module, further down, adds its
550      // questions to these and sends them together (jev-call.ts).
551      const junior = (() => {
552        const slot = juniorSlot(crew(), router() !== null)
553        return !!slot && slotUsable(slot)
554      })()
555      // What the project deploys with, from its file names (once per folder):
556      // a platform's skill fits only a project on that platform.
557      let platforms = platformCache
558      try {
559        const cwd = await $.session.cwd()
560        if (platforms?.cwd !== cwd) {
561          const paths: string[] = []
562          const top = await $.fs.list(cwd)
563          for (const entry of top) paths.push(entry.kind === 'dir' ? `${entry.name}/` : entry.name)
564          const dirs = top.filter((entry) => entry.kind === 'dir' && !SKIPPED_DIRS.has(entry.name)).slice(0, 60)
565          for (const dir of dirs) {
566            for (const entry of await $.fs.list(`${cwd}/${dir.name}`).catch(() => [])) {
567              paths.push(`${dir.name}/${entry.kind === 'dir' ? `${entry.name}/` : entry.name}`)
568            }
569          }
570          platforms = { cwd, text: platformsOf(paths) }
571          platformCache = platforms
572        }
573      } catch {
574        platforms = null
575      }
576      const state = { prompt: e.text, recent_context: recent, signals: signalsOf(e.text, messages), ...(platforms ? { project_platforms: platforms.text } : {}) }
577      let settled = false
578      let settledBlock: string | null = null
579      const part: RouterPart = {
580        prompt: e.text,
581        ask: async (extra) => {
582          const asked = await askJev(io, backend, state, planning, 'the turn', false, [], junior, extra, feature('quality'))
583          return { ...asked, ms: (await $.clock.now()) - startedAt }
584        },
585        // Once: a second settle (never expected) gets the first one's block.
586        settle: async (answer) => {
587          if (settled) return settledBlock
588          settled = true
589          const decision = answer.text === null ? null : readDecision(answer.text)
590          settledBlock = await finish(decision, answer.text !== null && !decision ? 'error' : answer.miss, answer.ms)
591          return settledBlock
592        },
593      }
594      offerPart(part)
595      const result = await next(withNote(e))
596      // Nobody took it (the skill module never ran): asked now, before the
597      // turn starts; too late for advice, in time for the effort.
598      stillOffered(part)
599      if (!settled) await part.settle(await part.ask({}))
600      return refused(result)
601    }
602
603    // No backend: the engine's own small-model classifier answers the same
604    // question, without the confidence the policy's threshold reads, and
605    // without a strategy (it answers one label).
606    let decision: Decision | null = null
607    try {
608      const input = recent ? `Recent conversation:\n${recent}\n\nLatest request:\n${e.text}` : e.text
609      const label = await $.model.classify(input, TIER_ORDER)
610      if (label) {
611        decision = {
612          tier: label as Tier,
613          confidence: null,
614          risky: null,
615          effort: null,
616          effortConfidence: null,
617        }
618      }
619    } catch (error) {
620      $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
621    }
622    const block = await finish(decision, null, (await $.clock.now()) - startedAt)
623    // Attached on the way down: one block after the prompt as typed, read by
624    // the model and never shown to the person.
625    return refused(await next(withNote(e, block)))
626  })
627
628  on('turn.step', async function* ($, e, next) {
629    const io: Io = {
630      fetch: (url, init) => $.http.fetch(url, init),
631      sleep: (ms) => $.clock.sleep(ms),
632      log: (text) => $.ui.log(text),
633      detail: (text) => (verbose ? $.ui.log(text) : undefined),
634      messages: () => $.session.messages(),
635    }
636    // Every request names the id the engine resolved for it, a subagent's
637    // included: that is where the main loop's switch finds its ids.
638    ids.learn(e.model)
639    // A custom model where no router serves it (a plain `claude` session
640    // whose default was set to jev-…): Sonnet instead of a request that fails.
641    if (/^jev-/.test(e.model ?? '') && router() === null) {
642      const id = requestModelId(policy.tiers.balanced, ids)
643      if (id) {
644        if (!warnedNoRouter) {
645          warnedNoRouter = true
646          $.ui.log(`jev · ${e.model} needs claude-jev (its router isn't running here), so this session uses ${id}. /model default sets your default back.`)
647        }
648        return yield* next({ ...e, model: id })
649      }
650    }
651    // A subagent's request: the effort it was routed to when it started,
652    // settled at its first request (from the effort the engine built it with)
653    // and kept for the rest of its run.
654    if (e.agentId) {
655      const agentId = e.agentId
656      let effort = subagentEffort.get(agentId)
657      const known = subagents.get(agentId)
658      // A model without effort (the engine left it unset) is left that way.
659      if (effort === undefined && known && routeSubagentEffort() && e.effort !== undefined) {
660        const routing = route(known.decision, { model: e.model, effort: e.effort }, policy)
661        effort = routing.effort
662        subagentEffort.set(agentId, effort)
663        if (effort && lines) {
664          $.ui.log(verbose ? `[jev-model-router] ${known.label} → effort ${effort}: ${routing.reason}` : `jev · subagent ${known.label} → effort ${effort}`)
665        }
666        if (effort && petOn()) {
667          say(`${known.label} → ${[known.model, effort].filter(Boolean).join(' · ')}`, 'focused')
668          $.ui.invalidate('ui.render')
669        }
670      }
671      return yield* next(effort && routeSubagentEffort() ? { ...e, effort } : e)
672    }
673    if (!routeMainLoop()) return yield* next(e)
674
675    // Every request after the first reuses what the turn settled on, so
676    // neither the model nor the effort changes under its own tool loop —
677    // unless the loop is visibly struggling, and then only the effort, up.
678    if (e.index > 0 && e.turnId === appliedTurnId) {
679      if (routeMainEffort() && feature('raise') && escalateAfterErrors > 0 && escalatedTurnId !== e.turnId && !holdEffort) {
680        const failed = failedInARow
681        // Going in circles counts as struggling too, however the calls ended.
682        const circling = spinning
683        if (failed >= escalateAfterErrors || circling) {
684          spinning = null
685          escalatedTurnId = e.turnId
686          const effort = applied?.effort ?? e.effort
687          const turn = current
688          // The turn's own prompt: without one (its decision was withheld),
689          // there is nothing to re-read, and the raise is the one rung.
690          const prompt = turnPrompt
691          const messages = prompt === null ? [] : await readMessages(io, contextLimits.messages > 0)
692          const reread =
693            prompt === null
694              ? null
695              : (
696                  await classify(
697                    io,
698                    backend,
699                    {
700                      prompt,
701                      recent_context: recentContext(messages, prompt, contextLimits),
702                      signals: signalsOf(prompt, messages),
703                      trouble: circling ?? `${failed} tool calls in a row have failed while working on this request`,
704                    },
705                    false,
706                    'the effort',
707                  )
708                ).decision
709          const level = reread ? effortScoreOf(reread, margin) : null
710          const raised = escalate(effort, Math.max(failed, circling ? escalateAfterErrors : 0), escalateAfterErrors, level, raisedCeiling)
711          if (raised) {
712            applied = { ...(applied ?? {}), effort: raised }
713            if (turn && turn.turnId === e.turnId) turn.raisedTo = raised
714            if (lines) {
715              $.ui.log(
716                verbose
717                  ? `[jev-model-router] main loop → effort ${raised}: ${circling ?? `${failed} tool calls failed in a row`}`
718                  : `jev · ${circling ? 'going in circles' : `${failed} failed in a row`} → effort ${raised}`,
719              )
720              $.ui.status(`jev · struggling → ${raised}`)
721            }
722            if (petOn()) {
723              // At max, the pilot pulls its goggles down for the rest of the turn.
724              if (raised === 'max') setBoost(true)
725              say(`${circling ? 'circles' : `${failed} fails`} → ${raised} ✈`, raised === 'max' ? 'boost' : moodOf(raised))
726              $.ui.invalidate('ui.render')
727            }
728          } else if (verbose) {
729            $.ui.log(`[jev-model-router] ${failed} tool calls failed in a row; effort ${String(effort)} kept`)
730          }
731        }
732      }
733      // The engine moved the turn to another model (its overload fallback):
734      // that model stands for the rest of the turn; the routed effort stays.
735      if (applied?.model && engineMoved(turnEngineModel, e.model)) {
736        const { model: _dropped, ...rest } = applied
737        applied = Object.keys(rest).length > 0 ? rest : null
738        if (verbose) $.ui.log(`[jev-model-router] main loop: the engine moved the turn to ${e.model}; that model stands`)
739        if (lines) $.ui.status(`jev · engine moved to ${e.model}${applied?.effort ? ` · effort ${applied.effort}` : ''}`)
740      }
741      return yield* next(applied ? { ...e, ...applied } : e)
742    }
743
744    // A new turn: failures of the last one say nothing about this one.
745    failedInARow = 0
746    edits.clear()
747    runs.clear()
748    spinning = null
749    spunTurnId = undefined
750    const taken = pending.take()
751    const decision = taken?.decision ?? null
752    turnPrompt = taken?.prompt ?? null
753    const routing = route(decision, { model: e.model, effort: e.effort }, policy, { noLowering: taken ? isFollowUp(taken.prompt) : false })
754    const change: { model?: string; effort?: Effort } = {}
755    // The main loop's `model` is sent to the API as written, so an alias
756    // becomes the id the engine was seen using for it; a subagent's
757    // (agent.spawn) may stay an alias.
758    if (routeMainModel() && routing.model) {
759      const id = requestModelId(routing.model, ids)
760      if (id) change.model = id
761      else if (verbose) {
762        $.ui.log(`[jev-model-router] no ${routing.model} model seen yet this session; model left as ${e.model}`)
763      }
764    }
765    if (holdEffort === null) {
766      holdEffort = effortClearsCache({
767        bedrock: await $.env.get('CLAUDE_CODE_USE_BEDROCK'),
768        vertex: await $.env.get('CLAUDE_CODE_USE_VERTEX'),
769        upstream: (await $.env.get('JEV_ANTHROPIC_UPSTREAM')) ?? (await $.env.get('ANTHROPIC_BASE_URL')),
770      })
771      if (holdEffort && lines) $.ui.log('jev · effort held for this session: here, changing it would clear the cached conversation')
772    }
773    if (holdEffort && heldEffort !== undefined) {
774      // Held: every turn keeps the effort the first one got.
775      if (heldEffort && heldEffort !== e.effort) change.effort = heldEffort
776    } else {
777      if (routeMainEffort() && routing.effort) change.effort = routing.effort
778      if (holdEffort) heldEffort = change.effort ?? (typeof e.effort === 'string' ? (e.effort as Effort) : null)
779    }
780
781    appliedTurnId = e.turnId
782    turnEngineModel = e.model
783    applied = Object.keys(change).length > 0 ? change : null
784    if (recordDecisions) {
785      // A decision withheld (two prompts waiting) is recorded as unanswered:
786      // the turn ran on the engine's own settings.
787      const known = decision ? (taken?.draft ?? null) : null
788      const startEffort = change.effort ?? e.effort
789      current = {
790        turnId: e.turnId,
791        at: await $.clock.now(),
792        answered: known?.answered ?? false,
793        ms: known?.ms ?? null,
794        tier: known?.tier ?? null,
795        tierConfidence: known?.tierConfidence ?? null,
796        effortLevel: known?.effortLevel ?? null,
797        effortConfidence: known?.effortConfidence ?? null,
798        startedFrom: typeof e.effort === 'string' ? e.effort : null,
799        started: typeof startEffort === 'string' ? startEffort : null,
800        raisedTo: null,
801        toolCalls: 0,
802        failures: 0,
803        strategy: known?.strategy ?? null,
804        strategyConfidence: known?.strategyConfidence ?? null,
805        advised: known?.advised ?? false,
806        outcome: null,
807        durationMs: null,
808        outputTokens: null,
809      }
810    }
811    // A row in the transcript scrolls away; this line stays on screen.
812    if (lines) $.ui.status(describeStatus(decision, applied))
813    // The skill module's pick for this turn's prompt, read (and so released)
814    // whatever the log mode.
815    const skillNote = takeSkill(taken?.prompt ?? null)
816    const known = decision ? (taken?.draft ?? null) : null
817    const level = decision ? effortScoreOf(decision, margin) : null
818    // A turn nobody typed (a subagent's or a background task's notification)
819    // was never put to the decision model: the bubble keeps what it said.
820    if (petOn() && taken) {
821      say(
822        turnSpeech({
823          answered: decision !== null,
824          miss: taken?.miss ?? null,
825          applied: change.effort ?? null,
826          current: typeof e.effort === 'string' ? e.effort : null,
827          wanted: level === null ? null : effortLevel(level),
828          confidence: decision?.effortConfidence ?? null,
829          skill: skillNote ? skillNote.skill : undefined,
830          advised: known?.advised ? (known.strategy ?? null) : null,
831        }).text,
832        decision ? moodOf(change.effort ?? (typeof e.effort === 'string' ? e.effort : null)) : 'alert',
833      )
834      $.ui.invalidate('ui.render')
835    }
836    if (lines && !verbose) {
837      // One line for the whole decision: effort, skill, advice, time.
838      $.ui.log(
839        turnLine({
840          answered: decision !== null,
841          confidence: decision?.effortConfidence ?? null,
842          applied: change.effort ?? null,
843          current: typeof e.effort === 'string' ? e.effort : null,
844          wanted: level === null ? null : effortLevel(level),
845          jevMs: known?.ms ?? null,
846          skill: skillNote,
847          advised: known?.advised ? (known.strategy ?? null) : null,
848        }),
849      )
850    }
851
852    if (!applied) {
853      // A turn left alone is the common case, and it used to be silent, which
854      // made a working mod look like one that never loaded. Say what happened.
855      if (verbose) {
856        const suppressed = routing.model && !routeMainModel() ? ' (main-loop model routing off)' : ''
857        $.ui.log(`[jev-model-router] main loop: ${routing.reason}${suppressed}`)
858      }
859      return yield* next(e)
860    }
861    if (verbose) {
862      const what = [change.model, change.effort && `effort ${change.effort}`]
863        .filter(Boolean)
864        .join(', ')
865      $.ui.log(`[jev-model-router] main loop → ${what}: ${routing.reason}`)
866    }
867    return yield* next({ ...e, ...change })
868  })
869
870  // Observation only: the main loop's tool calls, counted as they finish, so
871  // a struggling turn is seen without re-reading the transcript every step,
872  // and so the ledger knows how much each turn did.
873  // A refusal (a denied permission) is the person's choice, not the task
874  // going wrong: it neither counts nor clears the run.
875  on('tool.call', async ($, e, next) => {
876    const result = await next(e)
877    if (!e.agentId && !result.deny) {
878      failedInARow = result.isError ? failedInARow + 1 : 0
879      if (current) {
880        current.toolCalls++
881        current.failures = Math.max(current.failures, failedInARow)
882      }
883      // Going in circles: the same file edited again and again, or the same
884      // command run again and failing. Said once a turn, to the model only.
885      const input = e as unknown as { tool?: string; file_path?: unknown; command?: unknown }
886      const circle = spinOf(input, !!result.isError, edits, runs)
887      if (circle && feature('quality') && current && spunTurnId !== current.turnId) {
888        spunTurnId = current.turnId
889        spinning = circle
890        if (verbose) $.ui.log(`[jev-model-router] going in circles: ${circle}`)
891        if (petOn()) {
892          say('going in circles · stepping back', 'alert')
893          $.ui.invalidate('ui.render')
894        }
895        const now = applied?.effort ?? current?.started ?? null
896        const atTop = now === 'xhigh' || now === 'max'
897        return { ...result, context: [...(result.context ?? []), stepBackNote(circle, atTop)] } as typeof result
898      }
899    }
900    return result
901  })
902
903  // The end of a main-loop turn closes its ledger entry: how it ended, how
904  // long it took, what it produced. Written after the engine's own handling.
905  on('turn.complete', async ($, e, next) => {
906    const result = await next(e)
907    // A subagent that finished: its routing is done with.
908    if (e.agentId) {
909      subagents.delete(e.agentId)
910      subagentEffort.delete(e.agentId)
911      recordSubagentUsage(e.agentId, e.usage as Record<string, unknown> | null | undefined)
912    }
913    // A custom model that failed during the turn: said, since Claude's answer hides it.
914    if (!e.agentId && crew().slots.length > 0 && router()) {
915      const failed = await newFallbacks({
916        fetch: (url, init) => $.http.fetch(url, init),
917        sleep: (ms) => $.clock.sleep(ms),
918      })
919      for (const line of failed) $.ui.log(`jev · ${line}. /jev status shows the details`)
920      if (failed.length > 0 && petOn()) {
921        say('custom model failed · Claude answered', 'alert')
922        $.ui.invalidate('ui.render')
923      }
924    }
925    if (!e.agentId && current && current.turnId === e.turnId) {
926      const { turnId, ...entry } = current
927      current = null
928      // Every tenth turn, the bubble says what the session came to.
929      const turns = recordTurn(entry.startedFrom, entry.started)
930      if (turns % 10 === 0 && petOn()) {
931        say(shortStats(), 'calm')
932        $.ui.invalidate('ui.render')
933      }
934      const finished: LedgerEntry = {
935        ...entry,
936        id: turnId,
937        outcome: e.reason,
938        durationMs: e.durationMs,
939        outputTokens: e.usage?.output_tokens ?? null,
940      }
941      try {
942        const entries = appendEntry(await $.store.get(LEDGER_KEY), finished)
943        await $.store.set(LEDGER_KEY, entries)
944        lastEntryId = turnId
945        // Every 20 turns: what the ledger now suggests, in one line.
946        const found = dueToPropose(entries)
947        if (found.length > 0) {
948          const first = found[0] as (typeof found)[number]
949          if (lines) $.ui.log(`jev · learned from your last turns: ${first.why}. /jev tune shows the change, /jev tune apply takes it`)
950          if (petOn()) {
951            say(`tune? ${first.option} ${first.from}→${first.to} · /jev tune`, 'ready')
952            $.ui.invalidate('ui.render')
953          }
954        }
955      } catch (error) {
956        $.ui.log(`[jev-model-router] could not record the turn: ${String(error)}`)
957      }
958    }
959    return result
960  })
961
962  // `/clear` or a resume starts another session in this worker: nothing
963  // waiting or in progress carries over. The learned model ids stay: they
964  // are the engine's own and still valid. (Under a match-all matcher: the
965  // skill module hooks session.end too, and one unmatched hook per plugin.)
966  on('session.end', { sessionId: /(?:)/ }, async ($, e, next) => {
967    pending.clear()
968    resetBriefing()
969    resetStats()
970    lastDecision = null
971    lastEntryId = null
972    heldEffort = undefined
973    spunTurnId = undefined
974    subagents.clear()
975    subagentEffort.clear()
976    clearSkillNotes()
977    current = null
978    appliedTurnId = undefined
979    applied = null
980    escalatedTurnId = undefined
981    turnPrompt = null
982    failedInARow = 0
983    return next(e)
984  })
985
986  // A compaction rewrites the cached conversation anyway: a held effort may
987  // be chosen again. (Under a matcher: the skill module hooks it too.)
988  on('session.compact', { trigger: /(?:)/ }, async ($, e, next) => {
989    heldEffort = undefined
990    return next(e)
991  })
992
993  // `/jev-pilot:report`: the ledger summarised, with suggested changes;
994  // `/jev-pilot:report reset` clears it. The command's markdown is a
995  // placeholder: the prompt the model reads is written here.
996  on('skill.prompt', { skill: 'jev-pilot:report' }, async ($, e, next) => {
997    try {
998      if (/\breset\b/i.test(e.text)) {
999        await $.store.delete(LEDGER_KEY)
1000        return next({ ...e, text: 'Tell the user the jev-pilot decision ledger was cleared. Change nothing else.' })
1001      }
1002      const entries = entriesOf(await $.store.get(LEDGER_KEY))
1003      const home = (await $.env.get('HOME')) ?? '~'
1004      const settingsPath = `${home}/.claude/settings.json`
1005      // The key jev-pilot's options live under depends on how it was
1006      // installed (marketplace, --mod, --plugin-dir): read which exist.
1007      let keys: string[] = []
1008      try {
1009        if (await $.fs.exists(settingsPath)) keys = configKeysOf(await $.fs.read(settingsPath))
1010      } catch {
1011        keys = []
1012      }
1013      const text = reportPrompt(summarize(entries, effective()), settingsPath, suggestions(entries, effective()).length > 0, keys)
1014      return next({ ...e, text })
1015    } catch (error) {
1016      return next({ ...e, text: `Tell the user the jev-pilot ledger could not be read: ${String(error)}. Change nothing.` })
1017    }
1018  })
1019
1020  on('agent.spawn', async ($, e, next) => {
1021    const io: Io = {
1022      fetch: (url, init) => $.http.fetch(url, init),
1023      sleep: (ms) => $.clock.sleep(ms),
1024      log: (text) => $.ui.log(text),
1025      detail: (text) => (verbose ? $.ui.log(text) : undefined),
1026      messages: () => $.session.messages(),
1027    }
1028    // Before the routing guards: a module whose switches are all off has still
1029    // loaded, and that is exactly when its silence is most misleading.
1030    if (!announced) {
1031      announced = true
1032      if (lines) $.ui.log(readyLine())
1033    }
1034
1035    // A fork inherits its parent's model; `model` is ignored for it.
1036    if (!routeSubagentModel() || e.fork) return next(e)
1037
1038    if (!unusableReported) {
1039      unusableReported = true
1040      $.ui.log(`[jev-model-router] provider "${forced}" has no key set; using the built-in classifier`)
1041    }
1042
1043    const crewIo: CrewIo = {
1044      fetch: (url, init) => $.http.fetch(url, init),
1045      home: () => $.env.get('HOME'),
1046      routerUrl: () => $.env.get('JEV_ROUTER_URL'),
1047      write: (path, text) => $.fs.write(path, text),
1048      read: async (path) => ((await $.fs.exists(path)) ? $.fs.read(path) : null),
1049      run: (argv, timeoutMs) => $.process.run(argv, { timeoutMs }),
1050      storeGet: (key) => $.store.get(key),
1051      sleep: (ms) => $.clock.sleep(ms),
1052      register: async (spec) => {
1053        await $.agent.register(spec)
1054      },
1055    }
1056    await ensureCrew(crewIo, openrouterKey || null)
1057
1058    // A reviewer only relays to its CLI: the smallest model does.
1059    if (REVIEWERS.some((r) => e.subagentType === reviewerAgent(r))) {
1060      if (petOn()) {
1061        say(`${subagentLabel(e.description, 'review')} → ${e.subagentType.replace(/^jev-pilot:|-review$/g, '')}`, 'focused')
1062        $.ui.invalidate('ui.render')
1063      }
1064      const reviewed = await next({ ...e, model: policy.tiers.fast })
1065      recordSubagent(reviewed.agentId, policy.tiers.fast, e.parentModel ?? e.model)
1066      return reviewed
1067    }
1068
1069    // The junior runs on its slot; with the slot down, on Sonnet instead.
1070    if (e.subagentType === JUNIOR_AGENT) {
1071      const slot = juniorSlot(crew(), router() !== null)
1072      const model = slot && slotUsable(slot) ? slotAlias(slot.name) : policy.tiers.balanced
1073      if (petOn()) {
1074        say(`${subagentLabel(e.description, 'junior')} → junior on ${model}`, 'focused')
1075        $.ui.invalidate('ui.render')
1076      }
1077      if (lines) $.ui.log(`jev · junior on ${model}`)
1078      const spawned = await next({ ...e, model })
1079      recordSubagent(spawned.agentId, model, e.parentModel ?? e.model)
1080      return spawned
1081    }
1082
1083    // A model the caller named is kept: asked for by the user, or chosen on
1084    // purpose by the main model. Jev still sets the subagent's effort.
1085    const named = e.model ?? null
1086
1087    // Quality mode: Opus for every subagent, no cheaper models.
1088    if (crew().mode === 'quality' && !named) {
1089      if (verbose) $.ui.log(`[jev-model-router] ${e.subagentType}: quality mode, ${policy.tiers.deep}`)
1090      return next({ ...e, model: policy.tiers.deep })
1091    }
1092
1093    // The custom models Jev may choose here: this mode's, the router up, each
1094    // one answering its last check.
1095    const offered = slotsOffered(crew(), router() !== null).filter(slotUsable)
1096    const startedAt = await $.clock.now()
1097    let decision: Decision | null = null
1098    if (active) {
1099      // A subagent's brief is self-contained by design: no conversation added.
1100      decision = (
1101        await classify(
1102          io,
1103          backend,
1104          { prompt: e.prompt, description: e.description, agentType: e.subagentType },
1105          false,
1106          'the subagent',
1107          true,
1108          offered,
1109        )
1110      ).decision
1111    } else {
1112      try {
1113        const label = await $.model.classify(e.prompt, TIER_ORDER)
1114        if (label) {
1115          decision = {
1116            tier: label as Tier,
1117            confidence: null,
1118            risky: null,
1119            effort: null,
1120            effortConfidence: null,
1121          }
1122        }
1123      } catch (error) {
1124        $.ui.log(`[jev-model-router] built-in classifier failed: ${String(error)}`)
1125      }
1126    }
1127
1128    if (logDecisions) {
1129      const ms = (await $.clock.now()) - startedAt
1130      if (verbose) $.ui.log(`[jev-model-router] jev (${e.subagentType}): ${describeDecision(decision, ms)}`)
1131    }
1132
1133    // The subagent's own model wins when the caller named one; otherwise it
1134    // would inherit the parent's, so that is what a change is measured from.
1135    // The Agent tool takes no effort: that is set at the subagent's requests
1136    // (turn.step), from this same decision, kept by the id it starts with.
1137    const current = e.model ?? e.parentModel
1138    let { model, reason } = route(decision, { model: current }, policy)
1139    // A custom model: moving down to it takes the same confidence as any
1140    // cheaper model; the router sends `jev-<slot>` to it.
1141    const slot = decision?.slot ? offered.find((choice) => choice.name === decision.slot) : undefined
1142    if (named) {
1143      model = null
1144      reason = `the caller named ${named}`
1145    } else if (slot) {
1146      const sure = (decision?.confidence ?? 0) >= policy.minDowngradeConfidence
1147      model = sure ? slotAlias(slot.name) : null
1148      reason = sure ? `${slot.name}: ${slot.model} (confidence ${decision?.confidence?.toFixed(2)})` : `${slot.name} wanted, confidence ${decision?.confidence?.toFixed(2)} < ${policy.minDowngradeConfidence}`
1149    }
1150    // Named by its task in the bubble and the log, not its generic type.
1151    const label = subagentLabel(e.description, e.subagentType)
1152    if (!model) {
1153      if (verbose) $.ui.log(`[jev-model-router] ${label} (${e.subagentType}): model kept (${reason})`)
1154    } else {
1155      if (lines) {
1156        $.ui.log(verbose ? `[jev-model-router] ${label} (${e.subagentType}) → ${model}: ${reason}` : `jev · subagent ${label} → ${model}`)
1157      }
1158      if (petOn()) {
1159        say(`${label} → ${model}`, 'focused')
1160        $.ui.invalidate('ui.render')
1161      }
1162    }
1163    // A brief that will run at low effort and may change code gets Anthropic's
1164    // line for low effort: at `low` the check that exercises a change can be skipped.
1165    const low =
1166      routeSubagentEffort() && decision !== null && decision.tier !== 'fast' && effortLevel(effortScoreOf(decision, margin) ?? 1) === 'low'
1167    const brief = low && !e.prompt.includes(CHECK_AT_LOW) ? { prompt: `${e.prompt}\n\n${CHECK_AT_LOW}` } : {}
1168    const result = await next({ ...e, ...(model ? { model } : {}), ...brief })
1169    recordSubagent(result.agentId, model ?? current ?? null, e.parentModel ?? current)
1170    if (decision && result.agentId && routeSubagentEffort()) {
1171      subagents.set(result.agentId, { decision, label, model: model ?? null })
1172      while (subagents.size > MAX_SUBAGENTS) subagents.delete(subagents.keys().next().value as string)
1173    }
1174    return result
1175  })
1176}
1177
hooks/jev-skill-suggestion.ts 708 lines
1/**
2 * jev-skill-suggestion — Claude Mod (EARLY ACCESS)
3 *
4 * Takes the skill listing out of the context window and has TypeSafe's Jev,
5 * a System One decision model, suggest at most one skill per prompt, going
6 * by the skills' descriptions. The skills stay installed and loadable; what
7 * goes away is the listing the engine sends the model every session, one
8 * line per skill, whether the prompt has anything to do with any of them.
9 *
10 * The decision follows TypeSafe's "Skill suggestion" cookbook: two requests
11 * per prompt, one to rank every skill and ask whether the prompt needs a
12 * skill at all, one to re-read the top few with their full text and let each
13 * be rejected on its own. Either may come back empty-handed.
14 *
15 * Three hooks:
16 *   prompt.attachment  — the engine's `skill_listing` attachment is answered
17 *                        with `{ text: null }` (left out) or trimmed to the
18 *                        names in `alwaysListed`. Its names are remembered:
19 *                        they are the engine's word on which skills the model
20 *                        may invoke.
21 *   prompt.submit      — the two requests run, and the winner (if any) is
22 *                        attached to the prompt as a `<skill_relevance>`
23 *                        block: with the skill's own SKILL.md inside it
24 *                        (`inject: "content"`, the default), so the skill
25 *                        loads even when `skillOverrides` hides it from the
26 *                        model, or with its name for the Skill tool
27 *                        (`inject: "suggest"`).
28 *   skill.prompt       — observation: whether the model took the suggestion,
29 *                        or loaded a skill on its own. Also writes the prompt
30 *                        of the plugin's own `/jev-pilot:setup`,
31 *                        which hides every skill from the engine's listing
32 *                        (user-invocable-only) once the person has seen the
33 *                        list and said yes; the model makes the edit with its
34 *                        own tools, so it shows and asks like any other.
35 *
36 * Jev is reached one of three ways, whichever key is configured: TypeSafe's
37 * own API (`typesafeApiKey`) or OpenRouter's Decisions API
38 * (`openrouterApiKey`), which report a calibrated confidence, or the Vercel
39 * AI Gateway (`gatewayApiKey`), which does not. With none, the
40 * engine's own `$.model.classify` stands in with a single request and no
41 * gate, so the mod is useful without any account.
42 *
43 * The candidates come from `$.command.list()`, not from the listing: the
44 * listing is only rendered at the turn's first request, after `prompt.submit`
45 * has run, so the first prompt of a session would otherwise have nothing to
46 * choose from. The listing, once seen, narrows the candidates to what the
47 * engine itself would have shown.
48 *
49 * The second request reads the opening of each shortlisted skill's SKILL.md,
50 * found on disk by how Claude Code lays skills out (project and user
51 * `.claude/skills` and `.claude/commands`, a plugin's install path from
52 * `~/.claude/plugins/installed_plugins.json`). A body that cannot be found
53 * leaves that skill with its one-line description; nothing fails over it.
54 *
55 * Only the main conversation is handled. A subagent's own listing is left as
56 * the engine renders it: its prompt is a tool call's argument, not a
57 * `prompt.submit`, so nothing here could suggest for it.
58 *
59 * Every failure path is fail-open: a request that errors or runs past the
60 * latency budget lets the prompt through with no suggestion, and the listing
61 * hook always answers the same way, so the model's prompt cache holds.
62 *
63 * The API key comes from the plugin's options (userConfig "typesafeApiKey"
64 * or "gatewayApiKey"). Never hardcode it in this file.
65 *
66 * Needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 and Claude Code >= 2.1.278: the
67 * `prompt.attachment` event is that release's. Typed against Anthropic's
68 * declarations: https://github.com/anthropics/claude-code/tree/main/mods
69 *
70 * Privacy: with a key set, the prompt text, every candidate skill's name and
71 * description, and the opening of each shortlisted skill's SKILL.md are sent
72 * to whichever backend the key belongs to.
73 */
74import type { Register } from 'claude-code'
75import { missOf } from './model-router.policy.ts'
76import { NOT_A_TASK, recentContext } from './context.ts'
77import { noteSkill, resetBriefing } from './summary.ts'
78import { feature } from './features.ts'
79import {
80  DEFAULT_BASE_URL,
81  DEFAULT_MODEL,
82  NONE,
83  SETUP_COMMAND,
84  builtinWide,
85  catalog,
86  classifyText,
87  commandLike,
88  decide,
89  pickSkill,
90  designQuestion,
91  readDesign,
92  DESIGN_BAR,
93  packSkills,
94  designBlock,
95  describeRerank,
96  describeSetup,
97  describeStatus,
98  describeStillListed,
99  describeWide,
100  detailOf,
101  displayIds,
102  canonical,
103  endpoint,
104  injectionBlock,
105  installPathsOf,
106  parseListing,
107  parseNames,
108  passesGate,
109  pluginFileCandidates,
110  readRerank,
111  readSkillSettings,
112  modelInvocable,
113  readWide,
114  rerankQuestions,
115  requestBody,
116  requestHeaders,
117  selectProvider,
118  setupAborted,
119  setupInstructions,
120  setupPlan,
121  shortlistOf,
122  validBackup,
123  skillFileCandidates,
124  suggestionBlock,
125  syncedFileCandidates,
126  trimListing,
127  wideQuestions,
128  batchesOf,
129  MAX_CHOICES,
130  mergeWide,
131} from './skill-suggestion.policy.ts'
132import type { Candidate, PolicyConfig, Provider, Rerank, Skill, Suggestion, Wide } from './skill-suggestion.policy.ts'
133import { isContinuation, takePart } from './jev-call.ts'
134
135export const register: Register = (on, options) => {
136  const text = (key: string, fallback: string) =>
137    typeof options[key] === 'string' && options[key] ? (options[key] as string) : fallback
138  const number = (key: string, fallback: number) =>
139    typeof options[key] === 'number' ? (options[key] as number) : fallback
140  const flag = (key: string, fallback: boolean) =>
141    typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback
142
143  // With several keys set, `auto` takes TypeSafe's own API, then OpenRouter,
144  // then the Gateway: the first two report a calibrated confidence.
145  // `provider` forces one, "builtin" uses none.
146  const typesafeKey = text('typesafeApiKey', '')
147  const gatewayKey = text('gatewayApiKey', '')
148  const openrouterKey = text('openrouterApiKey', '')
149  const forced = text('provider', 'auto')
150  const active: Provider | null = selectProvider(forced, typesafeKey, gatewayKey, openrouterKey)
151
152  // Each backend keeps its own URL and model, so an override written for one
153  // can never be sent to the other when `auto` picks differently than expected.
154  const apiKey = !active ? '' : { typesafe: typesafeKey, gateway: gatewayKey, openrouter: openrouterKey }[active]
155  const modelId = !active ? '' : text(`${active}Model`, DEFAULT_MODEL[active])
156  const url = !active ? '' : endpoint(active, text(`${active}BaseUrl`, DEFAULT_BASE_URL[active]))
157
158  // A backend named in the options but missing its key degrades to the
159  // built-in classifier, which is silent; say so once, when a hook first runs.
160  let unusableReported = forced === 'auto' || forced === 'builtin' || active !== null
161
162  const hideListing = flag('hideListing', true)
163  // "content": the mod reads the chosen skill's SKILL.md and attaches it, so
164  // the skill loads even when the engine lists it as user-invocable-only or
165  // off. "suggest": the cookbook's block alone, and the model loads the skill
166  // with the Skill tool, which honours the engine's skillOverrides.
167  const injectContent = text('inject', 'content') !== 'suggest'
168  const alwaysListed = parseNames(text('alwaysListed', ''))
169  const neverSuggested = parseNames(text('neverSuggested', ''))
170  // Off (the default): the one request decides. On: a second request re-reads the shortlist.
171  const rerankEnabled = flag('rerank', false)
172  const excerptChars = number('excerptChars', 300)
173  const timeoutMs = number('timeoutMs', 1500)
174  const logDecisions = flag('logDecisions', true)
175  // One line per turn (written by the router) by default; every step here with verboseLog.
176  const verbose = logDecisions && flag('verboseLog', false)
177  // The setup tip goes to the transcript only where jev-pilot talks there.
178  const display = text('display', 'pet')
179  const lines = logDecisions && (verbose || display === 'transcript' || display === 'both')
180  // Shared with the model router: how much of the conversation Jev reads.
181  const contextLimits = {
182    messages: Math.max(0, Math.round(number('contextMessages', 4))),
183    chars: Math.max(0, number('contextChars', 2000)),
184  }
185  const policy: PolicyConfig = {
186    // The rerank is one Choice too, under the same 255-option limit.
187    shortlist: Math.min(MAX_CHOICES, Math.max(1, Math.round(number('shortlist', 3)))),
188    gateThreshold: number('gateThreshold', 0.3),
189    fitsThreshold: number('fitsThreshold', 0.3),
190  }
191
192  // The names every skill_listing attachment carried so far. Once non-empty,
193  // only these are offered to the decision model: the listing is the engine's
194  // word on which skills the model is allowed to invoke, and `$.command.list()`
195  // also names commands the model may not.
196  const listed = new Set<string>()
197  // The skill suggested for the current prompt, so a skill.prompt that loads
198  // it can be told apart from one the model reached for on its own.
199  let suggested: string | null = null
200  // Each skill's SKILL.md as first found, with where, or null when nowhere:
201  // read once per session, since the second request wants it on every prompt
202  // it is on, and the injection wants it whole.
203  const files = new Map<string, { path: string; markdown: string } | null>()
204  // The skills whose instructions were already attached this session: a
205  // second time, the block only names the skill again.
206  const injected = new Set<string>()
207  // Said once, the first time a hook runs. A mod that loaded and one that
208  // never loaded are otherwise told apart only by the absence of later lines,
209  // and absence is not evidence: with the listing gone, silence is the norm.
210  let announced = false
211  // The setup hint, once per session.
212  let hintedSetup = false
213  // The design pack's skills, and the tool names (for the screen libraries), read once.
214  const designSkillsOption = text('designSkills', 'design-taste-frontend, impeccable, web-design-guidelines')
215  let toolNames: string[] | null = null
216
217  // A skill whose frontmatter `name:` has spaces ("PocketBase API Rules") is
218  // reported by `$.command.list()` under that name, but the engine lists,
219  // runs and overrides it by its directory name (`pb-api-rules`). The map
220  // from one to the other is read from disk once per session, in whichever
221  // hook first needs it.
222  let displayToId: Map<string, string> | null = null
223
224  on('prompt.attachment', { type: 'skill_listing' }, async ($, e, next) => {
225    if (!announced) {
226      announced = true
227      if (verbose) $.ui.log(`[jev-skill-suggestion] ${describeSetup(active, url, hideListing, forced === 'builtin')}`)
228    }
229
230    // A subagent's listing is not ours: nothing here suggests for a subagent,
231    // so hiding it would leave the subagent with no skills, and adding its
232    // names would narrow the main conversation's roster to the subagent's.
233    if (e.agentId) return next(e)
234    // Skills switched off (/jev skills off): the listing reaches the model as
235    // the engine built it, and nothing is picked.
236    if (!feature('skills')) return next(e)
237
238    const skills = parseListing(e.text)
239    for (const skill of skills) listed.add(skill.name)
240
241    // With the mod loading skills itself, a listing that still names any is
242    // context the setup command would have saved: say so once.
243    if (injectContent && !hintedSetup && skills.length > 0) {
244      hintedSetup = true
245      if (lines) {
246        $.ui.log(
247          verbose
248            ? `[jev-skill-suggestion] ${describeStillListed(skills.length)}`
249            : `jev · tip: /jev-pilot:setup takes the ${skills.length} listed skills out of context`,
250        )
251      }
252    }
253
254    if (!hideListing) return next(e)
255
256    const kept = trimListing(e.text, alwaysListed)
257    if (verbose) {
258      const keptNames = kept
259        ? parseListing(kept)
260            .map((skill) => skill.name)
261            .join(', ')
262        : 'none'
263      $.ui.log(
264        `[jev-skill-suggestion] withheld the skill listing (${skills.length} skills, ${e.text.length} characters); kept listed: ${keptNames}`,
265      )
266    }
267    // Answered without `next`: the engine's text never reaches the model.
268    return { text: kept }
269  })
270
271  // Every prompt, under a matcher: jev-pilot.ts registers the model
272  // router first, and a plugin may hook an event once without a matcher. The
273  // two nest, router outermost, so this runs inside it on the same prompt.
274  on('prompt.submit', { text: /(?:)/ }, async ($, e, next) => {
275    if (!announced) {
276      announced = true
277      if (verbose) $.ui.log(`[jev-skill-suggestion] ${describeSetup(active, url, hideListing, forced === 'builtin')}`)
278    }
279    suggested = null
280    // The router's questions for this prompt, to send with ours in one
281    // request (jev-call.ts). Whatever happens here, they are sent and settled.
282    const part = takePart(e.text)
283    /** Passes the prompt on, the router's part asked alone when no ranking carried it. */
284    const alone = async (input: typeof e) => {
285      if (!part) return next(input)
286      const block = await part.settle(await part.ask({}))
287      return next(block ? { ...input, context: [...(input.context ?? []), block] } : input)
288    }
289    if (!feature('skills')) return alone(e)
290
291    // Notifications and peer messages are not tasks; a typed `/name` already
292    // names its skill. Neither gets a suggestion. Nor does "continue": the
293    // work in progress already has what it loaded.
294    if (!e.text.trim() || /^\/\S/.test(e.text.trim())) return alone(e)
295    if (e.origin && NOT_A_TASK.has(e.origin.kind)) return alone(e)
296    if (isContinuation(e.text)) return alone(e)
297
298    // The conversation before the prompt, so a follow-up is read as the work
299    // it continues. Read only when a backend will receive it, and not when
300    // the router's part already carries it.
301    let recent = ''
302    if (active && !part && contextLimits.messages > 0) {
303      try {
304        recent = recentContext(await $.session.messages(), e.text, contextLimits)
305      } catch (error) {
306        $.ui.log(`[jev-skill-suggestion] could not read the conversation: ${String(error)}`)
307      }
308    }
309
310    /** One request to the active backend, or null on timeout, error or a non-2xx. */
311    const ask = async (
312      prompt: string,
313      questions: Record<string, unknown>,
314      what: string,
315    ): Promise<string | null> => {
316      if (!active) return null
317      try {
318        const response = await Promise.race([
319          $.http.fetch(url, {
320            method: 'POST',
321            headers: requestHeaders(active, apiKey, modelId),
322            body: requestBody(active, prompt, questions, modelId, recent),
323          }),
324          $.clock.sleep(timeoutMs),
325        ])
326        if (response && response.ok) return response.text
327        // A timeout or a busy backend is routine (the pet says so): verbose
328        // only. Any other status is an error, and its body says why (a limit,
329        // a bad field): worth the one line.
330        const miss = missOf(response ? response.status : null)
331        const tell = (text: string) => (miss === 'error' || verbose ? $.ui.log(text) : undefined)
332        if (response) {
333          tell(`[jev-skill-suggestion] ${active} responded ${response.status} to the ${what}: ${response.text.slice(0, 200)}`)
334        } else tell(`[jev-skill-suggestion] ${what} passed ${timeoutMs}ms; no suggestion`)
335      } catch (error) {
336        $.ui.log(`[jev-skill-suggestion] ${what} failed: ${String(error)}`)
337      }
338      return null
339    }
340
341    /** A skill's file, found on disk by Claude Code's layout, or null. */
342    const fileOf = async (
343      skill: Skill,
344      plugin: string | undefined,
345    ): Promise<{ path: string; markdown: string } | null> => {
346      const cached = files.get(skill.name)
347      if (cached !== undefined) return cached
348      let found: { path: string; markdown: string } | null = null
349      try {
350        const home = (await $.env.get('HOME')) ?? ''
351        const relative = skillFileCandidates(skill.name, plugin)
352        // The engine reads the project's `.claude/` (the working directory
353        // only, not its ancestors) and the user's.
354        const candidates = [...relative, ...(home ? relative.map((file) => `${home}/${file}`) : [])]
355        if (plugin && home) {
356          const installed = `${home}/.claude/plugins/installed_plugins.json`
357          if (await $.fs.exists(installed)) {
358            for (const path of installPathsOf(await $.fs.read(installed), plugin)) {
359              candidates.push(...pluginFileCandidates(path, skill.name, plugin))
360            }
361          }
362        }
363        // A claude.ai-synced skill sits under an account directory only
364        // `$.fs.list` can name.
365        if (home) {
366          const synced = `${home}/.claude/skills/synced`
367          if (await $.fs.exists(synced)) {
368            const accounts = (await $.fs.list(synced)).filter((entry) => entry.kind === 'dir').map((entry) => entry.name)
369            candidates.push(...syncedFileCandidates(home, accounts, skill.name))
370          }
371        }
372        for (const file of candidates) {
373          if (await $.fs.exists(file)) {
374            found = { path: file, markdown: await $.fs.read(file) }
375            break
376          }
377        }
378      } catch (error) {
379        $.ui.log(`[jev-skill-suggestion] could not read /${skill.name}: ${String(error)}`)
380      }
381      files.set(skill.name, found)
382      return found
383    }
384    /** The opening of a skill's body, or null when its file is nowhere. */
385    const bodyOf = async (skill: Skill, plugin: string | undefined): Promise<string | null> =>
386      (await fileOf(skill, plugin))?.markdown ?? null
387
388    if (!unusableReported) {
389      unusableReported = true
390      $.ui.log(`[jev-skill-suggestion] provider "${forced}" has no key set; using the built-in classifier`)
391    }
392
393    let commands: Awaited<ReturnType<typeof $.command.list>>
394    try {
395      commands = await $.command.list()
396      if (!displayToId && commands.some((command) => !commandLike(command.name))) {
397        const found: { dir: string; markdown: string }[] = []
398        for (const root of [await $.session.cwd(), (await $.env.get('HOME')) ?? '']) {
399          const dir = root && `${root}/.claude/skills`
400          if (!dir || !(await $.fs.exists(dir))) continue
401          for (const entry of await $.fs.list(dir)) {
402            const file = `${dir}/${entry.name}/SKILL.md`
403            if (entry.kind === 'dir' && (await $.fs.exists(file))) found.push({ dir: entry.name, markdown: await $.fs.read(file) })
404          }
405        }
406        displayToId = displayIds(found)
407      }
408      commands = canonical(commands, displayToId ?? new Map())
409    } catch (error) {
410      $.ui.log(`[jev-skill-suggestion] could not list the skills: ${String(error)}`)
411      return alone(e)
412    }
413    // Loading the skill itself, the mod is not bound to what the engine would
414    // list: a skill hidden with skillOverrides is still a candidate.
415    const listedSkills = catalog(commands, injectContent ? new Set() : listed, neverSuggested)
416    if (listedSkills.length === 0) {
417      if (verbose) $.ui.log('[jev-skill-suggestion] no candidate skills; nothing to suggest')
418      return alone(e)
419    }
420    const pluginOf = new Map(commands.map((command) => [command.name, command.plugin]))
421    // Each skill as Jev reads it, in the one request: its description and the
422    // opening of its SKILL.md (read once per session), without the skills
423    // whose own frontmatter says the model may not invoke them.
424    const barred: string[] = []
425    const skills: Skill[] = []
426    for (const skill of listedSkills) {
427      const body = active && !rerankEnabled ? await bodyOf(skill, pluginOf.get(skill.name)) : null
428      if (body !== null && !modelInvocable(body)) {
429        barred.push(skill.name)
430        continue
431      }
432      skills.push(active && !rerankEnabled ? { ...skill, description: detailOf(skill, body, excerptChars) } : skill)
433    }
434    if (skills.length === 0) return alone(e)
435
436    // Request 1: rank everything, and ask whether the prompt wants a skill at
437    // all. The first batch carries the router's questions too: one request.
438    const startedAt = await $.clock.now()
439    let wide: Wide | null = null
440    let parts: (Wide | null)[] = []
441    let routerBlock: string | null = null
442    let routerSettled = false
443    let designRead: number | null = null
444    // The shortlist grows with the batches, so every batch's leaders reach
445    // the rerank (their scores do not compare across batches).
446    let picked: PolicyConfig = policy
447    if (active) {
448      // One Choice takes at most MAX_CHOICES options: a larger catalog is
449      // ranked in batches, side by side, and the gate asked once.
450      const batches = batchesOf(skills, MAX_CHOICES - 1)
451      picked = { ...policy, shortlist: Math.min(MAX_CHOICES, policy.shortlist * batches.length) }
452      const withNone = !rerankEnabled
453      const answers = await Promise.all(
454        batches.map(async (batch, index) => {
455          const questions = wideQuestions(active, batch, index === 0, withNone)
456          // UI design work? Asked in the same request, for the design pack.
457          if (index === 0 && feature('design')) questions.ui_design = designQuestion(active)
458          const what = batches.length > 1 ? `ranking ${index + 1}/${batches.length}` : 'ranking'
459          if (index > 0 || !part) {
460            const text = await ask(e.text, questions, what)
461            if (index === 0) designRead = readDesign(text)
462            return text
463          }
464          const answer = await part.ask(questions)
465          routerSettled = true
466          routerBlock = await part.settle(answer)
467          designRead = readDesign(answer.text)
468          return answer.text
469        }),
470      )
471      parts = answers.map((answer) => (answer ? readWide(answer) : null))
472      wide = mergeWide(parts)
473    } else {
474      // No backend: the engine's own small-model classifier answers the same
475      // question, with the descriptions folded into the text it reads. One
476      // label, no gate, no rerank.
477      try {
478        const label = await $.model.classify(classifyText(e.text, skills), [
479          NONE,
480          ...skills.map((skill) => skill.name),
481        ])
482        wide = builtinWide(label)
483        parts = [wide]
484      } catch (error) {
485        $.ui.log(`[jev-skill-suggestion] built-in classifier failed: ${String(error)}`)
486      }
487    }
488    if (part && !routerSettled) routerBlock = await part.settle(await part.ask({}))
489    // What the decision model actually answered, whatever the policy then
490    // does with it. This is the line that proves the ranking ran.
491    if (verbose) {
492      const ms = (await $.clock.now()) - startedAt
493      $.ui.log(`[jev-skill-suggestion] jev: ${describeWide(wide, skills.length, ms)}`)
494    }
495
496    let decision: Suggestion
497    if (active && rerankEnabled) {
498      // Request 2 (the `rerank` option): re-read the shortlist with each
499      // skill's full text, and let every candidate be rejected on its own.
500      let rerank: Rerank | null = null
501      let rerankAttempted = false
502      if (wide && passesGate(wide, policy)) {
503        const candidates: Candidate[] = []
504        for (const skill of shortlistOf(wide, skills, picked.shortlist)) {
505          const body = await bodyOf(skill, pluginOf.get(skill.name))
506          if (!modelInvocable(body)) {
507            barred.push(skill.name)
508            continue
509          }
510          candidates.push({ ...skill, detail: detailOf(skill, body, excerptChars) })
511        }
512        if (candidates.length > 0) {
513          const rerankStartedAt = await $.clock.now()
514          rerankAttempted = true
515          const answer = await ask(e.text, rerankQuestions(active, candidates), 'rerank')
516          if (answer) rerank = readRerank(answer)
517          if (verbose) {
518            const ms = (await $.clock.now()) - rerankStartedAt
519            const read = candidates.filter((candidate) => files.get(candidate.name)).length
520            $.ui.log(`[jev-skill-suggestion] jev: ${describeRerank(rerank, ms)} · ${read}/${candidates.length} bodies read`)
521          }
522        }
523      }
524      const offered = barred.length > 0 ? skills.filter((skill) => !barred.includes(skill.name)) : skills
525      decision = decide(wide, rerank, offered, picked, rerankAttempted)
526    } else if (active) {
527      decision = pickSkill(parts, skills, policy)
528    } else {
529      decision = decide(wide, null, skills, picked, false)
530    }
531    let pick = decision.name ? (skills.find((skill) => skill.name === decision.name) ?? null) : null
532    // The winner's own frontmatter has the last word, whichever path picked it.
533    if (pick && !barred.includes(pick.name) && !modelInvocable(await bodyOf(pick, pluginOf.get(pick.name)))) {
534      barred.push(pick.name)
535      decision = { name: null, reason: `/${pick.name} has disable-model-invocation` }
536      pick = null
537    }
538    if (verbose && barred.length > 0) {
539      $.ui.log(
540        `[jev-skill-suggestion] not model-invocable, left out: ${barred.map((name) => `/${name}`).join(', ')}`,
541      )
542    }
543    // The router writes the turn's one line (and its status) from this note.
544    noteSkill(e.text, { skill: pick?.name ?? null, ms: active ? Math.round((await $.clock.now()) - startedAt) : null })
545    if (verbose) $.ui.status(describeStatus(pick?.name ?? null))
546    if (verbose) {
547      $.ui.log(
548        pick
549          ? `[jev-skill-suggestion] suggesting /${pick.name}: ${decision.reason}`
550          : `[jev-skill-suggestion] no suggestion: ${decision.reason}`,
551      )
552    }
553
554    suggested = pick?.name ?? null
555    let block: string | null
556    if (injectContent && pick) {
557      const file = await fileOf(pick, pluginOf.get(pick.name))
558      const projectDir = await $.session.cwd()
559      block = injectionBlock(pick, file?.markdown ?? null, file?.path ?? null, projectDir, injected.has(pick.name))
560      if (verbose) {
561        $.ui.log(
562          file
563            ? injected.has(pick.name)
564              ? `[jev-skill-suggestion] /${pick.name} already injected this session; named again`
565              : `[jev-skill-suggestion] injected /${pick.name} from ${file.path} (${file.markdown.length} characters)`
566            : `[jev-skill-suggestion] no file found for /${pick.name}; suggested by name only`,
567        )
568      }
569      if (file) injected.add(pick.name)
570    } else {
571      block = suggestionBlock(pick, hideListing)
572    }
573    // Attached on the way down, the router's advice first: blocks after the
574    // prompt as typed, read by the model and never shown to the person.
575    // UI design work: the design pack, with the design skills installed and
576    // the screen libraries connected.
577    let design: string | null = null
578    // (Set inside the request's callbacks above, which the compiler can't follow.)
579    const designScore = designRead as number | null
580    if (feature('design') && designScore !== null && designScore >= DESIGN_BAR) {
581      const skills = packSkills([...parseNames(designSkillsOption)], listedSkills.map((skill) => skill.name))
582      if (!toolNames) toolNames = (await $.tool.list().catch(() => [])).map((tool) => tool.name)
583      design = designBlock(skills, {
584        mobbin: toolNames.some((name) => name.startsWith('mcp__mobbin__')),
585        inspo: toolNames.some((name) => name.startsWith('mcp__inspo__')),
586      })
587      if (verbose) $.ui.log(`[jev-skill-suggestion] design work (${designScore.toFixed(2)}): ${skills.map((name) => `/${name}`).join(', ') || 'no design skills installed'}`)
588    }
589    const blocks = [routerBlock, block, design].filter((b): b is string => b !== null)
590    if (blocks.length === 0) return next(e)
591    return next({ ...e, context: [...(e.context ?? []), ...blocks] })
592  })
593
594  // An injected skill lives in the conversation, not the process: `/clear`
595  // or a resume starts another under the same worker, and a compaction may
596  // summarize the block away. Either way the next pick goes in whole again.
597  on('session.end', async ($, e, next) => {
598    injected.clear()
599    suggested = null
600    // Everything learned about this session's skills goes with it: the next
601    // one may have another roster, and its SKILL.md files may have changed.
602    listed.clear()
603    files.clear()
604    displayToId = null
605    announced = false
606    hintedSetup = false
607    return next(e)
608  })
609  on('session.compact', async ($, e, next) => {
610    if (!e.agentId) {
611      injected.clear()
612      // The compacted conversation no longer holds the router's note to the
613      // model about what jev-pilot does: it is given again on the next prompt.
614      resetBriefing()
615    }
616    return next(e)
617  })
618
619  on('skill.prompt', { skill: 'jev-pilot:setup' }, async ($, e, next) => {
620    // The plugin's own setup command: its markdown is a placeholder, and the
621    // prompt the model reads is written here, from the roster as the engine
622    // has it and the user settings as they are. The model does the editing
623    // with its own tools, so the change shows as a diff and asks permission.
624    const mode = /\brestore\b/i.test(e.text) ? 'restore' : 'apply'
625    // No plan from a partial roster or unreadable settings: the edit would
626    // hide too little, and the backup would save the wrong values.
627    let commands: Awaited<ReturnType<typeof $.command.list>> = []
628    try {
629      commands = await $.command.list()
630      if (!displayToId && commands.some((command) => !commandLike(command.name))) {
631        const found: { dir: string; markdown: string }[] = []
632        for (const root of [await $.session.cwd(), (await $.env.get('HOME')) ?? '']) {
633          const dir = root && `${root}/.claude/skills`
634          if (!dir || !(await $.fs.exists(dir))) continue
635          for (const entry of await $.fs.list(dir)) {
636            const file = `${dir}/${entry.name}/SKILL.md`
637            if (entry.kind === 'dir' && (await $.fs.exists(file))) found.push({ dir: entry.name, markdown: await $.fs.read(file) })
638          }
639        }
640        displayToId = displayIds(found)
641      }
642      commands = canonical(commands, displayToId ?? new Map())
643    } catch (error) {
644      $.ui.log(`[jev-skill-suggestion] setup: could not list the skills: ${String(error)}`)
645      return next({ ...e, text: setupAborted(`the skills could not be listed (${String(error)})`) })
646    }
647    const home = (await $.env.get('HOME')) ?? '~'
648    const settingsPath = `${home}/.claude/settings.json`
649    const backupPath = `${home}/.claude/jev-pilot.skill-overrides.backup.json`
650    let json: string | null = null
651    try {
652      if (await $.fs.exists(settingsPath)) json = await $.fs.read(settingsPath)
653    } catch (error) {
654      $.ui.log(`[jev-skill-suggestion] setup: could not read ${settingsPath}: ${String(error)}`)
655      return next({ ...e, text: setupAborted(`${settingsPath} exists but could not be read (${String(error)})`) })
656    }
657    const settings = readSkillSettings(json)
658    // Never plan edits on a file that could not be parsed: the backup would
659    // miss what it holds, and the edit could destroy it.
660    if (settings.invalid) {
661      $.ui.log(`[jev-skill-suggestion] setup: ${settingsPath} is not valid JSON`)
662      return next({ ...e, text: setupAborted(`${settingsPath} is not a valid JSON object; ask the user to fix it first`) })
663    }
664    const plan = setupPlan(commands, settings, new Set([SETUP_COMMAND]))
665    // An earlier run's backup is reused only if it is one: a file that is
666    // not this mod's, or is corrupt, is nothing restore could apply, so no
667    // setup is built on top of it.
668    let backupExists = false
669    try {
670      backupExists = await $.fs.exists(backupPath)
671      if (backupExists && !validBackup(await $.fs.read(backupPath))) {
672        $.ui.log(`[jev-skill-suggestion] setup: ${backupPath} is not a valid backup`)
673        return next({
674          ...e,
675          text: setupAborted(
676            `${backupPath} exists but is not a backup this mod wrote (expected {"skillOverrides": {...}, "disableBundledSkills": true|false|null}); ask the user to inspect it and move it away, or fix it, before running the setup again`,
677          ),
678        })
679      }
680    } catch (error) {
681      $.ui.log(`[jev-skill-suggestion] setup: could not read ${backupPath}: ${String(error)}`)
682      return next({ ...e, text: setupAborted(`${backupPath} could not be read (${String(error)})`) })
683    }
684    if (logDecisions) {
685      $.ui.log(
686        `[jev-skill-suggestion] setup (${mode}): ${plan.hide.length} to hide, ${plan.alreadyHidden.length} already hidden, ${plan.locked.length} locked by a plugin`,
687      )
688    }
689    return next({ ...e, text: setupInstructions(mode, plan, settings, settingsPath, backupPath, backupExists) })
690  })
691
692  on('skill.prompt', async ($, e, next) => {
693    // Observation only: whether the model took the suggestion, or reached for
694    // a skill it was never told about, is the one measure of this mod's worth.
695    // The plugin's own commands (setup, report) are not skills to measure.
696    if (verbose && !e.skill.startsWith('jev-pilot:')) {
697      const how =
698        suggested === e.skill
699          ? 'as suggested'
700          : suggested
701            ? `suggested was /${suggested}`
702            : 'nothing was suggested'
703      $.ui.log(`[jev-skill-suggestion] skill /${e.skill} loaded (${how})`)
704    }
705    return next(e)
706  })
707}
708
hooks/jev-pet.tsx 408 lines
1/**
2 * jev-pet — Jev as a companion above the prompt, at the right: Claude the
3 * pilot (Claude Code's character in pilot gear), with a speech bubble saying what Jev just
4 * decided, so the decisions stay out of the conversation. It also owns
5 * `/jev`, the switches for every part of jev-pilot.
6 *
7 * While a turn runs the pilot shows what Claude is doing, and the bubble says
8 * it beside a spinner: thinking (a thought cloud), reading (a book),
9 * searching (a magnifier), writing (paper and a pencil), running a command
10 * (a terminal), anything else (a subagent) flying, goggles down. With
11 * subagents still working in the background after a turn, it cruises,
12 * goggles down, until they finish; at max effort the goggles stay down for
13 * the turn. Idle, it hovers, its scarf's end dipping every few seconds; it
14 * blinks, and every few seconds plays for a moment: jumps rope, waves, looks
15 * around. Resting, it redraws only when something moves.
16 *
17 * The router and the skill module set what it says (pet-art.ts `say`) and ask
18 * for a redraw; this module draws it, in the `AbovePrompt` band on the
19 * terminal. Nothing draws in `claude -p`, the desktop app or mobile.
20 */
21import type { Register, Timer } from 'claude-code'
22import {
23  describeFeatures,
24  feature,
25  FEATURE_INFO,
26  featureOverrides,
27  loadFeatureOverrides,
28  parseJevCommand,
29  setFeature,
30} from './features.ts'
31import { applyCrewCommand, describeChoice, describeCrew, describeSlot, MAX_MODELS, MODELS_PAGE, parseCrewCommand, resolveCodexChoice, resolveModel, resolveOpencodeChoice, slotAlias, type CrewCommand, type ReviewerResolved } from './crew.ts'
32import { resetBriefing } from './summary.ts'
33import { describeStats } from './session-stats.ts'
34import { entriesOf, LEDGER_KEY } from './ledger.ts'
35import { applied, currentTuning, describeTuning, proposals, setTuning, TUNING_KEY, tuningLoaded, tuningOf } from './tuning.ts'
36import { codexModels, crew, crewOverrides, describeHealth, router, setCrewOverrides } from './crew-state.ts'
37import { CREW_KEY, checkCrew, checkSlot, ensureCrew, modelCatalog, opencodeModels, publishSlots, refreshCodexModels, registerCrew, type CrewIo } from './crew-run.ts'
38import {
39  type Act,
40  ACT_LABEL,
41  actOfTool,
42  CANVAS_W,
43  currentSpeech,
44  isBoosted,
45  MOOD_COLOR,
46  PLAY_FRAMES,
47  PLAYS,
48  sceneRows,
49  setBoost,
50  WORK_ACTS,
51} from './pet-art.ts'
52
53const FEATURES_KEY = 'features'
54const FLY_MS = 200
55// Background subagents still working, no turn running: the pilot cruises,
56// goggles down, at a calmer rate; running agents are checked this often.
57const CRUISE_MS = 450
58const AGENTS_EVERY_MS = 1500
59const BLINK_EVERY_MS = 4600
60const BLINK_MS = 170
61const PLAY_EVERY_MS = 9000
62const PLAY_STEP_MS = 180
63/** How many loops each idle play runs for: about three seconds each. */
64const PLAY_LOOPS = { rope: 4, wave: 4, look: 2 } as const
65const SPINNER = '⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏'
66
67export const register: Register = (on, options) => {
68  let act: Act = 'rest'
69  let frame = 0
70  let blink = false
71  let workingTurn: string | null = null
72  let flying: Timer | null = null
73  let blinker: Timer | null = null
74  let player: Timer | null = null
75  let playing: Timer | null = null
76  let plays = 0
77  let cruising: Timer | null = null
78  let watcher: Timer | null = null
79  // The main loop's tool calls in flight, each with the act it shows.
80  const running = new Map<string, Act>()
81  let calls = 0
82
83  const stopPlay = () => {
84    playing?.cancel()
85    playing = null
86  }
87  const stopCruise = () => {
88    cruising?.cancel()
89    cruising = null
90  }
91
92  // Session setup: the saved switches, the /jev command, the blink. (Under a
93  // match-all matcher: the router hooks session.start too, one unmatched
94  // registration per plugin.)
95  on('session.start', { cwd: /(?:)/ }, async ($, e, next) => {
96    const result = await next(e)
97    loadFeatureOverrides(await $.store.get(FEATURES_KEY).catch(() => undefined))
98    await $.command
99      .register({
100        name: 'jev',
101        description: 'jev-pilot switches: /jev shows them, /jev <effort|raise|subagents|skills|strategy|model|pet> on|off',
102        argumentHint: '[<feature> on|off | all on|off | reset]',
103        immediate: true,
104      })
105      .catch((error) => $.ui.log(`[jev-pilot] /jev not registered: ${String(error)}`))
106    blinker?.cancel()
107    blinker = $.clock.every(BLINK_EVERY_MS, () => {
108      if (!feature('pet') || workingTurn || playing || cruising) return
109      blink = true
110      $.ui.invalidate('ui.render')
111      $.clock.after(BLINK_MS, () => {
112        blink = false
113        $.ui.invalidate('ui.render')
114      })
115    })
116    // Idle play: every few seconds one of the plays, in turn, for a moment.
117    player?.cancel()
118    player = $.clock.every(PLAY_EVERY_MS, () => {
119      if (!feature('pet') || workingTurn || playing || cruising) return
120      const play = PLAYS[plays++ % PLAYS.length] as (typeof PLAYS)[number]
121      const steps = PLAY_FRAMES[play] * PLAY_LOOPS[play]
122      act = play
123      frame = 0
124      playing = $.clock.every(PLAY_STEP_MS, () => {
125        frame++
126        if (frame >= steps) {
127          stopPlay()
128          act = 'rest'
129          frame = 0
130        }
131        $.ui.invalidate('ui.render')
132      })
133      $.ui.invalidate('ui.render')
134    })
135    // Subagents working in the background, with no turn running: the pilot
136    // cruises until they are all done. Asked of the engine every few seconds,
137    // so one that was stopped or failed never leaves it flying.
138    watcher?.cancel()
139    watcher = $.clock.every(AGENTS_EVERY_MS, async () => {
140      if (!feature('pet')) return
141      let busy = false
142      try {
143        busy = (await $.agent.list()).some((agent) => agent.status === 'running')
144      } catch {
145        busy = false
146      }
147      if (busy && !workingTurn && !cruising) {
148        stopPlay()
149        act = 'fly'
150        frame = 0
151        cruising = $.clock.every(CRUISE_MS, () => {
152          frame++
153          $.ui.invalidate('ui.render')
154        })
155      } else if (!busy && cruising) {
156        stopCruise()
157        if (!workingTurn) {
158          act = 'rest'
159          frame = 0
160          $.ui.invalidate('ui.render')
161        }
162      }
163    })
164    return result
165  })
166
167  // A new session in this worker (/clear, resume): the old timers stop.
168  on('session.end', { reason: /(?:)/ }, async ($, e, next) => {
169    flying?.cancel()
170    flying = null
171    stopPlay()
172    stopCruise()
173    setBoost(false)
174    act = 'rest'
175    workingTurn = null
176    running.clear()
177    return next(e)
178  })
179
180  on('command.run', { command: 'jev' }, async ($, e) => {
181    // `/jev tune [apply|reset]`: the changes the ledger suggests.
182    const tune = /^\s*tune(?:\s+(apply|reset))?\s*$/i.exec(e.args)
183    if (tune) {
184      if (!tuningLoaded()) setTuning(tuningOf(await $.store.get(TUNING_KEY).catch(() => undefined)))
185      const entries = entriesOf(await $.store.get(LEDGER_KEY).catch(() => undefined))
186      const action = tune[1]?.toLowerCase()
187      if (action === 'reset') {
188        setTuning({ values: {}, since: Date.now() })
189        await $.store.delete(TUNING_KEY).catch(() => undefined)
190        return { text: `tuning cleared; back to your settings.\n\n${describeTuning(entries)}` }
191      }
192      if (action === 'apply') {
193        const found = proposals(entries)
194        if (found.length === 0) return { text: `nothing to apply.\n\n${describeTuning(entries)}` }
195        const next = applied(currentTuning(), found, Date.now())
196        setTuning(next)
197        await $.store.set(TUNING_KEY, next).catch((error) => $.ui.log(`[jev-pilot] tuning not saved: ${String(error)}`))
198        return {
199          text: `applied ${found.map((f) => `${f.option} ${f.from} → ${f.to}`).join(', ')}. The next suggestion waits for 20 new turns.\n\n${describeTuning(entries)}`,
200        }
201      }
202      return { text: describeTuning(entries) }
203    }
204    // The crew: `/jev status`, `/jev mode ...`, `/jev alpha <model>|off`,
205    // `/jev junior <slot>`, `/jev reviewer <codex|opencode>`.
206    const crewCommand = /^\s*status\s*$/i.test(e.args) ? ({ kind: 'show' } as const) : parseCrewCommand(e.args)
207    if (crewCommand) {
208      const crewIo: CrewIo = {
209        fetch: (url, init) => $.http.fetch(url, init),
210        home: () => $.env.get('HOME'),
211        routerUrl: () => $.env.get('JEV_ROUTER_URL'),
212        write: (path, text) => $.fs.write(path, text),
213        read: async (path) => ((await $.fs.exists(path)) ? $.fs.read(path) : null),
214        run: (argv, timeoutMs) => $.process.run(argv, { timeoutMs }),
215        storeGet: (key) => $.store.get(key),
216        sleep: (ms) => $.clock.sleep(ms),
217        register: async (spec) => {
218          await $.agent.register(spec)
219        },
220      }
221      const openrouterKey = typeof options.openrouterApiKey === 'string' && options.openrouterApiKey ? options.openrouterApiKey : null
222      await ensureCrew(crewIo, openrouterKey)
223      if (crewCommand.kind === 'unknown') return { text: `jev-pilot: unknown "${crewCommand.text}".\n${describeCrew(crew(), router() !== null, codexModels())}` }
224      if (crewCommand.kind === 'show') {
225        await checkCrew(crewIo, openrouterKey)
226        await registerCrew(crewIo)
227        return { text: `${describeCrew(crew(), router() !== null, codexModels())}\n${describeHealth()}\n\n${describeStats()}` }
228      }
229      if (crewCommand.kind === 'slot') return { text: describeSlot(crew(), crewCommand.slot, crewOverrides().recent ?? []) }
230      // What was pasted, checked against OpenRouter's list before it's set.
231      let change: CrewCommand = crewCommand
232      let about: string | null = null
233      // A reviewer's model and effort, checked against the CLI's own list before it's set.
234      if (crewCommand.kind === 'reviewer-paste') {
235        const current = crew().reviewerChoices[crewCommand.reviewer] ?? {}
236        let resolved: ReviewerResolved
237        if (crewCommand.reviewer === 'codex') {
238          if (codexModels().length === 0) await refreshCodexModels(crewIo)
239          resolved = resolveCodexChoice(crewCommand.input, crewCommand.effort, codexModels(), current)
240        } else {
241          resolved = resolveOpencodeChoice(crewCommand.input, crewCommand.effort, crewCommand.input ? await opencodeModels(crewIo) : [], current)
242        }
243        if (!resolved.ok) {
244          const list = resolved.suggestions.length > 0 ? `\nAvailable:\n${resolved.suggestions.map((line) => `  ${line}`).join('\n')}` : ''
245          return { text: `The ${crewCommand.reviewer} reviewer is not changed: ${resolved.why}.${list}` }
246        }
247        change = { kind: 'reviewer-choice', reviewer: crewCommand.reviewer, choice: resolved.choice }
248        about = resolved.about || null
249      }
250      if (crewCommand.kind === 'paste') {
251        const resolved = resolveModel(crewCommand.input, await modelCatalog(crewIo))
252        if (!resolved.ok) {
253          const tries = resolved.suggestions.length > 0 ? `\nDid you mean:\n${resolved.suggestions.map((id) => `  /jev ${crewCommand.slot} ${id}`).join('\n')}` : ''
254          return { text: `${crewCommand.slot} not changed: ${resolved.why}.${tries}\nFind one at ${MODELS_PAGE} and paste its id, link or name.` }
255        }
256        // A new name past the limit: each custom model is an option Jev weighs.
257        if (!crew().slots.some((slot) => slot.name === crewCommand.slot) && crew().slots.length >= MAX_MODELS) {
258          return { text: `${crewCommand.slot} not added: ${MAX_MODELS} custom models is the most at once. Remove one first: /jev remove <name>` }
259        }
260        change = { kind: 'model', slot: crewCommand.slot, model: resolved.id, about: resolved.about }
261        about = resolved.about
262      }
263      if (change.kind === 'model' && !change.model && !crew().slots.some((slot) => slot.name === change.slot)) {
264        return { text: `No custom model is called ${change.slot}.\n\n${describeCrew(crew(), router() !== null, codexModels())}` }
265      }
266      setCrewOverrides(applyCrewCommand(crewOverrides(), change))
267      await $.store.set(CREW_KEY, crewOverrides()).catch((error) => $.ui.log(`[jev-pilot] crew not saved: ${String(error)}`))
268      await publishSlots(crewIo).catch((error) => $.ui.log(`[jev-pilot] custom models not published: ${String(error)}`))
269      if (change.kind === 'model' && change.model) await checkSlot(crewIo, openrouterKey, change.slot, change.model)
270      // A new junior, say: its agent type from the next turn, and the model
271      // told about the change with its next prompt.
272      await registerCrew(crewIo)
273      resetBriefing()
274      const headline =
275        change.kind === 'model'
276          ? change.model
277            ? `${change.slot} is now ${change.model}${about ? ` (${about})` : ''}. Saved for every session.\n` +
278              `It's in /model from your next claude-jev session (press s there to use it for that session only). To use it as the main model right now: /model ${slotAlias(change.slot)}. That also makes it your default for new sessions, and /model default undoes that.\n\n`
279            : `${change.slot} is removed, from every session and from /model.\n\n`
280          : change.kind === 'reviewer-choice'
281            ? `${change.reviewer} reviews now run on ${describeChoice(change.choice, codexModels())}${about ? ` (${about})` : ''}. Saved for every session. To use another model for one review, just say so ("review it with codex, luna, high").\n\n`
282            : ''
283      // A mode that hands work to custom models, with none set: say how to set one.
284      const unset =
285        change.kind === 'mode' && (change.mode === 'budget' || change.mode === 'junior-lead') && crew().slots.length === 0
286          ? `\n\nNo custom model is set yet, so this mode has nothing to hand work to. Add one from ${MODELS_PAGE}:\n  /jev <name you choose> <model>`
287          : ''
288      return { text: `${headline}${describeCrew(crew(), router() !== null, codexModels())}\n${describeHealth()}${unset}` }
289    }
290    const command = parseJevCommand(e.args)
291    if (command.kind === 'unknown') {
292      return { text: `jev-pilot: unknown "${command.text}".\n${describeFeatures()}` }
293    }
294    if (command.kind === 'reset') loadFeatureOverrides({})
295    if (command.kind === 'set') for (const name of command.features) setFeature(name, command.on)
296    if (command.kind !== 'status') {
297      await $.store.set(FEATURES_KEY, featureOverrides()).catch((error) => $.ui.log(`[jev-pilot] switches not saved: ${String(error)}`))
298      $.ui.invalidate('ui.render')
299    }
300    if (command.kind === 'set' && command.features.length === 1) {
301      const name = command.features[0] as keyof typeof FEATURE_INFO
302      return { text: `jev-pilot: ${name} ${command.on ? 'on' : 'off'} (${FEATURE_INFO[name]}) · /jev ${name} ${command.on ? 'off' : 'on'} undoes it` }
303    }
304    return { text: `${describeFeatures()}\n\n${describeStats()}` }
305  })
306
307  // While a turn runs, the pilot shows what Claude is doing: thinking first.
308  on('turn.start', async ($, e, next) => {
309    const result = await next(e)
310    workingTurn = e.turnId
311    flying?.cancel()
312    stopPlay()
313    stopCruise()
314    setBoost(false)
315    running.clear()
316    act = 'think'
317    frame = 0
318    flying = $.clock.every(FLY_MS, () => {
319      if (!feature('pet')) return
320      frame++
321      $.ui.invalidate('ui.render')
322    })
323    return result
324  })
325
326  // The model's response as it streams: thinking shows as thinking, the
327  // answer's text as writing. Observed only: every chunk passes on as it came.
328  on('turn.step', { turnId: /(?:)/ }, async function* ($, e, next) {
329    if (e.agentId || e.turnId !== workingTurn) return yield* next(e)
330    if (running.size === 0) act = 'think'
331    for await (const chunk of next(e)) {
332      if (running.size === 0) {
333        if (chunk.kind === 'thinking') act = 'think'
334        else if (chunk.kind === 'text') act = 'write'
335      }
336      yield chunk
337    }
338  })
339
340  // A main-loop tool call shows what the tool does: reading, searching,
341  // editing, running a command; anything else (a subagent) flies. When it is
342  // done, the next call still running shows, or thinking.
343  on('tool.call', { tool: /(?:)/ }, async ($, e, next) => {
344    if (e.agentId || !workingTurn) return next(e)
345    const id = e.tool_use_id ?? `call-${++calls}`
346    act = actOfTool(e.tool)
347    running.set(id, act)
348    try {
349      return await next(e)
350    } finally {
351      running.delete(id)
352      if (workingTurn) act = [...running.values()].at(-1) ?? 'think'
353    }
354  })
355
356  on('turn.complete', { turnId: /(?:)/ }, async ($, e, next) => {
357    const result = await next(e)
358    if (e.turnId === workingTurn) {
359      running.clear()
360      workingTurn = null
361      flying?.cancel()
362      flying = null
363      setBoost(false)
364      act = 'rest'
365      frame = 0
366      $.ui.invalidate('ui.render')
367    }
368    return result
369  })
370
371  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
372    if (!feature('pet') || e.props.hasSurvey || e.surface !== 'terminal') return next(e)
373    const { Box, Text } = $.ui.resolve(e)
374    // The band's own width: the transcript column's while a Pane is docked.
375    const columns = e.props.bodyColumns || (e.viewport?.columns ?? 80)
376    const speech = currentSpeech()
377    const color = MOOD_COLOR[speech.mood]
378    // Goggles down when flying, and at full power (effort raised to max).
379    const rows = sceneRows(act, frame, blink, frame, act === 'fly' || isBoosted())
380    const working = (WORK_ACTS as readonly Act[]).includes(act)
381    const text = working ? `${SPINNER[frame % SPINNER.length]} ${ACT_LABEL[act as (typeof WORK_ACTS)[number]]} · ${speech.text}` : speech.text
382    return (
383      <Box flexDirection="column">
384        <Box key="jev:pet" flexDirection="row" justifyContent="flex-end" alignItems="center" columnGap={1} width={columns} paddingRight={4}>
385          {/* The bubble gives way: a long line is cut, the pilot is never squeezed. */}
386          <Box key="jev:bubble" borderStyle="round" borderColor={color} paddingX={1} flexShrink={1}>
387            <Text key="jev:say" color={color} wrap="truncate-end">
388              {text}
389            </Text>
390          </Box>
391          <Box key="jev:sprite" flexDirection="column" flexShrink={0} width={CANVAS_W} minWidth={CANVAS_W}>
392            {rows.map((row, y) => (
393              <Text key={`jev:row${y}`} wrap="truncate-end">
394                {row.map((cell, x) => (
395                  <Text key={`jev:${y}:${x}`} color={cell.fg} backgroundColor={cell.bg}>
396                    {cell.ch}
397                  </Text>
398                ))}
399              </Text>
400            ))}
401          </Box>
402        </Box>
403        {await next(e)}
404      </Box>
405    )
406  })
407}
408
hooks/features.ts 109 lines
1/**
2 * jev-pilot — what is switched on. Every part can be turned off on its own,
3 * from the settings (defaults) or live with `/jev <feature> on|off`, which is
4 * remembered in the plugin's store across sessions.
5 *
6 * Shared: the modules run in the plugin's one worker, so they all read these.
7 */
8
9export type Feature = 'effort' | 'raise' | 'subagents' | 'skills' | 'strategy' | 'quality' | 'design' | 'model' | 'pet'
10
11export const FEATURES: readonly Feature[] = ['effort', 'raise', 'subagents', 'skills', 'strategy', 'quality', 'design', 'model', 'pet']
12
13/** What each switch does, for `/jev`. */
14export const FEATURE_INFO: Record<Feature, string> = {
15  effort: 'sets the reasoning effort of each turn',
16  raise: 'raises the effort when tool calls keep failing',
17  subagents: 'picks each subagent’s model and effort',
18  skills: 'picks the one skill a prompt needs (off: the full skill list stays)',
19  strategy: 'advises splitting big work across subagents',
20  quality: 'asks before guessing, tests bugs first, checks costly changes, and steps back when a turn goes in circles',
21  design: 'gives UI design work your design skills, a real product\u2019s DESIGN.md or Mobbin screens for direction, and the web guidelines to check against',
22  model: 'switches the main conversation’s model (resets the prompt cache)',
23  pet: 'shows Claude the pilot above the prompt',
24}
25
26let defaults: Record<Feature, boolean> = {
27  effort: true,
28  raise: true,
29  subagents: true,
30  skills: true,
31  strategy: true,
32  quality: true,
33  design: true,
34  model: false,
35  pet: true,
36}
37let overrides: Partial<Record<Feature, boolean>> = {}
38
39/** The defaults, from the plugin's options (once, when the plugin registers). */
40export function initFeatures(from: Record<Feature, boolean>): void {
41  defaults = { ...from }
42  overrides = {}
43}
44
45export function feature(name: Feature): boolean {
46  return overrides[name] ?? defaults[name]
47}
48
49export function setFeature(name: Feature, on: boolean): void {
50  if (on === defaults[name]) delete overrides[name]
51  else overrides[name] = on
52}
53
54/** What `/jev` changed, to keep in the store. */
55export function featureOverrides(): Partial<Record<Feature, boolean>> {
56  return { ...overrides }
57}
58
59/** Overrides read back from the store; anything that is not one is ignored. */
60export function loadFeatureOverrides(stored: unknown): void {
61  overrides = {}
62  if (!stored || typeof stored !== 'object' || Array.isArray(stored)) return
63  for (const [name, on] of Object.entries(stored as Record<string, unknown>)) {
64    if ((FEATURES as readonly string[]).includes(name) && typeof on === 'boolean') {
65      setFeature(name as Feature, on)
66    }
67  }
68}
69
70export type JevCommand =
71  | { kind: 'status' }
72  | { kind: 'set'; features: Feature[]; on: boolean }
73  | { kind: 'reset' }
74  | { kind: 'unknown'; text: string }
75
76/**
77 * `/jev` arguments:
78 *   (none)             what is on
79 *   <feature> on|off   one switch
80 *   all on|off         every switch
81 *   on | off           the pet (a shortcut)
82 *   reset              back to the settings' defaults
83 */
84export function parseJevCommand(args: string): JevCommand {
85  const words = args.trim().toLowerCase().split(/\s+/).filter(Boolean)
86  if (words.length === 0) return { kind: 'status' }
87  if (words.length === 1 && words[0] === 'reset') return { kind: 'reset' }
88  const onOff = (word: string | undefined) => (word === 'on' ? true : word === 'off' ? false : null)
89  if (words.length === 1) {
90    const on = onOff(words[0])
91    return on === null ? { kind: 'unknown', text: words[0] as string } : { kind: 'set', features: ['pet'], on }
92  }
93  const on = onOff(words[1])
94  if (on === null || words.length > 2) return { kind: 'unknown', text: args.trim() }
95  if (words[0] === 'all') return { kind: 'set', features: [...FEATURES], on }
96  if ((FEATURES as readonly string[]).includes(words[0] as string)) return { kind: 'set', features: [words[0] as Feature], on }
97  return { kind: 'unknown', text: words[0] as string }
98}
99
100/** `/jev`'s answer: every switch with its state. */
101export function describeFeatures(): string {
102  const width = Math.max(...FEATURES.map((name) => name.length))
103  return [
104    'jev-pilot switches (/jev <name> on|off, /jev all on|off, /jev reset):',
105    ...FEATURES.map((name) => `  ${feature(name) ? 'on ' : 'off'}  ${name.padEnd(width)}  ${FEATURE_INFO[name]}`),
106    'also: /jev status (the crew and its health) · /jev mode <name> · /jev tune (changes learned from your turns)',
107  ].join('\n')
108}
109
hooks/crew-state.ts 124 lines
1/**
2 * jev-pilot — the crew as this session has it: the options, what `/jev`
3 * changed, whether the router is up, and each worker's last health check.
4 * Shared, like features.ts: the router module routes with it, the pet module's
5 * `/jev` changes and shows it.
6 */
7import { crewOf, type CodexModel, type Crew, type CrewOverrides, type Reviewer, type Slot } from './crew.ts'
8
9export type Health = { ok: boolean; detail: string; at: number }
10
11let options: Record<string, unknown> = {}
12let overrides: CrewOverrides = {}
13let routerUrl: string | null = null
14let started = false
15const health = new Map<string, Health>()
16// Codex's model list as it was last read (`codex debug models`): what a tier such as `luna` means now.
17let codexList: CodexModel[] = []
18
19export function codexModels(): CodexModel[] {
20  return codexList
21}
22
23export function setCodexModels(list: CodexModel[]): void {
24  codexList = [...list]
25}
26
27export function initCrew(from: Record<string, unknown>): void {
28  options = { ...from }
29  overrides = {}
30  routerUrl = null
31  started = false
32  health.clear()
33}
34
35/** Whether this worker has set the crew up: after a reload it hasn't, until a hook does. */
36export function crewStarted(): boolean {
37  return started
38}
39
40export function markCrewStarted(): void {
41  started = true
42}
43
44export function crew(): Crew {
45  return crewOf(options, overrides)
46}
47
48export function crewOverrides(): CrewOverrides {
49  return { ...overrides }
50}
51
52export function setCrewOverrides(next: CrewOverrides): void {
53  overrides = { ...next }
54}
55
56/** The router's address when `claude-jev` started one and it answered; null otherwise. */
57export function router(): string | null {
58  return routerUrl
59}
60
61export function setRouter(url: string | null): void {
62  routerUrl = url
63}
64
65/** Health keys: `router`, `slot:alpha`, `agent:codex`, ... */
66export function setHealth(key: string, value: Health): void {
67  health.set(key, value)
68}
69
70export function healthOf(key: string): Health | null {
71  return health.get(key) ?? null
72}
73
74/** A slot is usable once its last check passed (unchecked counts as not yet). */
75export function slotHealthy(slot: Slot): boolean {
76  const h = health.get(`slot:${slot.name}`)
77  return !!h && h.ok && h.detail.startsWith(slot.model)
78}
79
80/**
81 * A slot that may be used: answering its last check, or not checked yet (a
82 * headless run's first prompt arrives before the check). An unchecked slot
83 * is safe to try: if it fails, the router gives the request to Claude.
84 * Only a slot that failed its check is left out.
85 */
86export function slotUsable(slot: Slot): boolean {
87  const h = health.get(`slot:${slot.name}`)
88  return !h || !h.detail.startsWith(slot.model) || h.ok
89}
90
91export function reviewerHealthy(reviewer: Reviewer): boolean {
92  return health.get(`agent:${reviewer}`)?.ok === true
93}
94
95/** The health lines for `/jev status`. */
96export function describeHealth(): string {
97  const c = crew()
98  const mark = (h: Health | null) => (h === null ? '·  not checked' : h.ok ? `✓  ${h.detail}` : `✗  ${h.detail}`)
99  const lines = ['health:']
100  lines.push(`  router     ${routerUrl ? mark(health.get('router') ?? null) : '✗  not running (start Claude Code with claude-jev)'}`)
101  for (const slot of c.slots) lines.push(`  ${slot.name.padEnd(10)} ${mark(health.get(`slot:${slot.name}`) ?? null)}`)
102  for (const agent of ['codex', 'opencode']) lines.push(`  ${agent.padEnd(10)} ${mark(health.get(`agent:${agent}`) ?? null)}`)
103  return lines.join('\n')
104}
105
106// ---- the checks' verdicts, pure: the hooks run them and hand the output here ----
107
108/** Codex is working when `codex login status` exits 0 and says it's logged in. */
109export function codexVerdict(exitCode: number, output: string): { ok: boolean; detail: string } {
110  if (exitCode !== 0) return { ok: false, detail: `codex login status failed: ${output.trim().slice(0, 80) || `exit ${exitCode}`}` }
111  const line = output.trim().split('\n')[0] ?? ''
112  return /^logged in\b/i.test(line) ? { ok: true, detail: line } : { ok: false, detail: `not logged in: ${line.slice(0, 80)}` }
113}
114
115/** OpenCode is working when it runs and has at least one provider logged in. */
116export function opencodeVerdict(versionExit: number, version: string, authOutput: string): { ok: boolean; detail: string } {
117  if (versionExit !== 0) return { ok: false, detail: 'opencode not found or not starting' }
118  const plain = authOutput.replace(/\x1b\[[0-9;]*m/g, '')
119  const providers = [...plain.matchAll(/●\s+([^\n]+?)\s+(?:oauth|api|wellknown)\b/g)].map((m) => (m[1] as string).trim())
120  return providers.length > 0
121    ? { ok: true, detail: `opencode ${version.trim()} · ${providers.slice(0, 3).join(', ')}${providers.length > 3 ? '…' : ''}` }
122    : { ok: false, detail: `opencode ${version.trim()} has no provider logged in (opencode auth login)` }
123}
124
hooks/context.ts 203 lines
1/**
2 * jev-pilot — what the decision model is shown of the conversation, and
3 * which prompts are tasks at all.
4 *
5 * Pure, like the policy modules: the hooks read `$.session.messages()` and
6 * hand the list here. Only message text and tool names travel, never a tool's
7 * input or output: those hold file contents and command output, which is more
8 * than a classifier needs and more than should leave the machine.
9 */
10
11/** Prompt origins that are not a task of the person's: nothing to plan or suggest for. */
12export const NOT_A_TASK: ReadonlySet<string> = new Set([
13  'task-notification',
14  'peer',
15  'peer-send-message',
16  'projects-relay',
17  'observer',
18  'observer-activity',
19  'scheduled-trigger',
20  'slack-ping',
21])
22
23/** The part of a transcript message this module reads (`SessionMessage`). */
24export interface ContextMessage {
25  role: 'user' | 'assistant'
26  text: string
27  toolUses?: readonly { tool: string; isError?: true }[]
28  toolResults?: readonly { isError: boolean }[]
29}
30
31export interface ContextLimits {
32  /** How many messages before the prompt to include; 0 sends none. */
33  messages: number
34  /** The most characters all of them may take together. */
35  chars: number
36}
37
38/** A message as one line: who, what they said, and which tools ran. */
39function lineOf(message: ContextMessage, cap: number, tail = 0): string | null {
40  const text = message.text.replace(/\s+/g, ' ').trim()
41  const tools = (message.toolUses ?? []).map((use) => (use.isError ? `${use.tool} (failed)` : use.tool))
42  if (!text && tools.length === 0) return null
43  // With a tail, a long message keeps its beginning and its end: where a
44  // reply asks its question ("Shall I start?") is usually the end.
45  const said =
46    text.length <= cap ? text : tail > 0 && tail < cap ? `${text.slice(0, cap - tail)} … ${text.slice(-tail)}` : `${text.slice(0, cap)}…`
47  const ran = tools.length > 0 ? ` [tools: ${tools.join(', ')}]` : ''
48  return `${message.role}: ${said}${ran}`
49}
50
51/**
52 * The conversation just before `prompt`, newest last, as the text the
53 * decision model reads beside it; '' when there is none or `messages` is 0.
54 *
55 * The prompt itself is dropped when the transcript already holds it, so it is
56 * never counted twice. Messages that carry only tool results (no text) are
57 * skipped: their outcome already shows on the tool call as "(failed)". The
58 * newest messages win the character budget; older ones are dropped, not
59 * squeezed. The newest is always sent, cut to the budget if it must be.
60 */
61export function recentContext(
62  messages: readonly ContextMessage[],
63  prompt: string,
64  limits: ContextLimits,
65): string {
66  if (limits.messages <= 0 || limits.chars <= 0) return ''
67  let list = messages
68  const last = list.at(-1)
69  if (last && last.role === 'user' && last.text.trim() === prompt.trim()) list = list.slice(0, -1)
70
71  // The newest assistant message is what a short reply answers ("yes", "1",
72  // "fix all and continue"), and its proposal is usually at its end: it gets
73  // about half the budget and keeps its beginning and its end; the others
74  // share the rest.
75  const cap = Math.max(200, Math.floor(limits.chars / limits.messages))
76  let newestAssistant = -1
77  for (let index = list.length - 1; index >= 0; index--) {
78    if ((list[index] as ContextMessage).role === 'assistant' && (list[index] as ContextMessage).text.trim()) {
79      newestAssistant = index
80      break
81    }
82  }
83  const bigCap = Math.min(limits.chars, Math.max(cap, Math.floor(limits.chars * 0.55)))
84  const otherCap = newestAssistant >= 0 && limits.messages > 1 ? Math.max(200, Math.floor((limits.chars - bigCap) / (limits.messages - 1))) : cap
85  const lines: string[] = []
86  let used = 0
87  for (let index = list.length - 1; index >= 0 && lines.length < limits.messages; index--) {
88    const message = list[index] as ContextMessage
89    const line = index === newestAssistant ? lineOf(message, bigCap, Math.floor(bigCap * 0.72)) : lineOf(message, otherCap)
90    if (!line) continue
91    if (used + line.length > limits.chars) {
92      // A budget smaller than one message still carries the newest one.
93      if (lines.length === 0) lines.unshift(`${line.slice(0, Math.max(0, limits.chars - 1))}…`)
94      break
95    }
96    lines.unshift(line)
97    used += line.length + 1
98  }
99  return lines.join('\n')
100}
101
102/**
103 * Plain facts about a request, sent beside it so the decision model does not
104 * have to infer them from prose: how long it is, how many files it names,
105 * whether it carries code or an error, whether it is phrased as a question,
106 * and what the recent turns did with their tools. Counts and flags only;
107 * nothing here quotes the conversation.
108 */
109export interface Signals {
110  prompt_chars: number
111  files_mentioned: number
112  has_code_or_error: boolean
113  is_question: boolean
114  /** Tool use over the last `window` messages, by kind. */
115  recent_tools: { edits: number; commands: number; reads: number; subagents: number; failed: number }
116}
117
118const TOOL_KINDS: Record<string, keyof Omit<Signals['recent_tools'], 'failed'>> = {
119  Edit: 'edits',
120  MultiEdit: 'edits',
121  Write: 'edits',
122  NotebookEdit: 'edits',
123  Bash: 'commands',
124  PowerShell: 'commands',
125  Read: 'reads',
126  Grep: 'reads',
127  Glob: 'reads',
128  LS: 'reads',
129  WebFetch: 'reads',
130  WebSearch: 'reads',
131  Agent: 'subagents',
132  Task: 'subagents',
133}
134
135const PATH = /(?:^|[\s`'"(])((?:[\w.-]+\/)+[\w.-]+|[\w-]+\.(?:tsx?|jsx?|mjs|py|go|rs|java|kt|swift|rb|php|cs|c|cc|cpp|h|hpp|sql|json|ya?ml|toml|md|css|scss|html|sh|lock))(?=$|[\s`'"),:;])/g
136const CODE_OR_ERROR = /```|Traceback|Exception|\berror\b|\bfailed\b|stack ?trace|\bat \S+:\d+/i
137const QUESTION = /^(what|why|how|when|where|which|who|can|could|does|do|did|is|are|should|would|will)\b/i
138
139export function signalsOf(prompt: string, messages: readonly ContextMessage[], window = 10): Signals {
140  const text = prompt.trim()
141  const files = new Set<string>()
142  for (const match of text.matchAll(PATH)) files.add(match[1] as string)
143  const tools = { edits: 0, commands: 0, reads: 0, subagents: 0, failed: 0 }
144  for (const message of messages.slice(-window)) {
145    for (const use of message.toolUses ?? []) {
146      const kind = TOOL_KINDS[use.tool]
147      if (kind) tools[kind]++
148      if (use.isError) tools.failed++
149    }
150  }
151  return {
152    prompt_chars: text.length,
153    files_mentioned: files.size,
154    has_code_or_error: CODE_OR_ERROR.test(text),
155    is_question: text.endsWith('?') || QUESTION.test(text),
156    recent_tools: tools,
157  }
158}
159
160// ---- what the project deploys with ------------------------------------------------------
161
162/**
163 * The files that show which platform a project deploys to or builds on, by
164 * what they're called. Read from the project's top folder and one level down
165 * (infra/cdk.json, apps/api/Dockerfile), never their contents.
166 */
167const PLATFORM_MARKERS: [RegExp, string][] = [
168  [/(^|\/)vercel\.json$|(^|\/)\.vercel\/$/, 'Vercel'],
169  [/(^|\/)cdk\.json$/, 'AWS CDK'],
170  [/(^|\/)serverless\.(yml|yaml|ts|js)$/, 'AWS Serverless'],
171  [/(^|\/)template\.ya?ml$|(^|\/)samconfig\.toml$/, 'AWS SAM'],
172  [/(^|\/)buildspec\.ya?ml$/, 'AWS CodeBuild'],
173  [/(^|\/)amplify\.ya?ml$|(^|\/)amplify\/$/, 'AWS Amplify'],
174  [/(^|\/)netlify\.toml$/, 'Netlify'],
175  [/(^|\/)firebase\.json$/, 'Firebase'],
176  [/(^|\/)supabase\/config\.toml$/, 'Supabase'],
177  [/(^|\/)fly\.toml$/, 'Fly.io'],
178  [/(^|\/)wrangler\.(toml|json|jsonc)$/, 'Cloudflare Workers'],
179  [/(^|\/)app\.ya?ml$|(^|\/)cloudbuild\.ya?ml$/, 'Google Cloud'],
180  [/(^|\/)render\.ya?ml$/, 'Render'],
181  [/(^|\/)railway\.(json|toml)$/, 'Railway'],
182  [/(^|\/)(docker-)?compose\.ya?ml$/, 'Docker Compose'],
183  [/(^|\/)Dockerfile$/, 'Docker'],
184  [/(^|\/)\.github\/workflows\/$/, 'GitHub Actions'],
185  [/(^|\/)bitbucket-pipelines\.yml$/, 'Bitbucket Pipelines'],
186]
187
188/** Folders never looked into for platform files: dependencies and build output. */
189export const SKIPPED_DIRS = new Set(['node_modules', '.git', 'dist', 'build', '.next', 'out', 'cdk.out', 'vendor', 'target', '.venv', 'venv', '__pycache__', 'coverage'])
190
191/**
192 * What a project deploys to or builds on, from its file names (folders end
193 * in "/"): "AWS CDK, Docker", or "none found". Jev reads it so a platform's
194 * skill (Vercel's, say) fits only a project that uses that platform.
195 */
196export function platformsOf(paths: readonly string[]): string {
197  const found: string[] = []
198  for (const [marker, name] of PLATFORM_MARKERS) {
199    if (!found.includes(name) && paths.some((path) => marker.test(path))) found.push(name)
200  }
201  return found.length > 0 ? found.join(', ') : 'none found'
202}
203
hooks/summary.ts 105 lines
1/**
2 * jev-pilot — the one line each turn gets in the transcript, and the note the
3 * skill module leaves for it.
4 *
5 * Both modules run in the plugin's one worker, so this module is shared: the
6 * skill module notes its pick per prompt at `prompt.submit`, and the router
7 * reads it when the turn starts and writes a single line for both, instead
8 * of each module logging every step (that detail is `verboseLog`).
9 */
10
11import { sure } from './pet-art.ts'
12
13/** What the skill module decided for one prompt. */
14export interface SkillNote {
15  /** The skill attached, or null for none. */
16  skill: string | null
17  /** How long its requests took, ms; null when none was made. */
18  ms: number | null
19}
20
21const notes = new Map<string, SkillNote>()
22const MAX_NOTES = 32
23
24/** Records the skill module's pick for `prompt`, bounded to the last few. */
25export function noteSkill(prompt: string, note: SkillNote): void {
26  notes.delete(prompt)
27  notes.set(prompt, note)
28  while (notes.size > MAX_NOTES) notes.delete(notes.keys().next().value as string)
29}
30
31/** The pick noted for `prompt`, removed as it is read; null when none. */
32export function takeSkill(prompt: string | null): SkillNote | null {
33  if (prompt === null) return null
34  const note = notes.get(prompt) ?? null
35  notes.delete(prompt)
36  return note
37}
38
39/** A new session: no pick carries over. */
40export function clearSkillNotes(): void {
41  notes.clear()
42}
43
44/** What the turn line says. */
45export interface TurnFacts {
46  /** Whether the decision model answered for this turn. */
47  answered: boolean
48  /** How sure the decision model was of the effort it picked. */
49  confidence: number | null
50  /** The effort this turn was set to, or null when left as built. */
51  applied: string | null
52  /** The effort the engine built the turn with. */
53  current: string | null
54  /** The level the decision model's answer pointed to. */
55  wanted: string | null
56  /** The router's request time, ms. */
57  jevMs: number | null
58  skill: SkillNote | null
59  /** The strategy advised to the model, if any. */
60  advised: string | null
61}
62
63/**
64 * One line for a turn, the effort with how sure the decision model was of it:
65 *   jev · low (93% sure) · no skill · 1.3s
66 *   jev · xhigh (88% sure) · skill /systematic-debugging · parallel advised · 1.4s
67 *   jev · high kept (wanted low, 42% sure) · no skill · 1.2s
68 */
69export function turnLine(facts: TurnFacts): string {
70  const parts = ['jev']
71  if (!facts.answered) {
72    parts.push('no answer, turn left as built')
73  } else {
74    const read = facts.confidence === null ? '' : sure(facts.confidence)
75    if (facts.applied) {
76      parts.push(`${facts.applied}${read ? ` (${read})` : ''}`)
77    } else if (facts.wanted && facts.current && facts.wanted !== facts.current) {
78      parts.push(`${facts.current} kept (wanted ${facts.wanted}${read ? `, ${read}` : ''})`)
79    } else {
80      parts.push(`${facts.current ?? 'default effort'}${read ? ` (${read})` : ''}`)
81    }
82  }
83  if (facts.skill) parts.push(facts.skill.skill ? `skill /${facts.skill.skill}` : 'no skill')
84  if (facts.advised) parts.push(`${facts.advised} advised`)
85  const ms = (facts.jevMs ?? 0) + (facts.skill?.ms ?? 0)
86  if (ms > 0) parts.push(`${(ms / 1000).toFixed(1)}s`)
87  return parts.join(' · ')
88}
89
90// ---- the note to the model: once per session, and again after a compaction ----
91
92let briefed = false
93
94/** Whether the model still needs jev-pilot's capability note; marks it given. */
95export function takeBriefing(): boolean {
96  if (briefed) return false
97  briefed = true
98  return true
99}
100
101/** A new session or a compacted conversation: the note is due again. */
102export function resetBriefing(): void {
103  briefed = false
104}
105
hooks/pet-art.ts 409 lines
1/**
2 * jev-pilot — the pet: Claude Code's character as a pilot, drawn in
3 * half-block characters at the bottom right, and what it says.
4 *
5 * Pure: the sprite as rows of cells, the speech texts, and the one shared
6 * speech state the hooks update and the pet's render hook reads (the modules
7 * run in the plugin's one worker, so this state is shared between them).
8 */
9
10/** How the pet feels about the last thing it did: sets the bubble's color. */
11export type Mood = 'ready' | 'calm' | 'focused' | 'boost' | 'alert'
12
13export const MOOD_COLOR: Record<Mood, string> = {
14  ready: '#8b949e',
15  calm: '#7cf0c4',
16  focused: '#c9d2ea',
17  boost: '#d4ff4f',
18  alert: '#ff8f8f',
19}
20
21export interface Speech {
22  text: string
23  mood: Mood
24}
25
26let speech: Speech = { text: 'ready', mood: 'ready' }
27
28export function say(text: string, mood: Mood): void {
29  speech = { text, mood }
30}
31
32export function currentSpeech(): Speech {
33  return speech
34}
35
36// Full power: the effort was raised to max this turn. The pilot wears its
37// goggles down until the turn ends.
38let boosted = false
39
40export function setBoost(on: boolean): void {
41  boosted = on
42}
43
44export function isBoosted(): boolean {
45  return boosted
46}
47
48/**
49 * A subagent's name in the bubble: its task's short description (the Agent
50 * tool's `description`, e.g. "Fix S2a Codex findings"), cut to fit; its type
51 * (`general-purpose`, `Explore`) when it has none.
52 */
53export function subagentLabel(description: string | undefined, type: string, max = 32): string {
54  const text = (description ?? '').replace(/\s+/g, ' ').trim()
55  if (!text) return type
56  return text.length <= max ? text : `${text.slice(0, max - 1).trimEnd()}…`
57}
58
59/** The mood an effort level reads as. */
60export function moodOf(effort: string | null): Mood {
61  if (effort === 'low' || effort === 'medium') return 'calm'
62  if (effort === 'high') return 'focused'
63  if (effort === 'xhigh' || effort === 'max') return 'boost'
64  return 'ready'
65}
66
67/** How sure the decision model was of the effort, as people read it: `89% sure`. */
68export function sure(confidence: number): string {
69  return `${Math.round(confidence * 100)}% sure`
70}
71
72/**
73 * What the pet says when a turn starts:
74 *   low · no skill · 89% sure
75 *   jev busy · left as is          (the backend overloaded; the turn runs as set)
76 *   xhigh · /systematic-debugging · parallel · 88% sure
77 *   high kept · wanted low · 42% sure
78 */
79export function turnSpeech(facts: {
80  answered: boolean
81  /** Why the decision model gave no answer, when it gave none. */
82  miss?: 'timeout' | 'busy' | 'error' | null
83  applied: string | null
84  current: string | null
85  wanted: string | null
86  confidence: number | null
87  skill: string | null | undefined
88  advised: string | null
89}): Speech {
90  if (!facts.answered) {
91    const why = facts.miss === 'busy' ? 'jev busy' : facts.miss === 'error' ? 'jev error' : 'no answer in time'
92    return { text: `${why} · left as is`, mood: 'alert' }
93  }
94  const parts: string[] = []
95  const effort = facts.applied ?? facts.current
96  if (facts.applied) parts.push(facts.applied)
97  else if (facts.wanted && facts.current && facts.wanted !== facts.current) parts.push(`${facts.current} kept`, `wanted ${facts.wanted}`)
98  else parts.push(facts.current ?? 'effort as set')
99  if (facts.skill !== undefined) parts.push(facts.skill ? `/${facts.skill}` : 'no skill')
100  if (facts.advised) parts.push(facts.advised)
101  if (facts.confidence !== null) parts.push(sure(facts.confidence))
102  return { text: parts.join(' · '), mood: moodOf(effort) }
103}
104
105// ---- the sprite -------------------------------------------------------------
106
107/** One character of the sprite: a glyph and its colors. */
108export interface Cell {
109  ch: string
110  fg?: string
111  bg?: string
112}
113
114const CORAL = '#D97757'
115export const PALETTE: Record<string, string> = {
116  C: CORAL, // the pilot's body
117  E: '#0b1020', // eyes
118  L: '#d4ff4f', // goggle lenses
119  W: '#ffffff', // the lenses' glint
120  G: '#2b3a67', // goggle strap
121  S: '#7cf0c4', // scarf
122  F: '#ffd166', // flame
123  f: '#ff7a59', // flame, outer
124  R: '#e6edf3', // jump rope, thought dots, cursor
125  B: '#5b8def', // book cover
126  P: '#f5f0e1', // book pages
127  M: '#56607d', // magnifier rim
128  H: '#a06a3f', // magnifier handle, pencil wood
129  r: '#ff5f57', // terminal: close
130  y: '#febc2e', // terminal: minimise
131  g: '#28c840', // terminal: zoom
132  l: '#9fe7ff', // magnifier lens
133  K: '#8b949e', // laptop
134  k: '#0b1020', // laptop screen
135}
136
137/**
138 * The pilot: Claude Code's character (the same pilot as the README banner and
139 * the demo video) at 3/4 of the banner's 16x12, every part kept: goggles
140 * pushed up on the forehead (pulled down over the eyes to fly), their lenses
141 * glinting, the head, two-pixel eyes, both rows of arms, a teal scarf, long
142 * legs, and the jet flames under them.
143 */
144const CLAWD = [
145  '.GWLGGGGWLG.',
146  '.CCCCCCCCCC.',
147  '.CCECCCCECC.',
148  'CCCECCCCECCC',
149  'CCCCCCCCCCCC',
150  '.SSSSSSSSSS.',
151  '..C.C..C.C..',
152  '..C.C..C.C..',
153]
154const FLAMES = ['..F.f..F.f..', '..f.F..f.F..']
155
156/**
157 * The canvas every scene draws on: fixed, so the band never changes size.
158 * The pilot's own part is the left BODY_W columns; to its right, what it holds.
159 */
160export const BODY_W = 14
161export const CANVAS_W = 24
162export const CANVAS_H = 10
163
164/**
165 * What the pet is doing. While a turn runs, what Claude is doing:
166 *   think   thinking: eyes up, thought dots rising
167 *   read    reading files or pages: a book, its pages turning
168 *   search  searching: a magnifier sweeping, the eyes following it
169 *   write   editing files: typing code on a laptop
170 *   run     running commands: a terminal prompt, the cursor blinking
171 *   fly     anything else (subagents, other tools): flying, goggles down
172 * Idle:
173 *   rest    hovering, the scarf flapping, blinking now and then
174 *   rope    play: jumping rope
175 *   wave    play: waving
176 *   look    play: looking around
177 */
178export type Act = 'think' | 'read' | 'search' | 'write' | 'run' | 'fly' | 'rest' | 'rope' | 'wave' | 'look'
179
180/** The acts a turn shows, by what Claude is doing. */
181export const WORK_ACTS = ['think', 'read', 'search', 'write', 'run', 'fly'] as const
182
183/** What the bubble says Claude is doing, for each working act. */
184export const ACT_LABEL: Record<(typeof WORK_ACTS)[number], string> = {
185  think: 'thinking',
186  read: 'reading',
187  search: 'searching',
188  write: 'writing',
189  run: 'running',
190  fly: 'working',
191}
192
193/** The act a tool call shows: what the tool does, by its name. */
194export function actOfTool(tool: string): Act {
195  if (/^(Read|NotebookRead|WebFetch|ReadMcpResource)/.test(tool)) return 'read'
196  if (/^(Grep|Glob|LS|WebSearch|ToolSearch)$/.test(tool)) return 'search'
197  if (/^(Edit|MultiEdit|Write|NotebookEdit)$/.test(tool)) return 'write'
198  if (/^(Bash|PowerShell|Monitor)$/.test(tool)) return 'run'
199  return 'fly'
200}
201
202/** The idle plays, in the order they take turns. */
203export const PLAYS = ['rope', 'wave', 'look'] as const
204export type Play = (typeof PLAYS)[number]
205
206/** How many frames one loop of each idle play has. */
207export const PLAY_FRAMES: Record<Play, number> = { rope: 4, wave: 2, look: 4 }
208
209type Grid = string[][]
210
211function blank(): Grid {
212  return Array.from({ length: CANVAS_H }, () => Array.from({ length: CANVAS_W }, () => '.'))
213}
214
215function paste(grid: Grid, rows: readonly string[], ox: number, oy: number): void {
216  rows.forEach((row, y) => {
217    for (let x = 0; x < row.length; x++) {
218      const ch = row[x] as string
219      if (ch !== '.') dot(grid, ox + x, oy + y, ch)
220    }
221  })
222}
223
224function dot(grid: Grid, x: number, y: number, ch: string): void {
225  if (y >= 0 && y < CANVAS_H && x >= 0 && x < CANVAS_W) (grid[y] as string[])[x] = ch
226}
227
228function setAt(row: string, x: number, ch: string): string {
229  return row.slice(0, x) + ch + row.slice(x + 1)
230}
231
232/** The pilot for one frame of one act, before it is placed. */
233function clawdFor(act: Act, frame: number, blink: boolean, goggles: boolean): string[] {
234  let rows = [...CLAWD]
235  if (act === 'look' && !blink) {
236    // The eyes glance left, back, right, back.
237    const shift = [-1, 0, 1, 0][frame % 4] as number
238    rows = rows.map((row, y) => {
239      if (y !== 2 && y !== 3) return row
240      return setAt(setAt(row.replace(/E/g, 'C'), 3 + shift, 'E'), 8 + shift, 'E')
241    })
242  }
243  if (act === 'think') rows[3] = (rows[3] as string).replace(/E/g, 'C') // eyes up
244  if (act === 'read' || act === 'search' || act === 'write' || act === 'run') {
245    // The eyes turn to what the pilot holds at its side (down, for the page).
246    rows = rows.map((row, y) => (y === 2 || y === 3 ? setAt(setAt(row.replace(/E/g, 'C'), 4, 'E'), 9, 'E') : row))
247    if (act !== 'search') rows[2] = (rows[2] as string).replace(/E/g, 'C')
248  }
249  if (blink) rows = rows.map((row) => row.replace(/E/g, 'C'))
250  if (act === 'wave' && frame % 2 === 0) {
251    // The right arm up beside the head.
252    for (const y of [3, 4]) rows[y] = setAt(rows[y] as string, 11, '.')
253    for (const y of [1, 2]) rows[y] = setAt(rows[y] as string, 11, 'C')
254  }
255  if (goggles) {
256    // Goggles down over the eyes, the strap round the head: flying, or at
257    // full power.
258    rows[0] = '.CCCCCCCCCC.'
259    rows[1] = '.CCCCCCCCCC.'
260    rows[2] = '.GWLGGGGWLG.'
261    rows[3] = 'CCLLCCCCLLCC'
262  }
263  if (act === 'fly') {
264    // The flames flicker long and short, their colors steady.
265    rows.push(...(frame % 2 === 1 ? [FLAMES[0] as string] : FLAMES))
266  } else if (act !== 'rope') {
267    // Hovering: the flames on, steady. (Jumping rope, the feet do the work.)
268    rows.push(FLAMES[0] as string)
269  }
270  return rows
271}
272
273/**
274 * The pixel canvas for one frame of one act. `wind` is no longer used (the
275 * scarf's end that flapped in it is gone) and is kept only so callers need
276 * not change. `goggles` pulls them down over the eyes: always when flying.
277 */
278export function scenePixels(act: Act, frame = 0, blink = false, wind = frame, goggles = act === 'fly'): string[] {
279  const grid = blank()
280  const clawd = clawdFor(act, frame, blink, goggles)
281  const X = 1
282  if (act === 'fly') {
283    // Flying: a bob, down a pixel on the short flame.
284    paste(grid, clawd, X, frame % 2)
285  } else if (act === 'rope') {
286    // Four beats: the rope overhead, coming down in front, under the feet
287    // (the pilot up in the air), coming round behind.
288    const beat = frame % 4
289    const lift = beat === 2 ? 2 : beat === 1 ? 1 : 0
290    const top = CANVAS_H - clawd.length - lift
291    paste(grid, clawd, X, top)
292    const hands = top + 3
293    const right = BODY_W - 1
294    if (beat === 0) {
295      for (let x = 1; x < right; x++) dot(grid, x, 0, 'R')
296      for (let y = 1; y < hands; y++) dot(grid, 0, y, 'R'), dot(grid, right, y, 'R')
297    } else if (beat === 2) {
298      for (let x = 1; x < right; x++) dot(grid, x, CANVAS_H - 1, 'R')
299      for (let y = hands + 1; y < CANVAS_H - 1; y++) dot(grid, 0, y, 'R'), dot(grid, right, y, 'R')
300    } else if (beat === 1) {
301      for (let y = hands; y < CANVAS_H; y++) dot(grid, 0, y, 'R'), dot(grid, right, y, 'R')
302    } else {
303      for (let y = 0; y <= hands; y++) dot(grid, 0, y, 'R'), dot(grid, right, y, 'R')
304    }
305  } else {
306    paste(grid, clawd, X, CANVAS_H - clawd.length)
307    drawProp(grid, act, frame)
308  }
309  return grid.map((row) => row.join(''))
310}
311
312/** Pixel art for what the pilot holds, each drawn to read as the thing at a glance. */
313const CLOUD = ['.RR.RRR..', 'RRRRRRRRR', 'RRRRRRRRR', '.RRRRRRR.']
314const BOOK = [
315  '.PPP.PPP.',
316  'BkkPMkkPB',
317  'BPPPMPPPB',
318  'BkkPMkPPB',
319  'BBBBBBBBB',
320]
321const LENS = ['.MMMM.', 'MRRllM', 'MRlllM', 'MllllM', 'MllllM', '.MMMM.']
322const PAPER = ['PPPPM.', 'PPPPPP', 'PPPPPP', 'PPPPPP', 'PPPPPP', 'PPPPPP', 'PPPPPP', 'PPPPPP']
323const TERMINAL = [
324  'KKKKKKKKK',
325  'KrKyKgKKK',
326  'KkkkkkkkK',
327  'KkkkkkkkK',
328  'KkkkkkkkK',
329  'KkkkkkkkK',
330  'KkkkkkkkK',
331  'KKKKKKKKK',
332]
333
334/**
335 * What the pilot holds at its side while working, right of its body, by the
336 * right hand (resting, the arm ends at column 12, rows 4 and 5).
337 */
338function drawProp(grid: Grid, act: Act, frame: number): void {
339  const art = (x: number, y: number, rows: readonly string[]) => paste(grid, rows, x, y)
340  const line = (x: number, y: number, pixels: string) => paste(grid, [pixels], x, y)
341  const X = BODY_W // the first column right of the pilot
342  if (act === 'think') {
343    // Thought bubbles rising from the head into a cloud, "..." filling in.
344    art(X, 0, CLOUD)
345    dot(grid, 12, 3, 'R')
346    dot(grid, X - 1, 2, 'R')
347    const dots = Math.floor(frame / 2) % 4
348    for (let i = 0; i < dots; i++) dot(grid, X + 2 + i * 2, 2, 'k')
349  } else if (act === 'read') {
350    // An open book: two pages of text, the fold between them; the line
351    // being read lights up, down the left page, then the right.
352    art(X - 1, 3, BOOK)
353    const lines: [number, number, number][] = [[X, 4, 2], [X, 6, 2], [X + 4, 4, 2], [X + 4, 6, 1]]
354    const [lx, ly, len] = lines[frame % lines.length] as [number, number, number]
355    line(lx, ly, 'S'.repeat(len))
356  } else if (act === 'search') {
357    // A magnifying glass held out by its handle, from the hand to the rim,
358    // sweeping in and out, the lens glinting.
359    const dx = [0, 1, 2, 1][frame % 4] as number
360    art(X + 2 + dx, 1, LENS)
361    for (let x = X - 1; x < X + 2 + dx; x++) dot(grid, x, 5, 'H')
362    if (frame % 4 === 1) dot(grid, X + 5 + dx, 2, 'R')
363  } else if (act === 'write') {
364    // A sheet of paper, its corner turned; lines of writing appear as a
365    // pencil moves along them.
366    art(X + 1, 1, PAPER)
367    const done = frame % 12
368    for (let i = 0; i < done; i++) dot(grid, X + 2 + (i % 4), 3 + Math.floor(i / 4) * 2, 'k')
369    const px = X + 2 + (done % 4)
370    const py = 3 + Math.floor(done / 4) * 2
371    // The pencil: point, wood, yellow body, eraser, leaning up and right.
372    dot(grid, px, py, 'G')
373    dot(grid, px + 1, py - 1, 'H')
374    dot(grid, px + 2, py - 2, 'F')
375    dot(grid, px + 3, py - 3, 'f')
376  } else if (act === 'run') {
377    // A terminal window: its three buttons, output scrolling, a > prompt
378    // and a blinking cursor.
379    art(X, 1, TERMINAL)
380    const outputs = ['LLL.LLL', 'LL.LLL.', 'LLLLL.L', 'L.LLLL.', 'LLL.LL.']
381    for (let i = 0; i < 2; i++) line(X + 1, 3 + i, outputs[(frame + i) % outputs.length] as string)
382    dot(grid, X + 1, 5, 'S')
383    dot(grid, X + 2, 6, 'S')
384    dot(grid, X + 1, 7, 'S')
385    if (frame % 2 === 0) line(X + 4, 7, 'RR')
386  }
387}
388
389/** The canvas as terminal lines of cells, two pixel rows per line. */
390export function sceneRows(act: Act, frame = 0, blink = false, wind = frame, goggles = act === 'fly'): Cell[][] {
391  const pixels = scenePixels(act, frame, blink, wind, goggles)
392  const rows: Cell[][] = []
393  for (let y = 0; y < pixels.length; y += 2) {
394    const top = pixels[y] as string
395    const bottom = pixels[y + 1] ?? ''
396    const row: Cell[] = []
397    for (let x = 0; x < top.length; x++) {
398      const up = PALETTE[top[x] as string]
399      const down = PALETTE[bottom[x] ?? '.']
400      if (up && down) row.push({ ch: '▀', fg: up, bg: down })
401      else if (up) row.push({ ch: '▀', fg: up })
402      else if (down) row.push({ ch: '▄', fg: down })
403      else row.push({ ch: ' ' })
404    }
405    rows.push(row)
406  }
407  return rows
408}
409
hooks/crew.ts 790 lines
1/**
2 * jev-pilot — the crew: which workers a session can use and how the user wants
3 * them used. Pure: the hooks read options and the store, and act on this.
4 *
5 * Workers:
6 *   Claude       haiku, sonnet, opus (the plan's own models)
7 *   model slots  alpha, beta, gamma: any OpenRouter model, reached through the
8 *                local router (router/jev-router.mjs) as `jev-<slot>`
9 *   agents       Codex and OpenCode, their own CLIs, as reviewers
10 *
11 * Modes (the user's choice, `/jev mode <name>`); Jev decides within one:
12 *   standard        Claude only (the default)
13 *   budget          model slots take subagent work that needs no judgment
14 *   junior-lead     a junior on a slot writes easy, well-specified code;
15 *                   the main conversation reviews it as the tech lead
16 *   second-opinion  an external agent reviews significant changes
17 *   quality         Opus for every subagent, and an external review
18 */
19
20import { FEATURES } from './features.ts'
21
22export type Mode = 'standard' | 'budget' | 'junior-lead' | 'second-opinion' | 'quality'
23export const MODES: readonly Mode[] = ['standard', 'budget', 'junior-lead', 'second-opinion', 'quality']
24
25export const MODE_INFO: Record<Mode, string> = {
26  standard: 'Claude only: Jev picks Haiku, Sonnet or Opus and the effort for each task',
27  budget: 'custom models take subagent work that needs no judgment; Opus keeps the judgment',
28  'junior-lead': 'a junior on a custom model writes easy, well-specified code; Opus reviews it as tech lead',
29  'second-opinion': 'an external agent (Codex or OpenCode) reviews significant changes before they are done',
30  quality: 'Opus for every subagent, and an external review of significant changes',
31}
32
33export type Reviewer = 'codex' | 'opencode'
34export const REVIEWERS: readonly Reviewer[] = ['codex', 'opencode']
35
36export interface Slot {
37  /** Its name, as you type it after /jev: `flash` in `/jev flash deepseek/...`. */
38  name: string
39  /** The OpenRouter model id, e.g. deepseek/deepseek-v4.1-flash. */
40  model: string
41  /** When to choose it: what the decision model reads. */
42  when: string
43  /** What OpenRouter says it is, when known: "DeepSeek: DeepSeek V4.1 Flash · 1M context · $0.14 in · …". */
44  about?: string
45}
46
47/**
48 * A custom model's name, as you type it after `/jev`: lowercase letters,
49 * digits and "-", starting with a letter. Claude Code sees it as `jev-<name>`.
50 */
51export const SLOT_NAME = /^[a-z][a-z0-9-]{0,23}$/
52
53/** Words `/jev` already means something by: never a custom model's name. */
54export const RESERVED_NAMES: ReadonlySet<string> = new Set([
55  ...FEATURES,
56  'all', 'reset', 'status', 'mode', 'models', 'crew', 'junior', 'reviewer', 'tune',
57  'remove', 'delete', 'on', 'off', 'help', 'list', 'default', 'set', 'when',
58])
59
60/** A name `/jev <name> <model>` may use. */
61export function validName(name: string): boolean {
62  return SLOT_NAME.test(name) && !RESERVED_NAMES.has(name)
63}
64
65/** At most this many custom models at once: each is an option in Jev's question. */
66export const MAX_MODELS = 8
67
68/** How many entries a record (the store, models.json) is read for, removed ones included. */
69export const RECORD_LIMIT = 64
70
71/** The names the `alphaModel`/`betaModel`/`gammaModel` settings fill (before names were yours to choose). */
72export const LEGACY_NAMES = ['alpha', 'beta', 'gamma'] as const
73
74/**
75 * Read-only bulk work only: code is written on the main model (Anthropic's
76 * own guidance for Opus 5.5, and our junior-mode runs agreed: a cheaper
77 * model writing code cost more once the lead's review was counted).
78 */
79export const DEFAULT_SLOT_WHEN =
80  'Choose for read-only bulk work where cost matters more than precision: searching or reading across many files and reporting what is there, summarizing logs or test output, listing. Never for writing or changing code.'
81
82/** The model name Claude Code uses for a slot; the router maps it to the slot's model. */
83export function slotAlias(name: string): string {
84  return `jev-${name}`
85}
86
87/** What `/jev` has changed. */
88export interface CrewOverrides {
89  mode?: Mode
90  /** Each custom model by its name; '' once removed (so a setting can't bring it back). */
91  models?: Record<string, string>
92  junior?: string
93  reviewer?: Reviewer
94  /** The models set before, newest first, so switching back is one command. */
95  recent?: string[]
96  /** What OpenRouter said each model is, by id. */
97  about?: Record<string, string>
98  /** The model and effort each reviewer runs with, when not its CLI's own default. */
99  reviewerChoices?: Partial<Record<Reviewer, ReviewerChoice>>
100}
101
102/** A reviewer's model and effort; either may be left to the CLI's own config. */
103export interface ReviewerChoice {
104  model?: string
105  effort?: string
106}
107
108/** How many models `/jev <name>` remembers. */
109export const RECENT_MODELS = 5
110
111export interface Crew {
112  mode: Mode
113  /** The custom models set, in the order they were added. */
114  slots: Slot[]
115  /** The junior's model name ('' when there is none). */
116  junior: string
117  reviewer: Reviewer
118  reviewerChoices: Partial<Record<Reviewer, ReviewerChoice>>
119}
120
121/** The crew from the options and what `/jev` changed. */
122export function crewOf(options: Record<string, unknown>, overrides: CrewOverrides = {}): Crew {
123  const text = (key: string, fallback: string) => (typeof options[key] === 'string' ? (options[key] as string).trim() : fallback)
124  const models: Record<string, string> = { ...(overrides.models ?? {}) }
125  for (const name of LEGACY_NAMES) if (!(name in models)) models[name] = text(`${name}Model`, '')
126  const slots: Slot[] = []
127  for (const [name, model] of Object.entries(models)) {
128    // A record written elsewhere with more than the limit: the first ones count.
129    if (slots.length >= MAX_MODELS) break
130    if (!model || !validName(name)) continue
131    const about = overrides.about?.[model]
132    slots.push({ name, model, when: text(`${name}When`, '') || DEFAULT_SLOT_WHEN, ...(about ? { about } : {}) })
133  }
134  const modeOption = text('mode', 'standard') as Mode
135  const reviewerOption = text('reviewer', 'codex') as Reviewer
136  const juniorWanted = overrides.junior ?? text('junior', '')
137  return {
138    mode: overrides.mode ?? (MODES.includes(modeOption) ? modeOption : 'standard'),
139    slots,
140    // The one named, while it's set; else the first model.
141    junior: slots.some((slot) => slot.name === juniorWanted) ? juniorWanted : (slots[0]?.name ?? ''),
142    reviewer: overrides.reviewer ?? (REVIEWERS.includes(reviewerOption) ? reviewerOption : 'codex'),
143    reviewerChoices: { ...(overrides.reviewerChoices ?? {}) },
144  }
145}
146
147/** A reviewer model id as the CLIs name them (checked against their lists when set). */
148const REVIEWER_MODEL_ID = /^[\w.:~-]+(?:\/[\w.:~-]+)?$/
149/** A reasoning effort word (the CLIs' own names: low, medium, high, xhigh, max, ultra…). */
150const EFFORT_WORD = /^[a-z]{2,12}$/
151
152/** A reviewer choice read back from a record; anything that isn't one is dropped. */
153function reviewerChoicesOf(raw: unknown): Partial<Record<Reviewer, ReviewerChoice>> | undefined {
154  if (!raw || typeof raw !== 'object' || Array.isArray(raw)) return undefined
155  const out: Partial<Record<Reviewer, ReviewerChoice>> = {}
156  for (const reviewer of REVIEWERS) {
157    const entry = (raw as Record<string, unknown>)[reviewer]
158    if (!entry || typeof entry !== 'object') continue
159    const { model, effort } = entry as { model?: unknown; effort?: unknown }
160    const choice: ReviewerChoice = {}
161    if (typeof model === 'string' && model.length <= 100 && REVIEWER_MODEL_ID.test(model)) choice.model = model
162    if (typeof effort === 'string' && EFFORT_WORD.test(effort)) choice.effort = effort
163    if (choice.model || choice.effort) out[reviewer] = choice
164  }
165  return out
166}
167
168/** Overrides read back from the store or models.json; anything that isn't one is dropped. */
169export function overridesOf(stored: unknown): CrewOverrides {
170  if (!stored || typeof stored !== 'object' || Array.isArray(stored)) return {}
171  const raw = stored as Record<string, unknown>
172  const out: CrewOverrides = {}
173  if (typeof raw.mode === 'string' && MODES.includes(raw.mode as Mode)) out.mode = raw.mode as Mode
174  if (typeof raw.junior === 'string' && validName(raw.junior)) out.junior = raw.junior
175  if (typeof raw.reviewer === 'string' && REVIEWERS.includes(raw.reviewer as Reviewer)) out.reviewer = raw.reviewer as Reviewer
176  if (Array.isArray(raw.recent)) {
177    out.recent = raw.recent.filter((id): id is string => typeof id === 'string' && MODEL_ID.test(id)).slice(0, RECENT_MODELS)
178  }
179  if (raw.models && typeof raw.models === 'object' && !Array.isArray(raw.models)) {
180    const models: Record<string, string> = {}
181    for (const [name, value] of Object.entries(raw.models as Record<string, unknown>).slice(0, RECORD_LIMIT)) {
182      if (validName(name) && typeof value === 'string' && (value.trim() === '' || MODEL_ID.test(value.trim()))) models[name] = value.trim()
183    }
184    out.models = models
185  }
186  const choices = reviewerChoicesOf(raw.reviewerChoices)
187  if (choices) out.reviewerChoices = choices
188  if (raw.about && typeof raw.about === 'object' && !Array.isArray(raw.about)) {
189    const about: Record<string, string> = {}
190    for (const [id, value] of Object.entries(raw.about as Record<string, unknown>).slice(0, RECORD_LIMIT)) {
191      if (MODEL_ID.test(id) && typeof value === 'string') about[id] = value.slice(0, 200)
192    }
193    out.about = about
194  }
195  return out
196}
197
198/**
199 * ~/.claude/jev-pilot/models.json: the router's table, and the record of
200 * what you set: every model by its name (a removed one as ""), what
201 * OpenRouter said it is, and the models set before. Every session, in any
202 * project and whichever way jev-pilot is installed, reads the same choice back.
203 */
204export function routerTable(crew: Crew, overrides: Pick<CrewOverrides, 'models' | 'recent' | 'reviewerChoices'> = {}): string {
205  const slots: Record<string, { model: string; about?: string }> = {}
206  for (const [name, model] of Object.entries(overrides.models ?? {})) if (model === '' && validName(name)) slots[name] = { model: '' }
207  for (const slot of crew.slots) slots[slot.name] = { model: slot.model, ...(slot.about ? { about: slot.about } : {}) }
208  const reviewers = overrides.reviewerChoices ?? {}
209  return JSON.stringify({ slots, recent: (overrides.recent ?? []).slice(0, RECENT_MODELS), ...(Object.keys(reviewers).length > 0 ? { reviewers } : {}) }, null, 2)
210}
211
212/**
213 * The models recorded in models.json, as overrides; null when the file
214 * isn't one (missing, or not jev-pilot's).
215 */
216export function recordedModels(fileText: string | null): Pick<CrewOverrides, 'models' | 'recent' | 'about' | 'reviewerChoices'> | null {
217  if (!fileText) return null
218  let parsed: unknown
219  try {
220    parsed = JSON.parse(fileText)
221  } catch {
222    return null
223  }
224  const slots = (parsed as { slots?: unknown })?.slots
225  if (!slots || typeof slots !== 'object' || Array.isArray(slots)) return null
226  const models: Record<string, string> = {}
227  const about: Record<string, string> = {}
228  for (const [name, entry] of Object.entries(slots as Record<string, { model?: unknown; about?: unknown } | undefined>).slice(0, RECORD_LIMIT)) {
229    const model = entry?.model
230    if (!validName(name) || typeof model !== 'string' || !(model === '' || MODEL_ID.test(model))) continue
231    models[name] = model
232    if (model && typeof entry?.about === 'string') about[model] = entry.about.slice(0, 200)
233  }
234  const recent = overridesOf({ recent: (parsed as { recent?: unknown }).recent }).recent
235  const reviewerChoices = reviewerChoicesOf((parsed as { reviewers?: unknown }).reviewers) ?? {}
236  return { models, ...(recent ? { recent } : {}), ...(Object.keys(about).length > 0 ? { about } : {}), reviewerChoices }
237}
238
239/**
240 * The rows jev-pilot adds to Claude Code's `/model` list, one per custom
241 * model (the `modelPicker` setting, in a file `claude-jev` passes with
242 * `--settings`, so they're there only where the router is). `behavesAs`
243 * lets Claude Code, which doesn't know the model, treat it like Sonnet on
244 * its side (prompt, capabilities, effort); the requests still go to the
245 * model itself.
246 */
247export function pickerSettings(crew: Crew): string {
248  const options = crew.slots.map((slot) => {
249    const [name, ...details] = (slot.about ?? '').split(' · ')
250    return {
251      model: slotAlias(slot.name),
252      label: `${slot.name}${name ? ` · ${name.replace(/^[^:]+:\s*/, '')}` : ''}`,
253      description: [`${slot.model} on OpenRouter`, ...details.filter((d) => !/context$/.test(d))].join(' · ') + ' · claude-jev only',
254      behavesAs: 'sonnet',
255    }
256  })
257  return JSON.stringify({ modelPicker: { options } }, null, 2)
258}
259
260/**
261 * The slots Jev may choose for a subagent. Only with the router running, and
262 * only in the modes that hand work to custom models.
263 */
264export function slotsOffered(crew: Crew, routerOn: boolean): Slot[] {
265  if (!routerOn) return []
266  return crew.mode === 'budget' || crew.mode === 'junior-lead' ? crew.slots : []
267}
268
269/** The junior's slot, when the junior can work: junior-lead mode, the router up, the slot set. */
270export function juniorSlot(crew: Crew, routerOn: boolean): Slot | null {
271  if (!routerOn || crew.mode !== 'junior-lead') return null
272  return crew.slots.find((slot) => slot.name === crew.junior) ?? null
273}
274
275/** Whether significant changes get an external review. */
276export function reviews(crew: Crew): boolean {
277  return crew.mode === 'second-opinion' || crew.mode === 'quality'
278}
279
280export type CrewCommand =
281  | { kind: 'show' }
282  | { kind: 'mode'; mode: Mode }
283  /** `model` is the resolved id ('' removes the model); `about` what OpenRouter says it is. */
284  | { kind: 'model'; slot: string; model: string; about?: string }
285  /** `/jev <name> <what you pasted>`: resolved against OpenRouter's list before it's set. */
286  | { kind: 'paste'; slot: string; input: string }
287  /** `/jev <name>`: that model, and the ones set before. */
288  | { kind: 'slot'; slot: string }
289  | { kind: 'junior'; slot: string }
290  | { kind: 'reviewer'; reviewer: Reviewer }
291  /** `/jev reviewer codex luna high`: resolved against the CLI's own model list before it's set. */
292  | { kind: 'reviewer-paste'; reviewer: Reviewer; input: string; effort?: string }
293  /** A reviewer's choice as set (after resolving); an empty choice goes back to the CLI's config. */
294  | { kind: 'reviewer-choice'; reviewer: Reviewer; choice: ReviewerChoice }
295  | { kind: 'unknown'; text: string }
296
297/**
298 * `/jev` arguments about the crew, or null when they're about something else
299 * (a switch such as `/jev skills off`):
300 *   models                      the crew: mode, custom models, junior, reviewer
301 *   mode <name>                 standard · budget · junior-lead · second-opinion · quality
302 *   <name> <openrouter model>   add a custom model under a name you choose, pasted as its
303 *                               id, page link or name; the same name again replaces it
304 *   <name>                      that model, and the ones set before
305 *   remove <name>               delete it (also: <name> off)
306 *   junior <name>               which custom model the junior runs on
307 *   reviewer <codex|opencode>   which external agent reviews
308 */
309export function parseCrewCommand(args: string): CrewCommand | null {
310  const words = args.trim().split(/\s+/).filter(Boolean)
311  const head = words[0]?.toLowerCase()
312  if (!head) return null
313  if (head === 'models' || head === 'crew') return words.length === 1 ? { kind: 'show' } : { kind: 'unknown', text: args.trim() }
314  if (head === 'mode') {
315    const mode = words[1]?.toLowerCase() as Mode | undefined
316    return mode && MODES.includes(mode) && words.length === 2 ? { kind: 'mode', mode } : { kind: 'unknown', text: args.trim() }
317  }
318  if (head === 'remove' || head === 'delete') {
319    const name = words[1]?.toLowerCase()
320    return name && validName(name) && words.length === 2 ? { kind: 'model', slot: name, model: '' } : { kind: 'unknown', text: args.trim() }
321  }
322  if (head === 'junior') {
323    const name = words[1]?.toLowerCase()
324    return name && validName(name) && words.length === 2 ? { kind: 'junior', slot: name } : { kind: 'unknown', text: args.trim() }
325  }
326  if (head === 'reviewer') {
327    const reviewer = words[1]?.toLowerCase() as Reviewer | undefined
328    if (!reviewer || !REVIEWERS.includes(reviewer)) return { kind: 'unknown', text: args.trim() }
329    if (words.length === 2) return { kind: 'reviewer', reviewer }
330    const third = (words[2] as string).toLowerCase()
331    //   reviewer codex default          back to the CLI's own model and effort
332    //   reviewer codex effort high      the effort alone
333    //   reviewer codex luna [high]      a model, and an effort
334    if (third === 'default' && words.length === 3) return { kind: 'reviewer-choice', reviewer, choice: {} }
335    if (third === 'effort' && words.length === 4 && EFFORT_WORD.test((words[3] as string).toLowerCase())) {
336      return { kind: 'reviewer-paste', reviewer, input: '', effort: (words[3] as string).toLowerCase() }
337    }
338    if (words.length === 3 || (words.length === 4 && EFFORT_WORD.test((words[3] as string).toLowerCase()))) {
339      return { kind: 'reviewer-paste', reviewer, input: words[2] as string, ...(words[3] ? { effort: (words[3] as string).toLowerCase() } : {}) }
340    }
341    return { kind: 'unknown', text: args.trim() }
342  }
343  if (validName(head)) {
344    const input = args.trim().slice(head.length).trim()
345    if (!input) return { kind: 'slot', slot: head }
346    if (/^(off|remove|delete)$/i.test(input)) return { kind: 'model', slot: head, model: '' }
347    return { kind: 'paste', slot: head, input }
348  }
349  return null
350}
351
352/** Applies a command to the overrides. */
353export function applyCrewCommand(overrides: CrewOverrides, command: CrewCommand): CrewOverrides {
354  if (command.kind === 'mode') return { ...overrides, mode: command.mode }
355  if (command.kind === 'model') {
356    const recent = command.model ? [command.model, ...(overrides.recent ?? []).filter((id) => id !== command.model)].slice(0, RECENT_MODELS) : overrides.recent
357    const about = command.model && command.about ? { ...(overrides.about ?? {}), [command.model]: command.about } : overrides.about
358    const next: CrewOverrides = { ...overrides, models: { ...(overrides.models ?? {}), [command.slot]: command.model } }
359    if (recent) next.recent = recent
360    if (about) next.about = about
361    // A removed model can't stay the junior.
362    if (!command.model && next.junior === command.slot) delete next.junior
363    return next
364  }
365  if (command.kind === 'junior') return { ...overrides, junior: command.slot }
366  if (command.kind === 'reviewer') return { ...overrides, reviewer: command.reviewer }
367  if (command.kind === 'reviewer-choice') {
368    const choices = { ...(overrides.reviewerChoices ?? {}) }
369    if (command.choice.model || command.choice.effort) choices[command.reviewer] = { ...command.choice }
370    else delete choices[command.reviewer]
371    return { ...overrides, reviewerChoices: choices }
372  }
373  return overrides
374}
375
376/** `/jev models`: the crew as it stands. */
377export function describeCrew(crew: Crew, routerOn: boolean, codexList: readonly CodexModel[] = []): string {
378  const lines = [
379    `jev-pilot mode: ${crew.mode} (${MODE_INFO[crew.mode]})`,
380    `  /jev mode <${MODES.join('|')}>`,
381    `custom models (${routerOn ? 'router running' : 'router not running: start Claude Code with claude-jev'}):`,
382  ]
383  const width = Math.max(6, ...crew.slots.map((slot) => slot.name.length))
384  for (const slot of crew.slots) lines.push(`  ${slot.name.padEnd(width)} ${slot.model}${slot.name === crew.junior ? '   (the junior)' : ''}`)
385  if (crew.slots.length === 0) lines.push('  none yet')
386  lines.push(`  add: /jev <name> <model from openrouter.ai/models> · remove: /jev remove <name> · junior: /jev junior <name>`)
387  const said = (r: Reviewer) => describeChoice(crew.reviewerChoices[r], r === 'codex' ? codexList : [])
388  lines.push(`reviewer: ${crew.reviewer}   /jev reviewer <codex|opencode>`)
389  lines.push(`  codex    ${said('codex')}   /jev reviewer codex <model> [effort] · default`)
390  lines.push(`  opencode ${said('opencode')}   /jev reviewer opencode <provider/model> [effort] · default`)
391  return lines.join('\n')
392}
393
394// ---- the junior: a coder on a custom model, reviewed by the lead ----------------
395
396/** An agent type jev-pilot registers, as `$.agent.register` takes it. */
397export interface AgentSpec {
398  name: string
399  description: string
400  prompt: string
401  tools: string[]
402  model: string
403  maxTurns: number
404}
405
406export const JUNIOR_AGENT = 'jev-pilot:junior'
407
408/** The junior's agent definition, registered as `jev-pilot:junior`. */
409export function juniorSpec(model: string): AgentSpec {
410  return {
411    name: 'junior',
412    description:
413      'A junior developer on a cheaper model. Give it one easy, well-specified coding change: the files to change, the exact behavior, and the command that proves it (usually the tests). It implements and reports; review its diff yourself before calling the work done.',
414    prompt: [
415      'You are the junior developer on a small team. Your lead gives you one well-specified coding task.',
416      'Do exactly that: change only what the brief names, follow the existing code style, and run the command the brief gives to prove it (usually the tests).',
417      "Don't refactor, rename or add anything that wasn't asked for.",
418      "If the brief is unclear, or the task turns out bigger than it says, stop and say so instead of guessing.",
419      'When you are done, reply with: what you changed (file by file, one line each), the command you ran and its result, and anything you were unsure about.',
420    ].join('\n'),
421    tools: ['Read', 'Edit', 'Write', 'Grep', 'Glob', 'Bash'],
422    model,
423    maxTurns: 40,
424  }
425}
426
427// ---- the reviewers: Codex and OpenCode, their own CLI agents --------------------
428
429export function reviewerAgent(reviewer: Reviewer): string {
430  return `jev-pilot:${reviewer}-review`
431}
432
433const REVIEWER_NAME: Record<Reviewer, string> = { codex: 'Codex', opencode: 'OpenCode' }
434
435/** What the external agent is asked, ahead of the lead's brief. */
436export const REVIEW_INSTRUCTIONS = [
437  "Review a code change in this repository, read-only: don't edit anything.",
438  'Read the changed files and `git diff` (or `git diff HEAD~1` when the change is already committed).',
439  'Check that it does what the brief says, and look for wrong behavior, input that is not checked, broken edge cases and missing tests.',
440  'Report each finding as P1 (wrong behavior or lost data), P2 (a likely bug or a missing check) or P3 (minor), with file:line and a one-line fix.',
441  'End with one verdict: PASS, PASS-WITH-FOLLOWUP or NEEDS FIXES. Keep it tight.',
442].join('\n')
443
444/** A model Codex offers, as `codex debug models` lists it. */
445export interface CodexModel {
446  slug: string
447  name: string
448  description: string
449  efforts: string[]
450}
451
452/** Codex's model list from `codex debug models` (the ones it shows; hidden ones left out). */
453export function codexCatalog(json: string): CodexModel[] {
454  let parsed: unknown
455  try {
456    parsed = JSON.parse(json)
457  } catch {
458    return []
459  }
460  const list = Array.isArray(parsed) ? parsed : (parsed as { models?: unknown })?.models
461  if (!Array.isArray(list)) return []
462  return list
463    .filter((m): m is Record<string, unknown> => !!m && typeof m === 'object' && typeof (m as { slug?: unknown }).slug === 'string')
464    .filter((m) => m.visibility === undefined || m.visibility === 'list')
465    .map((m) => ({
466      slug: m.slug as string,
467      name: typeof m.display_name === 'string' ? m.display_name : (m.slug as string),
468      description: typeof m.description === 'string' ? m.description : '',
469      efforts: Array.isArray(m.supported_reasoning_levels)
470        ? (m.supported_reasoning_levels as { effort?: unknown }[]).map((level) => level?.effort).filter((e): e is string => typeof e === 'string')
471        : [],
472    }))
473}
474
475/** A model's short name, its tier, as people say it: `gpt-5.6-luna` → `luna`, `gpt-6-astra` → `astra`. */
476export function shortName(slug: string): string {
477  return slug.split('-').pop() ?? slug
478}
479
480/** The version numbers in a model id, to order a tier's models: `gpt-5.6-luna` → [5, 6]. */
481function versionOf(slug: string): number[] {
482  return (slug.match(/\d+/g) ?? []).map(Number)
483}
484
485function newer(a: number[], b: number[]): number {
486  for (let i = 0; i < Math.max(a.length, b.length); i++) {
487    const diff = (a[i] ?? 0) - (b[i] ?? 0)
488    if (diff !== 0) return diff
489  }
490  return 0
491}
492
493/** Codex's tiers (astra, sol, terra, luna…), each with its newest model now. */
494export function codexTiers(catalog: readonly CodexModel[]): Map<string, CodexModel> {
495  const tiers = new Map<string, CodexModel>()
496  for (const model of catalog) {
497    const tier = shortName(model.slug).toLowerCase()
498    if (!/^[a-z]+$/.test(tier)) continue
499    const held = tiers.get(tier)
500    if (!held || newer(versionOf(model.slug), versionOf(held.slug)) > 0) tiers.set(tier, model)
501  }
502  return tiers
503}
504
505/**
506 * The model a choice runs on now: a tier (`luna`) is its newest model in
507 * Codex's list at this moment, so a new generation is taken up by itself;
508 * an id (`gpt-5.6-luna`) stays that model. Unknown to the list: as given.
509 */
510export function resolvedChoice(reviewer: Reviewer, choice: ReviewerChoice = {}, catalog: readonly CodexModel[] = []): ReviewerChoice {
511  if (reviewer !== 'codex' || !choice.model || choice.model.includes('-')) return choice
512  const latest = codexTiers(catalog).get(choice.model.toLowerCase())
513  return latest ? { ...choice, model: latest.slug } : choice
514}
515
516export type ReviewerResolved = { ok: true; choice: ReviewerChoice; about: string } | { ok: false; why: string; suggestions: string[] }
517
518/**
519 * `/jev reviewer codex <model> [effort]`, checked against Codex's own list:
520 * the id (`gpt-5.6-luna`), its name (`GPT-5.6-Luna`) or its short name
521 * (`luna`); the effort must be one that model takes.
522 */
523export function resolveCodexChoice(input: string, effort: string | undefined, catalog: readonly CodexModel[], current: ReviewerChoice = {}): ReviewerResolved {
524  const wanted = input.trim().toLowerCase()
525  const tiers = codexTiers(catalog)
526  // A tier (`luna`) is kept as the tier: its newest model is used each time.
527  // An id or a full name (`gpt-5.6-luna`) pins that exact model.
528  let kept: string | undefined = current.model
529  let model: CodexModel | undefined
530  if (wanted) {
531    const tier = tiers.get(wanted)
532    const pinned = catalog.find((m) => m.slug.toLowerCase() === wanted) ?? catalog.find((m) => m.name.toLowerCase() === wanted)
533    model = tier ?? pinned
534    kept = tier ? wanted : pinned?.slug
535    if (!model || !kept) {
536      return {
537        ok: false,
538        why: catalog.length > 0 ? `Codex has no model "${input}"` : "Codex's model list couldn't be read (codex debug models)",
539        suggestions: [...tiers].map(([name, m]) => `${name} (now ${m.slug}): ${m.description}`),
540      }
541    }
542  } else if (kept) {
543    model = tiers.get(kept.toLowerCase()) ?? catalog.find((m) => m.slug === kept)
544  }
545  const efforts = model?.efforts ?? []
546  if (effort && efforts.length > 0 && !efforts.includes(effort)) {
547    return { ok: false, why: `${model?.slug ?? 'that model'} doesn't take effort "${effort}"`, suggestions: efforts }
548  }
549  const keptEffort = effort ?? (wanted ? undefined : current.effort)
550  const choice: ReviewerChoice = { ...(kept ? { model: kept } : {}), ...(keptEffort ? { effort: keptEffort } : {}) }
551  const about = model ? (wanted && tiers.get(wanted) ? `the newest ${wanted}, now ${model.slug}: ${model.description}` : `${model.name}: ${model.description}`) : ''
552  return { ok: true, choice, about }
553}
554
555/**
556 * `/jev reviewer opencode <model> [effort]`, checked against `opencode
557 * models` (provider/model): the full id, or a model name only one provider has.
558 */
559export function resolveOpencodeChoice(input: string, effort: string | undefined, models: readonly string[], current: ReviewerChoice = {}): ReviewerResolved {
560  const wanted = input.trim().toLowerCase()
561  let id: string | undefined = current.model
562  if (wanted) {
563    const exact = models.find((m) => m.toLowerCase() === wanted)
564    const byName = models.filter((m) => m.toLowerCase().endsWith(`/${wanted}`))
565    id = exact ?? (byName.length === 1 ? byName[0] : undefined)
566    if (!id) {
567      const close = byName.length > 1 ? byName : models.filter((m) => m.toLowerCase().includes(wanted))
568      return { ok: false, why: byName.length > 1 ? `more than one provider has "${input}"` : `OpenCode has no model "${input}"`, suggestions: close.slice(0, 5) }
569    }
570  }
571  return { ok: true, choice: { ...(id ? { model: id } : {}), ...((effort ?? (wanted ? undefined : current.effort)) ? { effort: effort ?? current.effort } : {}) }, about: '' }
572}
573
574/** "luna (newest, now gpt-5.6-luna) · effort high", or the CLI's own when nothing is chosen. */
575export function describeChoice(choice: ReviewerChoice | undefined, catalog: readonly CodexModel[] = []): string {
576  const model = choice?.model
577  const now = model && !model.includes('-') ? codexTiers(catalog).get(model.toLowerCase())?.slug : undefined
578  const parts = [model ? (now ? `${model} (newest, now ${now})` : model) : undefined, choice?.effort ? `effort ${choice.effort}` : undefined].filter(Boolean)
579  return parts.length > 0 ? parts.join(' · ') : "the CLI's own model and effort"
580}
581
582const shellQuote = (value: string) => `'${value.replace(/'/g, `'\\''`)}'`
583
584/** The model and effort flags for a reviewer's CLI, as `set --` arguments. */
585export function reviewerArgs(reviewer: Reviewer, choice: ReviewerChoice = {}): string {
586  const args: string[] = []
587  if (choice.model) args.push('-m', shellQuote(choice.model))
588  if (choice.effort) args.push(...(reviewer === 'codex' ? ['-c', shellQuote(`model_reasoning_effort="${choice.effort}"`)] : ['--variant', shellQuote(choice.effort)]))
589  return `set --${args.length > 0 ? ` ${args.join(' ')}` : ''}`
590}
591
592/** The command that runs the review, reading the brief from `$brief` and the flags from `set --`. */
593function reviewCommand(reviewer: Reviewer): string {
594  return reviewer === 'codex'
595    ? 'codex exec "$@" -s read-only --skip-git-repo-check -C "$PWD" -o "$brief.out" - < "$brief" > /dev/null 2> "$brief.err"; echo "exit $?"; cat "$brief.out" 2>/dev/null || tail -20 "$brief.err"'
596    : 'opencode run "$@" --agent plan --dir "$PWD" "$(cat "$brief")" < /dev/null 2> "$brief.err"; echo "exit $?"; [ -s "$brief.err" ] && tail -5 "$brief.err"'
597}
598
599/** How the reviewer is told which models it may switch to, when the brief asks for one. */
600function modelMenu(reviewer: Reviewer, codexModels: readonly CodexModel[]): string[] {
601  if (reviewer === 'codex') {
602    if (codexModels.length === 0) return ['If the brief asks for a Codex model or effort, pass it with -m <model id> and -c model_reasoning_effort="<effort>" in the set -- line.']
603    return [
604      'If the brief asks for a particular Codex model or effort for this review, change the set -- line to it (and only then). A model named by its tier means its newest model:',
605      ...[...codexTiers(codexModels)].map(([tier, m]) => `  ${tier} = -m '${m.slug}' (${m.description}; efforts: ${m.efforts.join(', ')})`),
606      `  an exact id the brief gives (e.g. ${codexModels[0]?.slug ?? 'gpt-…'}) = -m '<that id>'`,
607      `  effort: -c 'model_reasoning_effort="<effort>"'`,
608      "If it names a model or effort that isn't listed, don't guess: reply that it isn't available, with the list.",
609    ]
610  }
611  return [
612    "If the brief asks for a particular OpenCode model or effort, change the set -- line to -m '<provider/model>' and --variant '<effort>'. If you're unsure of the id, run `opencode models | grep -i <name>` first and use the exact line it prints; if it prints none or several, reply with them instead of guessing.",
613  ]
614}
615
616/**
617 * A reviewer's agent definition, registered as `jev-pilot:<reviewer>-review`:
618 * a small Claude model that hands the brief to the external CLI and brings
619 * its findings back, so the long review stays out of the main conversation.
620 */
621export function reviewerSpec(reviewer: Reviewer, model: string, choice: ReviewerChoice = {}, codexModels: readonly CodexModel[] = []): AgentSpec {
622  const name = REVIEWER_NAME[reviewer]
623  return {
624    name: `${reviewer}-review`,
625    description: `Gets a code review from ${name}, an external coding agent (its own CLI, not Claude), on ${describeChoice(choice)}. Brief it with what changed and why, the files, and what to check; to use another ${name} model or effort for this review, say which in the brief. It returns ${name}'s findings (P1/P2/P3) and verdict. Takes a few minutes: run it in the background when there is other work.`,
626    prompt: [
627      `You hand a code review to ${name}, an external coding agent, and bring back what it finds. You don't review the code yourself.`,
628      'Run one Bash command (timeout 600000), with the brief you were given pasted between the JEV_BRIEF lines exactly as given:',
629      '',
630      reviewerArgs(reviewer, choice),
631      'brief=$(mktemp /tmp/jev-review-XXXXXX); cat > "$brief" <<\'JEV_BRIEF\'',
632      REVIEW_INSTRUCTIONS,
633      '',
634      'The change:',
635      '<the brief>',
636      'JEV_BRIEF',
637      reviewCommand(reviewer),
638      '',
639      ...modelMenu(reviewer, codexModels),
640      '',
641      `Then reply with ${name}'s findings and verdict as it gave them, without adding your own, and say which model and effort it ran on.`,
642      `If the command fails or times out, reply with the exit code and the error lines instead, and say the review didn't run. Retry at most once, and only on a timeout.`,
643    ].join('\n'),
644    tools: ['Bash'],
645    model,
646    maxTurns: 6,
647  }
648}
649
650/**
651 * The crew's lines for the note to the main model: the mode the user chose,
652 * the junior, and the reviewers that are working.
653 */
654export function crewNote(crew: Crew, junior: Slot | null, working: Reviewer[], offered: readonly Slot[] = [], workflows = false): string[] {
655  const lines: string[] = []
656  if (crew.mode !== 'standard') lines.push(`The user chose the ${crew.mode} mode: ${MODE_INFO[crew.mode]}.`)
657  // Workflow agents never pass the Agent tool, so jev-pilot can't route them:
658  // the script sets each one's model (agent(prompt, { model })).
659  if (workflows) {
660    lines.push(
661      "Agents a Workflow script starts (agent()) don't pass through jev-pilot, so choose each one's model in the script with opts.model: 'haiku' for searching, reading and reporting, 'sonnet' for ordinary well-specified work, and leave it out for work that needs judgment.",
662    )
663    if (offered.length > 0) {
664      lines.push(
665        `In this mode, the user wants bulk work on their custom model: for a workflow agent searching, reading and reporting, or doing other work with nothing to judge, use ${offered.map((slot) => `opts.model: '${slotAlias(slot.name)}' (${slot.model})`).join(' or ')} rather than 'haiku'.`,
666      )
667    }
668  }
669  if (junior) {
670    lines.push(
671      `${JUNIOR_AGENT} is a junior developer on ${junior.model}. Give it easy, well-specified coding changes (the files, the exact behavior, the command that proves it); then read its diff. Rerun the tests only if its report doesn't show them passing or the diff goes beyond the brief. If it falls short, send your findings back once or fix small things yourself.`,
672    )
673  }
674  if (working.length > 0) {
675    lines.push(
676      `External reviewers (their own CLI agents, not Claude): ${working.map((r) => `${reviewerAgent(r)} (${REVIEWER_NAME[r]}, on ${describeChoice(crew.reviewerChoices[r])})`).join(', ')}. Use one when the user asks for a review by it. When the user names a model or effort for the review (for Codex: astra, sol, terra, luna…), put it in the brief as "Model: <name>, effort: <level>".`,
677    )
678  }
679  if (reviews(crew)) {
680    const chosen = working.includes(crew.reviewer) ? crew.reviewer : (working[0] ?? null)
681    lines.push(
682      chosen
683        ? `Before calling done a change that adds a feature or touches security, money or stored data, spawn ${reviewerAgent(chosen)} with a brief: what changed and why, the files, what to check. Fix the findings you agree with; mention briefly any you set aside.`
684        : `No external reviewer is working (${REVIEWER_NAME[crew.reviewer]} failed its check; /jev status shows why), so there is no external review: say so once when one would have been due.`,
685    )
686  }
687  return lines
688}
689
690// ---- setting a slot: what you paste from OpenRouter, checked against its list -----
691
692/** An OpenRouter model id: provider/model, optionally with a :variant or a ~ alias. */
693export const MODEL_ID = /^~?[\w.-]+\/[\w.:-]+$/
694
695/** Where to find a model to paste: OpenRouter's list, filtered to models that can call tools. */
696export const MODELS_PAGE = 'https://openrouter.ai/models?supported_parameters=tools'
697
698/** A model as OpenRouter's list (GET /api/v1/models) describes it. */
699export interface OpenRouterModel {
700  id: string
701  name?: string
702  canonical_slug?: string
703  context_length?: number
704  /** When OpenRouter added it, seconds since the epoch: the newest are suggested first. */
705  created?: number
706  pricing?: { prompt?: string; completion?: string }
707  supported_parameters?: string[]
708}
709
710export type Resolved =
711  | { ok: true; id: string; about: string }
712  | { ok: false; why: string; suggestions: string[] }
713
714const clean = (text: string) => text.trim().replace(/^[`'"<]+|[`'">]+$/g, '').trim()
715
716/** Per million tokens, as OpenRouter shows it: "$0.14 in · $0.42 out". */
717function price(model: OpenRouterModel): string {
718  const perMillion = (value: string | undefined) => {
719    const n = Number(value)
720    return Number.isFinite(n) ? `$${(n * 1e6).toFixed(n * 1e6 < 1 ? 3 : 2).replace(/0+$/, '').replace(/\.$/, '')}` : '?'
721  }
722  return `${perMillion(model.pricing?.prompt)} in · ${perMillion(model.pricing?.completion)} out per million tokens`
723}
724
725/** "DeepSeek: DeepSeek V4.1 Flash · 1M context · $0.14 in · $0.42 out per million tokens" */
726export function aboutModel(model: OpenRouterModel): string {
727  const context = model.context_length
728    ? ` · ${model.context_length >= 1e6 ? `${Math.round(model.context_length / 1e5) / 10}M` : `${Math.round(model.context_length / 1000)}k`} context`
729    : ''
730  return `${model.name ?? model.id}${context} · ${price(model)}`
731}
732
733/**
734 * What was pasted, as a model OpenRouter serves and a subagent can use.
735 * Accepted: the id (`deepseek/deepseek-v4.1-flash`), its page link
736 * (`https://openrouter.ai/deepseek/deepseek-v4.1-flash`), its dated slug, or
737 * its name as the list shows it (`DeepSeek: DeepSeek V4.1 Flash`, or without
738 * the `DeepSeek: ` prefix). Refused: anything not in the list (with the
739 * closest ids to try), and models that can't call tools: a subagent works
740 * through tools, so one without them could do nothing.
741 *
742 * With no list (OpenRouter unreachable), an id-shaped paste is taken as is;
743 * the slot's own check (a 1-token request) then says whether it answers.
744 */
745export function resolveModel(pasted: string, catalog: readonly OpenRouterModel[] | null): Resolved {
746  let text = clean(pasted)
747  const link = /^(?:https?:\/\/)?(?:www\.)?openrouter\.ai\/(?:models\/)?([^?#\s]+)/i.exec(text)
748  if (link) text = (link[1] as string).split('/').slice(0, 2).join('/')
749  if (!text) return { ok: false, why: 'nothing pasted', suggestions: [] }
750  if (!catalog) {
751    return MODEL_ID.test(text)
752      ? { ok: true, id: text, about: `${text} (OpenRouter's list couldn't be read to check it)` }
753      : { ok: false, why: `"${text}" isn't a model id (provider/model), and OpenRouter's list couldn't be read to look it up`, suggestions: [] }
754  }
755  const lower = text.toLowerCase()
756  const bare = (name: string | undefined) => (name ?? '').toLowerCase().replace(/^[^:]+:\s*/, '')
757  const found =
758    catalog.find((m) => m.id.toLowerCase() === lower) ??
759    catalog.find((m) => (m.canonical_slug ?? '').toLowerCase() === lower) ??
760    catalog.find((m) => (m.name ?? '').toLowerCase() === lower) ??
761    catalog.find((m) => bare(m.name) === lower.replace(/^[^:]+:\s*/, ''))
762  if (!found) {
763    const words = lower.split(/[^a-z0-9.]+/).filter((w) => w.length > 1)
764    const score = (m: OpenRouterModel) => words.filter((w) => `${m.id} ${m.name ?? ''}`.toLowerCase().includes(w)).length
765    const suggestions = catalog
766      .filter((m) => (m.supported_parameters ?? []).includes('tools') && score(m) > 0)
767      .sort((a, b) => score(b) - score(a) || (b.created ?? 0) - (a.created ?? 0) || a.id.length - b.id.length)
768      .slice(0, 3)
769      .map((m) => m.id)
770    return { ok: false, why: `OpenRouter has no model "${text}"`, suggestions }
771  }
772  if (!(found.supported_parameters ?? []).includes('tools')) {
773    return { ok: false, why: `${found.id} can't call tools on OpenRouter, so it can't work as a subagent`, suggestions: [] }
774  }
775  return { ok: true, id: found.id, about: aboutModel(found) }
776}
777
778/** `/jev alpha`: the slot, its model, and the models set before (to switch back). */
779export function describeSlot(crew: Crew, name: string, recent: readonly string[] = []): string {
780  const slot = crew.slots.find((s) => s.name === name)
781  const lines = [
782    slot ? `${name}: ${slot.model}${slot.about ? ` (${slot.about})` : ''}${name === crew.junior ? ', the junior' : ''}` : `${name}: no model by that name yet`,
783    `  ${slot ? 'replace it' : 'add it'}: /jev ${name} <model>, pasting its id, page link or name from ${MODELS_PAGE}`,
784  ]
785  if (slot) lines.push(`  use it as the main model: /model ${slotAlias(name)} (in claude-jev) · delete it: /jev remove ${name}`)
786  const others = recent.filter((id) => id !== slot?.model)
787  if (others.length > 0) lines.push('  set before:', ...others.map((id) => `    /jev ${name} ${id}`))
788  return lines.join('\n')
789}
790
hooks/crew-run.ts 286 lines
1/**
2 * jev-pilot — starting the crew for a session and checking every worker.
3 *
4 * The hooks hand in the engine calls this needs (`CrewIo`); `$` never leaves
5 * the hook. Checks are cheap: a 1-token call per custom model on OpenRouter,
6 * `codex login status`, `opencode --version` and `opencode auth list`.
7 */
8import { aboutModel, codexCatalog, juniorSlot, juniorSpec, overridesOf, pickerSettings, resolvedChoice, recordedModels, REVIEWERS, reviewerSpec, routerTable, slotAlias, type AgentSpec, type OpenRouterModel } from './crew.ts'
9import { codexModels, codexVerdict, crew, crewOverrides, crewStarted, setCodexModels, markCrewStarted, opencodeVerdict, reviewerHealthy, router, setCrewOverrides, setHealth, setRouter } from './crew-state.ts'
10
11export interface CrewIo {
12  fetch: (url: string, init: { method?: string; headers?: Record<string, string>; body?: string }) => Promise<{ ok: boolean; status: number; text: string }>
13  /** $HOME and $JEV_ROUTER_URL: the only variables this reads. */
14  home: () => Promise<string | undefined>
15  routerUrl: () => Promise<string | undefined>
16  write: (path: string, text: string) => Promise<void>
17  /** A file's text, or null when it isn't there. */
18  read: (path: string) => Promise<string | null>
19  run: (argv: string[], timeoutMs: number) => Promise<{ exitCode: number; stdout: string; stderr: string }>
20  storeGet: (key: string) => Promise<unknown>
21  sleep: (ms: number) => Promise<void>
22  /** Registers an agent type, `jev-pilot:<name>`. */
23  register: (spec: AgentSpec) => Promise<void>
24}
25
26export const CREW_KEY = 'crew'
27
28/** The router's address as shown to you: without the secret in its path. */
29export function shownUrl(url: string): string {
30  return url.replace(/\/[0-9a-f]{16,}$/i, '')
31}
32
33/** Where the router reads the slots from. */
34export async function modelsFile(io: CrewIo): Promise<string> {
35  const home = (await io.home()) ?? '~'
36  return `${home}/.claude/jev-pilot/models.json`
37}
38
39/** The `/model` rows `claude-jev` passes to Claude Code with `--settings`. */
40export async function pickerFile(io: CrewIo): Promise<string> {
41  const home = (await io.home()) ?? '~'
42  return `${home}/.claude/jev-pilot/picker.json`
43}
44
45/**
46 * Writes the router's table and the record of what's set (models.json), and
47 * the `/model` rows (picker.json), from the crew as it stands now.
48 */
49export async function publishSlots(io: CrewIo): Promise<void> {
50  await io.write(await modelsFile(io), routerTable(crew(), crewOverrides()))
51  await io.write(await pickerFile(io), pickerSettings(crew()))
52}
53
54async function within<T>(io: Pick<CrewIo, 'sleep'>, ms: number, work: Promise<T>): Promise<T | null> {
55  return Promise.race([work, io.sleep(ms).then(() => null)])
56}
57
58/** A session's crew: the saved `/jev` changes, the router, the slots file. */
59export async function startCrew(io: CrewIo): Promise<void> {
60  markCrewStarted()
61  // The mode, junior and reviewer from the store; the models from
62  // models.json, which every install and project shares, and which holds
63  // the latest change made anywhere.
64  const stored = overridesOf(await io.storeGet(CREW_KEY).catch(() => undefined))
65  const recorded = recordedModels(await io.read(await modelsFile(io)).catch(() => null))
66  setCrewOverrides(recorded ? { ...stored, ...recorded } : stored)
67  const url = (await io.routerUrl())?.replace(/\/$/, '') || null
68  setRouter(null)
69  if (url) {
70    const answer = await within(io, 1500, io.fetch(`${url}/jev-router/health`, { method: 'GET' }).catch(() => null))
71    const ok = !!answer && answer.ok
72    setHealth('router', { ok, detail: ok ? shownUrl(url) : `no answer at ${shownUrl(url)}`, at: Date.now() })
73    if (ok) setRouter(url)
74  }
75  await publishSlots(io).catch(() => undefined)
76}
77
78/** Checks one custom model: a 1-token request on OpenRouter. */
79export async function checkSlot(io: CrewIo, key: string | null, name: string, model: string): Promise<void> {
80  if (!key) {
81    setHealth(`slot:${name}`, { ok: false, detail: `${model} · no OpenRouter key`, at: Date.now() })
82    return
83  }
84  const answer = await within(
85    io,
86    8000,
87    io
88      .fetch('https://openrouter.ai/api/v1/messages', {
89        method: 'POST',
90        headers: { 'content-type': 'application/json', authorization: `Bearer ${key}`, 'anthropic-version': '2023-06-01' },
91        body: JSON.stringify({ model, max_tokens: 1, messages: [{ role: 'user', content: 'ok' }] }),
92      })
93      .catch(() => null),
94  )
95  const ok = !!answer && answer.ok
96  const why = answer ? `HTTP ${answer.status}: ${answer.text.slice(0, 80)}` : 'no answer in 8s'
97  setHealth(`slot:${name}`, { ok, detail: ok ? `${model} · answering` : `${model} · ${why}`, at: Date.now() })
98}
99
100/** Checks the external agents: installed, and logged in. */
101export async function checkAgents(io: CrewIo): Promise<void> {
102  try {
103    const codex = await io.run(['codex', 'login', 'status'], 15_000)
104    const verdict = codexVerdict(codex.exitCode, `${codex.stdout}\n${codex.stderr}`)
105    setHealth('agent:codex', { ...verdict, at: Date.now() })
106    // Its models now: what a tier such as `luna` resolves to this session.
107    if (verdict.ok) await refreshCodexModels(io)
108  } catch {
109    setHealth('agent:codex', { ok: false, detail: 'codex not installed', at: Date.now() })
110  }
111  try {
112    const version = await io.run(['opencode', '--version'], 15_000)
113    const auth = version.exitCode === 0 ? await io.run(['opencode', 'auth', 'list'], 15_000) : { stdout: '', stderr: '', exitCode: 1 }
114    setHealth('agent:opencode', { ...opencodeVerdict(version.exitCode, version.stdout, `${auth.stdout}\n${auth.stderr}`), at: Date.now() })
115  } catch {
116    setHealth('agent:opencode', { ok: false, detail: 'opencode not installed', at: Date.now() })
117  }
118}
119
120/** Reads Codex's model list (`codex debug models`); kept as it was when it can't be read. */
121export async function refreshCodexModels(io: CrewIo): Promise<void> {
122  const listed = await io.run(['codex', 'debug', 'models'], 15_000).catch(() => null)
123  const models = listed && listed.exitCode === 0 ? codexCatalog(listed.stdout) : []
124  if (models.length > 0) setCodexModels(models)
125}
126
127/** OpenCode's models (`opencode models`), one provider/model id per line; [] when it can't be read. */
128export async function opencodeModels(io: CrewIo): Promise<string[]> {
129  const listed = await io.run(['opencode', 'models'], 30_000).catch(() => null)
130  if (!listed || listed.exitCode !== 0) return []
131  return listed.stdout
132    .split('\n')
133    .map((line) => line.replace(/\x1b\[[0-9;]*m/g, '').trim())
134    .filter((line) => /^[\w.:~-]+\/[\w.:~\/-]+$/.test(line))
135}
136
137/** The router still answering, and any custom model it had to hand to Claude. */
138export async function checkRouter(io: CrewIo): Promise<void> {
139  const url = router()
140  if (!url) return
141  const answer = await within(io, 1500, io.fetch(`${url}/jev-router/health`, { method: 'GET' }).catch(() => null))
142  if (!answer || !answer.ok) {
143    setHealth('router', { ok: false, detail: `no answer at ${shownUrl(url)}`, at: Date.now() })
144    return
145  }
146  setHealth('router', { ok: true, detail: `${shownUrl(url)}${fallbackNote(answer.text)}`, at: Date.now() })
147}
148
149/** "· alpha fell back to claude-sonnet-5 2× (OpenRouter 503)", from the router's health answer. */
150export function fallbackNote(healthText: string): string {
151  try {
152    const fallbacks = (JSON.parse(healthText) as { fallbacks?: Record<string, { count?: number; to?: string; why?: string }> }).fallbacks ?? {}
153    const notes = Object.entries(fallbacks)
154      .filter(([, f]) => typeof f?.count === 'number' && f.count > 0)
155      .map(([slot, f]) => `${slot} fell back to ${f.to ?? 'Claude'} ${f.count}× (last: ${String(f.why ?? '').slice(0, 60)})`)
156    return notes.length > 0 ? ` · ${notes.join('; ')}` : ''
157  } catch {
158    return ''
159  }
160}
161
162/** Every check, in parallel; then what OpenRouter says each model is, where that's missing. */
163export async function checkCrew(io: CrewIo, key: string | null): Promise<void> {
164  await Promise.all([...crew().slots.map((slot) => checkSlot(io, key, slot.name, slot.model)), checkAgents(io), checkRouter(io)])
165  await fillAbout(io).catch(() => undefined)
166}
167
168/**
169 * A model set without its description (before descriptions were kept, or
170 * while OpenRouter's list was out of reach): looked up once it's there, for
171 * `/model` and `/jev status`.
172 */
173export async function fillAbout(io: CrewIo): Promise<void> {
174  const missing = crew().slots.filter((slot) => !slot.about)
175  if (missing.length === 0) return
176  const catalog = await modelCatalog(io)
177  if (!catalog) return
178  const about = { ...(crewOverrides().about ?? {}) }
179  for (const slot of missing) {
180    const found = catalog.find((model) => model.id === slot.model)
181    if (found) about[slot.model] = aboutModel(found)
182  }
183  if (Object.keys(about).length === Object.keys(crewOverrides().about ?? {}).length) return
184  setCrewOverrides({ ...crewOverrides(), about })
185  await publishSlots(io)
186}
187
188/** The small Claude model a reviewer runs on: it only relays. */
189export const REVIEWER_MODEL = 'haiku'
190
191/**
192 * The crew's agent types: the junior in junior-lead mode (with the router
193 * up), and each reviewer whose CLI passed its check.
194 */
195export async function registerCrew(io: CrewIo): Promise<void> {
196  const junior = juniorSlot(crew(), router() !== null)
197  const specs = [
198    ...(junior ? [juniorSpec(slotAlias(junior.name))] : []),
199    // Each reviewer on its chosen model, a tier resolved to its newest now.
200    ...REVIEWERS.filter(reviewerHealthy).map((r) => reviewerSpec(r, REVIEWER_MODEL, resolvedChoice(r, crew().reviewerChoices[r], codexModels()), codexModels())),
201  ]
202  for (const spec of specs) await io.register(spec).catch(() => undefined)
203}
204
205/** A session's crew, start to finish: set up, agents registered, then checked and registered again. */
206export async function startSession(io: CrewIo, key: string | null): Promise<void> {
207  await startCrew(io).catch(() => undefined)
208  await registerCrew(io)
209  void checkCrew(io, key)
210    .then(() => registerCrew(io))
211    .catch(() => undefined)
212}
213
214/**
215 * The crew, set up if this worker hasn't yet: a reload mid-session (a plugin
216 * update) starts a fresh worker without a session start, and routing must
217 * not think the router is gone. Checks run in the background.
218 */
219export async function ensureCrew(io: CrewIo, key: string | null): Promise<void> {
220  if (!crewStarted()) await startSession(io, key)
221  else await refreshModels(io, key)
222}
223
224/**
225 * The models set in another session since this one started (models.json
226 * changed): taken up here too, and a newly set model checked. One small
227 * file read per prompt.
228 */
229export async function refreshModels(io: CrewIo, key: string | null): Promise<void> {
230  const recorded = recordedModels(await io.read(await modelsFile(io)).catch(() => null))
231  if (!recorded) return
232  const current = crewOverrides()
233  const before = crew().slots
234  const same = (a: unknown, b: unknown) => JSON.stringify(a ?? null) === JSON.stringify(b ?? null)
235  if (same(recorded.models, current.models) && same(recorded.recent, current.recent) && same(recorded.about, current.about) && same(recorded.reviewerChoices, current.reviewerChoices)) return
236  setCrewOverrides({ ...current, ...recorded })
237  for (const slot of crew().slots) {
238    if (!before.some((b) => b.name === slot.name && b.model === slot.model)) void checkSlot(io, key, slot.name, slot.model).catch(() => undefined)
239  }
240  await registerCrew(io)
241}
242
243let catalogCache: { at: number; models: OpenRouterModel[] } | null = null
244
245/**
246 * OpenRouter's model list, to check what `/jev <slot>` was given against
247 * (public, no key). Kept 10 minutes; null when it can't be read in 6 s.
248 */
249export async function modelCatalog(io: CrewIo): Promise<OpenRouterModel[] | null> {
250  if (catalogCache && Date.now() - catalogCache.at < 600_000) return catalogCache.models
251  const answer = await within(io, 6000, io.fetch('https://openrouter.ai/api/v1/models', { method: 'GET' }).catch(() => null))
252  if (!answer || !answer.ok) return null
253  try {
254    const models = (JSON.parse(answer.text) as { data?: unknown }).data
255    if (!Array.isArray(models)) return null
256    catalogCache = { at: Date.now(), models: models.filter((m): m is OpenRouterModel => !!m && typeof (m as OpenRouterModel).id === 'string') }
257    return catalogCache.models
258  } catch {
259    return null
260  }
261}
262
263let fallbacksSeen: number | null = null
264
265/**
266 * New fallbacks since the last look: each custom model that failed and was
267 * answered by Claude instead, with why. Empty on the first look (it only
268 * counts from there) and when the router isn't answering.
269 */
270export async function newFallbacks(io: Pick<CrewIo, 'fetch' | 'sleep'>): Promise<string[]> {
271  const url = router()
272  if (!url) return []
273  const answer = await within(io, 800, io.fetch(`${url}/jev-router/health`, { method: 'GET' }).catch(() => null))
274  if (!answer || !answer.ok) return []
275  try {
276    const fallbacks = (JSON.parse(answer.text) as { fallbacks?: Record<string, { count?: number; to?: string; why?: string }> }).fallbacks ?? {}
277    const total = Object.values(fallbacks).reduce((sum, f) => sum + (typeof f?.count === 'number' ? f.count : 0), 0)
278    const before = fallbacksSeen
279    fallbacksSeen = total
280    if (before === null || total <= before) return []
281    return Object.entries(fallbacks).map(([name, f]) => `${name} failed (${String(f.why ?? '').slice(0, 90)}), so ${f.to ?? 'Claude'} answered instead`)
282  } catch {
283    return []
284  }
285}
286
hooks/jev-call.ts 92 lines
1/**
2 * jev-pilot — one request to Jev per prompt.
3 *
4 * Two modules ask Jev about the same prompt: the router (effort, model,
5 * strategy) and the skill picker (which skill, if any). They used to ask one
6 * after the other, three requests in a row (the router's, the skill ranking,
7 * and a re-read of the shortlist): about 1.5 s before each turn. Now the
8 * router hands its questions down, and the skill module, which runs inside
9 * it on the same prompt, sends them with its own in one request. One request
10 * with both sets of questions takes as long as the slower of the two alone,
11 * and answers them the same (measured: same effort, tier and strategy on
12 * every prompt tried).
13 *
14 * The hand-off: the router's prompt.submit offers its part, then passes the
15 * prompt on; the skill module's takes it, asks (`ask`, with its own
16 * questions added), and gives the answer back (`settle`), which returns the
17 * router's block for the prompt (strategy advice). A part nobody took is
18 * asked for by the router itself once the prompt comes back up.
19 */
20import type { Miss } from './model-router.policy.ts'
21
22export interface Answer {
23  /** The response body, or null with why (`miss`). */
24  text: string | null
25  miss: Miss | null
26  ms: number
27}
28
29export interface RouterPart {
30  /** The prompt this part is for: a part is only ever taken for its own prompt. */
31  prompt: string
32  /**
33   * One request with the router's questions and `extra` (the skill
34   * module's), in the router's state (the prompt, recent context, signals).
35   */
36  ask: (extra: Record<string, unknown>) => Promise<Answer>
37  /** Hands the answer to the router; returns its block for the prompt, if any. */
38  settle: (answer: Answer) => Promise<string | null>
39}
40
41/**
42 * The parts waiting, by prompt text, oldest first. The engine gives a
43 * prompt no id of its own, so its text is the key; two prompts with the same
44 * text (a quick double submit) can only swap parts built for identical text,
45 * and a take for one prompt never touches another's part.
46 */
47const offered = new Map<string, RouterPart[]>()
48const MAX_WAITING = 16
49
50/** The router's part for this prompt, waiting for the skill module. */
51export function offerPart(part: RouterPart): void {
52  const queue = offered.get(part.prompt) ?? []
53  queue.push(part)
54  offered.set(part.prompt, queue)
55  // Bounded: parts nobody takes (a module that never ran) don't pile up.
56  while ([...offered.values()].reduce((n, q) => n + q.length, 0) > MAX_WAITING) {
57    const [oldest] = offered.keys()
58    if (oldest === undefined) break
59    const q = offered.get(oldest) as RouterPart[]
60    q.shift()
61    if (q.length === 0) offered.delete(oldest)
62  }
63}
64
65/** Takes the oldest part waiting for this prompt, once: null when there is none for it. */
66export function takePart(prompt: string): RouterPart | null {
67  const queue = offered.get(prompt)
68  const part = queue?.shift() ?? null
69  if (queue && queue.length === 0) offered.delete(prompt)
70  return part
71}
72
73/** Whether this part is still waiting (nobody took it); it's withdrawn either way. */
74export function stillOffered(part: RouterPart): boolean {
75  const queue = offered.get(part.prompt)
76  const at = queue ? queue.indexOf(part) : -1
77  if (!queue || at < 0) return false
78  queue.splice(at, 1)
79  if (queue.length === 0) offered.delete(part.prompt)
80  return true
81}
82
83/**
84 * A prompt that only says to go on with the work in progress ("continue",
85 * "keep going"): the turn before already decided how to do that work, so
86 * Jev isn't asked again. An approval ("yes", "go ahead", "do it") is not one:
87 * it often starts the very work the turn before only proposed.
88 */
89export function isContinuation(prompt: string): boolean {
90  return /^(?:please\s+)?(?:continue|go on|keep going|carry on|resume)(?:\s+please)?[.!]*$/i.test(prompt.trim())
91}
92