SLOPSHOPPER

jev-skill-scout

Before each prompt reaches the model, TypeSafe's Jev ranks your installed skills against it and attaches one line naming the skill to load. Sends the prompt…

newstatuspromptnetwork
v0.1.0MITupdated 2026-09-23karanb192/jev-skill-scout
A shopper browsing a rack in a slop shop
README

jev-skill-scout

Claude Code picks skills on its own, from a list of one-line descriptions that sits in its context next to everything else. Sometimes it does not pick. You find out later: the new screen ignores the design system you wrote a skill for, the endpoint ships with no tests even though your testing skill asks for them, the commit message skips the format you set. Every one of those skills was installed the whole time.

On my own transcripts, 221 sessions over nine weeks, 73% of the turns that needed a skill loaded none. When I checked a sample by hand, about half of those held up. Numbers and method below.

This repo does two things about that.

  1. The audit. npx jev-skill-scout audit replays every prompt in your Claude Code transcripts through TypeSafe's Jev and counts the turns where a skill should have loaded and did not. Each prompt is judged against the skill list its own session showed the model, which the transcript records. One command, one key, one HTML report you can label.
  2. The mod. A Claude Code function-hook plugin that runs the same judgment live, before each prompt reaches the model, and attaches one line: Relevant to this request: frontend-design. The model still decides. Your skill list does not change, so prompt caching over it still holds.

Both use the same code in lib/. The audit is the mod's brain run offline, so its numbers are what the mod would have done on your history. Once the mod is on, the audit also reads its trace in later transcripts and reports whether the agent followed each suggestion, and the miss rate with the mod against without.

What the audit found on my transcripts

<!-- audit:start --> 221 sessions, 3,431 human prompts, 3,083 judged (348 were under 12 characters). 3,411 of the prompts were judged against the skill list their own session showed the model (56 to 77 skills, depending on the day). 5,817 Jev calls, 21.4M input tokens, $0.90, 3 minutes 22 seconds at 12 requests in parallel from India.

Fit thresholdTurns where Jev saw a skill needLoaded nothingMiss rate
0.3 (default, the cookbook's)1,5511,12972.8%
0.51,02170268.8%
0.739424762.7%

The rate moves a little with the threshold; the count moves a lot. Most missed at fit 0.5: a past-session search skill 86, Claude Code's own code-review 46, an open-source contribution checklist 44, the bundled update-config 42, my report-writing skill 38, browser automation 37. Two of the top six ship with Claude Code itself.

I then read 30 random misses at fit 0.5 and labelled each one myself: 15 right, 15 wrong, so about 50% precision (an earlier sample of 37 on a disk-only roster came out at 57%). Take the 702 down to roughly 350 real misses across nine weeks. Right: "do you remember that I applied to [a company], any details?" (past-session search), "can you check why CI was failing on that PR?" (the contribution checklist), "is the PR all solid to merge, are we sure?" (code-review). Wrong: "allowed the key.." went to update-config, and "which one is your recommendation? top 3" went to a design skill because its description promises options. Jev reads descriptions literally, so the wrong half is mostly descriptions that overclaim; the doctor below is for those.

The other direction exists too. At fit 0.3 the agent loaded a skill Jev did not pick 43 times, and only 60 turns were a clean hit. Jev is a second opinion, not an oracle. <!-- audit:end -->

Jev is the judge here, not ground truth. Every row in the report shows the pick, its fit probability and what the turn actually loaded, so you can tick right or wrong on a sample and the page turns your ticks into a precision number.

Run the audit

export TYPESAFE_API_KEY=...        # https://console.typesafe.ai/settings/keys
npx jev-skill-scout audit          # counts prompts, shows the cost, asks before spending

What comes back, from my run:

jev-skill-scout audit: 3431 prompts, 56 skills in the roster

    1129  miss               Jev picked a skill; the turn loaded none, and it was not already loaded
      60  hit                Jev picked the skill the turn loaded
     302  already-loaded     Jev picked a skill that an earlier turn had loaded
      60  disagree           Jev picked one skill; the turn loaded a different one
      43  unsuggested-load   The turn loaded a skill; Jev picked none
    1489  quiet              Neither picked a skill
     348  trivial            Too short to judge; skipped without a call
       0  error              The request failed

  Turns where Jev saw a skill need: 1551. Missed by the agent: 1129 (72.8%).
  Most missed skills:
     165  (your skills, by name)
     ...

  3411 of 3431 prompts were judged against the skill list their own session showed the model; the rest against what is installed now.
  5817 Jev calls, 21,407,337 input tokens, about $0.899, 1392 ms per judged prompt on average.

  Report: ./skill-audit/report.html

Useful flags:

--dry-run          count prompts and estimate cost, no requests
--days 30          only sessions touched in the last 30 days
--limit 200        stop after 200 judged prompts
--project name     only projects whose folder contains this text
--out dir          where report.html, cases.json and cache.json go (default ./skill-audit)
--gate 0.3         gate threshold; --fit 0.3 the winner's fit threshold
--yes              skip the confirmation

It reads ~/.claude/projects/*/*.jsonl and every SKILL.md under ~/.claude/skills, ~/.claude/plugins/cache and ./.claude/skills. Nothing is written outside the output directory. Judgments are cached, so a second run with new thresholds is free.

Cost: about $0.0003 per judged prompt at the listed Jev price. My 3,083 prompts cost $0.90 and took under four minutes at 12 in parallel. The median prompt is 850 ms for both calls from India over one kept-alive HTTP/1.1 connection. (An earlier run averaged 7 seconds and wedged twice at 200 prompts: Node's fetch negotiates HTTP/2 with the API and the session spins the event loop, so the CLI now uses node:https directly.)

What counts as a miss

Every human prompt is one turn. A turn is a miss only when all three hold:

  • Jev picked a skill after both stages (ranking the whole roster, then re-reading the top three skills' actual instructions and being allowed to reject all of them).
  • The turn loaded no skill: no Skill tool call, no Launching skill result, no slash command typed.
  • That skill was not already loaded earlier in the same session. Skills stay in context once loaded, so suggesting one again would be noise.

The roster for each prompt is the one Claude Code listed to the model in that session (transcripts carry a skill_listing record at session start and whenever it changes), so a skill you installed last week is not held against prompts from last month. Sessions with no such record fall back to what is installed now. The verify stage reads today's SKILL.md body for skills that still exist.

The other buckets are reported too: hit (Jev and the turn agree), already-loaded, disagree (Jev picked one, the turn loaded another), unsuggested-load (the turn loaded a skill Jev did not pick), quiet (neither), and trivial (prompts under 12 characters, skipped without a call). Subagent transcripts and harness notifications are excluded.

Install the mod

Function hooks are early access. Nothing loads unless the flag is set, and the API can change between Claude Code releases.

claude plugin marketplace add karanb192/jev-skill-scout
claude plugin install jev-skill-scout@jev-skill-scout

Then in ~/.claude/settings.json:

{ "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1", "TYPESAFE_API_KEY": "..." } }

Or for one session from a checkout: CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir /path/to/jev-skill-scout.

Every option (key, thresholds, timeout, model, quiet, shadow, on/off) is a plugin setting under /config, so there is no extra command to learn. With no key the mod loads and does nothing.

Shadow mode judges every prompt and shows the pick in the status line but attaches nothing, so you can watch what it would do before letting it. Turn it on under /config, or for one session with JEV_SKILL_SCOUT_SHADOW=1.

Per prompt it adds one status line and, from India, about 2 to 3 seconds before the model starts (two round trips to a West Coast API). From the US it is well under a second. Set timeoutMs lower if that bothers you; on timeout the turn runs untouched.

What it can reach

Validated on Claude Code 2.1.278:

❯ ./register.ts hooks: session.start, prompt.submit ❯ ./register.ts calls: $.clock.now, $.clock.sleep, $.env.get, $.fs.exists, $.fs.list, $.fs.read, $.http.fetch, $.session.cwd, $.store.get, $.store.set, $.ui.log, $.ui.status ❯ ./register.ts env reads: HOME, JEV_SKILL_SCOUT_SHADOW, TYPESAFE_API_KEY, TYPESAFE_KEY

Reach L3, network. Sees every prompt you type.

  • Reads: SKILL.md files under your home and project skill directories and the enabled plugins' caches, ~/.claude/settings.json for which plugins are enabled, and four environment variables.
  • Runs: nothing. No shell.
  • Sends: your prompt text, the skill names and descriptions, and on the second call the first 700 characters of three skills' instructions, to api.typesafe.ai. Nothing else leaves the machine.
  • Persists: the skill roster in the plugin's own store, refreshed every ten minutes.
  • Hostile input: a prompt that tries to steer Jev can at most cause a wrong or missing suggestion line, which the model is told to ignore if it does not fit. The mod never loads a skill itself and never blocks a prompt; on any failure it enters the prompt untouched, exactly once.

How the judgment works

It follows TypeSafe's skill suggestion cookbook, which measured wrong-skill loads on Haiku dropping from 16.8% to 7.3% with one suggestion line.

  1. Rank. One request with a Choice over every skill by its description plus a none option, and three Nouls that gate the turn: does it act on the user's files or accounts, would a written procedure help, would prose suffice. Below the gate, nothing is suggested.
  2. Verify. The top three skills are re-read with the opening of their instructions. A second Choice picks among them or rejects all; a Noul per candidate asks whether loading it would change the work. The winner needs its fit above the threshold.
  3. Attach. One <skill_relevance> block after the prompt, invisible to you, telling the model which skill to load first and to ignore the hint if it does not fit.

Jev returns typed answers with probabilities in one parallel pass, so a 58-skill roster is one request, not 58. A Choice takes at most 255 options; a roster past 250 is ranked in parallel chunks, each chunk keeps its top three, and the verify stage settles it with real excerpts.

Fix the description, not the symptom

Most misses trace back to a description that does not say when the skill applies. npx jev-skill-scout doctor <skill> pulls the real prompts from your last audit that involve that skill, in three groups: the ones that loaded it, the ones where Jev picked it and nothing loaded, and the ones where the turn chose a different skill. It scores the current description against all of them in one request, and any rewrite you pass with --desc "..." or --desc-file beside it:

jev-skill-scout doctor: frontend-design
  9 prompts loaded it (9 scored), 12 where Jev picked it and nothing loaded (12 scored), 1 where the turn loaded another skill (1 scored).

  description     loaded it  Jev missed  suspect   what you want
                       high        high      low
  current               54%         72%      74%   2,323 tokens, 1064 ms
  rewrite               61%         68%      71%   2,333 tokens, 399 ms

It then lists the prompts the current description matches least among those that really used the skill, and the ones it still matches among those that used another. Anthropic's /skill-doctor tests a description against prompts it invents; this tests it against yours, in about a second, for a fraction of a cent.

Does the agent obey it?

The line the mod attaches is recorded in the transcript, so the audit can see it. For every turn where the mod spoke, the report shows the suggestion and whether the agent loaded that skill, and it splits the miss rate into sessions where the mod was active and sessions where it was not. Run the mod for a few days, run the audit again, and that paragraph fills in with your own before and after. Nothing else in this space measures that on real sessions; TypeSafe's cookbook number below is from a synthetic set on Haiku.

Why one line works when the list does not

Claude Code already puts every skill's name and description in context. Three things differ:

  • Menu vs verdict. The default is a 58-item menu Claude has to match against your prompt on the side, while it plans the answer. The mod hands it a decision: load this one. Following an instruction is a far easier task for a model than noticing a match.
  • Where it sits. The list lives in the static prefix, tens of thousands of tokens above your prompt. The line is attached to the prompt itself, the last thing Claude reads before it starts.
  • Who decided. The list is judged on descriptions alone. The pick here was made after re-reading the skills' actual instructions, and Claude is told to drop it if it does not fit.

Limits

  • Claude Code's bundled skills (code-review, deep-research, simplify and the rest) appear in the transcript listing but not on disk, so the audit can judge them and the mod cannot suggest them.
  • Skills that were compacted out of context still count as already loaded.
  • Jev 1.13 reads literally; a skill with a vague description gets ranked on that vague description. The audit's disagree and unsuggested-load rows are where to look for descriptions worth rewriting.
  • Precision is yours to measure. Label a sample in the report before quoting the miss rate anywhere.
  • The obedience numbers only exist once you have run the mod for a while; a fresh audit reports Jev's opinion of your history, not what Claude did with a suggestion.

Related

  • typesafe-mod ranks skills on prompt.submit too, in one request, with the router off by default and a shell scan for the roster.
  • skillranker is the most complete live picker: a Rust CLI wired in as a classic UserPromptSubmit shell hook, with the same two-stage Jev judgment, local feedback records, a TUI, and Cursor and Pi support. Use it if you want a picker across harnesses. This repo is the same judgment as a mod (no process spawn, footprint printed by the validator) plus the audit, which skillranker does not have.
  • skill-router picks skills from the shell at session start.
  • awesome-claude-code-mods scans every mod on GitHub nightly and prints what each one can reach.

License

MIT.

Source 3 files
hooks/register.ts 111 lines
1import type { Register } from 'claude-code'
2
3import { readRoster } from '../lib/roster.js'
4import { contextLine, suggest } from '../lib/scout.js'
5
6type Skill = { name: string; description: string; excerpt: string; path: string; source: string }
7type Init = { method?: string; headers?: Record<string, string>; body?: string }
8type Suggestion = {
9  suggestion: string | null
10  verify: { fits: Record<string, number> } | null
11  ms: number
12}
13type Options = {
14  apiKey?: string
15  enabled?: boolean
16  gateThreshold?: number
17  fitsThreshold?: number
18  timeoutMs?: number
19  model?: string
20  quiet?: boolean
21  shadow?: boolean
22}
23
24const ROSTER_KEY = 'roster.v1'
25const ROSTER_TTL_MS = 10 * 60 * 1000
26
27export const register: Register = (on, options) => {
28  const opt = (options ?? {}) as Options
29  const enabled = opt.enabled !== false
30
31  on('session.start', async ($, e, next) => {
32    const r = await next(e)
33    if (!enabled) return r
34    const home = await $.env.get('HOME')
35    if (!home) return r
36    const roster = await readRoster(
37      { list: p => $.fs.list(p), read: p => $.fs.read(p), exists: p => $.fs.exists(p) },
38      { home, cwd: e.cwd },
39    )
40    await $.store.set(ROSTER_KEY, { at: Date.now(), cwd: e.cwd, skills: roster })
41    if (!opt.quiet) $.ui.log(`jev-skill-scout: ${roster.length} skills indexed`)
42    return r
43  })
44
45  on('prompt.submit', async ($, e, next) => {
46    const typed = e.origin.kind === 'composer' || e.origin.kind === 'bridge'
47    if (!enabled || !typed || e.text.length < 12 || e.text.startsWith('/')) return next(e)
48    const key = opt.apiKey || (await $.env.get('TYPESAFE_API_KEY')) || (await $.env.get('TYPESAFE_KEY'))
49    if (!key) return next(e)
50    const cached = (await $.store.get(ROSTER_KEY)) as { at: number; cwd: string; skills: Skill[] } | undefined
51    let roster = cached?.skills ?? []
52    if (!cached || Date.now() - cached.at > ROSTER_TTL_MS) {
53      const home = await $.env.get('HOME')
54      const cwd = await $.session.cwd()
55      if (home) {
56        roster = await readRoster(
57          { list: p => $.fs.list(p), read: p => $.fs.read(p), exists: p => $.fs.exists(p) },
58          { home, cwd },
59        )
60        await $.store.set(ROSTER_KEY, { at: Date.now(), cwd, skills: roster })
61      }
62    }
63    if (!roster.length) return next(e)
64
65    const started = await $.clock.now()
66    let result: Suggestion | null = null
67    try {
68      result = (await Promise.race([
69        suggest({
70          // The engine's init has no abort signal; the race below is the timeout.
71          fetchImpl: (url: string, init: Init) => $.http.fetch(url, { method: init.method, headers: init.headers, body: init.body }),
72          key,
73          roster,
74          request: e.text,
75          options: {
76            model: opt.model,
77            gateThreshold: opt.gateThreshold,
78            fitsThreshold: opt.fitsThreshold,
79            timeoutMs: opt.timeoutMs ?? 4000,
80          },
81        }),
82        $.clock.sleep(opt.timeoutMs ?? 4000).then(() => null),
83      ])) as Suggestion | null
84    } catch (err) {
85      if (!opt.quiet) $.ui.log(`jev-skill-scout: off for this turn (${String(err).slice(0, 120)})`)
86      return next(e)
87    }
88    const ms = (await $.clock.now()) - started
89    if (!result) {
90      if (!opt.quiet) $.ui.status(`jev-skill-scout: no answer in ${ms} ms, turn left alone`)
91      return next(e)
92    }
93    if (!result.suggestion) {
94      if (!opt.quiet) $.ui.status(`jev-skill-scout: no skill (${ms} ms)`)
95      return next(e)
96    }
97    const fit = result.verify?.fits?.[result.suggestion] ?? 0
98    const shadow = opt.shadow || (await $.env.get('JEV_SKILL_SCOUT_SHADOW')) === '1'
99    if (shadow) {
100      $.ui.status(`jev-skill-scout (shadow): would suggest ${result.suggestion} (fit ${fit.toFixed(2)}, ${ms} ms)`)
101      return next(e)
102    }
103    if (!opt.quiet) $.ui.status(`jev-skill-scout: ${result.suggestion} (fit ${fit.toFixed(2)}, ${ms} ms)`)
104    return next({ ...e, context: [...(e.context ?? []), contextLine(result.suggestion)] })
105  }).catch(($, e, next) => {
106    // Core has a side effect on prompt.submit: enter the prompt exactly once.
107    if (next.called) return undefined
108    return next(e)
109  })
110}
111
lib/roster.js 126 lines
1// Finds every SKILL.md Claude Code can load and reads its frontmatter.
2// Pure over an injected fs so the mod (engine $.fs) and the CLI (node:fs) share it.
3
4const PLUGIN_ROOT = ['.claude', 'plugins', 'cache'];
5
6export function parseFrontmatter(text) {
7  const lines = text.split('\n');
8  let i = 0;
9  while (i < lines.length && lines[i].trim() === '') i++;
10  if ((lines[i] ?? '').trim() !== '---') return { fields: {}, body: text };
11  i++;
12  const fields = {};
13  let key = '';
14  let buf = [];
15  const flush = () => {
16    if (key) fields[key] = buf.join(' ').replace(/\s+/g, ' ').trim();
17    key = '';
18    buf = [];
19  };
20  for (; i < lines.length; i++) {
21    const line = lines[i];
22    if (line.trim() === '---') { i++; break; }
23    const m = /^([A-Za-z_][\w-]*):[ \t]*(.*)$/.exec(line);
24    if (m && !/^[ \t]/.test(line)) {
25      flush();
26      key = m[1];
27      const value = m[2].trim();
28      if (value && !/^[>|][-+]?$/.test(value)) buf.push(value.replace(/^["']|["']$/g, ''));
29    } else if (key && line.trim()) {
30      buf.push(line.trim());
31    }
32  }
33  flush();
34  return { fields, body: lines.slice(i).join('\n') };
35}
36
37// Version directories sort as 1.10.0 > 1.9.0, not as strings.
38function versionKey(v) {
39  return v.split(/[.-]/).map(p => (/^\d+$/.test(p) ? Number(p) : -1));
40}
41function newestVersion(names) {
42  return [...names].sort((a, b) => {
43    const ka = versionKey(a), kb = versionKey(b);
44    for (let i = 0; i < Math.max(ka.length, kb.length); i++) {
45      const d = (kb[i] ?? -1) - (ka[i] ?? -1);
46      if (d) return d;
47    }
48    return 0;
49  })[0];
50}
51
52/**
53 * @param {{ list(p:string):Promise<{name:string,kind:string}[]>, read(p:string):Promise<string>, exists(p:string):Promise<boolean> }} fs
54 * @param {{ home:string, cwd?:string, bodyChars?:number }} where
55 * @returns {Promise<{ name:string, description:string, excerpt:string, path:string, source:string }[]>}
56 */
57export async function readRoster(fs, { home, cwd, bodyChars = 700 }) {
58  const out = new Map();
59  const add = async (skillDir, name, source) => {
60    const p = `${skillDir}/SKILL.md`;
61    if (!(await fs.exists(p))) return;
62    let text;
63    try { text = await fs.read(p); } catch { return; }
64    const { fields, body } = parseFrontmatter(text.slice(0, 12000));
65    const description = fields.description ?? '';
66    if (!description) return;
67    const key = name;
68    if (out.has(key)) return;
69    out.set(key, {
70      name: key,
71      description,
72      excerpt: body.replace(/\s+/g, ' ').trim().slice(0, bodyChars),
73      path: p,
74      source,
75    });
76  };
77  const dirs = async p => {
78    try { return (await fs.list(p)).filter(e => !e.name.startsWith('.')); } catch { return []; }
79  };
80
81  if (cwd) for (const e of await dirs(`${cwd}/.claude/skills`)) await add(`${cwd}/.claude/skills/${e.name}`, e.name, 'project');
82  for (const e of await dirs(`${home}/.claude/skills`)) await add(`${home}/.claude/skills/${e.name}`, e.name, 'user');
83
84  // Only enabled plugins reach the model's skill list; a cached but disabled
85  // one must not be suggested. Without readable settings, every plugin counts.
86  let enabled = null;
87  try {
88    const settings = JSON.parse(await fs.read(`${home}/.claude/settings.json`));
89    if (settings && typeof settings.enabledPlugins === 'object') enabled = settings.enabledPlugins;
90  } catch { enabled = null; }
91
92  const cache = [home, ...PLUGIN_ROOT].join('/');
93  for (const market of await dirs(cache)) {
94    for (const plugin of await dirs(`${cache}/${market.name}`)) {
95      if (enabled && enabled[`${plugin.name}@${market.name}`] !== true) continue;
96      const versions = (await dirs(`${cache}/${market.name}/${plugin.name}`)).map(v => v.name);
97      if (!versions.length) continue;
98      const v = newestVersion(versions);
99      const skillsDir = `${cache}/${market.name}/${plugin.name}/${v}/skills`;
100      for (const s of await dirs(skillsDir)) await add(`${skillsDir}/${s.name}`, `${plugin.name}:${s.name}`, `plugin ${market.name}`);
101    }
102  }
103  return [...out.values()];
104}
105
106/**
107 * The roster as Claude Code itself listed it in a transcript (`skill_listing`
108 * attachments): one `- name: description` per skill, long descriptions wrapped.
109 */
110export function parseListing(content) {
111  const out = [];
112  for (const raw of String(content ?? '').split('\n')) {
113    const m = /^- ([^\s:]+(?::[^\s:]+)?): ?(.*)$/.exec(raw);
114    if (m) out.push({ name: m[1], description: m[2].trim() });
115    else if (out.length && raw.trim()) out[out.length - 1].description += ` ${raw.trim()}`;
116  }
117  return out;
118}
119
120// Skill tool calls name plugin skills as plugin:skill, sometimes as just skill.
121export function sameSkill(a, b) {
122  if (!a || !b) return false;
123  if (a === b) return true;
124  return a.split(':').pop() === b.split(':').pop();
125}
126
lib/scout.js 202 lines
1// The judgment. Two TypeSafe requests at most, following the skill-suggestion
2// cookbook (https://docs.typesafe.ai/cookbooks/skill_suggestion): rank the whole
3// roster and gate the turn, then re-read the top few properly and allow a reject.
4// Pure: the caller supplies fetch, so the mod and the audit make the same calls.
5
6export const NONE = '__none__';
7export const ENDPOINT = 'https://api.typesafe.ai/v1/systemone';
8// A Choice takes at most 255 options; rosters past this are ranked in chunks.
9export const CHUNK = 250;
10
11export const DEFAULTS = {
12  model: 'jev-latest',
13  shortlist: 3,
14  gateThreshold: 0.3,   // mean of the oriented gate nouls; below it, no skill
15  fitsThreshold: 0.3,   // the winner's "really fits" noul; below it, no skill
16  descriptionChars: 400,
17  excerptChars: 700,
18  contextChars: 400,
19  timeoutMs: 4000,
20  minPromptChars: 12,
21};
22
23const GATES = {
24  acts: 'Does `request` ask the agent to do something with the user\'s files, repositories, accounts, sites, documents or data, rather than only answer from general knowledge?',
25  procedure: 'Would a careful agent answer `request` better by following a specific written procedure, checklist or house style, rather than by general skill alone?',
26  prose: 'Can `request` be fully satisfied by a short reply in prose, with no tools, no files and no procedure?',
27};
28const INVERTED = new Set(['prose']);
29
30const cut = (s, n) => {
31  const flat = String(s ?? '').replace(/\s+/g, ' ').trim();
32  return flat.length <= n ? flat : `${flat.slice(0, n - 3).trimEnd()}...`;
33};
34
35export function buildRank(roster, request, recent, o = DEFAULTS) {
36  const criteria = {};
37  for (const s of roster) criteria[s.name] = cut(s.description, o.descriptionChars);
38  criteria[NONE] = 'None of the skills above is what this request needs. Also the answer for small talk, a question answerable from what is already on screen, or a small direct edit that needs no procedure.';
39  const questions = {
40    which: {
41      type: 'choice',
42      instructions: 'Which single skill should the agent read before answering `request`? Judge each skill only by what its description says it is for, and prefer the skill whose trigger conditions the request matches most specifically.',
43      criteria,
44    },
45  };
46  for (const [k, text] of Object.entries(GATES)) questions[`gate_${k}`] = { type: 'noul', instructions: text };
47  return {
48    model: o.model,
49    state: { request: cut(request, 4000), recent_context: cut(recent, o.contextChars) },
50    questions,
51  };
52}
53
54export function readRank(answers) {
55  const which = answers.which;
56  const ranked = Object.entries(which.probabilities ?? {}).sort((a, b) => b[1] - a[1]);
57  const gates = {};
58  let sum = 0, n = 0;
59  for (const [k] of Object.entries(GATES)) {
60    const v = answers[`gate_${k}`]?.noul;
61    if (typeof v !== 'number') continue;
62    gates[k] = v;
63    sum += INVERTED.has(k) ? 1 - v : v;
64    n++;
65  }
66  return { ranked, gate: n ? sum / n : 0, gates, confidence: which.confidence ?? null };
67}
68
69export function buildVerify(candidates, request, recent, o = DEFAULTS) {
70  const criteria = {};
71  for (const c of candidates) criteria[c.name] = `${cut(c.description, o.descriptionChars)} Instructions begin: ${cut(c.excerpt, o.excerptChars)}`;
72  criteria[NONE] = 'None of these skills should be loaded for this request.';
73  const questions = {
74    which: {
75      type: 'choice',
76      instructions: 'Now that each candidate skill\'s real instructions are visible, which one should the agent load before answering `request`?',
77      criteria,
78    },
79  };
80  candidates.forEach((c, i) => {
81    questions[`fits_${i}`] = {
82      type: 'noul',
83      instructions: `Would loading skill \`candidates[${i}].name\` change how the agent handles \`request\` for the better, judged by the skill's own instructions, not its name?`,
84      criteria: { true: 'The skill\'s instructions cover this request and following them would change the work.', false: 'Wrong domain, or the request would be handled the same way without it.' },
85    };
86  });
87  return {
88    model: o.model,
89    state: {
90      request: cut(request, 4000),
91      recent_context: cut(recent, o.contextChars),
92      candidates: candidates.map(c => ({ name: c.name, description: cut(c.description, o.descriptionChars), instructions: cut(c.excerpt, o.excerptChars) })),
93    },
94    questions,
95  };
96}
97
98export function readVerify(answers, candidates) {
99  const which = answers.which;
100  const fits = {};
101  candidates.forEach((c, i) => { fits[c.name] = answers[`fits_${i}`]?.noul ?? 0; });
102  return { winner: which.choice, probabilities: which.probabilities ?? {}, confidence: which.confidence ?? null, fits };
103}
104
105/**
106 * Ask Jev. `fetchImpl(url, init)` resolves `{ ok, status, text }` (the engine's
107 * $.http.fetch and a wrapper over global fetch both fit).
108 */
109export async function ask(fetchImpl, key, body, timeoutMs = DEFAULTS.timeoutMs, retries = 0) {
110  for (let attempt = 0; ; attempt++) {
111    try {
112      return await askOnce(fetchImpl, key, body, timeoutMs);
113    } catch (e) {
114      const msg = String(e?.message ?? e);
115      const retryable = /aborted|429|50\d|fetch failed|ECONN|ETIMEDOUT/i.test(msg);
116      if (!retryable || attempt >= retries) throw e;
117      await new Promise(r => setTimeout(r, 500 * (attempt + 1)));
118    }
119  }
120}
121
122async function askOnce(fetchImpl, key, body, timeoutMs) {
123  const controller = typeof AbortController === 'function' ? new AbortController() : null;
124  let timer = null;
125  // Two guards: the abort for transports that honour it, the race for those that do not.
126  const deadline = new Promise((_, reject) => { timer = setTimeout(() => { controller?.abort(); reject(new Error('aborted: timeout')); }, timeoutMs); });
127  const work = (async () => {
128    const r = await fetchImpl(ENDPOINT, {
129      method: 'POST',
130      headers: { authorization: `Bearer ${key}`, 'content-type': 'application/json' },
131      body: JSON.stringify(body),
132      signal: controller?.signal,
133    });
134    const text = typeof r.text === 'function' ? await r.text() : r.text;
135    if (!r.ok) throw new Error(`typesafe ${r.status}: ${String(text).slice(0, 200)}`);
136    return JSON.parse(text);
137  })();
138  try {
139    return await Promise.race([work, deadline]);
140  } finally {
141    if (timer) clearTimeout(timer);
142  }
143}
144
145/** Re-applies the thresholds to a stored judgment, so a rerun with new ones costs nothing. */
146export function decide(res, options = {}) {
147  const o = { ...DEFAULTS, ...options };
148  if (!res || !res.rank) return null;
149  if (res.rank.gate < o.gateThreshold) return null;
150  const v = res.verify;
151  if (!v || !v.winner || v.winner === NONE) return null;
152  return (v.fits[v.winner] ?? 0) >= o.fitsThreshold ? v.winner : null;
153}
154
155/**
156 * The whole judgment for one request. Returns what the mod would attach and
157 * everything the audit needs to explain it.
158 */
159export async function suggest({ fetchImpl, key, roster, request, recent = '', options = {} }) {
160  const o = { ...DEFAULTS, ...options };
161  const t0 = Date.now();
162  const chunks = [];
163  for (let i = 0; i < roster.length; i += CHUNK) chunks.push(roster.slice(i, i + CHUNK));
164  const replies = await Promise.all(chunks.map(c => ask(fetchImpl, key, buildRank(c, request, recent, o), o.timeoutMs, o.retries ?? 0)));
165  const usage = { input: 0, output: 0, calls: replies.length };
166  for (const r of replies) { usage.input += r.usage?.input_tokens ?? 0; usage.output += r.usage?.output_tokens ?? 0; }
167  const rank = mergeRanks(replies.map(r => readRank(r.answers)), o.shortlist);
168  const out = { suggestion: null, rank, verify: null, usage, ms: 0, stage: 'gate' };
169  if (rank.gate < o.gateThreshold) { out.ms = Date.now() - t0; return out; }
170  const short = rank.ranked.filter(([n]) => n !== NONE).slice(0, o.shortlist).map(([n]) => roster.find(s => s.name === n)).filter(Boolean);
171  if (!short.length) { out.ms = Date.now() - t0; return out; }
172  const r2 = await ask(fetchImpl, key, buildVerify(short, request, recent, o), o.timeoutMs, o.retries ?? 0);
173  const verify = readVerify(r2.answers, short);
174  usage.input += r2.usage?.input_tokens ?? 0;
175  usage.output += r2.usage?.output_tokens ?? 0;
176  usage.calls += 1;
177  out.verify = verify;
178  out.stage = 'verify';
179  out.ms = Date.now() - t0;
180  out.suggestion = decide(out, o);
181  return out;
182}
183
184// Probabilities from different chunks are not comparable, so each chunk keeps
185// only its own top few and the verify stage settles it with real excerpts.
186function mergeRanks(ranks, shortlist) {
187  if (ranks.length === 1) return ranks[0];
188  const ranked = [];
189  let gate = 0;
190  for (const r of ranks) {
191    ranked.push(...r.ranked.filter(([n]) => n !== NONE).slice(0, shortlist));
192    gate += r.gate;
193  }
194  ranked.sort((a, b) => b[1] - a[1]);
195  ranked.push([NONE, Math.max(...ranks.map(r => r.ranked.find(([n]) => n === NONE)?.[1] ?? 0))]);
196  return { ranked, gate: gate / ranks.length, gates: ranks[0].gates, confidence: null };
197}
198
199export function contextLine(name) {
200  return `<skill_relevance>\nRelevant to this request: ${name}. Load it with the Skill tool before answering. Ignore this if it does not fit what the user actually asked for.\n</skill_relevance>`;
201}
202