SLOPSHOPPER

typesafe-mod

Routes decisions to TypeSafe's Jev model: ranks installed skills for each prompt, and answers the agent's own this-or-that questions when confident.

newrowsguardpromptprocessnetwork
★ 4v0.5.1MITupdated 2026-09-17BeLazy167/typesafe-mod
A shopper browsing a rack in a slop shop
README

typesafe-mod

A Claude Code function hook that sends two kinds of decision to TypeSafe's Jev model.

Jev returns typed answers (Choice, Noul, Score) with probabilities instead of prose, so your code can branch on both the answer and how confident it is.

This is a function hook, not a shell hook. It is a TypeScript module that runs inside the agent loop, so it can answer a tool call itself rather than only allowing or blocking one.

What it does

Decision router, on tool.call for AskUserQuestion. When the agent stops to ask you a this-or-that question, the hook sends it to Jev and puts the probability for each option in the transcript beside the dialog. You still choose. TYPESAFE_AUTO_ANSWER=1 lets Jev answer instead, so the dialog never opens. This runs by default.

A dialog may carry several questions, and they all ride one request. Jev answers independent questions in parallel, so two steps cost one call. Each gets its own transcript line, and each is judged on its own: it can be confident about one step and hand the next back to you.

typesafe-mod: (1/2) Jev picks Tag v0.3.0 as-is (Tag v0.3.0 as-is 0.89,
Bump to 0.4.0 0.10, Backfill v0.2.0 0.01), confidence 0.84
typesafe-mod: (2/2) Jev leans Everything since the first commit
(Everything 0.51, Only since the bump 0.49), but confidence 0.02 is under
the floor, so this one is yours

Skill router, on prompt.submit. One request ranks every installed skill against your prompt and asks whether the turn needs a procedure at all. A confident winner becomes one advisory line of context. The skill list the agent already has does not change, so prefix caching over it still works. This is off by default. See the cost section for why.

Install

claude plugin marketplace add BeLazy167/typesafe-mod
claude plugin install typesafe-mod@typesafe-mod

To run it from a checkout without installing:

claude --plugin-dir /path/to/typesafe-mod

Two things are required. Without either one, the mod loads and does nothing.

1. Function hooks turned on. They are early access. In ~/.claude/settings.json:

{ "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" } }

2. A TypeSafe API key in the environment Claude Code runs in. Get one at <https://console.typesafe.ai/keys>.

export TYPESAFE_API_KEY="your-key-here"

Switches

VariableEffect
TYPESAFE_API_KEYUnset, and neither router runs.
TYPESAFE_SKILL_ROUTERSet to 1 to turn the skill router on. Off otherwise.
TYPESAFE_DECIDE_OFFSet to any value to stop Jev answering your questions.
TYPESAFE_AUTO_ANSWERSet to 1 to let Jev answer the dialog outright instead of advising. See the warning below.

By default the router advises. The dialog opens as usual, with a panel above it drawing a bar per option, and the same numbers go to the transcript.

  TypeSafe decision router

  ███████████████████████░░░░░░░ 0.73  Fix forward
  ████████░░░░░░░░░░░░░░░░░░░░░░ 0.27  Roll back

  confidence 0.47, under the 0.75 floor, so this one is yours

A dialog with several questions gets one row each instead, because the render event never says which step is on screen:

  TypeSafe decision router  (4 questions, one row each)

  ████████████████ 0.97  No pattern       0.95 ✓
  ███████████░░░░░ 0.71  Depends on size  0.57
  ██████████░░░░░░ 0.64  Both             0.46
  ████████░░░░░░░░ 0.50  Rebase           0.25

  1 of 4 cleared the 0.75 floor, marked with a tick

Drawing one question's options would pin the panel to that question while you page through the rest, showing step one's numbers above step three.

The panel wraps core's dialog rather than replacing it. AskUserQuestion is drawn by exactly one engine node, so a tree without that node is refused and core draws its own. The engine also caps how much a hook may add around a dialog, so the panel draws at most four rows.

The transcript line carries the same information in one row:

typesafe-mod: Jev leans Fix forward (Fix forward 0.73, Roll back 0.27),
but confidence 0.47 is under the floor, so this one is yours

TYPESAFE_AUTO_ANSWER=1 makes the router answer the call instead, so the dialog never opens. Know the cost before you turn it on. Answering a tool call from a hook requires deny, and the engine defines deny as "the model receives the text as an error result". So a decision that worked renders in red as a failure, and the agent may argue with it rather than proceed. Answering with { result } would avoid that, but it needs this tool's output schema, which the generated types do not declare.

Both thresholds live in hooks/suggest.ts. DEFAULTS.minConfidence is 0.6 for the skill hint. DECISION_DEFAULTS.minConfidence is 0.75 for answering instead of you.

The decision floor is higher for a reason. When the skill router says nothing, you lose one hint. When the decision router is confidently wrong, it answers a question that was yours to answer. Tune both against your own turns rather than these numbers.

What it costs

Jev bills input tokens only, at $0.042 per million. It does no token-by-token decoding, so output is not billed.

Input tokensCost
One decision374$0.000016
One skill-router prompt6,772$0.00028

The decision router fires only when the agent stops to ask you something, so it costs almost nothing.

The skill router runs on every prompt. At 100 prompts a day that is about $0.85 a month, which is not the reason it is off by default. It adds 166 to 420 ms to every prompt you type, and that is the reason.

Measured on a 194-skill install

The scan takes about 1.1 s, once per session.

PromptAnswer
"the merge is conflicted on three files"resolving-merge-conflicts at 1.00
"what time is it in Tokyo right now"nothing, needs_skill 0.02
"review the changes before I open a PR"code-review at 0.96

When it fails

Every path fails open. A missing key, an HTTP error, a timeout, a hook that throws, a body that will not parse, and a skill list that never built all fall through to next(e), and the turn carries on untouched.

A hook cannot catch the engine's 10 second dispatch budget, because the engine enforces it from outside. So each call races $.clock.sleep(4000) and gives up first.

The decision router also checks that the answer is one of the options you offered. If it is not, the router ignores it and you get asked.

Design notes

Every call writes $ out in full, as in $.http.fetch(...). The loader rejects a module that binds, passes, spreads, or destructures $. So every engine call sits inline in register.ts, and the logic lives in pure functions in suggest.ts. Those functions take plain data and return plain data, which is why the tests need no mocks.

One $.process.run reads all the skills, rather than about 194 $.fs.read calls. $.fs compares paths against the project directory as strings, so it cannot reach ~/.claude.

The scan runs find -L. Most of ~/.claude/skills is symlinks, and without -L the scan found 4 of 46 entries.

In auto-answer mode the hook returns { deny: text }, which is the only way a hook can answer a tool call it is given. The engine defines deny as "the model receives the text as an error result", so the answer renders in red. { result } would render normally, but core validates it against the tool's output schema and the generated types declare none for AskUserQuestion. That is why advising is the default.

The router skips multi-select steps, because a multi-select answer is not a Choice and approximating one would answer a question you did not ask. The other steps in the same dialog still route.

A skipped step says which step it was and what was wrong with it, because "did not parse" does not tell you whether to change anything:

typesafe-mod: not routed, step 1 is multi-select, which is not a Choice
typesafe-mod: skipped 1 of 3, step 2 offers 1, so there is nothing to
choose between

The panel draws bars for the first answered step only. The engine caps what a hook may add around a dialog, so a second set of bars would be refused and core would draw its own. Every step still gets a transcript line, where text costs nothing.

Known limits

  • The decision router reads the question and its options, not the conversation. It also passes the working directory. Questions that depend on anything else will route worse.
  • Ranking 194 skills costs 6,772 input tokens per prompt. A shorter list costs less. Cutting the list blindly is what broke the first version: a cap of 120 dropped resolving-merge-conflicts, and a missing skill looks exactly like "nothing fits".
  • 54 skill names here appear in more than one plugin. The scan keeps the first one it finds, so the router may name a different plugin's copy.
claude plugin validate typesafe-skill-mod
claude plugin test typesafe-skill-mod     # 48 tests
Source 2 files
hooks/register.ts 292 lines
1import type { Register } from 'claude-code';
2import {
3  DEFAULTS,
4  DECISION_DEFAULTS,
5  DECISION_KEY,
6  ROSTER_KEY,
7  MAX_PANEL_ROWS,
8  SCAN_COMMAND,
9  bar,
10  rankOptions,
11  buildRequest,
12  contextBlock,
13  decisionNote,
14  isSkillEntry,
15  parseRoster,
16  pickDecision,
17  pickWinner,
18  scanAskQuestions,
19  buildBatchRequest,
20  readDecisions,
21  isDecisionViewList,
22  decisionLine,
23  summaryRow,
24  pairDecisions,
25} from './suggest';
26
27/**
28 * Two decision points, both answered by TypeSafe's Jev model.
29 *
30 * `prompt.submit` ranks the installed skills for the turn and attaches one
31 * advisory line. `tool.call` on AskUserQuestion answers the agent's own
32 * this-or-that question when Jev is confident enough, so the turn continues
33 * without stopping the user.
34 *
35 * Every path fails open. Routing is an optimisation, and no turn should wait
36 * on, or break because of, a third-party service.
37 *
38 * `$` is written out at every call site: the loader refuses a module that
39 * binds, passes, spreads or destructures it.
40 */
41export const register: Register = (on) => {
42  // The roster is built once per session. The gotchas list says to do work
43  // like this in session.start: a prompt.submit hook has a 10 s dispatch
44  // budget, enforced outside the hook, that a cold scan could blow.
45  on('session.start', async ($, e, next) => {
46    // The skill router is opt-in. Off, there is nothing to scan for.
47    const enabled = await $.env.get('TYPESAFE_SKILL_ROUTER');
48    if (!enabled) return next(e);
49    try {
50      const scan = await $.process.run(SCAN_COMMAND, { timeoutMs: 8000 });
51      const roster = parseRoster(scan.stdout);
52      await $.store.set(ROSTER_KEY, roster);
53      // A cap that is too low drops real skills, and the router then says
54      // nothing about them. That looks exactly like "nothing fits", so report
55      // it instead of staying silent.
56      if (roster.length >= DEFAULTS.maxSkills) {
57        $.ui.log(`typesafe-mod: roster hit the ${DEFAULTS.maxSkills} cap; raise maxSkills`);
58      } else {
59        $.ui.log(`typesafe-mod: ${roster.length} skills indexed`);
60      }
61    } catch (err) {
62      // A hook that throws is skipped silently, so catch and say so.
63      $.ui.log(`typesafe-mod: skill scan failed, router off (${String(err)})`);
64    }
65    return next(e);
66  });
67
68  on('prompt.submit', async ($, e, next) => {
69    // Opt-in: this hook runs on every prompt, so it costs 160-420 ms of the
70    // turn each time. The money is negligible (Jev bills input only, at
71    // $0.042/M, so about $1.35 a month at 100 prompts a day); the latency is
72    // not. Off by default, on with TYPESAFE_SKILL_ROUTER=1.
73    const enabled = await $.env.get('TYPESAFE_SKILL_ROUTER');
74    if (!enabled) return next(e);
75
76    const text = typeof e.text === 'string' ? e.text.trim() : '';
77    // A slash command already names what it wants, and a very short prompt
78    // carries too little to route on.
79    if (text.length < 12 || text.startsWith('/')) return next(e);
80
81    const key = await $.env.get('TYPESAFE_API_KEY');
82    if (!key) return next(e);
83
84    const cached = await $.store.get(ROSTER_KEY);
85    const roster = Array.isArray(cached) ? cached.filter(isSkillEntry) : [];
86    if (roster.length === 0) return next(e);
87
88    let block = '';
89    try {
90      // The dispatch budget cannot be caught from in here, so race it and
91      // give up first.
92      const res = await Promise.race([
93        $.http.fetch(DEFAULTS.endpoint, {
94          method: 'POST',
95          headers: { 'Content-Type': 'application/json', Authorization: `Bearer ${key}` },
96          body: JSON.stringify(buildRequest(roster, text)),
97        }),
98        $.clock.sleep(DEFAULTS.budgetMs).then(() => null),
99      ]);
100
101      if (res === null) {
102        $.ui.log('typesafe-mod: skill router timed out, continuing');
103      } else if (res.ok) {
104        const winner = pickWinner(JSON.parse(res.text));
105        if (winner) {
106          block = contextBlock(winner);
107          $.ui.log(`typesafe-mod: skill hint ${winner.name} (${winner.confidence.toFixed(2)})`);
108        }
109      } else {
110        $.ui.log(`typesafe-mod: skill router HTTP ${res.status}`);
111      }
112    } catch (err) {
113      $.ui.log(`typesafe-mod: skill router unavailable (${String(err)})`);
114    }
115
116    if (!block) return next(e);
117    return next({ ...e, context: [...(e.context ?? []), block] });
118  });
119
120  // The agent's own this-or-that question is the general decision point. A
121  // dialog may carry several, and they ride one request together.
122  on('tool.call', { tool: 'AskUserQuestion' }, async ($, e, next) => {
123    const scan = scanAskQuestions((e as { questions?: unknown }).questions);
124    const questions = scan.routable;
125    // Nothing routable here. A multi-select step is not a Choice, and
126    // approximating one would answer a question the agent did not ask.
127    if (questions.length === 0) {
128      $.ui.log(`typesafe-mod: not routed, ${scan.skipped.join('; ')}`);
129      return next(e);
130    }
131    // Some steps routed and some did not. Say which, so a dialog that is only
132    // half covered does not look like a dialog the router ignored.
133    if (scan.skipped.length > 0) {
134      $.ui.log(`typesafe-mod: skipped ${scan.skipped.length} of ${scan.skipped.length + questions.length}, ${scan.skipped.join('; ')}`);
135    }
136
137    const off = await $.env.get('TYPESAFE_DECIDE_OFF');
138    if (off) return next(e);
139
140    const key = await $.env.get('TYPESAFE_API_KEY');
141    if (!key) return next(e);
142
143    try {
144      const cwd = await $.session.cwd();
145      const res = await Promise.race([
146        $.http.fetch(DEFAULTS.endpoint, {
147          method: 'POST',
148          headers: { 'Content-Type': 'application/json', Authorization: `Bearer ${key}` },
149          body: JSON.stringify(
150            buildBatchRequest(
151              questions,
152              `A coding agent paused mid-task in ${cwd} to ask this. ` +
153                `Judge only from the questions and the options as written.`
154            )
155          ),
156        }),
157        $.clock.sleep(DEFAULTS.budgetMs).then(() => null),
158      ]);
159
160      if (res !== null && res.ok) {
161        const payload: unknown = JSON.parse(res.text);
162        const views = readDecisions(payload, questions, DECISION_DEFAULTS.minConfidence);
163
164        // Record them whatever happens next, so the render hook can show why
165        // the router stood aside on the questions it did not answer.
166        if (views.length > 0) await $.store.set(DECISION_KEY, views);
167        else await $.store.delete(DECISION_KEY);
168
169        // Answering the call outright needs `deny`, which the engine defines as
170        // "the model receives the text as an error result", so a working
171        // decision renders red. It is opt-in, and only for a lone question: a
172        // single deny string cannot answer a dialog of several.
173        const auto = await $.env.get('TYPESAFE_AUTO_ANSWER');
174        const only = questions.length === 1 ? questions[0] : undefined;
175        if (auto && only) {
176          const decision = pickDecision(payload, only, DECISION_DEFAULTS.minConfidence);
177          if (decision) {
178            $.ui.log(
179              `typesafe-mod: decided "${decision.label}" (${decision.confidence.toFixed(2)}) without asking`
180            );
181            return { deny: decisionNote(only, decision) };
182          }
183        }
184
185        // Text has no element budget, so every question gets a line even when
186        // the panel can only draw one.
187        views.forEach((view, i) => {
188          const q = questions[i];
189          if (!q) return;
190          const prefix = questions.length > 1 ? `(${i + 1}/${questions.length}) ` : '';
191          $.ui.log(decisionLine(view, q, prefix));
192        });
193      }
194    } catch (err) {
195      $.ui.log(`typesafe-mod: decision router unavailable (${String(err)})`);
196    }
197
198    // Not confident, unreachable, or switched off, so the user gets asked.
199    return next(e);
200  });
201
202  // Draws Jev's distribution over the question dialog. Only fires in
203  // show-your-work mode, because otherwise the hook answers the tool call
204  // and no dialog is ever drawn.
205  on('ui.render', { component: 'AskUserQuestion' }, async ($, e, next) => {
206    const stored = await $.store.get(DECISION_KEY);
207    if (!isDecisionViewList(stored)) {
208      $.ui.log('typesafe-mod: render skipped, no stored decision');
209      return next(e);
210    }
211
212    const questions = scanAskQuestions((e.props as { questions?: unknown }).questions);
213    if (questions.routable.length === 0) {
214      $.ui.log(`typesafe-mod: render skipped, ${questions.skipped.join('; ')}`);
215      return next(e);
216    }
217
218    const steps = questions.routable;
219    const paired = pairDecisions(steps, stored);
220    if (paired.length === 0) {
221      $.ui.log('typesafe-mod: render skipped, no decision matches this dialog');
222      return next(e);
223    }
224
225    const el = $.ui.resolve(e);
226    const columns = e.viewport?.columns ?? 80;
227    const multi = paired.length > 1;
228
229    // The render event never says which step the dialog is showing, so a
230    // batched dialog gets one row per question rather than one question's
231    // option bars. Pinning the panel to step one would misread every later
232    // step as belonging to the first.
233    let rows;
234    let title;
235    let footer;
236    if (multi) {
237      const shown = paired.slice(0, MAX_PANEL_ROWS);
238      const labelWidth = Math.min(
239        shown.reduce((w, p) => Math.max(w, p.view.choice.length), 0),
240        Math.max(8, columns - 34)
241      );
242      const barWidth = Math.max(8, Math.min(columns - labelWidth - 28, 16));
243      rows = shown.map((p, i) =>
244        el.Text({
245          key: `q${i}`,
246          color: p.view.wouldAnswer ? 'green' : 'gray',
247          bold: p.view.wouldAnswer,
248          children: summaryRow(p.view, barWidth, labelWidth),
249        })
250      );
251      const over = paired.filter((p) => p.view.wouldAnswer).length;
252      title = `TypeSafe decision router  (${paired.length} questions, one row each)`;
253      footer =
254        over === 0
255          ? 'none cleared the 0.75 floor, so every one of these is yours'
256          : `${over} of ${paired.length} cleared the 0.75 floor, marked with a tick`;
257    } else {
258      const only = paired[0]!;
259      const width = Math.max(16, Math.min(columns - 28, 32));
260      rows = rankOptions(only.view, only.question)
261        .slice(0, MAX_PANEL_ROWS)
262        .map((row) =>
263          el.Text({
264            key: row.label,
265            color: row.label === only.view.choice ? 'green' : 'gray',
266            bold: row.label === only.view.choice,
267            children: `${bar(row.p, width)} ${row.p.toFixed(2)}  ${row.label}`,
268          })
269        );
270      title = 'TypeSafe decision router';
271      footer = only.view.wouldAnswer
272        ? `confidence ${only.view.confidence.toFixed(2)}, over the 0.75 floor`
273        : `confidence ${only.view.confidence.toFixed(2)}, under the 0.75 floor, so this one is yours`;
274    }
275
276    // The dialog is drawn by exactly one engine node, so this wraps core's own
277    // tree rather than replacing it. A tree with no engine node is refused.
278    const core = await next(e);
279
280    return el.Box({
281      flexDirection: 'column',
282      paddingX: 1,
283      children: [
284        el.Text({ bold: true, color: 'cyan', children: title }),
285        ...rows,
286        el.Text({ dimColor: true, children: footer }),
287        core,
288      ],
289    });
290  });
291};
292
hooks/suggest.ts 593 lines
1/**
2 * Pure helpers for the skill-suggestion mod.
3 *
4 * Nothing here touches `$`. The loader refuses a module that binds, passes or
5 * destructures it, so every engine call stays written out in register.ts and
6 * these functions move plain data only. That constraint also makes them
7 * unit-testable under `claude plugin test`.
8 *
9 * Design follows https://docs.typesafe.ai/cookbooks/skill_suggestion.md :
10 * one request ranks the whole roster and asks whether the turn needs a skill
11 * at all, and the winner becomes one advisory line of context.
12 */
13
14/** One skill as the router sees it: the two frontmatter fields that matter. */
15export type SkillEntry = { name: string; description: string };
16
17/** What the router decided for one turn. */
18export type Suggestion = { name: string; confidence: number; needsSkill: number };
19
20/** The option meaning "none of these fit". */
21export const NONE = '__none__';
22
23/** Key under which the session's roster is cached in `$.store`. */
24export const ROSTER_KEY = 'skill-roster.v1';
25
26export const DEFAULTS = {
27  endpoint: 'https://api.typesafe.ai/v1/systemone',
28  model: 'jev-latest',
29  /** Below this Choice confidence, suggest nothing. Tune on your own turns. */
30  minConfidence: 0.6,
31  /** Below this Noul, the turn does not want a procedure at all. */
32  minNeedsSkill: 0.5,
33  // A cap that is too low drops real skills without saying so. At 120 this
34  // list lost `resolving-merge-conflicts`, and the router then said nothing
35  // about merge conflicts. This sits well above the 194 skills seen here.
36  // Raise it if a scan reports hitting it.
37  maxSkills: 500,
38  // 80 measured better than 220. It used 37% fewer input tokens and gave
39  // higher confidence on the same answer. The rest of a description was noise.
40  descriptionChars: 80,
41  /** Give up before the engine's uncatchable 10 s dispatch budget bites. */
42  budgetMs: 4000,
43};
44
45/**
46 * Lists every SKILL.md the session could load, with its frontmatter.
47 *
48 * One process call rather than ~50 `$.fs.read` calls, because `$.fs` is fenced
49 * to the project by a string compare and cannot reach `~/.claude`.
50 */
51export const SCAN_COMMAND: readonly string[] = [
52  'sh',
53  '-c',
54  // -L follows symlinks. Most of ~/.claude/skills is symlinked, and without
55  // it the scan saw 4 of 46 entries and could never suggest those skills.
56  'find -L "$HOME/.claude/skills" "$HOME/.claude/plugins" -type f -name SKILL.md 2>/dev/null ' +
57    '| head -2000 ' +
58    '| while IFS= read -r f; do printf "===SKILL===%s\\n" "$f"; sed -n "1,60p" "$f"; done',
59];
60
61const asRecord = (v: unknown): Record<string, unknown> | null =>
62  typeof v === 'object' && v !== null && !Array.isArray(v) ? (v as Record<string, unknown>) : null;
63
64const asNumber = (v: unknown): number | null =>
65  typeof v === 'number' && Number.isFinite(v) ? v : null;
66
67const asString = (v: unknown): string | null => (typeof v === 'string' ? v : null);
68
69/** Narrows a value recovered from `$.store`, whose contents are untyped. */
70export function isSkillEntry(v: unknown): v is SkillEntry {
71  const r = asRecord(v);
72  return r !== null && typeof r.name === 'string' && typeof r.description === 'string';
73}
74
75/**
76 * Flatten to one line and cap the length, ellipsis included.
77 *
78 * The ellipsis counts against the budget. Slicing to `max - 1` and then adding
79 * three characters returns `max + 2`, which overspends the token budget on
80 * every skill in the roster at once.
81 */
82function truncate(s: string, max: number): string {
83  const flat = s.replace(/\s+/g, ' ').trim();
84  if (flat.length <= max) return flat;
85  return `${flat.slice(0, Math.max(0, max - 3)).trimEnd()}...`;
86}
87
88/**
89 * Read the `key: value` pairs out of a leading YAML frontmatter block.
90 *
91 * Only `name` and `description` matter, and the hooks environment has no YAML
92 * library, so this reads the block directly. Folded scalars (`description: >`)
93 * and plain multi-line continuations both join into one line.
94 *
95 * @param lines The file's opening lines, frontmatter fence included.
96 * @returns The top-level scalar keys found, values flattened to one line.
97 */
98export function readFrontmatter(lines: readonly string[]): Record<string, string> {
99  const out: Record<string, string> = {};
100  let i = 0;
101  while (i < lines.length && (lines[i] ?? '').trim() === '') i++;
102  if ((lines[i] ?? '').trim() !== '---') return out;
103  i++;
104
105  let key = '';
106  let buf: string[] = [];
107  const flush = () => {
108    if (key) out[key] = buf.join(' ').trim();
109    key = '';
110    buf = [];
111  };
112
113  for (; i < lines.length; i++) {
114    const line = lines[i] ?? '';
115    if (line.trim() === '---') break;
116    const m = /^([A-Za-z_][\w-]*):[ \t]*(.*)$/.exec(line);
117    if (m && !/^[ \t]/.test(line)) {
118      flush();
119      key = m[1] ?? '';
120      const value = (m[2] ?? '').trim();
121      // `>` and `|` introduce a block scalar; the text is on the lines below.
122      if (value && !/^[>|][-+]?$/.test(value)) buf.push(value);
123    } else if (key && line.trim()) {
124      buf.push(line.trim());
125    }
126  }
127  flush();
128  return out;
129}
130
131/** The directory name that holds a SKILL.md, used when frontmatter has no name. */
132function dirName(path: string): string {
133  const parts = path.split('/').filter(Boolean);
134  return parts.length >= 2 ? (parts[parts.length - 2] ?? '') : '';
135}
136
137/**
138 * Turn the scan command's stdout into a deduplicated roster.
139 *
140 * @param stdout Output of SCAN_COMMAND: `===SKILL===<path>` then that file's head.
141 * @param limits Caps on roster size and description length.
142 * @returns One entry per uniquely named skill that declares a description.
143 */
144export function parseRoster(
145  stdout: string,
146  limits: { maxSkills: number; descriptionChars: number } = DEFAULTS
147): SkillEntry[] {
148  const out: SkillEntry[] = [];
149  const seen = new Set<string>();
150
151  for (const chunk of stdout.split('===SKILL===').slice(1)) {
152    const lines = chunk.split('\n');
153    const path = (lines.shift() ?? '').trim();
154    const fm = readFrontmatter(lines);
155    const name = (fm.name ?? '').trim() || dirName(path);
156    const description = (fm.description ?? '').trim();
157    // A skill with no description tells the router nothing, so it cannot be ranked.
158    if (!name || !description || seen.has(name)) continue;
159    seen.add(name);
160    out.push({ name, description: truncate(description, limits.descriptionChars) });
161    if (out.length >= limits.maxSkills) break;
162  }
163  return out;
164}
165
166/**
167 * Build the TypeSafe v1 request that ranks the roster for one turn.
168 *
169 * The two questions are independent, so they ride one request and run in
170 * parallel: the Choice picks a skill, the Noul decides whether the turn wants
171 * a procedure at all. Neither can see the other's answer, which is why the
172 * Choice also carries its own no-match option.
173 */
174export function buildRequest(
175  roster: readonly SkillEntry[],
176  prompt: string,
177  model: string = DEFAULTS.model
178): Record<string, unknown> {
179  const criteria: Record<string, string> = {};
180  for (const s of roster) criteria[s.name] = s.description;
181  criteria[NONE] =
182    'None of the skills above fits this request, or the request is ordinary conversation, ' +
183    'a question about something already on screen, or a small direct edit that needs no procedure.';
184
185  return {
186    model,
187    state: { user_request: prompt },
188    questions: {
189      skill: {
190        type: 'choice',
191        instructions:
192          'Which single skill should the agent read before answering `state.user_request`? ' +
193          `Judge each skill only by what its description says it is for. Answer ${NONE} when none of them fits.`,
194        criteria,
195      },
196      needs_skill: {
197        type: 'noul',
198        instructions:
199          'Does answering `state.user_request` call for following a written procedure, ' +
200          'rather than just replying or making one direct edit?',
201        criteria: {
202          true: 'The request asks for work with steps worth following: a workflow, a review, a migration, a build, a debugging loop.',
203          false: 'The request is conversation, a factual question, or a small direct change that needs no procedure.',
204        },
205      },
206    },
207  };
208}
209
210/**
211 * Read a winner out of a TypeSafe v1 response, or decide to stay quiet.
212 *
213 * Returns null on any doubt: an unparseable body, the no-match option, a turn
214 * the Noul says wants no procedure, or a Choice below the confidence floor.
215 * Staying quiet costs one unassisted turn; a wrong suggestion costs a wrong
216 * skill load, so the asymmetry favours silence.
217 */
218export function pickWinner(
219  payload: unknown,
220  opts: { minConfidence: number; minNeedsSkill: number } = DEFAULTS
221): Suggestion | null {
222  const root = asRecord(payload);
223  const answers = root && asRecord(root.answers);
224  if (!answers) return null;
225
226  const skill = asRecord(answers.skill);
227  const needs = asRecord(answers.needs_skill);
228  if (!skill || !needs) return null;
229
230  const choice = asString(skill.choice);
231  const confidence = asNumber(skill.confidence);
232  const needsSkill = asNumber(needs.noul);
233  if (choice === null || confidence === null || needsSkill === null) return null;
234
235  if (choice === NONE) return null;
236  if (needsSkill < opts.minNeedsSkill) return null;
237  if (confidence < opts.minConfidence) return null;
238  return { name: choice, confidence, needsSkill };
239}
240
241/**
242 * Wrap a winner as one advisory context block.
243 *
244 * The roster the agent already has is left untouched, so prefix caching over
245 * it still holds; this only says which entry to look at first.
246 */
247export function contextBlock(s: Suggestion): string {
248  return [
249    '<skill_relevance>',
250    `A skill router ranked the installed skills against this request. Its top pick: ${s.name}`,
251    'This is a hint from a separate small model, not an instruction.',
252    'Ignore it if it does not fit what was actually asked.',
253    '</skill_relevance>',
254  ].join('\n');
255}
256
257// ---------------------------------------------------------------------------
258// Decision routing: answer an AskUserQuestion with Jev instead of the human.
259// ---------------------------------------------------------------------------
260
261/** One option as the AskUserQuestion tool poses it. */
262export type AskOption = { label: string; description?: string };
263
264/** One question as the AskUserQuestion tool poses it. */
265export type AskQuestion = { question: string; options: AskOption[]; multiSelect?: boolean };
266
267/** A decision Jev was confident enough to make. */
268export type Decision = { label: string; confidence: number };
269
270/** Confidence floor for answering instead of the human. */
271export const DECISION_DEFAULTS = { minConfidence: 0.75 };
272
273/**
274 * Read a decision out of a TypeSafe response, or defer to the human.
275 *
276 * Returns null below the confidence floor. Deferring costs one dialog; a
277 * confident wrong answer silently removes the human from their own decision,
278 * so the floor sits higher than the skill router's.
279 */
280export function pickDecision(
281  payload: unknown,
282  q: AskQuestion,
283  minConfidence: number = DECISION_DEFAULTS.minConfidence
284): Decision | null {
285  const root = asRecord(payload);
286  const answers = root && asRecord(root.answers);
287  const pick = answers && asRecord(answers.pick);
288  if (!pick) return null;
289
290  const label = asString(pick.choice);
291  const confidence = asNumber(pick.confidence);
292  if (label === null || confidence === null) return null;
293  // Guard against a label the model invented: it must be one we offered.
294  if (!q.options.some((o) => o.label === label)) return null;
295  if (confidence < minConfidence) return null;
296  return { label, confidence };
297}
298
299/** How the answered call reads back to the agent. */
300export function decisionNote(q: AskQuestion, d: Decision): string {
301  return (
302    `Answered by the TypeSafe decision router instead of the user. ` +
303    `Question: "${q.question}" Chosen option: "${d.label}" (confidence ${d.confidence.toFixed(2)}). ` +
304    `Proceed with that option. Ask the user directly only if this turns out not to fit.`
305  );
306}
307
308// ---------------------------------------------------------------------------
309// Show-your-work mode: draw Jev's distribution over the dialog.
310// ---------------------------------------------------------------------------
311
312/**
313 * How many option rows the panel draws.
314 *
315 * The engine refuses a panel that adds too much around the dialog. Measured on
316 * 2.1.274: four rows draw, six are refused, with the title, footer, wrapper and
317 * core node also counting against the budget. Four leaves a margin, and a
318 * question with more options than that is rare.
319 */
320export const MAX_PANEL_ROWS = 4;
321
322/** Slot holding the last decision, for the render hook to read. */
323export const DECISION_KEY = 'last-decision.v1';
324
325/** A decision plus the full distribution behind it. */
326export type DecisionView = {
327  question: string;
328  probabilities: Record<string, number>;
329  choice: string;
330  confidence: number;
331  /** True when the router would have answered without asking. */
332  wouldAnswer: boolean;
333};
334
335/**
336 * Draw a proportional bar.
337 *
338 * @param p Probability from 0 to 1. Values outside that range are clamped.
339 * @param width Total cells the bar occupies.
340 */
341export function bar(p: number, width: number): string {
342  const safe = Math.max(0, Math.min(1, Number.isFinite(p) ? p : 0));
343  const filled = Math.round(safe * width);
344  return '█'.repeat(filled) + '░'.repeat(Math.max(width - filled, 0));
345}
346
347/** Options ordered by probability, highest first, for display. */
348export function rankOptions(view: DecisionView, q: AskQuestion): Array<{ label: string; p: number }> {
349  return q.options
350    .map((o) => ({ label: o.label, p: view.probabilities[o.label] ?? 0 }))
351    .sort((a, b) => b.p - a.p);
352}
353
354/** Narrows a DecisionView recovered from `$.store`. */
355export function isDecisionView(v: unknown): v is DecisionView {
356  const r = asRecord(v);
357  return (
358    r !== null &&
359    typeof r.question === 'string' &&
360    typeof r.choice === 'string' &&
361    typeof r.confidence === 'number' &&
362    asRecord(r.probabilities) !== null
363  );
364}
365
366// ---------------------------------------------------------------------------
367// Batched dialogs: several questions in one AskUserQuestion call.
368// ---------------------------------------------------------------------------
369
370/** A plain-English name for a value, for diagnostics. */
371const describe = (v: unknown): string =>
372  v === null ? 'null' : Array.isArray(v) ? 'an array' : typeof v;
373
374/**
375 * What a scan of a dialog's steps found, and why anything was dropped.
376 *
377 * A bare list of routable questions cannot say why a dialog produced none, so
378 * a skipped step reports itself. "step 2 is multi-select" and "questions is
379 * undefined" are different faults, and only one of them is worth changing.
380 */
381export type QuestionScan = {
382  /** Steps the router can turn into a Choice, in the order the dialog draws them. */
383  routable: AskQuestion[];
384  /** One line per dropped step, naming which step and what was wrong. */
385  skipped: string[];
386};
387
388/**
389 * Read every routable question out of an AskUserQuestion call, and report the
390 * rest.
391 *
392 * A dialog may carry several steps. Each single-select step with two or more
393 * options is its own Choice. Anything else is dropped with a reason rather
394 * than approximated.
395 *
396 * @param input The tool call's `questions` value, typed `unknown[]`.
397 */
398export function scanAskQuestions(input: unknown): QuestionScan {
399  if (!Array.isArray(input)) {
400    return { routable: [], skipped: [`questions is ${describe(input)}, not an array`] };
401  }
402  if (input.length === 0) {
403    return { routable: [], skipped: ['the call carried no questions'] };
404  }
405
406  const routable: AskQuestion[] = [];
407  const skipped: string[] = [];
408
409  for (let i = 0; i < input.length; i++) {
410    const at = `step ${i + 1}`;
411    const raw = input[i];
412    const q = asRecord(raw);
413    if (!q) {
414      skipped.push(`${at} is ${describe(raw)}, not an object`);
415      continue;
416    }
417    if (q.multiSelect === true) {
418      skipped.push(`${at} is multi-select, which is not a Choice`);
419      continue;
420    }
421    const question = asString(q.question);
422    if (!question) {
423      skipped.push(`${at} has no question text`);
424      continue;
425    }
426    if (!Array.isArray(q.options)) {
427      skipped.push(`${at} has ${describe(q.options)} where its options should be`);
428      continue;
429    }
430    if (q.options.length < 2) {
431      skipped.push(`${at} offers ${q.options.length}, so there is nothing to choose between`);
432      continue;
433    }
434
435    const options: AskOption[] = [];
436    let fault = '';
437    for (let j = 0; j < q.options.length; j++) {
438      const o = asRecord(q.options[j]);
439      const label = o && asString(o.label);
440      if (!label) {
441        // A label that is not a string cannot be matched against an answer.
442        fault = `${at} option ${j + 1} has no string label`;
443        break;
444      }
445      options.push({ label, description: (o && asString(o.description)) || undefined });
446    }
447    if (fault) {
448      skipped.push(fault);
449      continue;
450    }
451    routable.push({ question, options });
452  }
453  return { routable, skipped };
454}
455
456/** The routable steps alone, for callers that do not report faults. */
457export function readAskQuestions(input: unknown): AskQuestion[] {
458  return scanAskQuestions(input).routable;
459}
460
461/** The id a question carries in a batched request. */
462const questionId = (index: number) => `q${index}`;
463
464/**
465 * Build one request covering every question in a batched dialog.
466 *
467 * The questions are independent, so they ride one request and Jev answers them
468 * in parallel. One call costs what a single question costs in latency, and the
469 * shared situation is sent once rather than per question.
470 */
471export function buildBatchRequest(
472  questions: readonly AskQuestion[],
473  situation: string,
474  model: string = DEFAULTS.model
475): Record<string, unknown> {
476  const asked: Record<string, unknown> = {};
477  const state: Record<string, unknown> = { situation };
478
479  questions.forEach((q, i) => {
480    const id = questionId(i);
481    const criteria: Record<string, string> = {};
482    for (const o of q.options) criteria[o.label] = o.description ?? o.label;
483    state[id] = q.question;
484    asked[id] = {
485      type: 'choice',
486      instructions:
487        `${q.question} Decide from \`state.${id}\` and \`state.situation\`, ` +
488        'judging each option only by its description.',
489      criteria,
490    };
491  });
492
493  return { model, state, questions: asked };
494}
495
496/**
497 * Read one view per question out of a batched response.
498 *
499 * A question Jev could not answer is skipped rather than guessed at, so the
500 * result may be shorter than the input. Order follows the questions given.
501 */
502export function readDecisions(
503  payload: unknown,
504  questions: readonly AskQuestion[],
505  minConfidence: number = DECISION_DEFAULTS.minConfidence
506): DecisionView[] {
507  const root = asRecord(payload);
508  const answers = root && asRecord(root.answers);
509  if (!answers) return [];
510
511  const out: DecisionView[] = [];
512  questions.forEach((q, i) => {
513    const pick = asRecord(answers[questionId(i)]);
514    if (!pick) return;
515
516    const choice = asString(pick.choice);
517    const confidence = asNumber(pick.confidence);
518    if (choice === null || confidence === null) return;
519
520    const probabilities: Record<string, number> = {};
521    const raw = asRecord(pick.probabilities);
522    if (raw) {
523      for (const [label, value] of Object.entries(raw)) {
524        const n = asNumber(value);
525        if (n !== null) probabilities[label] = n;
526      }
527    }
528    const offered = q.options.some((o) => o.label === choice);
529    out.push({
530      question: q.question,
531      probabilities,
532      choice,
533      confidence,
534      wouldAnswer: offered && confidence >= minConfidence,
535    });
536  });
537  return out;
538}
539
540/** Narrows a list of views recovered from `$.store`. */
541export function isDecisionViewList(v: unknown): v is DecisionView[] {
542  return Array.isArray(v) && v.length > 0 && v.every(isDecisionView);
543}
544
545/** One transcript line describing what Jev said about one question. */
546export function decisionLine(view: DecisionView, q: AskQuestion, prefix = ''): string {
547  const spread = rankOptions(view, q)
548    .map((r) => `${r.label} ${r.p.toFixed(2)}`)
549    .join(', ');
550  return view.wouldAnswer
551    ? `typesafe-mod: ${prefix}Jev picks ${view.choice} (${spread}), confidence ${view.confidence.toFixed(2)}`
552    : `typesafe-mod: ${prefix}Jev leans ${view.choice} (${spread}), but confidence ${view.confidence.toFixed(2)} is under the floor, so this one is yours`;
553}
554
555/**
556 * One compact row summarising what Jev said about one question.
557 *
558 * A batched dialog draws one of these per question. The render event does not
559 * say which step the dialog is currently showing, so drawing one question's
560 * option bars would pin the panel to that question while the reader pages
561 * through the others. A row per question is true whatever step is on screen.
562 *
563 * @param view The decision for this question.
564 * @param barWidth Cells for the winner's probability bar.
565 * @param labelWidth Cells the label is padded to, so the columns line up.
566 */
567export function summaryRow(view: DecisionView, barWidth: number, labelWidth: number): string {
568  const top = Object.values(view.probabilities).reduce((a, b) => Math.max(a, b), 0);
569  const label =
570    view.choice.length > labelWidth
571      ? `${view.choice.slice(0, Math.max(0, labelWidth - 1))}…`
572      : view.choice.padEnd(labelWidth);
573  const mark = view.wouldAnswer ? ' ✓' : '';
574  return `${bar(top, barWidth)} ${top.toFixed(2)}  ${label}  ${view.confidence.toFixed(2)}${mark}`;
575}
576
577/**
578 * Pair each question in a dialog with its decision, dropping the unanswered.
579 *
580 * Order follows the dialog, so row one is step one.
581 */
582export function pairDecisions(
583  questions: readonly AskQuestion[],
584  views: readonly DecisionView[]
585): Array<{ question: AskQuestion; view: DecisionView }> {
586  const out: Array<{ question: AskQuestion; view: DecisionView }> = [];
587  for (const question of questions) {
588    const view = views.find((v) => v.question === question.question);
589    if (view) out.push({ question, view });
590  }
591  return out;
592}
593