SLOPSHOPPER

jeffort

Per-turn effort routing with TypeSafe's Jev, applied per request at turn.step. Off until /jev on.

newcommandstatuspromptnetworkagents
v0.1.0MITupdated 2026-10-05vntrungld/jeffort
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · jeffort
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /jev ⎿ jeffort: Jev routing is ON, but no TypeSafe key was found, so prompts pass through unrouted. Set it in /plugin configur ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ jeffort: jev
README

jeffort

A Claude Code mod that asks Jev (TypeSafe) how much effort each prompt needs, then sends every request of that turn at that effort. The main loop's model is never changed. Off by default; turn it on with /jev on.

This is a fork of jjjjjjjjjjjjjjjjacob/jev-router (MIT), taken from commit 50d7e40 on 2026-09-27. The questions sent to Jev (hooks/lib/questions.ts), the level policy (hooks/lib/policy.ts) and the eval set are kept from upstream. The parts that apply effort and filter data were rewritten.

How it differs from upstream

upstream jev-routerjeffort
How effort is appliedTells Claude to load a jev-<level> skill and blocks tool calls until it doesA turn.step mod writes effort into each request of the turn
Turns that call no toolsMay run at the session level (skill skipped)Still routed
Turns from notifications, peers, pluginsRouted like ordinary promptsSkipped; only prompts you type are routed
Data sent to TypeSafePrompt (pasted blocks removed), head/tail truncatedSame as upstream, plus masking of code blocks, secrets, URLs, e-mails, IPs and absolute paths
Highest effortmaxxhigh (configurable)
CacheTrusts the docsWatches real cache_read/cache_creation, warns and then pauses when an effort change loses the cache
SubagentsInherit the turn's effort; model routed by judgment/delegatedKeep their own effort; model routed as upstream, via agent.spawn
RuntimeBun/Node, shell scripts, tool-call hooksRuns inside Claude Code's engine; no Bun/Node needed

Install

Requires Claude Code 2.1.289 or later (the tested version) and a TypeSafe API key.

# 1. Add the vntrungld marketplace and install (no clone needed):
claude plugin marketplace add vntrungld/jeffort
claude plugin install jeffort@vntrungld

# 2. Key: store it in the keychain via /plugin configure (step 3), or set the env var
export TYPESAFE_API_KEY=...

The vntrungld marketplace lists both jeffort and tightlip. This repo and the tightlip repo carry the same catalog, so adding either one (vntrungld/jeffort or vntrungld/tightlip) gives you both plugins under @vntrungld; adding the other later just replaces it with the same list. Update with claude plugin marketplace update vntrungld and claude plugin update jeffort@vntrungld.

To work on the code, clone it and load it for one session instead: claude --plugin-dir <path-to-clone>.

  1. In Claude Code: /plugin configure jeffort@vntrungld to enter the key and adjust options. The non-sensitive options are also in /config.

If your organization sets allowManagedModsOnly or allowManagedHooksOnly in managed settings, the mod will not load. If your Claude Code is older and reports that function hooks are disabled, set CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1.

Usage

/jev on       turn on for this session
/jev status   show state and the last 8 decisions
/jev off      turn off

Each routed turn shows a dim line such as jev → high · verification decides success (conf 0.72). The status line shows jev · <level> while routing is on. The next turn goes back to the session level unless Jev picks another level.

Options

OptionDefaultMeaning
typesafe_api_key(empty)The key. When empty, reads TYPESAFE_API_KEY or JEV_API_KEY
enabled_by_defaultfalseStart every session with routing on
max_effortxhighHighest level the router may pick
route_subagentstrueRoute subagent models
judgment_model / delegated_modelopus / sonnetModel for review/debug/design subagents, and for subagents doing delegated work
redacttrueMask data before sending (see below)
cache_safe_onlytrueOnly change effort on Opus 5.5, Sonnet 5.5, Fable 5.1
cache_guardtrueWarn on the first cache miss, pause on the second
timeout_ms2500How long to wait for Jev before the turn runs at the session level
show_decisionstrueShow the jev → … line
base_urlhttps://api.typesafe.aihttps only, or http to localhost
jev_modeljev-1.13.0Pinned because the policy thresholds were tuned on this version

What gets sent

Only while routing is on, and only for prompts you type (no slash commands, notifications, or messages from other sessions). Each turn makes one request to api.typesafe.ai/v1/systemone, containing:

  • the current prompt after filtering, at most the first 3,000 and last 1,000 characters;
  • the first 1,500 characters of the previous prompt (filtered), so Jev recognizes "ok, go ahead";
  • for subagents: the Agent call's prompt and description, also filtered.

The filter replaces:

ContentWith
pasted block[pasted text: N chars]
fenced ``` code block[code block: N lines]
inline code of 40+ characters[code]
PEM keys, sk-…, ghp_…, xox…, AKIA…, JWTs, password=…, long hex or base64<secret>
URLs, DSNs (postgres://…)<url>
e-mail<email>
IPv4<ip>
absolute paths, ~/…, C:\…<path>

Kept: short function names in backticks (getUser) and relative paths (app/Http/Kernel.php), since they tell Jev what kind of work is being asked for. The filter is regex-based, not DLP: unusual secret formats or lowercase customer names can slip through. For company repositories, check your internal policy before turning it on.

Locally, decisions are kept in the mod's $.store (at most 200 entries, each the first 80 characters of the filtered prompt) for tuning thresholds later.

Test results

  • bun test ./test: 79 unit tests for the policy (ported from upstream), the filter, the Jev client and the cache guard.
  • claude plugin test .: 18 tests running the mod on Claude Code's real engine (per-turn routing, subagents, notifications, timeouts, base URL, filtering, cache guard, /jev). Three key behaviours were deliberately broken (skipping subagent steps, filtering, routing only cache-safe models) to confirm the tests catch each one.
  • claude plugin validate . and tsc are clean.
  • Real run in a headless Claude Code 2.1.289 session (Sonnet 5.5), against a mock Jev server on localhost:
  • The transcript records effort: high / low correctly for each routed turn, so the effort really reaches the API.
  • In the test environment (a cloud container, requests going through an ANTHROPIC_BASE_URL proxy), each effort change kept the cache for the system prompt and tools (~13k tokens) but rewrote the conversation part (~4.5k tokens). Claude Code's built-in /effort command costs exactly the same, so this cost comes from the environment, not the mod. The cache guard caught it and paused routing after the second time.

Not yet verified:

  1. TypeSafe has not been called for real (no key), so the 97% accuracy is upstream's number, measured on unfiltered English prompts. The filter changes none of upstream's 42 fixtures, since they contain no code or sensitive data; your real prompts will differ. Run TYPESAFE_API_KEY=… bun eval/run.ts and --no-redact to compare, then add your own real prompts to eval/fixtures.local.jsonl.
  2. Not tested on a machine with a subscription calling the API directly. Per the docs, effort changes keep the cache there; check the Prompt cache (main) line in /usage, or let the cache guard check for you.
  3. Subagent routing has not been tested in a real session, only through the engine test suite.

Development

bun install
bun test ./test            # unit tests
claude plugin test .       # mod tests on the engine
claude plugin validate .
tsc -p tsconfig.bun.json && tsc -p .   # tsc -p . needs .claude-plugin/types, generated by the engine when loaded with --plugin-dir
TYPESAFE_API_KEY=… bun eval/run.ts     # live eval; --no-redact; --replay

hooks/lib/questions.ts is kept word for word, because Jev reads the questions literally. Changing the wording means rerunning the live eval.

License

MIT. See LICENSE: original copyright by jjjjjjjjjjjjjjjjacob, modifications by this fork.

Source 6 files
hooks/register.ts 350 lines
1// jeffort: a mod that asks TypeSafe's Jev how much effort each prompt needs and sends
2// that turn's model requests at that effort. Forked from jjjjjjjjjjjjjjjjacob/jev-router
3// (MIT): the questions, policy and eval are theirs; how the level is applied is not.
4//
5// Upstream had no way to set effort from a hook, so it told Claude to load a jev-<level>
6// skill and gated tool calls until it did; a turn that answered without tools skipped it.
7// A mod can rewrite `effort` on every `turn.step`, so here:
8//
9//   prompt.submit  marks prompts the person typed (not notifications, peers or plugins)
10//   turn.start     asks Jev once per marked prompt, before the turn's first request
11//   turn.step      sends each main-loop request of that turn at the routed level
12//   turn.complete  forgets the turn; the next prompt starts at the session level again
13//   agent.spawn    picks a subagent's model from the kind of work it is given
14//
15// Everything fails open: no key, a timeout, an error or an unsure answer leaves the turn at
16// the session's own effort. Subagents keep their own effort; only their model is routed.
17
18import type { EngineInterface, Register } from "claude-code";
19import { CacheGuard, isCacheSafeModel } from "./lib/cache-guard.ts";
20import { buildRequest, DEFAULT_BASE_URL, DEFAULT_MODEL, DEFAULT_TIMEOUT_MS, isAllowedBaseUrl, parseResponse, type JevAnswers, type Question } from "./lib/jev.ts";
21import { capLevel, decideEffort, decideSubagentModel, isLevel, readEffortAnswers, SKIPPED_SUBAGENT_TYPES, type Level } from "./lib/policy.ts";
22import { effortQuestions, subagentQuestions, type EffortState, type SubagentState } from "./lib/questions.ts";
23import { preview, promptKey, redact, shapePrompt, truncate } from "./lib/redact.ts";
24
25const PREVIOUS_REQUEST_CHARS = 1500;
26const MAX_MARKED_PROMPTS = 20;
27const RECENT_DECISIONS = 8;
28const STORED_DECISIONS = 200;
29const ROUTED_ORIGINS = new Set(["composer", "bridge", "sdk"]);
30
31type Decision = {
32  at: number;
33  kind: "effort" | "subagent";
34  preview: string;
35  level?: Level | null;
36  sessionLevel?: string | null;
37  applied?: boolean;
38  model?: string | null;
39  reasons: string[];
40  confidence?: number;
41  latencyMs?: number;
42  redactions?: number;
43};
44
45type Route = { level: Level; decision: Decision; logged: boolean };
46
47type Config = {
48  apiKey: string | undefined;
49  enabledByDefault: boolean;
50  maxEffort: Level;
51  routeSubagents: boolean;
52  judgmentModel: string;
53  delegatedModel: string;
54  redact: boolean;
55  cacheSafeOnly: boolean;
56  cacheGuard: boolean;
57  timeoutMs: number;
58  showDecisions: boolean;
59  baseUrl: string;
60  jevModel: string;
61};
62
63// Session state lives in the module: a reload of the plugin starts it over, which only
64// happens on /reload-plugins for an installed plugin. `register` fills it in.
65let config: Config;
66const session = {
67  enabled: false,
68  pausedReason: null as string | null,
69  lastLevel: null as Level | null,
70  lastRequest: undefined as string | undefined,
71  marked: new Map<string, number>(),
72  routed: new Map<string, Route>(),
73  recent: [] as Decision[],
74  guard: new CacheGuard(),
75};
76
77export const register: Register = (on, options) => {
78  config = {
79    apiKey: text(options.typesafe_api_key),
80    enabledByDefault: options.enabled_by_default === true,
81    maxEffort: (isLevel(options.max_effort) ? options.max_effort : "xhigh") as Level,
82    routeSubagents: options.route_subagents !== false,
83    judgmentModel: text(options.judgment_model) ?? "opus",
84    delegatedModel: text(options.delegated_model) ?? "sonnet",
85    redact: options.redact !== false,
86    cacheSafeOnly: options.cache_safe_only !== false,
87    cacheGuard: options.cache_guard !== false,
88    timeoutMs: typeof options.timeout_ms === "number" ? options.timeout_ms : DEFAULT_TIMEOUT_MS,
89    showDecisions: options.show_decisions !== false,
90    baseUrl: text(options.base_url) ?? DEFAULT_BASE_URL,
91    jevModel: text(options.jev_model) ?? DEFAULT_MODEL,
92  };
93  session.enabled = config.enabledByDefault;
94
95  on("session.start", async ($, e, next) => {
96    await $.command.register({
97      name: "jev",
98      description: "Jev effort routing for this session: on, off, or status",
99      argumentHint: "[on|off|status]",
100    });
101    showStatus($);
102    return next(e);
103  });
104
105  on("command.run", { command: "jev" }, async ($, e) => {
106    const arg = (e.args.trim().split(/\s+/)[0] || "on").toLowerCase();
107    if (arg === "on") {
108      setEnabled($, true);
109      if (!(await apiKey($))) {
110        return { text: "Jev routing is ON, but no TypeSafe key was found, so prompts pass through unrouted. Set it in /plugin configure, or export TYPESAFE_API_KEY." };
111      }
112      if (!isAllowedBaseUrl(config.baseUrl)) {
113        return { text: `Jev routing is ON, but base_url ${config.baseUrl} is not https (or http to localhost), so nothing is sent.` };
114      }
115      const subagents = config.routeSubagents ? ` Subagents: ${config.judgmentModel} for judgment, ${config.delegatedModel} for delegated work.` : "";
116      return { text: `Jev routing is ON for this session (up to ${config.maxEffort}).${subagents} /jev off to stop.` };
117    }
118    if (arg === "off") {
119      setEnabled($, false);
120      session.pausedReason = null;
121      return { text: "Jev routing is OFF. Turns run at the session's effort." };
122    }
123    if (arg === "status") return { text: await statusText($) };
124    return { text: `Unknown /jev argument "${arg}". Use /jev on, /jev off or /jev status.` };
125  });
126
127  // Only prompts the person sent start a routed turn. One typed while a turn runs is folded
128  // into that turn (e.turnId set) and is left alone, as are notifications and peers.
129  on("prompt.submit", async ($, e, next) => {
130    if (session.enabled && !e.turnId && ROUTED_ORIGINS.has(e.origin.kind)) {
131      const key = promptKey(e.text);
132      if (key) {
133        session.marked.set(key, (session.marked.get(key) ?? 0) + 1);
134        while (session.marked.size > MAX_MARKED_PROMPTS) session.marked.delete(session.marked.keys().next().value as string);
135      }
136    }
137    return next(e);
138  });
139
140  on("turn.start", async ($, e, next) => {
141    const markedKey = promptKey(e.text);
142    const count = session.marked.get(markedKey);
143    if (!session.enabled || !count) return next(e);
144    if (count > 1) session.marked.set(markedKey, count - 1);
145    else session.marked.delete(markedKey);
146
147    const shaped = shapePrompt(e.text, { redact: config.redact });
148    if (shaped.skip) return next(e);
149    const key = await apiKey($);
150    if (!key) return next(e);
151
152    const state: EffortState = { current_request: shaped.text };
153    if (session.lastRequest) state.previous_request = session.lastRequest;
154    const result = await askJev($, key, state, effortQuestions());
155    const answers = result && readEffortAnswers(result.answers);
156    if (!result || !answers) return next(e);
157
158    const decided = decideEffort(answers, { previousLevel: session.lastLevel });
159    const level = decided.level ? capLevel(decided.level, config.maxEffort) : null;
160    const reasons = level && decided.level !== level ? [...decided.reasons, `capped at ${level}`] : decided.reasons;
161    session.lastLevel = level ?? session.lastLevel;
162    session.lastRequest = shaped.text.slice(0, PREVIOUS_REQUEST_CHARS);
163
164    const decision: Decision = {
165      at: await $.clock.now(),
166      kind: "effort",
167      preview: preview(shaped.text),
168      level,
169      reasons,
170      confidence: round(answers.confidence),
171      latencyMs: result.latencyMs,
172      redactions: shaped.redactions,
173    };
174    await remember($, decision);
175    if (level) session.routed.set(e.turnId, { level, decision, logged: false });
176    showStatus($);
177    return next(e);
178  });
179
180  on("turn.step", async function* ($, e, next) {
181    // Subagent loops keep their own effort: their steps carry agentId and their own turnId.
182    if (e.agentId) return yield* next(e);
183
184    const route = session.enabled ? session.routed.get(e.turnId) : undefined;
185    const current = e.effort;
186    const routable =
187      route !== undefined &&
188      typeof current === "string" &&
189      isLevel(current) &&
190      (!config.cacheSafeOnly || isCacheSafeModel(e.model));
191    const sent = routable && current !== route.level ? { ...e, effort: route.level } : e;
192
193    if (route && !route.logged) {
194      route.logged = true;
195      route.decision.applied = sent !== e;
196      route.decision.sessionLevel = typeof current === "string" ? current : null;
197      if (sent !== e && config.showDecisions) {
198        $.ui.log(`jev → ${route.level} · ${route.decision.reasons.join(", ")} (conf ${route.decision.confidence?.toFixed(2)})`);
199      }
200    }
201
202    const result = yield* next(sent);
203
204    if (session.enabled && config.cacheGuard && result.usage) {
205      const verdict = session.guard.observe({
206        at: await $.clock.now(),
207        model: e.model,
208        effort: sent.effort,
209        cacheRead: result.usage.cache_read_input_tokens,
210        cacheWrite: result.usage.cache_creation_input_tokens,
211      });
212      if (verdict.kind === "miss") {
213        $.ui.log(`jev: the prompt cache missed right after an effort change on ${e.model} (${verdict.written} tokens re-written). Routing pauses if it happens again.`);
214      } else if (verdict.kind === "pause") {
215        setEnabled($, false);
216        session.pausedReason = `cache missed ${verdict.misses}× after effort changes on ${e.model}`;
217        $.ui.log(`jev: paused for this session. Effort changes keep re-reading the prompt cache on ${e.model}. /jev on to resume.`);
218      }
219    }
220    return result;
221  });
222
223  on("turn.complete", async ($, e, next) => {
224    if (!e.agentId) session.routed.delete(e.turnId);
225    return next(e);
226  });
227
228  on("agent.spawn", async ($, e, next) => {
229    if (!session.enabled || !config.routeSubagents || e.fork || SKIPPED_SUBAGENT_TYPES.has(e.subagentType)) return next(e);
230    const task = e.prompt.trim();
231    if (!task) return next(e);
232    const key = await apiKey($);
233    if (!key) return next(e);
234
235    const description = config.redact ? redact(e.description).text : e.description;
236    const body = config.redact ? redact(task).text : task;
237    const state: SubagentState = { subagent_task: truncate(body) };
238    if (description) state.description = description;
239
240    const result = await askJev($, key, state, subagentQuestions());
241    if (!result) return next(e);
242    const answer = result.answers.work;
243    const decision = decideSubagentModel(answer, e.model, {
244      judgment: config.judgmentModel,
245      delegated: config.delegatedModel,
246    });
247    await remember($, {
248      at: await $.clock.now(),
249      kind: "subagent",
250      preview: preview(description || body),
251      model: decision.model,
252      reasons: [decision.reason],
253      confidence: answer?.type === "choice" ? round(answer.confidence) : undefined,
254      latencyMs: result.latencyMs,
255    });
256    if (!decision.model) return next(e);
257    if (config.showDecisions) $.ui.log(`jev → subagent on ${decision.model} (${decision.reason})`);
258    return next({ ...e, model: decision.model });
259  });
260};
261
262// --- helpers: top-level functions, the only places a hook may hand `$` to ---------------
263
264async function apiKey($: EngineInterface): Promise<string | null> {
265  if (config.apiKey) return config.apiKey;
266  return text(await $.env.get("TYPESAFE_API_KEY")) ?? text(await $.env.get("JEV_API_KEY")) ?? null;
267}
268
269// One request, a hard timeout, never throws: every failure is null.
270async function askJev(
271  $: EngineInterface,
272  key: string,
273  state: unknown,
274  questions: Record<string, Question>,
275): Promise<(JevAnswers & { latencyMs: number }) | null> {
276  if (!isAllowedBaseUrl(config.baseUrl)) return null;
277  const request = buildRequest(state, questions, { apiKey: key, model: config.jevModel, baseUrl: config.baseUrl });
278  const timer = new AbortController();
279  try {
280    const started = await $.clock.now();
281    const call = $.http.fetch(request.url, request.init).then(
282      (response) => parseResponse(response, questions, config.jevModel),
283      () => null,
284    );
285    const timeout = $.clock.sleep(config.timeoutMs, { signal: timer.signal }).then(
286      () => null,
287      () => null,
288    );
289    const answers = await Promise.race([call, timeout]);
290    if (!answers) return null;
291    return { ...answers, latencyMs: Math.round((await $.clock.now()) - started) };
292  } catch {
293    return null;
294  } finally {
295    timer.abort();
296  }
297}
298
299async function remember($: EngineInterface, decision: Decision): Promise<void> {
300  session.recent.push(decision);
301  if (session.recent.length > RECENT_DECISIONS) session.recent.shift();
302  try {
303    const stored = await $.store.get("decisions");
304    const list = Array.isArray(stored) ? stored : [];
305    list.push(decision);
306    await $.store.set("decisions", list.slice(-STORED_DECISIONS));
307  } catch {
308    // the log is for tuning only
309  }
310}
311
312function showStatus($: EngineInterface): void {
313  $.ui.status(session.enabled ? (session.lastLevel ? `jev · ${session.lastLevel}` : "jev") : undefined);
314}
315
316function setEnabled($: EngineInterface, value: boolean): void {
317  session.enabled = value;
318  if (value) {
319    session.pausedReason = null;
320    session.guard.reset();
321  } else {
322    session.routed.clear();
323    session.marked.clear();
324  }
325  showStatus($);
326}
327
328async function statusText($: EngineInterface): Promise<string> {
329  const key = (await apiKey($)) ? "key set" : "no key";
330  const state = session.enabled ? "ON" : session.pausedReason ? `PAUSED (${session.pausedReason})` : "OFF";
331  const head = `Jev routing is ${state} · ${key} · up to ${config.maxEffort} · redaction ${config.redact ? "on" : "off"} · cache-safe models only: ${config.cacheSafeOnly ? "yes" : "no"}`;
332  if (session.recent.length === 0) return `${head}\nNo decisions yet.`;
333  const lines = session.recent.map((d) => {
334    if (d.kind === "subagent") return `- subagent → ${d.model ?? "unchanged"} (${d.reasons.join(", ")}) · "${d.preview}"`;
335    const level = d.level ?? "unchanged";
336    const applied = d.applied === undefined ? "" : d.applied ? ` (session ${d.sessionLevel ?? "?"})` : " (not applied)";
337    const conf = d.confidence === undefined ? "" : ` · conf ${d.confidence.toFixed(2)}`;
338    return `- ${level}${applied} · ${d.reasons.join(", ")}${conf} · ${d.latencyMs ?? "?"} ms · "${d.preview}"`;
339  });
340  return `${head}\nLast ${session.recent.length}:\n${lines.join("\n")}`;
341}
342
343function text(value: unknown): string | undefined {
344  return typeof value === "string" && value.trim() ? value.trim() : undefined;
345}
346
347function round(value: number): number {
348  return Math.round(value * 100) / 100;
349}
350
hooks/lib/cache-guard.ts 76 lines
1// Watches whether changing effort between requests costs the prompt cache. On Opus 5.5,
2// Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription it should not; on other
3// models, providers or gateways each change re-reads the whole conversation. The guard turns
4// that suspicion into evidence from the API's own usage numbers and pauses routing when the
5// evidence shows up, instead of trusting a model-name allowlist alone.
6
7export type StepRecord = {
8  at: number; // ms since epoch, when the response finished
9  model: string;
10  effort: string | number | undefined;
11  cacheRead: number;
12  cacheWrite: number;
13};
14
15export type GuardVerdict =
16  | { kind: "ok" }
17  | { kind: "miss"; written: number } // an effort change coincided with a cache miss
18  | { kind: "pause"; written: number; misses: number };
19
20export const GUARD = {
21  // How much of the previous request's cached prefix must go missing to count. Small drops
22  // are noise; a real invalidation loses at least the conversation part of the prefix.
23  minLost: 2048,
24  // Past this, a miss is the cache TTL (5 minutes by default), not the effort change.
25  maxGapMs: 4 * 60 * 1000,
26  pauseAfterMisses: 2,
27};
28
29export class CacheGuard {
30  private last: StepRecord | null = null;
31  private misses = 0;
32
33  constructor(private readonly limits = GUARD) {}
34
35  // Feed every main-loop step, routed or not, in order.
36  observe(step: StepRecord): GuardVerdict {
37    const previous = this.last;
38    this.last = step;
39    if (!previous) return { kind: "ok" };
40    const effortChanged = previous.effort !== step.effort;
41    const sameModel = previous.model === step.model;
42    const recent = step.at - previous.at <= this.limits.maxGapMs;
43    // A miss is often partial: the system prompt and tools stay cached and only the messages
44    // are re-written (seen on a gateway with Sonnet 5.5, for /effort and this mod alike), so
45    // compare against everything the previous request left cached, not against zero.
46    const previousPrefix = previous.cacheRead + previous.cacheWrite;
47    const lost = previousPrefix - step.cacheRead;
48    const missed = lost >= this.limits.minLost && step.cacheWrite >= this.limits.minLost / 2;
49    if (!(effortChanged && sameModel && recent)) return { kind: "ok" };
50    if (!missed) {
51      // An effort change that kept the cache is evidence the setup keeps it: forget older misses.
52      this.misses = 0;
53      return { kind: "ok" };
54    }
55    this.misses++;
56    if (this.misses >= this.limits.pauseAfterMisses) {
57      return { kind: "pause", written: step.cacheWrite, misses: this.misses };
58    }
59    return { kind: "miss", written: step.cacheWrite };
60  }
61
62  reset(): void {
63    this.last = null;
64    this.misses = 0;
65  }
66}
67
68// Models whose per-request effort change keeps the cache, per the Claude Code prompt-caching
69// docs (2.1.260+). Provider-prefixed ids (Bedrock, Vertex) still match here; the guard above
70// is what catches those, since the docs say the cache is not kept there.
71const CACHE_SAFE_MODEL = /claude-(?:opus-5-5|sonnet-5-5|fable-5-1)\b/i;
72
73export function isCacheSafeModel(model: string): boolean {
74  return CACHE_SAFE_MODEL.test(model);
75}
76
hooks/lib/jev.ts 90 lines
1// TypeSafe System One wire format. Pure: no network here. The mod sends the request with
2// $.http.fetch and the eval harness with fetch, so both share one request and one parser.
3//
4// Adapted from jjjjjjjjjjjjjjjjacob/jev-router (MIT), skills/jev/scripts/lib/jev.ts.
5
6export const DEFAULT_BASE_URL = "https://api.typesafe.ai";
7export const DEFAULT_MODEL = "jev-1.13.0"; // pinned: policy thresholds are tuned against this version
8export const DEFAULT_TIMEOUT_MS = 2500;
9
10export type NoulQuestion = {
11  type: "noul";
12  instructions: unknown;
13  criteria?: { true?: unknown; false?: unknown };
14};
15export type ChoiceQuestion = { type: "choice"; instructions: unknown; criteria: Record<string, unknown> };
16export type ScoreQuestion = { type: "score"; instructions: unknown; criteria: unknown[] };
17export type Question = NoulQuestion | ChoiceQuestion | ScoreQuestion;
18
19export type NoulAnswer = { type: "noul"; noul: number };
20export type ChoiceAnswer = {
21  type: "choice";
22  choice: string;
23  probabilities: Record<string, number>;
24  confidence: number;
25};
26export type ScoreAnswer = {
27  type: "score";
28  score: number;
29  probabilities: Record<string, number>;
30  legend: Record<string, string>;
31  confidence: number;
32};
33export type Answer = NoulAnswer | ChoiceAnswer | ScoreAnswer;
34
35export type JevAnswers = { model: string; answers: Record<string, Answer> };
36
37export type JevRequest = {
38  url: string;
39  init: { method: "POST"; headers: Record<string, string>; body: string };
40};
41
42export function buildRequest(
43  state: unknown,
44  questions: Record<string, Question>,
45  options: { apiKey: string; model?: string; baseUrl?: string },
46): JevRequest {
47  const base = (options.baseUrl || DEFAULT_BASE_URL).replace(/\/+$/, "");
48  return {
49    url: `${base}/v1/systemone`,
50    init: {
51      method: "POST",
52      headers: { Authorization: `Bearer ${options.apiKey}`, "Content-Type": "application/json" },
53      body: JSON.stringify({ model: options.model || DEFAULT_MODEL, state, questions }),
54    },
55  };
56}
57
58// Never throws: a response that is not ok, not JSON, or missing any asked question is null.
59export function parseResponse(
60  response: { ok: boolean; text: string },
61  questions: Record<string, Question>,
62  requestedModel: string = DEFAULT_MODEL,
63): JevAnswers | null {
64  if (!response.ok) return null;
65  try {
66    const body = JSON.parse(response.text) as { model?: unknown; answers?: Record<string, Answer> };
67    if (!body || typeof body !== "object" || !body.answers || typeof body.answers !== "object") return null;
68    for (const id of Object.keys(questions)) {
69      if (!body.answers[id]) return null;
70    }
71    return { model: typeof body.model === "string" ? body.model : requestedModel, answers: body.answers };
72  } catch {
73    return null;
74  }
75}
76
77// Only https to a real host, or http to loopback (a local proxy). Anything else would send
78// the API key and the prompt somewhere unintended, so the mod refuses it and stays off.
79export function isAllowedBaseUrl(raw: string): boolean {
80  try {
81    const url = new URL(raw);
82    if (url.username || url.password) return false;
83    if (url.protocol === "https:") return url.hostname.length > 0;
84    if (url.protocol === "http:") return ["localhost", "127.0.0.1", "[::1]"].includes(url.hostname);
85    return false;
86  } catch {
87    return false;
88  }
89}
90
hooks/lib/policy.ts 167 lines
1// Adapted from jjjjjjjjjjjjjjjjacob/jev-router (MIT), skills/jev/scripts/lib/policy.ts.
2// Unchanged except for levelIndex/capLevel at the end.
3
4// Turns Jev's raw judgments into a routing decision. Pure: no I/O, so thresholds can be
5// table-tested and re-tuned against eval/fixtures.jsonl without calling the API.
6
7import type { Answer, ChoiceAnswer, NoulAnswer, ScoreAnswer } from "./jev.ts";
8import { SIGNALS, SUBAGENT_WORK_CRITERIA, type Signal, type SubagentWork } from "./questions.ts";
9
10export const LEVELS = ["low", "medium", "high", "xhigh", "max"] as const;
11export type Level = (typeof LEVELS)[number];
12
13export function isLevel(value: unknown): value is Level {
14  return typeof value === "string" && (LEVELS as readonly string[]).includes(value);
15}
16
17export const THRESHOLDS = {
18  continuesPrevious: 0.6,
19  quickPassCap: 0.7,
20  verificationFloor: 0.7,
21  floorMaxConfidence: 0.85,
22  autonomy: 0.7,
23  autonomyCompanion: 0.6,
24  maxScoreForMax: 3,
25  minScoreConfidence: 0.35,
26  subagentConfidence: 0.6,
27};
28
29export type EffortAnswers = {
30  score: number;
31  confidence: number;
32  signals: Record<Signal, number>;
33};
34
35export type EffortDecision = {
36  level: Level | null; // null: leave the session's effort alone this turn
37  reasons: string[];
38  rules: string[];
39};
40
41const LEVEL_REASON: Record<Level, string> = {
42  low: "quick, in the loop",
43  medium: "regular feature work",
44  high: "verification decides success",
45  xhigh: "edge-case-dense",
46  max: "autonomous, mission-critical",
47};
48
49export function readEffortAnswers(answers: Record<string, Answer>): EffortAnswers | null {
50  const effort = answers.effort as ScoreAnswer | undefined;
51  if (effort?.type !== "score" || typeof effort.score !== "number") return null;
52  const signals = {} as Record<Signal, number>;
53  for (const signal of SIGNALS) {
54    const answer = answers[signal] as NoulAnswer | undefined;
55    if (answer?.type !== "noul" || typeof answer.noul !== "number") return null;
56    signals[signal] = answer.noul;
57  }
58  return { score: effort.score, confidence: effort.confidence ?? 0, signals };
59}
60
61export function decideEffort(
62  answers: EffortAnswers,
63  context: { previousLevel?: Level | null } = {},
64  t = THRESHOLDS,
65): EffortDecision {
66  const { score, confidence, signals } = answers;
67
68  if (signals.continues_previous > t.continuesPrevious && context.previousLevel) {
69    return {
70      level: context.previousLevel,
71      reasons: ["continues previous request"],
72      rules: ["continue"],
73    };
74  }
75
76  let index = clampIndex(Math.round(score));
77  const rules: string[] = [];
78  const reasons: string[] = [];
79
80  // Signals only overrule the Score when the Score itself is unsure: on the eval set a
81  // confident Score was right wherever the floor would have raised it.
82  const verification = signals.verification_central > t.verificationFloor;
83  const edgeCases = signals.hidden_edge_cases > t.verificationFloor;
84  if ((verification || edgeCases) && confidence < t.floorMaxConfidence) {
85    if (index < 2) index = 2;
86    rules.push("verification-floor");
87    if (verification) reasons.push("verification");
88    if (edgeCases) reasons.push("edge cases");
89  }
90
91  const autonomyCompanion =
92    signals.hidden_edge_cases > t.autonomyCompanion || signals.verification_central > t.autonomyCompanion;
93  if (signals.wants_autonomy > t.autonomy && autonomyCompanion) {
94    const floor = score >= t.maxScoreForMax ? 4 : 3;
95    if (index < floor) index = floor;
96    rules.push("autonomy-floor");
97    reasons.push("autonomous");
98  }
99
100  // Applied last: an explicit ask for a quick pass beats inferred difficulty.
101  if (signals.wants_quick_pass > t.quickPassCap && index > 1) {
102    index = 1;
103    rules.push("quick-pass-cap");
104    reasons.push("quick pass requested");
105  }
106
107  if (rules.length === 0 && confidence < t.minScoreConfidence) {
108    return { level: null, reasons: ["low confidence"], rules: ["abstain"] };
109  }
110
111  const level = LEVELS[index]!;
112  if (reasons.length === 0) reasons.push(LEVEL_REASON[level]);
113  return { level, reasons, rules };
114}
115
116function clampIndex(index: number): number {
117  return Math.min(LEVELS.length - 1, Math.max(0, index));
118}
119
120// --- subagents -------------------------------------------------------------------------
121
122// Lookup agents keep their own defaults; fork ignores `model` entirely.
123export const SKIPPED_SUBAGENT_TYPES = new Set(["fork", "Explore", "claude-code-guide", "statusline-setup"]);
124
125export type SubagentModels = { judgment: string; delegated: string };
126export type SubagentDecision = { model: string | null; work: SubagentWork | null; reason: string };
127
128export function decideSubagentModel(
129  answer: Answer | undefined,
130  requestedModel: string | undefined,
131  models: SubagentModels,
132  t = THRESHOLDS,
133): SubagentDecision {
134  const choice = answer as ChoiceAnswer | undefined;
135  if (choice?.type !== "choice" || !(choice.choice in SUBAGENT_WORK_CRITERIA)) {
136    return { model: null, work: null, reason: "no usable answer" };
137  }
138  const work = choice.choice as SubagentWork;
139  const target = models[work];
140  const requested = normalize(requestedModel);
141  const reason = work === "judgment" ? "judgment work" : "delegated work";
142  if (requested === normalize(target)) return { model: null, work, reason: "already on routed model" };
143  // The router owns the two configured tiers: a model outside them is always replaced
144  // (so a pair that leaves out haiku means haiku never runs); otherwise only a confident
145  // answer overrides what Claude asked for.
146  const tiers = [normalize(models.judgment), normalize(models.delegated)];
147  if (requested && !tiers.includes(requested)) return { model: target, work, reason: `${reason}, replaces ${requested}` };
148  if (choice.confidence < t.subagentConfidence) return { model: null, work, reason: "low confidence" };
149  return { model: target, work, reason };
150}
151
152function normalize(model: string | undefined): string | undefined {
153  return model?.trim().toLowerCase() || undefined;
154}
155
156// --- added in this fork ------------------------------------------------------------------
157
158export function levelIndex(level: Level): number {
159  return LEVELS.indexOf(level);
160}
161
162// Never route above `max`: max effort is costly and prone to overthinking, so the mod's
163// default ceiling is xhigh and max is opt-in.
164export function capLevel(level: Level, max: Level): Level {
165  return levelIndex(level) > levelIndex(max) ? max : level;
166}
167
hooks/lib/questions.ts 90 lines
1// Copied unchanged from jjjjjjjjjjjjjjjjacob/jev-router (MIT), skills/jev/scripts/lib/questions.ts.
2// Jev reads these literally: any rewording needs a live `bun run eval` before it ships.
3
4// The judgments Jev makes. Wording is literal on purpose: Jev answers the question as
5// written, so each Noul names its exact condition and each Score level is a concrete situation.
6
7import type { ChoiceQuestion, NoulQuestion, Question, ScoreQuestion } from "./jev.ts";
8
9export const EFFORT_LEVEL_CRITERIA = [
10  "Quick and in the loop: a short question or explanation, a brainstorm, a rough sketch, or a small mechanical edit such as a rename, formatting, a config value, a commit message, or a git command.",
11  "Regular work with a clear goal that the user will review: implement a feature or change, build a component, write tests, write or revise documents, copy, or plans, or do a focused analysis.",
12  "Work where verification decides success: debug or fix a problem in existing code or systems, review code, tune performance, do an in-depth investigation, or handle edge cases that must be reproduced and tested.",
13  "Edge-case-dense work where a first attempt usually fails: security-sensitive code such as sanitizers or auth, parsers, concurrency, storage engines, or numerical and scientific analysis.",
14  "Fully autonomous, mission-critical work: build and verify a whole app end to end, audit critical software for vulnerabilities, or solve a very hard problem with no user in the loop.",
15];
16
17export const SIGNALS = [
18  "continues_previous",
19  "wants_quick_pass",
20  "wants_autonomy",
21  "hidden_edge_cases",
22  "verification_central",
23] as const;
24export type Signal = (typeof SIGNALS)[number];
25
26const SIGNAL_QUESTIONS: Record<Signal, NoulQuestion> = {
27  continues_previous: {
28    type: "noul",
29    instructions:
30      "Is `current_request` only an approval or continuation of earlier work, such as 'yes', 'go ahead', 'continue', or 'do it', without describing a new task?",
31    criteria: {
32      true: "Approves or continues earlier work without describing a new task",
33      false: "Describes a task or question of its own",
34    },
35  },
36  wants_quick_pass: {
37    type: "noul",
38    instructions:
39      "Does `current_request` ask for a quick, rough, or first-draft result, or ask Claude to be fast or brief?",
40  },
41  wants_autonomy: {
42    type: "noul",
43    instructions:
44      "Does `current_request` ask Claude to keep working to completion without checking in with the user, for example overnight, end to end, or until everything passes?",
45  },
46  hidden_edge_cases: {
47    type: "noul",
48    instructions:
49      "Does the task in `current_request` have many hidden edge cases, where more testing would change whether the result is correct?",
50  },
51  verification_central: {
52    type: "noul",
53    instructions:
54      "Is checking correctness, by reproducing a bug, running or writing tests, or reviewing code or results, a central part of what `current_request` asks for?",
55  },
56};
57
58export function effortQuestions(): Record<string, Question> {
59  const effort: ScoreQuestion = {
60    type: "score",
61    instructions:
62      "How much independent verification and edge-case testing does the work that `current_request` asks for need?",
63    criteria: EFFORT_LEVEL_CRITERIA,
64  };
65  return { effort, ...SIGNAL_QUESTIONS };
66}
67
68export type EffortState = { current_request: string; previous_request?: string };
69
70// Subagent routing asks what kind of work the task is; code maps each kind to the model
71// configured for it (judgmentModel / delegatedModel in config.ts).
72export const SUBAGENT_WORK_CRITERIA = {
73  delegated:
74    "Well-scoped delegated work whose result the parent agent will review: implementing a specified change, research, searching or reading code, running commands, or mechanical edits.",
75  judgment:
76    "Judgment the parent agent will rely on without redoing: reviewing code or plans, adversarial verification, design or architecture decisions, root-cause debugging, user-facing copy or UI, or redoing work that failed review.",
77} as const;
78export type SubagentWork = keyof typeof SUBAGENT_WORK_CRITERIA;
79
80export function subagentQuestions(): Record<string, Question> {
81  const work: ChoiceQuestion = {
82    type: "choice",
83    instructions: "Which kind of work does the task in `subagent_task` ask the subagent to do?",
84    criteria: SUBAGENT_WORK_CRITERIA,
85  };
86  return { work };
87}
88
89export type SubagentState = { subagent_task: string; description?: string };
90
hooks/lib/redact.ts 114 lines
1// What a prompt looks like before it leaves the machine for TypeSafe. Jev only has to judge
2// how hard the task is, so it never needs the code, the secrets or where things live:
3//
4//   pasted blocks         -> [pasted text: N chars]          (upstream behavior)
5//   fenced code blocks    -> [code block: N lines]
6//   long inline code      -> [code]
7//   PEM keys, API tokens, JWTs, key=value secrets, long hex/base64 -> <secret>
8//   URLs and DSNs         -> <url>
9//   e-mail addresses      -> <email>
10//   IPv4 addresses        -> <ip>
11//   absolute/home paths   -> <path>
12//
13// then long prompts keep only their head and tail. Short identifiers in backticks (`getUser`)
14// and relative paths without a leading slash stay: they carry the task's meaning.
15//
16// shapePrompt/truncate adapted from jjjjjjjjjjjjjjjjacob/jev-router (MIT), lib/prompt.ts.
17
18const PASTED_BLOCK = /<pasted_content id="([^"]*)">\n?([\s\S]*?)\n?<\/pasted_content id="\1">/g;
19
20export const HEAD_CHARS = 3000;
21export const TAIL_CHARS = 1000;
22
23export type ShapedPrompt =
24  | { skip: false; text: string; redactions: number }
25  | { skip: true; reason: "empty" | "slash-command" };
26
27export function shapePrompt(raw: string | undefined | null, options: { redact?: boolean } = {}): ShapedPrompt {
28  const trimmed = (raw ?? "").trim();
29  if (!trimmed) return { skip: true, reason: "empty" };
30  // Slash commands and skills carry their own effort; /jev itself is a command.
31  if (trimmed.startsWith("/")) return { skip: true, reason: "slash-command" };
32
33  let redactions = 0;
34  let text = trimmed.replace(PASTED_BLOCK, (_match, _id, body: string) => {
35    redactions++;
36    return `[pasted text: ${body.length} chars]`;
37  });
38  if (options.redact !== false) {
39    const redacted = redact(text);
40    text = redacted.text;
41    redactions += redacted.count;
42  }
43  text = text.trim();
44  if (!text) return { skip: true, reason: "empty" };
45  return { skip: false, text: truncate(text), redactions };
46}
47
48type Rule = { pattern: RegExp; replace: (match: string, ...groups: string[]) => string };
49
50// Order matters: whole blocks first, then secrets (which can sit inside URLs or paths), then
51// the location-like things.
52const RULES: Rule[] = [
53  {
54    pattern: /(```|~~~)[^\n]*\n[\s\S]*?(?:\n\1[^\n]*(?=\n|$)|$)/g,
55    replace: (match) => `[code block: ${Math.max(1, match.split("\n").length - 2)} lines]`,
56  },
57  { pattern: /-----BEGIN [A-Z0-9 ]+-----[\s\S]*?(?:-----END [A-Z0-9 ]+-----|$)/g, replace: () => "<secret>" },
58  { pattern: /`[^`\n]{40,}`/g, replace: () => "[code]" },
59  {
60    pattern: /\b(password|passwd|pwd|secret|token|api[_-]?key|access[_-]?key|client[_-]?secret|private[_-]?key)(\s*[:=]\s*)("[^"\n]*"|'[^'\n]*'|\S+)/gi,
61    replace: (_match, key: string, sep: string) => `${key}${sep}<secret>`,
62  },
63  { pattern: /\b(?:sk|pk|rk)-[A-Za-z0-9_-]{16,}/g, replace: () => "<secret>" },
64  { pattern: /\b(?:gh[pousr]_[A-Za-z0-9]{20,}|github_pat_[A-Za-z0-9_]{20,})\b/g, replace: () => "<secret>" },
65  { pattern: /\bxox[abprs]-[A-Za-z0-9-]{10,}/g, replace: () => "<secret>" },
66  { pattern: /\b(?:AKIA|ASIA)[0-9A-Z]{16}\b/g, replace: () => "<secret>" },
67  { pattern: /\bAIza[0-9A-Za-z_-]{30,}/g, replace: () => "<secret>" },
68  { pattern: /\beyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{4,}/g, replace: () => "<secret>" },
69  {
70    pattern: /\b(?:https?|wss?|ftp|ssh|git|postgres(?:ql)?|mysql|mariadb|redis|rediss|mongodb(?:\+srv)?|amqps?|s3):\/\/[^\s<>"'`)\]]+/gi,
71    replace: () => "<url>",
72  },
73  { pattern: /\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b/g, replace: () => "<email>" },
74  { pattern: /\b(?:\d{1,3}\.){3}\d{1,3}(?::\d{1,5})?\b/g, replace: () => "<ip>" },
75  { pattern: /\b[A-Za-z]:\\(?:[^\\\s]+\\)*[^\\\s]*/g, replace: () => "<path>" },
76  { pattern: /(?<![\w<>/.~@-])(?:~|\.{1,2})?\/(?:[\w.@+-]+\/)+[\w.@+-]*/g, replace: () => "<path>" },
77  { pattern: /(?<![\w<>/.~-])~\/[\w.@+-]+/g, replace: () => "<path>" },
78  { pattern: /\b[a-f0-9]{32,}\b/gi, replace: () => "<secret>" },
79  {
80    // Long base64-ish runs with both letters and digits: tokens, keys, signatures.
81    pattern: /(?<![\w+-])(?=[A-Za-z0-9+_-]*\d)(?=[A-Za-z0-9+_-]*[A-Za-z])[A-Za-z0-9+_-]{40,}={0,2}(?![\w+-])/g,
82    replace: () => "<secret>",
83  },
84];
85
86export function redact(input: string): { text: string; count: number } {
87  let count = 0;
88  let text = input;
89  for (const rule of RULES) {
90    text = text.replace(rule.pattern, (match: string, ...groups: unknown[]) => {
91      count++;
92      return rule.replace(match, ...(groups.filter((g) => typeof g === "string") as string[]));
93    });
94  }
95  return { text, count };
96}
97
98export function truncate(text: string, head = HEAD_CHARS, tail = TAIL_CHARS): string {
99  if (text.length <= head + tail) return text;
100  const omitted = text.length - head - tail;
101  return `${text.slice(0, head)}\n[... ${omitted} chars omitted ...]\n${text.slice(-tail)}`;
102}
103
104// A one-line preview for logs that stay on this machine.
105export function preview(text: string, length = 80): string {
106  const flat = text.replace(/\s+/g, " ").trim();
107  return flat.length > length ? `${flat.slice(0, length - 1)}…` : flat;
108}
109
110// The key a prompt is matched on between prompt.submit and turn.start.
111export function promptKey(text: string): string {
112  return text.replace(/\s+/g, " ").trim().slice(0, 500);
113}
114