Per-turn effort routing with TypeSafe's Jev, applied per request at turn.step. Off until /jev on.

A Claude Code mod that asks Jev (TypeSafe) how much effort each prompt needs, then sends every request of that turn at that effort. The main loop's model is never changed. Off by default; turn it on with /jev on.
This is a fork of jjjjjjjjjjjjjjjjacob/jev-router (MIT), taken from commit 50d7e40 on 2026-09-27. The questions sent to Jev (hooks/lib/questions.ts), the level policy (hooks/lib/policy.ts) and the eval set are kept from upstream. The parts that apply effort and filter data were rewritten.
| upstream jev-router | jeffort | |
|---|---|---|
| How effort is applied | Tells Claude to load a jev-<level> skill and blocks tool calls until it does | A turn.step mod writes effort into each request of the turn |
| Turns that call no tools | May run at the session level (skill skipped) | Still routed |
| Turns from notifications, peers, plugins | Routed like ordinary prompts | Skipped; only prompts you type are routed |
| Data sent to TypeSafe | Prompt (pasted blocks removed), head/tail truncated | Same as upstream, plus masking of code blocks, secrets, URLs, e-mails, IPs and absolute paths |
| Highest effort | max | xhigh (configurable) |
| Cache | Trusts the docs | Watches real cache_read/cache_creation, warns and then pauses when an effort change loses the cache |
| Subagents | Inherit the turn's effort; model routed by judgment/delegated | Keep their own effort; model routed as upstream, via agent.spawn |
| Runtime | Bun/Node, shell scripts, tool-call hooks | Runs inside Claude Code's engine; no Bun/Node needed |
Requires Claude Code 2.1.289 or later (the tested version) and a TypeSafe API key.
# 1. Add the vntrungld marketplace and install (no clone needed):
claude plugin marketplace add vntrungld/jeffort
claude plugin install jeffort@vntrungld
# 2. Key: store it in the keychain via /plugin configure (step 3), or set the env var
export TYPESAFE_API_KEY=...
The vntrungld marketplace lists both jeffort and tightlip. This repo and the tightlip repo carry the same catalog, so adding either one (vntrungld/jeffort or vntrungld/tightlip) gives you both plugins under @vntrungld; adding the other later just replaces it with the same list. Update with claude plugin marketplace update vntrungld and claude plugin update jeffort@vntrungld.
To work on the code, clone it and load it for one session instead: claude --plugin-dir <path-to-clone>.
/plugin configure jeffort@vntrungld to enter the key and adjust options. The non-sensitive options are also in /config.If your organization sets allowManagedModsOnly or allowManagedHooksOnly in managed settings, the mod will not load. If your Claude Code is older and reports that function hooks are disabled, set CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1.
/jev on turn on for this session
/jev status show state and the last 8 decisions
/jev off turn off
Each routed turn shows a dim line such as jev → high · verification decides success (conf 0.72). The status line shows jev · <level> while routing is on. The next turn goes back to the session level unless Jev picks another level.
| Option | Default | Meaning |
|---|---|---|
typesafe_api_key | (empty) | The key. When empty, reads TYPESAFE_API_KEY or JEV_API_KEY |
enabled_by_default | false | Start every session with routing on |
max_effort | xhigh | Highest level the router may pick |
route_subagents | true | Route subagent models |
judgment_model / delegated_model | opus / sonnet | Model for review/debug/design subagents, and for subagents doing delegated work |
redact | true | Mask data before sending (see below) |
cache_safe_only | true | Only change effort on Opus 5.5, Sonnet 5.5, Fable 5.1 |
cache_guard | true | Warn on the first cache miss, pause on the second |
timeout_ms | 2500 | How long to wait for Jev before the turn runs at the session level |
show_decisions | true | Show the jev → … line |
base_url | https://api.typesafe.ai | https only, or http to localhost |
jev_model | jev-1.13.0 | Pinned because the policy thresholds were tuned on this version |
Only while routing is on, and only for prompts you type (no slash commands, notifications, or messages from other sessions). Each turn makes one request to api.typesafe.ai/v1/systemone, containing:
prompt and description, also filtered.The filter replaces:
| Content | With |
|---|---|
| pasted block | [pasted text: N chars] |
| fenced ``` code block | [code block: N lines] |
| inline code of 40+ characters | [code] |
PEM keys, sk-…, ghp_…, xox…, AKIA…, JWTs, password=…, long hex or base64 | <secret> |
URLs, DSNs (postgres://…) | <url> |
<email> | |
| IPv4 | <ip> |
absolute paths, ~/…, C:\… | <path> |
Kept: short function names in backticks (getUser) and relative paths (app/Http/Kernel.php), since they tell Jev what kind of work is being asked for. The filter is regex-based, not DLP: unusual secret formats or lowercase customer names can slip through. For company repositories, check your internal policy before turning it on.
Locally, decisions are kept in the mod's $.store (at most 200 entries, each the first 80 characters of the filtered prompt) for tuning thresholds later.
bun test ./test: 79 unit tests for the policy (ported from upstream), the filter, the Jev client and the cache guard.claude plugin test .: 18 tests running the mod on Claude Code's real engine (per-turn routing, subagents, notifications, timeouts, base URL, filtering, cache guard, /jev). Three key behaviours were deliberately broken (skipping subagent steps, filtering, routing only cache-safe models) to confirm the tests catch each one.claude plugin validate . and tsc are clean.effort: high / low correctly for each routed turn, so the effort really reaches the API.ANTHROPIC_BASE_URL proxy), each effort change kept the cache for the system prompt and tools (~13k tokens) but rewrote the conversation part (~4.5k tokens). Claude Code's built-in /effort command costs exactly the same, so this cost comes from the environment, not the mod. The cache guard caught it and paused routing after the second time.Not yet verified:
TYPESAFE_API_KEY=… bun eval/run.ts and --no-redact to compare, then add your own real prompts to eval/fixtures.local.jsonl.Prompt cache (main) line in /usage, or let the cache guard check for you.bun install
bun test ./test # unit tests
claude plugin test . # mod tests on the engine
claude plugin validate .
tsc -p tsconfig.bun.json && tsc -p . # tsc -p . needs .claude-plugin/types, generated by the engine when loaded with --plugin-dir
TYPESAFE_API_KEY=… bun eval/run.ts # live eval; --no-redact; --replay
hooks/lib/questions.ts is kept word for word, because Jev reads the questions literally. Changing the wording means rerunning the live eval.
MIT. See LICENSE: original copyright by jjjjjjjjjjjjjjjjacob, modifications by this fork.
hooks/register.ts 350 lines1// jeffort: a mod that asks TypeSafe's Jev how much effort each prompt needs and sends
2// that turn's model requests at that effort. Forked from jjjjjjjjjjjjjjjjacob/jev-router
3// (MIT): the questions, policy and eval are theirs; how the level is applied is not.
4//
5// Upstream had no way to set effort from a hook, so it told Claude to load a jev-<level>
6// skill and gated tool calls until it did; a turn that answered without tools skipped it.
7// A mod can rewrite `effort` on every `turn.step`, so here:
8//
9// prompt.submit marks prompts the person typed (not notifications, peers or plugins)
10// turn.start asks Jev once per marked prompt, before the turn's first request
11// turn.step sends each main-loop request of that turn at the routed level
12// turn.complete forgets the turn; the next prompt starts at the session level again
13// agent.spawn picks a subagent's model from the kind of work it is given
14//
15// Everything fails open: no key, a timeout, an error or an unsure answer leaves the turn at
16// the session's own effort. Subagents keep their own effort; only their model is routed.
17
18import type { EngineInterface, Register } from "claude-code";
19import { CacheGuard, isCacheSafeModel } from "./lib/cache-guard.ts";
20import { buildRequest, DEFAULT_BASE_URL, DEFAULT_MODEL, DEFAULT_TIMEOUT_MS, isAllowedBaseUrl, parseResponse, type JevAnswers, type Question } from "./lib/jev.ts";
21import { capLevel, decideEffort, decideSubagentModel, isLevel, readEffortAnswers, SKIPPED_SUBAGENT_TYPES, type Level } from "./lib/policy.ts";
22import { effortQuestions, subagentQuestions, type EffortState, type SubagentState } from "./lib/questions.ts";
23import { preview, promptKey, redact, shapePrompt, truncate } from "./lib/redact.ts";
24
25const PREVIOUS_REQUEST_CHARS = 1500;
26const MAX_MARKED_PROMPTS = 20;
27const RECENT_DECISIONS = 8;
28const STORED_DECISIONS = 200;
29const ROUTED_ORIGINS = new Set(["composer", "bridge", "sdk"]);
30
31type Decision = {
32 at: number;
33 kind: "effort" | "subagent";
34 preview: string;
35 level?: Level | null;
36 sessionLevel?: string | null;
37 applied?: boolean;
38 model?: string | null;
39 reasons: string[];
40 confidence?: number;
41 latencyMs?: number;
42 redactions?: number;
43};
44
45type Route = { level: Level; decision: Decision; logged: boolean };
46
47type Config = {
48 apiKey: string | undefined;
49 enabledByDefault: boolean;
50 maxEffort: Level;
51 routeSubagents: boolean;
52 judgmentModel: string;
53 delegatedModel: string;
54 redact: boolean;
55 cacheSafeOnly: boolean;
56 cacheGuard: boolean;
57 timeoutMs: number;
58 showDecisions: boolean;
59 baseUrl: string;
60 jevModel: string;
61};
62
63// Session state lives in the module: a reload of the plugin starts it over, which only
64// happens on /reload-plugins for an installed plugin. `register` fills it in.
65let config: Config;
66const session = {
67 enabled: false,
68 pausedReason: null as string | null,
69 lastLevel: null as Level | null,
70 lastRequest: undefined as string | undefined,
71 marked: new Map<string, number>(),
72 routed: new Map<string, Route>(),
73 recent: [] as Decision[],
74 guard: new CacheGuard(),
75};
76
77export const register: Register = (on, options) => {
78 config = {
79 apiKey: text(options.typesafe_api_key),
80 enabledByDefault: options.enabled_by_default === true,
81 maxEffort: (isLevel(options.max_effort) ? options.max_effort : "xhigh") as Level,
82 routeSubagents: options.route_subagents !== false,
83 judgmentModel: text(options.judgment_model) ?? "opus",
84 delegatedModel: text(options.delegated_model) ?? "sonnet",
85 redact: options.redact !== false,
86 cacheSafeOnly: options.cache_safe_only !== false,
87 cacheGuard: options.cache_guard !== false,
88 timeoutMs: typeof options.timeout_ms === "number" ? options.timeout_ms : DEFAULT_TIMEOUT_MS,
89 showDecisions: options.show_decisions !== false,
90 baseUrl: text(options.base_url) ?? DEFAULT_BASE_URL,
91 jevModel: text(options.jev_model) ?? DEFAULT_MODEL,
92 };
93 session.enabled = config.enabledByDefault;
94
95 on("session.start", async ($, e, next) => {
96 await $.command.register({
97 name: "jev",
98 description: "Jev effort routing for this session: on, off, or status",
99 argumentHint: "[on|off|status]",
100 });
101 showStatus($);
102 return next(e);
103 });
104
105 on("command.run", { command: "jev" }, async ($, e) => {
106 const arg = (e.args.trim().split(/\s+/)[0] || "on").toLowerCase();
107 if (arg === "on") {
108 setEnabled($, true);
109 if (!(await apiKey($))) {
110 return { text: "Jev routing is ON, but no TypeSafe key was found, so prompts pass through unrouted. Set it in /plugin configure, or export TYPESAFE_API_KEY." };
111 }
112 if (!isAllowedBaseUrl(config.baseUrl)) {
113 return { text: `Jev routing is ON, but base_url ${config.baseUrl} is not https (or http to localhost), so nothing is sent.` };
114 }
115 const subagents = config.routeSubagents ? ` Subagents: ${config.judgmentModel} for judgment, ${config.delegatedModel} for delegated work.` : "";
116 return { text: `Jev routing is ON for this session (up to ${config.maxEffort}).${subagents} /jev off to stop.` };
117 }
118 if (arg === "off") {
119 setEnabled($, false);
120 session.pausedReason = null;
121 return { text: "Jev routing is OFF. Turns run at the session's effort." };
122 }
123 if (arg === "status") return { text: await statusText($) };
124 return { text: `Unknown /jev argument "${arg}". Use /jev on, /jev off or /jev status.` };
125 });
126
127 // Only prompts the person sent start a routed turn. One typed while a turn runs is folded
128 // into that turn (e.turnId set) and is left alone, as are notifications and peers.
129 on("prompt.submit", async ($, e, next) => {
130 if (session.enabled && !e.turnId && ROUTED_ORIGINS.has(e.origin.kind)) {
131 const key = promptKey(e.text);
132 if (key) {
133 session.marked.set(key, (session.marked.get(key) ?? 0) + 1);
134 while (session.marked.size > MAX_MARKED_PROMPTS) session.marked.delete(session.marked.keys().next().value as string);
135 }
136 }
137 return next(e);
138 });
139
140 on("turn.start", async ($, e, next) => {
141 const markedKey = promptKey(e.text);
142 const count = session.marked.get(markedKey);
143 if (!session.enabled || !count) return next(e);
144 if (count > 1) session.marked.set(markedKey, count - 1);
145 else session.marked.delete(markedKey);
146
147 const shaped = shapePrompt(e.text, { redact: config.redact });
148 if (shaped.skip) return next(e);
149 const key = await apiKey($);
150 if (!key) return next(e);
151
152 const state: EffortState = { current_request: shaped.text };
153 if (session.lastRequest) state.previous_request = session.lastRequest;
154 const result = await askJev($, key, state, effortQuestions());
155 const answers = result && readEffortAnswers(result.answers);
156 if (!result || !answers) return next(e);
157
158 const decided = decideEffort(answers, { previousLevel: session.lastLevel });
159 const level = decided.level ? capLevel(decided.level, config.maxEffort) : null;
160 const reasons = level && decided.level !== level ? [...decided.reasons, `capped at ${level}`] : decided.reasons;
161 session.lastLevel = level ?? session.lastLevel;
162 session.lastRequest = shaped.text.slice(0, PREVIOUS_REQUEST_CHARS);
163
164 const decision: Decision = {
165 at: await $.clock.now(),
166 kind: "effort",
167 preview: preview(shaped.text),
168 level,
169 reasons,
170 confidence: round(answers.confidence),
171 latencyMs: result.latencyMs,
172 redactions: shaped.redactions,
173 };
174 await remember($, decision);
175 if (level) session.routed.set(e.turnId, { level, decision, logged: false });
176 showStatus($);
177 return next(e);
178 });
179
180 on("turn.step", async function* ($, e, next) {
181 // Subagent loops keep their own effort: their steps carry agentId and their own turnId.
182 if (e.agentId) return yield* next(e);
183
184 const route = session.enabled ? session.routed.get(e.turnId) : undefined;
185 const current = e.effort;
186 const routable =
187 route !== undefined &&
188 typeof current === "string" &&
189 isLevel(current) &&
190 (!config.cacheSafeOnly || isCacheSafeModel(e.model));
191 const sent = routable && current !== route.level ? { ...e, effort: route.level } : e;
192
193 if (route && !route.logged) {
194 route.logged = true;
195 route.decision.applied = sent !== e;
196 route.decision.sessionLevel = typeof current === "string" ? current : null;
197 if (sent !== e && config.showDecisions) {
198 $.ui.log(`jev → ${route.level} · ${route.decision.reasons.join(", ")} (conf ${route.decision.confidence?.toFixed(2)})`);
199 }
200 }
201
202 const result = yield* next(sent);
203
204 if (session.enabled && config.cacheGuard && result.usage) {
205 const verdict = session.guard.observe({
206 at: await $.clock.now(),
207 model: e.model,
208 effort: sent.effort,
209 cacheRead: result.usage.cache_read_input_tokens,
210 cacheWrite: result.usage.cache_creation_input_tokens,
211 });
212 if (verdict.kind === "miss") {
213 $.ui.log(`jev: the prompt cache missed right after an effort change on ${e.model} (${verdict.written} tokens re-written). Routing pauses if it happens again.`);
214 } else if (verdict.kind === "pause") {
215 setEnabled($, false);
216 session.pausedReason = `cache missed ${verdict.misses}× after effort changes on ${e.model}`;
217 $.ui.log(`jev: paused for this session. Effort changes keep re-reading the prompt cache on ${e.model}. /jev on to resume.`);
218 }
219 }
220 return result;
221 });
222
223 on("turn.complete", async ($, e, next) => {
224 if (!e.agentId) session.routed.delete(e.turnId);
225 return next(e);
226 });
227
228 on("agent.spawn", async ($, e, next) => {
229 if (!session.enabled || !config.routeSubagents || e.fork || SKIPPED_SUBAGENT_TYPES.has(e.subagentType)) return next(e);
230 const task = e.prompt.trim();
231 if (!task) return next(e);
232 const key = await apiKey($);
233 if (!key) return next(e);
234
235 const description = config.redact ? redact(e.description).text : e.description;
236 const body = config.redact ? redact(task).text : task;
237 const state: SubagentState = { subagent_task: truncate(body) };
238 if (description) state.description = description;
239
240 const result = await askJev($, key, state, subagentQuestions());
241 if (!result) return next(e);
242 const answer = result.answers.work;
243 const decision = decideSubagentModel(answer, e.model, {
244 judgment: config.judgmentModel,
245 delegated: config.delegatedModel,
246 });
247 await remember($, {
248 at: await $.clock.now(),
249 kind: "subagent",
250 preview: preview(description || body),
251 model: decision.model,
252 reasons: [decision.reason],
253 confidence: answer?.type === "choice" ? round(answer.confidence) : undefined,
254 latencyMs: result.latencyMs,
255 });
256 if (!decision.model) return next(e);
257 if (config.showDecisions) $.ui.log(`jev → subagent on ${decision.model} (${decision.reason})`);
258 return next({ ...e, model: decision.model });
259 });
260};
261
262// --- helpers: top-level functions, the only places a hook may hand `$` to ---------------
263
264async function apiKey($: EngineInterface): Promise<string | null> {
265 if (config.apiKey) return config.apiKey;
266 return text(await $.env.get("TYPESAFE_API_KEY")) ?? text(await $.env.get("JEV_API_KEY")) ?? null;
267}
268
269// One request, a hard timeout, never throws: every failure is null.
270async function askJev(
271 $: EngineInterface,
272 key: string,
273 state: unknown,
274 questions: Record<string, Question>,
275): Promise<(JevAnswers & { latencyMs: number }) | null> {
276 if (!isAllowedBaseUrl(config.baseUrl)) return null;
277 const request = buildRequest(state, questions, { apiKey: key, model: config.jevModel, baseUrl: config.baseUrl });
278 const timer = new AbortController();
279 try {
280 const started = await $.clock.now();
281 const call = $.http.fetch(request.url, request.init).then(
282 (response) => parseResponse(response, questions, config.jevModel),
283 () => null,
284 );
285 const timeout = $.clock.sleep(config.timeoutMs, { signal: timer.signal }).then(
286 () => null,
287 () => null,
288 );
289 const answers = await Promise.race([call, timeout]);
290 if (!answers) return null;
291 return { ...answers, latencyMs: Math.round((await $.clock.now()) - started) };
292 } catch {
293 return null;
294 } finally {
295 timer.abort();
296 }
297}
298
299async function remember($: EngineInterface, decision: Decision): Promise<void> {
300 session.recent.push(decision);
301 if (session.recent.length > RECENT_DECISIONS) session.recent.shift();
302 try {
303 const stored = await $.store.get("decisions");
304 const list = Array.isArray(stored) ? stored : [];
305 list.push(decision);
306 await $.store.set("decisions", list.slice(-STORED_DECISIONS));
307 } catch {
308 // the log is for tuning only
309 }
310}
311
312function showStatus($: EngineInterface): void {
313 $.ui.status(session.enabled ? (session.lastLevel ? `jev · ${session.lastLevel}` : "jev") : undefined);
314}
315
316function setEnabled($: EngineInterface, value: boolean): void {
317 session.enabled = value;
318 if (value) {
319 session.pausedReason = null;
320 session.guard.reset();
321 } else {
322 session.routed.clear();
323 session.marked.clear();
324 }
325 showStatus($);
326}
327
328async function statusText($: EngineInterface): Promise<string> {
329 const key = (await apiKey($)) ? "key set" : "no key";
330 const state = session.enabled ? "ON" : session.pausedReason ? `PAUSED (${session.pausedReason})` : "OFF";
331 const head = `Jev routing is ${state} · ${key} · up to ${config.maxEffort} · redaction ${config.redact ? "on" : "off"} · cache-safe models only: ${config.cacheSafeOnly ? "yes" : "no"}`;
332 if (session.recent.length === 0) return `${head}\nNo decisions yet.`;
333 const lines = session.recent.map((d) => {
334 if (d.kind === "subagent") return `- subagent → ${d.model ?? "unchanged"} (${d.reasons.join(", ")}) · "${d.preview}"`;
335 const level = d.level ?? "unchanged";
336 const applied = d.applied === undefined ? "" : d.applied ? ` (session ${d.sessionLevel ?? "?"})` : " (not applied)";
337 const conf = d.confidence === undefined ? "" : ` · conf ${d.confidence.toFixed(2)}`;
338 return `- ${level}${applied} · ${d.reasons.join(", ")}${conf} · ${d.latencyMs ?? "?"} ms · "${d.preview}"`;
339 });
340 return `${head}\nLast ${session.recent.length}:\n${lines.join("\n")}`;
341}
342
343function text(value: unknown): string | undefined {
344 return typeof value === "string" && value.trim() ? value.trim() : undefined;
345}
346
347function round(value: number): number {
348 return Math.round(value * 100) / 100;
349}
350hooks/lib/cache-guard.ts 76 lines1// Watches whether changing effort between requests costs the prompt cache. On Opus 5.5,
2// Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription it should not; on other
3// models, providers or gateways each change re-reads the whole conversation. The guard turns
4// that suspicion into evidence from the API's own usage numbers and pauses routing when the
5// evidence shows up, instead of trusting a model-name allowlist alone.
6
7export type StepRecord = {
8 at: number; // ms since epoch, when the response finished
9 model: string;
10 effort: string | number | undefined;
11 cacheRead: number;
12 cacheWrite: number;
13};
14
15export type GuardVerdict =
16 | { kind: "ok" }
17 | { kind: "miss"; written: number } // an effort change coincided with a cache miss
18 | { kind: "pause"; written: number; misses: number };
19
20export const GUARD = {
21 // How much of the previous request's cached prefix must go missing to count. Small drops
22 // are noise; a real invalidation loses at least the conversation part of the prefix.
23 minLost: 2048,
24 // Past this, a miss is the cache TTL (5 minutes by default), not the effort change.
25 maxGapMs: 4 * 60 * 1000,
26 pauseAfterMisses: 2,
27};
28
29export class CacheGuard {
30 private last: StepRecord | null = null;
31 private misses = 0;
32
33 constructor(private readonly limits = GUARD) {}
34
35 // Feed every main-loop step, routed or not, in order.
36 observe(step: StepRecord): GuardVerdict {
37 const previous = this.last;
38 this.last = step;
39 if (!previous) return { kind: "ok" };
40 const effortChanged = previous.effort !== step.effort;
41 const sameModel = previous.model === step.model;
42 const recent = step.at - previous.at <= this.limits.maxGapMs;
43 // A miss is often partial: the system prompt and tools stay cached and only the messages
44 // are re-written (seen on a gateway with Sonnet 5.5, for /effort and this mod alike), so
45 // compare against everything the previous request left cached, not against zero.
46 const previousPrefix = previous.cacheRead + previous.cacheWrite;
47 const lost = previousPrefix - step.cacheRead;
48 const missed = lost >= this.limits.minLost && step.cacheWrite >= this.limits.minLost / 2;
49 if (!(effortChanged && sameModel && recent)) return { kind: "ok" };
50 if (!missed) {
51 // An effort change that kept the cache is evidence the setup keeps it: forget older misses.
52 this.misses = 0;
53 return { kind: "ok" };
54 }
55 this.misses++;
56 if (this.misses >= this.limits.pauseAfterMisses) {
57 return { kind: "pause", written: step.cacheWrite, misses: this.misses };
58 }
59 return { kind: "miss", written: step.cacheWrite };
60 }
61
62 reset(): void {
63 this.last = null;
64 this.misses = 0;
65 }
66}
67
68// Models whose per-request effort change keeps the cache, per the Claude Code prompt-caching
69// docs (2.1.260+). Provider-prefixed ids (Bedrock, Vertex) still match here; the guard above
70// is what catches those, since the docs say the cache is not kept there.
71const CACHE_SAFE_MODEL = /claude-(?:opus-5-5|sonnet-5-5|fable-5-1)\b/i;
72
73export function isCacheSafeModel(model: string): boolean {
74 return CACHE_SAFE_MODEL.test(model);
75}
76hooks/lib/jev.ts 90 lines1// TypeSafe System One wire format. Pure: no network here. The mod sends the request with
2// $.http.fetch and the eval harness with fetch, so both share one request and one parser.
3//
4// Adapted from jjjjjjjjjjjjjjjjacob/jev-router (MIT), skills/jev/scripts/lib/jev.ts.
5
6export const DEFAULT_BASE_URL = "https://api.typesafe.ai";
7export const DEFAULT_MODEL = "jev-1.13.0"; // pinned: policy thresholds are tuned against this version
8export const DEFAULT_TIMEOUT_MS = 2500;
9
10export type NoulQuestion = {
11 type: "noul";
12 instructions: unknown;
13 criteria?: { true?: unknown; false?: unknown };
14};
15export type ChoiceQuestion = { type: "choice"; instructions: unknown; criteria: Record<string, unknown> };
16export type ScoreQuestion = { type: "score"; instructions: unknown; criteria: unknown[] };
17export type Question = NoulQuestion | ChoiceQuestion | ScoreQuestion;
18
19export type NoulAnswer = { type: "noul"; noul: number };
20export type ChoiceAnswer = {
21 type: "choice";
22 choice: string;
23 probabilities: Record<string, number>;
24 confidence: number;
25};
26export type ScoreAnswer = {
27 type: "score";
28 score: number;
29 probabilities: Record<string, number>;
30 legend: Record<string, string>;
31 confidence: number;
32};
33export type Answer = NoulAnswer | ChoiceAnswer | ScoreAnswer;
34
35export type JevAnswers = { model: string; answers: Record<string, Answer> };
36
37export type JevRequest = {
38 url: string;
39 init: { method: "POST"; headers: Record<string, string>; body: string };
40};
41
42export function buildRequest(
43 state: unknown,
44 questions: Record<string, Question>,
45 options: { apiKey: string; model?: string; baseUrl?: string },
46): JevRequest {
47 const base = (options.baseUrl || DEFAULT_BASE_URL).replace(/\/+$/, "");
48 return {
49 url: `${base}/v1/systemone`,
50 init: {
51 method: "POST",
52 headers: { Authorization: `Bearer ${options.apiKey}`, "Content-Type": "application/json" },
53 body: JSON.stringify({ model: options.model || DEFAULT_MODEL, state, questions }),
54 },
55 };
56}
57
58// Never throws: a response that is not ok, not JSON, or missing any asked question is null.
59export function parseResponse(
60 response: { ok: boolean; text: string },
61 questions: Record<string, Question>,
62 requestedModel: string = DEFAULT_MODEL,
63): JevAnswers | null {
64 if (!response.ok) return null;
65 try {
66 const body = JSON.parse(response.text) as { model?: unknown; answers?: Record<string, Answer> };
67 if (!body || typeof body !== "object" || !body.answers || typeof body.answers !== "object") return null;
68 for (const id of Object.keys(questions)) {
69 if (!body.answers[id]) return null;
70 }
71 return { model: typeof body.model === "string" ? body.model : requestedModel, answers: body.answers };
72 } catch {
73 return null;
74 }
75}
76
77// Only https to a real host, or http to loopback (a local proxy). Anything else would send
78// the API key and the prompt somewhere unintended, so the mod refuses it and stays off.
79export function isAllowedBaseUrl(raw: string): boolean {
80 try {
81 const url = new URL(raw);
82 if (url.username || url.password) return false;
83 if (url.protocol === "https:") return url.hostname.length > 0;
84 if (url.protocol === "http:") return ["localhost", "127.0.0.1", "[::1]"].includes(url.hostname);
85 return false;
86 } catch {
87 return false;
88 }
89}
90hooks/lib/policy.ts 167 lines1// Adapted from jjjjjjjjjjjjjjjjacob/jev-router (MIT), skills/jev/scripts/lib/policy.ts.
2// Unchanged except for levelIndex/capLevel at the end.
3
4// Turns Jev's raw judgments into a routing decision. Pure: no I/O, so thresholds can be
5// table-tested and re-tuned against eval/fixtures.jsonl without calling the API.
6
7import type { Answer, ChoiceAnswer, NoulAnswer, ScoreAnswer } from "./jev.ts";
8import { SIGNALS, SUBAGENT_WORK_CRITERIA, type Signal, type SubagentWork } from "./questions.ts";
9
10export const LEVELS = ["low", "medium", "high", "xhigh", "max"] as const;
11export type Level = (typeof LEVELS)[number];
12
13export function isLevel(value: unknown): value is Level {
14 return typeof value === "string" && (LEVELS as readonly string[]).includes(value);
15}
16
17export const THRESHOLDS = {
18 continuesPrevious: 0.6,
19 quickPassCap: 0.7,
20 verificationFloor: 0.7,
21 floorMaxConfidence: 0.85,
22 autonomy: 0.7,
23 autonomyCompanion: 0.6,
24 maxScoreForMax: 3,
25 minScoreConfidence: 0.35,
26 subagentConfidence: 0.6,
27};
28
29export type EffortAnswers = {
30 score: number;
31 confidence: number;
32 signals: Record<Signal, number>;
33};
34
35export type EffortDecision = {
36 level: Level | null; // null: leave the session's effort alone this turn
37 reasons: string[];
38 rules: string[];
39};
40
41const LEVEL_REASON: Record<Level, string> = {
42 low: "quick, in the loop",
43 medium: "regular feature work",
44 high: "verification decides success",
45 xhigh: "edge-case-dense",
46 max: "autonomous, mission-critical",
47};
48
49export function readEffortAnswers(answers: Record<string, Answer>): EffortAnswers | null {
50 const effort = answers.effort as ScoreAnswer | undefined;
51 if (effort?.type !== "score" || typeof effort.score !== "number") return null;
52 const signals = {} as Record<Signal, number>;
53 for (const signal of SIGNALS) {
54 const answer = answers[signal] as NoulAnswer | undefined;
55 if (answer?.type !== "noul" || typeof answer.noul !== "number") return null;
56 signals[signal] = answer.noul;
57 }
58 return { score: effort.score, confidence: effort.confidence ?? 0, signals };
59}
60
61export function decideEffort(
62 answers: EffortAnswers,
63 context: { previousLevel?: Level | null } = {},
64 t = THRESHOLDS,
65): EffortDecision {
66 const { score, confidence, signals } = answers;
67
68 if (signals.continues_previous > t.continuesPrevious && context.previousLevel) {
69 return {
70 level: context.previousLevel,
71 reasons: ["continues previous request"],
72 rules: ["continue"],
73 };
74 }
75
76 let index = clampIndex(Math.round(score));
77 const rules: string[] = [];
78 const reasons: string[] = [];
79
80 // Signals only overrule the Score when the Score itself is unsure: on the eval set a
81 // confident Score was right wherever the floor would have raised it.
82 const verification = signals.verification_central > t.verificationFloor;
83 const edgeCases = signals.hidden_edge_cases > t.verificationFloor;
84 if ((verification || edgeCases) && confidence < t.floorMaxConfidence) {
85 if (index < 2) index = 2;
86 rules.push("verification-floor");
87 if (verification) reasons.push("verification");
88 if (edgeCases) reasons.push("edge cases");
89 }
90
91 const autonomyCompanion =
92 signals.hidden_edge_cases > t.autonomyCompanion || signals.verification_central > t.autonomyCompanion;
93 if (signals.wants_autonomy > t.autonomy && autonomyCompanion) {
94 const floor = score >= t.maxScoreForMax ? 4 : 3;
95 if (index < floor) index = floor;
96 rules.push("autonomy-floor");
97 reasons.push("autonomous");
98 }
99
100 // Applied last: an explicit ask for a quick pass beats inferred difficulty.
101 if (signals.wants_quick_pass > t.quickPassCap && index > 1) {
102 index = 1;
103 rules.push("quick-pass-cap");
104 reasons.push("quick pass requested");
105 }
106
107 if (rules.length === 0 && confidence < t.minScoreConfidence) {
108 return { level: null, reasons: ["low confidence"], rules: ["abstain"] };
109 }
110
111 const level = LEVELS[index]!;
112 if (reasons.length === 0) reasons.push(LEVEL_REASON[level]);
113 return { level, reasons, rules };
114}
115
116function clampIndex(index: number): number {
117 return Math.min(LEVELS.length - 1, Math.max(0, index));
118}
119
120// --- subagents -------------------------------------------------------------------------
121
122// Lookup agents keep their own defaults; fork ignores `model` entirely.
123export const SKIPPED_SUBAGENT_TYPES = new Set(["fork", "Explore", "claude-code-guide", "statusline-setup"]);
124
125export type SubagentModels = { judgment: string; delegated: string };
126export type SubagentDecision = { model: string | null; work: SubagentWork | null; reason: string };
127
128export function decideSubagentModel(
129 answer: Answer | undefined,
130 requestedModel: string | undefined,
131 models: SubagentModels,
132 t = THRESHOLDS,
133): SubagentDecision {
134 const choice = answer as ChoiceAnswer | undefined;
135 if (choice?.type !== "choice" || !(choice.choice in SUBAGENT_WORK_CRITERIA)) {
136 return { model: null, work: null, reason: "no usable answer" };
137 }
138 const work = choice.choice as SubagentWork;
139 const target = models[work];
140 const requested = normalize(requestedModel);
141 const reason = work === "judgment" ? "judgment work" : "delegated work";
142 if (requested === normalize(target)) return { model: null, work, reason: "already on routed model" };
143 // The router owns the two configured tiers: a model outside them is always replaced
144 // (so a pair that leaves out haiku means haiku never runs); otherwise only a confident
145 // answer overrides what Claude asked for.
146 const tiers = [normalize(models.judgment), normalize(models.delegated)];
147 if (requested && !tiers.includes(requested)) return { model: target, work, reason: `${reason}, replaces ${requested}` };
148 if (choice.confidence < t.subagentConfidence) return { model: null, work, reason: "low confidence" };
149 return { model: target, work, reason };
150}
151
152function normalize(model: string | undefined): string | undefined {
153 return model?.trim().toLowerCase() || undefined;
154}
155
156// --- added in this fork ------------------------------------------------------------------
157
158export function levelIndex(level: Level): number {
159 return LEVELS.indexOf(level);
160}
161
162// Never route above `max`: max effort is costly and prone to overthinking, so the mod's
163// default ceiling is xhigh and max is opt-in.
164export function capLevel(level: Level, max: Level): Level {
165 return levelIndex(level) > levelIndex(max) ? max : level;
166}
167hooks/lib/questions.ts 90 lines1// Copied unchanged from jjjjjjjjjjjjjjjjacob/jev-router (MIT), skills/jev/scripts/lib/questions.ts.
2// Jev reads these literally: any rewording needs a live `bun run eval` before it ships.
3
4// The judgments Jev makes. Wording is literal on purpose: Jev answers the question as
5// written, so each Noul names its exact condition and each Score level is a concrete situation.
6
7import type { ChoiceQuestion, NoulQuestion, Question, ScoreQuestion } from "./jev.ts";
8
9export const EFFORT_LEVEL_CRITERIA = [
10 "Quick and in the loop: a short question or explanation, a brainstorm, a rough sketch, or a small mechanical edit such as a rename, formatting, a config value, a commit message, or a git command.",
11 "Regular work with a clear goal that the user will review: implement a feature or change, build a component, write tests, write or revise documents, copy, or plans, or do a focused analysis.",
12 "Work where verification decides success: debug or fix a problem in existing code or systems, review code, tune performance, do an in-depth investigation, or handle edge cases that must be reproduced and tested.",
13 "Edge-case-dense work where a first attempt usually fails: security-sensitive code such as sanitizers or auth, parsers, concurrency, storage engines, or numerical and scientific analysis.",
14 "Fully autonomous, mission-critical work: build and verify a whole app end to end, audit critical software for vulnerabilities, or solve a very hard problem with no user in the loop.",
15];
16
17export const SIGNALS = [
18 "continues_previous",
19 "wants_quick_pass",
20 "wants_autonomy",
21 "hidden_edge_cases",
22 "verification_central",
23] as const;
24export type Signal = (typeof SIGNALS)[number];
25
26const SIGNAL_QUESTIONS: Record<Signal, NoulQuestion> = {
27 continues_previous: {
28 type: "noul",
29 instructions:
30 "Is `current_request` only an approval or continuation of earlier work, such as 'yes', 'go ahead', 'continue', or 'do it', without describing a new task?",
31 criteria: {
32 true: "Approves or continues earlier work without describing a new task",
33 false: "Describes a task or question of its own",
34 },
35 },
36 wants_quick_pass: {
37 type: "noul",
38 instructions:
39 "Does `current_request` ask for a quick, rough, or first-draft result, or ask Claude to be fast or brief?",
40 },
41 wants_autonomy: {
42 type: "noul",
43 instructions:
44 "Does `current_request` ask Claude to keep working to completion without checking in with the user, for example overnight, end to end, or until everything passes?",
45 },
46 hidden_edge_cases: {
47 type: "noul",
48 instructions:
49 "Does the task in `current_request` have many hidden edge cases, where more testing would change whether the result is correct?",
50 },
51 verification_central: {
52 type: "noul",
53 instructions:
54 "Is checking correctness, by reproducing a bug, running or writing tests, or reviewing code or results, a central part of what `current_request` asks for?",
55 },
56};
57
58export function effortQuestions(): Record<string, Question> {
59 const effort: ScoreQuestion = {
60 type: "score",
61 instructions:
62 "How much independent verification and edge-case testing does the work that `current_request` asks for need?",
63 criteria: EFFORT_LEVEL_CRITERIA,
64 };
65 return { effort, ...SIGNAL_QUESTIONS };
66}
67
68export type EffortState = { current_request: string; previous_request?: string };
69
70// Subagent routing asks what kind of work the task is; code maps each kind to the model
71// configured for it (judgmentModel / delegatedModel in config.ts).
72export const SUBAGENT_WORK_CRITERIA = {
73 delegated:
74 "Well-scoped delegated work whose result the parent agent will review: implementing a specified change, research, searching or reading code, running commands, or mechanical edits.",
75 judgment:
76 "Judgment the parent agent will rely on without redoing: reviewing code or plans, adversarial verification, design or architecture decisions, root-cause debugging, user-facing copy or UI, or redoing work that failed review.",
77} as const;
78export type SubagentWork = keyof typeof SUBAGENT_WORK_CRITERIA;
79
80export function subagentQuestions(): Record<string, Question> {
81 const work: ChoiceQuestion = {
82 type: "choice",
83 instructions: "Which kind of work does the task in `subagent_task` ask the subagent to do?",
84 criteria: SUBAGENT_WORK_CRITERIA,
85 };
86 return { work };
87}
88
89export type SubagentState = { subagent_task: string; description?: string };
90hooks/lib/redact.ts 114 lines1// What a prompt looks like before it leaves the machine for TypeSafe. Jev only has to judge
2// how hard the task is, so it never needs the code, the secrets or where things live:
3//
4// pasted blocks -> [pasted text: N chars] (upstream behavior)
5// fenced code blocks -> [code block: N lines]
6// long inline code -> [code]
7// PEM keys, API tokens, JWTs, key=value secrets, long hex/base64 -> <secret>
8// URLs and DSNs -> <url>
9// e-mail addresses -> <email>
10// IPv4 addresses -> <ip>
11// absolute/home paths -> <path>
12//
13// then long prompts keep only their head and tail. Short identifiers in backticks (`getUser`)
14// and relative paths without a leading slash stay: they carry the task's meaning.
15//
16// shapePrompt/truncate adapted from jjjjjjjjjjjjjjjjacob/jev-router (MIT), lib/prompt.ts.
17
18const PASTED_BLOCK = /<pasted_content id="([^"]*)">\n?([\s\S]*?)\n?<\/pasted_content id="\1">/g;
19
20export const HEAD_CHARS = 3000;
21export const TAIL_CHARS = 1000;
22
23export type ShapedPrompt =
24 | { skip: false; text: string; redactions: number }
25 | { skip: true; reason: "empty" | "slash-command" };
26
27export function shapePrompt(raw: string | undefined | null, options: { redact?: boolean } = {}): ShapedPrompt {
28 const trimmed = (raw ?? "").trim();
29 if (!trimmed) return { skip: true, reason: "empty" };
30 // Slash commands and skills carry their own effort; /jev itself is a command.
31 if (trimmed.startsWith("/")) return { skip: true, reason: "slash-command" };
32
33 let redactions = 0;
34 let text = trimmed.replace(PASTED_BLOCK, (_match, _id, body: string) => {
35 redactions++;
36 return `[pasted text: ${body.length} chars]`;
37 });
38 if (options.redact !== false) {
39 const redacted = redact(text);
40 text = redacted.text;
41 redactions += redacted.count;
42 }
43 text = text.trim();
44 if (!text) return { skip: true, reason: "empty" };
45 return { skip: false, text: truncate(text), redactions };
46}
47
48type Rule = { pattern: RegExp; replace: (match: string, ...groups: string[]) => string };
49
50// Order matters: whole blocks first, then secrets (which can sit inside URLs or paths), then
51// the location-like things.
52const RULES: Rule[] = [
53 {
54 pattern: /(```|~~~)[^\n]*\n[\s\S]*?(?:\n\1[^\n]*(?=\n|$)|$)/g,
55 replace: (match) => `[code block: ${Math.max(1, match.split("\n").length - 2)} lines]`,
56 },
57 { pattern: /-----BEGIN [A-Z0-9 ]+-----[\s\S]*?(?:-----END [A-Z0-9 ]+-----|$)/g, replace: () => "<secret>" },
58 { pattern: /`[^`\n]{40,}`/g, replace: () => "[code]" },
59 {
60 pattern: /\b(password|passwd|pwd|secret|token|api[_-]?key|access[_-]?key|client[_-]?secret|private[_-]?key)(\s*[:=]\s*)("[^"\n]*"|'[^'\n]*'|\S+)/gi,
61 replace: (_match, key: string, sep: string) => `${key}${sep}<secret>`,
62 },
63 { pattern: /\b(?:sk|pk|rk)-[A-Za-z0-9_-]{16,}/g, replace: () => "<secret>" },
64 { pattern: /\b(?:gh[pousr]_[A-Za-z0-9]{20,}|github_pat_[A-Za-z0-9_]{20,})\b/g, replace: () => "<secret>" },
65 { pattern: /\bxox[abprs]-[A-Za-z0-9-]{10,}/g, replace: () => "<secret>" },
66 { pattern: /\b(?:AKIA|ASIA)[0-9A-Z]{16}\b/g, replace: () => "<secret>" },
67 { pattern: /\bAIza[0-9A-Za-z_-]{30,}/g, replace: () => "<secret>" },
68 { pattern: /\beyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{4,}/g, replace: () => "<secret>" },
69 {
70 pattern: /\b(?:https?|wss?|ftp|ssh|git|postgres(?:ql)?|mysql|mariadb|redis|rediss|mongodb(?:\+srv)?|amqps?|s3):\/\/[^\s<>"'`)\]]+/gi,
71 replace: () => "<url>",
72 },
73 { pattern: /\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b/g, replace: () => "<email>" },
74 { pattern: /\b(?:\d{1,3}\.){3}\d{1,3}(?::\d{1,5})?\b/g, replace: () => "<ip>" },
75 { pattern: /\b[A-Za-z]:\\(?:[^\\\s]+\\)*[^\\\s]*/g, replace: () => "<path>" },
76 { pattern: /(?<![\w<>/.~@-])(?:~|\.{1,2})?\/(?:[\w.@+-]+\/)+[\w.@+-]*/g, replace: () => "<path>" },
77 { pattern: /(?<![\w<>/.~-])~\/[\w.@+-]+/g, replace: () => "<path>" },
78 { pattern: /\b[a-f0-9]{32,}\b/gi, replace: () => "<secret>" },
79 {
80 // Long base64-ish runs with both letters and digits: tokens, keys, signatures.
81 pattern: /(?<![\w+-])(?=[A-Za-z0-9+_-]*\d)(?=[A-Za-z0-9+_-]*[A-Za-z])[A-Za-z0-9+_-]{40,}={0,2}(?![\w+-])/g,
82 replace: () => "<secret>",
83 },
84];
85
86export function redact(input: string): { text: string; count: number } {
87 let count = 0;
88 let text = input;
89 for (const rule of RULES) {
90 text = text.replace(rule.pattern, (match: string, ...groups: unknown[]) => {
91 count++;
92 return rule.replace(match, ...(groups.filter((g) => typeof g === "string") as string[]));
93 });
94 }
95 return { text, count };
96}
97
98export function truncate(text: string, head = HEAD_CHARS, tail = TAIL_CHARS): string {
99 if (text.length <= head + tail) return text;
100 const omitted = text.length - head - tail;
101 return `${text.slice(0, head)}\n[... ${omitted} chars omitted ...]\n${text.slice(-tail)}`;
102}
103
104// A one-line preview for logs that stay on this machine.
105export function preview(text: string, length = 80): string {
106 const flat = text.replace(/\s+/g, " ").trim();
107 return flat.length > length ? `${flat.slice(0, length - 1)}…` : flat;
108}
109
110// The key a prompt is matched on between prompt.submit and turn.start.
111export function promptKey(text: string): string {
112 return text.replace(/\s+/g, " ").trim().slice(0, 500);
113}
114