Jev-guided model, effort and subagent routing for Claude Code (Opus 5.5 / Sonnet 5.5), auto-created specialist subagents and Jev-pruned compaction.

A mod for Claude Code that cuts what a long session costs without lowering the quality of the work. It uses Jev, a fast judge model on OpenRouter ($0.042 per million input tokens, ~300 ms), to decide per step how much model each step needs, and to prune old context verbatim instead of summarising it.
Mods shipped in Claude Code 2.1.287. This one uses only that official mechanism: function hooks inside the unmodified client. No proxy, no ANTHROPIC_BASE_URL, no OAuth token handling (the mods API has no access to it). Jev is called with your own OpenRouter key.
low…max) is picked on every turn. Switching to Sonnet 5.5 is priced in dollars: rewriting the context into the other model's cache against what Sonnet saves on output over the next turns (cache reads cost the same on both). Short continuations ("yes", "go on") are decided locally, without asking Jev.compaction.onReturn) also while you are away, once the cache has expired, so that the rewrite on your return is smaller. The mod also sets Claude Code's auto-compact window (compaction.autoWindowTokens, 250k) so the engine starts a compaction in the middle of a long turn and inside subagents; those go through Jev too.compaction.fold). Jev keeps the ones the current work still relies on. On a long real chat this removed 35–67% of the dialog text per compaction./jevg getctx builds a compact "capsule" of the chat (brief, work steps, changed files, latest checks); /jevg fresh does it in the same window: capsule, /clear, capsule attached to your next message.[secret] before every request. Projects in the "Projects without Jev" list (excludeProjects) send nothing.Details: docs/ARCHITECTURE.md, docs/DATA.md, docs/CONTEXT.md, docs/MONITORING.md (the last two are in Russian).
The UI shows two numbers: exact (computed from the journal: tokens a compaction or trim removed × the requests that would have re-read them from the cache, minus the cache rewrite it caused, plus model price differences) and estimated (the effort lever, a rule of thumb). In our own use over one day of mixed work it came to roughly 7–13% of the API-equivalent spend; on long single-task sessions with big contexts it was far higher in A/B runs (~60%). It does not know what you would have spent without it, so treat the numbers as a lower bound, and run your own A/B on one task with and without the mod before you trust them.
Requires Claude Code 2.1.287 or newer (2.1.286 with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1). A recent Node (it runs the TypeScript files directly; developed on Node 26) for the UI and scripts. macOS only for now: the projects board opens Terminal.app and uses pbcopy/osascript; the core (routing, compaction, trimming) does not depend on the OS but is untested elsewhere.
git clone https://github.com/roma-vibe/jev-governor.git
cd jev-governor
npm install
~/.claude/jev-governor/openrouter.key, or in OPENROUTER_API_KEY, or save it from the UI (Settings).env in ~/.claude/settings.json: {
"env": {
"CLAUDE_CODE_PLUGIN_DIRS": "/absolute/path/to/jev-governor"
}
}
For a one-off check: claude --plugin-dir /absolute/path/to/jev-governor.
jev ▸ model·effort │ 5h N% · 7d M%. No such line means the mod did not load; npm run validate shows why. jev ✕ means Jev has not answered several times in a row; the mod then changes nothing and Claude Code behaves as without it.Mods run with your permissions and are not sandboxed: read the code before you enable one. Everything this mod sends out goes to openrouter.ai with your key.
/jevg: status (key, last decision, limits, agents)/jevg on / /jevg off: enable or disable the mod/jevg chat on|off: switch the mod off (or back on) for this chat only; nothing is routed, trimmed, compacted or sent to Jev there. Remembered for the chat, also after a restart or resume/jevg idle on|off: compact this chat after a pause (cache expired) even though the setting compaction.onReturn is off/jevg ui: start the settings UI in the background (http://127.0.0.1:4777)/jevg reload: re-read config, key and the agent registry/jevg fresh [--brief|--nobrief] [focus]: continue in this window with a clean chat/jevg getctx [--brief|--nobrief] [focus]: a compact context of this chat for a new one (prompt copied to the clipboard)/jevg ctx [list|<id>]: in a new chat, attach a project capsule to the next messageThe mod's own messages (chat notices, toasts, command replies) and the UI come in English and Russian. Setting ui.language: auto (default) follows the language set in Claude Code (language in ~/.claude/settings.json), then the system language; en or ru forces one.
npm run ui:install
npm run ui:build
npm run ui
~/.claude/jev-governor/ (override with JEV_GOVERNOR_HOME): config.json, agents/, skills/, drafts/, ledger/ (decisions and token usage), archives of removed output. Formats: docs/DATA.md.
node scripts/monitor.mjs --since v0.2.6
Prints a report from the mod's journal and Claude Code's transcripts: context per step, compaction counts and how often the model needed something that was removed, routing, cost, and what to change. See docs/MONITORING.md.
npm test
npm run typecheck
npm run validate
sh scripts/install-hooks.sh
The pre-commit hook checks, on what is staged, that the engine loads the mod, the types and the tests.
hooks/lib/providers.ts) but is not implemented.MIT. The compaction library is adapted from fast-jev-compaction (MIT).
hooks/register.ts 2621 lines1// jev-governor: Jev decides how much model each step of a Claude Code session
2// needs. Main conversation: effort every turn, model only where the prompt
3// cache allows. Subagents: model + effort at spawn, mapped onto (or creating)
4// a specialist agent. Compaction: Jev prunes stale tool calls instead of a
5// lossy summary. The policies live in ./lib (pure, unit-tested); this file
6// wires them to the engine. Data: ~/.claude/jev-governor (see docs/DATA.md).
7
8import type { EngineInterface, Register } from 'claude-code';
9
10import type { Message } from './lib/compaction/types.ts';
11
12import { MOD_VERSION } from './lib/version.ts';
13import { budgetPressure, describePressure, type Pressure } from './lib/budget.ts';
14import { archiveText, editedPath, indexLine, isPointerText, rerunKey, withPointers, type PrunedCall } from './lib/archive.ts';
15import {
16 attachedPrompt,
17 BRIEF_MARKER,
18 briefPrompt,
19 buildSkeleton,
20 candidateTurns,
21 capsuleId,
22 CTX_REF,
23 parseGetctxArgs,
24 pastePrompt,
25 planTurns,
26 projectSlug,
27 renderCapsule,
28 turnQuestions,
29 turnState,
30 type CapsuleMessage,
31 type CapsuleMeta,
32} from './lib/capsule.ts';
33import { planFolds, runCompaction, summarizeCompaction } from './lib/compaction/adapter.ts';
34import { messageChars } from './lib/compaction/compact.ts';
35import { applyFolds, DEFAULT_FOLD_OPTIONS, foldArchiveText, foldFileName, foldKeeps, foldStats, type FoldOptions } from './lib/compaction/fold.ts';
36import { resolveLang, systemLocales } from './lib/lang.ts';
37import { expandHome, isExcluded, resolveConfig } from './lib/config.ts';
38import { clip, estimateContextTokens, recentHistory, type HistoryMessage } from './lib/history.ts';
39import { choice, jevAsker, noul, withTimeout, type JevAsker, type JevQuestions, type JevResponse } from './lib/jev.ts';
40import { displayModel } from './lib/providers.ts';
41import { redact } from './lib/redact.ts';
42import { chunkQuestions, isClaudeSavedOutput, isLogCandidate, isRunnerCommand, LOG_KEY_LINES, logQuestion, persistedOutputPath, MIN_TRIM_GAIN, planTrim, renderTrim, trimKind, worthTrimming, type Chunk } from './lib/trim.ts';
43import { commandDir, describePrompt, describeSystem, normalizeObserved, parseDescriptions } from './lib/commands.ts';
44import {
45 agentQuestion,
46 agentType,
47 composePrompt,
48 DRAFT_SYSTEM,
49 draftPrompt,
50 NONE,
51 parseDraft,
52 PLUGIN,
53 rankCandidates,
54 retireCandidate,
55 sanitizeAgent,
56 sanitizeSkill,
57 withRolePreamble,
58 withWaitNote,
59} from './lib/registry.ts';
60import {
61 decideMain,
62 decideSubagent,
63 escalate,
64 fallbackDecision,
65 isShortFollowUp,
66 MAIN_STATE_CONTEXT,
67 mainQuestions,
68 readSignals,
69 SUBAGENT_STATE_CONTEXT,
70 subagentQuestions,
71 tierOf,
72 type Decision,
73} from './lib/router.ts';
74import type {
75 AgentRecord,
76 DescribeRequest,
77 DraftRecord,
78 LedgerEntry,
79 SkillRecord,
80} from './lib/types.ts';
81
82import {
83 basename,
84 cacheWarm,
85 type CapsuleRecord,
86 errorText,
87 type Handoff,
88 handoffDir,
89 active,
90 idleCompaction,
91 jevKey,
92 kTokens,
93 learnTurn,
94 type MainOutcome,
95 nowIso,
96 outputsDir,
97 type PendingSpawn,
98 promptKey,
99 prunedOf,
100 type Rate,
101 rememberPruned,
102 resultText,
103 type Route,
104 L,
105 S,
106 shadow,
107 type SpawnPlan,
108 statusText,
109 type SubState,
110 usageBase,
111 addStepUsage,
112 noteToolResult,
113 archiveOf,
114 stepKey,
115 turnUsageParts,
116} from './mod/state.ts';
117
118// How specialists reach a subagent (verified against the engine, 2.1.286):
119// a spawn may only be rewritten to an agent type that was offered to the
120// model when the turn began; a hidden or freshly registered type is refused,
121// the Agent call fails, and `next` cannot be retried. So the spawn's type is
122// never rewritten: the specialist's role and skills ride in front of the
123// task of the original (generic) subagent, which also keeps the model's agent
124// listing, and its prompt cache, untouched. With `agents.exposeToModel` the
125// specialists are additionally registered and listed, so the model may call
126// one by name; such a call is routed like any other spawn.
127/** Most specialists offered to Jev per spawn (pre-ranked by word overlap). */
128const MAX_CANDIDATES = 12;
129const LEDGER_MAX_LINES = 5000;
130/** Jev failures in a row after which the status line shows `jev ✕`. */
131const JEV_FAILS_SHOWN = 3;
132/** Pause before the one retry of a routing request to Jev. */
133const JEV_RETRY_PAUSE_MS = 600;
134
135// ---------------------------------------------------------------- storage --
136
137async function readJson($: EngineInterface, path: string): Promise<unknown> {
138 try {
139 return JSON.parse(String(await $.fs.read(path)));
140 } catch {
141 return undefined;
142 }
143}
144
145/** The language of the mod's messages: the setting, or (automatic) Claude Code's language, then the system's. */
146async function setLang($: EngineInterface): Promise<void> {
147 const setting = S.cfg.ui.language;
148 const claude = setting === 'auto' ? (await readJson($, `${S.data.replace(/[\\/][^\\/]+[\\/]?$/, '')}/settings.json`) as { language?: unknown } | undefined)?.language : undefined;
149 S.lang = resolveLang(setting, claude, systemLocales());
150}
151
152async function loadConfig($: EngineInterface, force: boolean): Promise<void> {
153 const path = `${S.data}/config.json`;
154 try {
155 const stat = await $.fs.stat(path);
156 if (!force && stat.mtimeMs === S.cfgMtime) return;
157 S.cfg = resolveConfig(await readJson($, path));
158 S.cfgMtime = stat.mtimeMs;
159 await setLang($);
160 // A window changed in the settings takes effect at once (bindSession logs it at start).
161 if (!force) await applyAutoWindow($, false);
162 } catch {
163 // First run: write the defaults so the settings UI has a file to edit.
164 S.cfg = resolveConfig(undefined);
165 await setLang($);
166 try {
167 await $.fs.write(path, `${JSON.stringify(S.cfg, null, 2)}\n`);
168 S.cfgMtime = (await $.fs.stat(path)).mtimeMs;
169 } catch {
170 S.cfgMtime = -1;
171 }
172 }
173}
174
175/**
176 * Claude Code compacts on its own when the context reaches its auto-compaction window, also in
177 * the middle of a turn and inside subagents, where the mod cannot start one. A smaller window
178 * (`compaction.autoWindowTokens`) makes those compactions come sooner, and the session.compact
179 * hook prunes them through Jev like ours. The variable is the process's: the value the mod set is
180 * remembered in JEV_GOVERNOR_AUTO_WINDOW, so a value the person set is never overwritten and a
181 * reload or a config change can update or clear the mod's own. Logged once per session with what
182 * the engine actually measures against, so the monitor can tell whether it took.
183 */
184async function applyAutoWindow($: EngineInterface, log: boolean): Promise<void> {
185 try {
186 // Only where the hook prunes: in shadow mode, an excluded project or without a key a smaller
187 // window would only bring Claude Code's paid summaries sooner.
188 const prunes = active() && S.cfg.compaction.enabled && jevKey() !== undefined && !shadow();
189 const want = prunes ? S.cfg.compaction.autoWindowTokens : 0;
190 const current = await $.env.get('CLAUDE_CODE_AUTO_COMPACT_WINDOW');
191 const mine = await $.env.get('JEV_GOVERNOR_AUTO_WINDOW');
192 const ours = current !== undefined && current === mine;
193 let by: 'mod' | 'user' | 'none' = 'none';
194 if (current && !ours) by = 'user';
195 else if (want > 0) {
196 if (current !== String(want)) {
197 await $.env.set('CLAUDE_CODE_AUTO_COMPACT_WINDOW', String(want));
198 await $.env.set('JEV_GOVERNOR_AUTO_WINDOW', String(want));
199 }
200 by = 'mod';
201 } else if (ours) {
202 await $.env.set('CLAUDE_CODE_AUTO_COMPACT_WINDOW', undefined);
203 await $.env.set('JEV_GOVERNOR_AUTO_WINDOW', undefined);
204 }
205 S.autoWindow = by === 'mod' ? want : undefined;
206 // Once per session and version: a reload into new code logs again, so the monitor sees
207 // which code a long session runs and whether its window took.
208 const logKey = `${S.session}@${MOD_VERSION}`;
209 if (!log || S.autoWindowLogged === logKey) return;
210 S.autoWindowLogged = logKey;
211 let tokens: number | undefined;
212 let source: string | undefined;
213 try {
214 const b = (await $.session.usage({ breakdown: 'summary' })).context.breakdown;
215 tokens = b?.rawMaxTokens;
216 source = b?.autocompactSource;
217 } catch {
218 // No breakdown before the first request in some hosts: the entry still says what was set.
219 }
220 await ledger($, {
221 kind: 'window',
222 text: by === 'mod' ? `auto-compaction window ${kTokens(want)}` : by === 'user' ? `auto-compaction window set outside the mod (${current})` : 'auto-compaction window: Claude Code default',
223 autoWindow: { by, tokens, source, ...(want > 0 ? { wanted: want } : {}) },
224 });
225 } catch (error) {
226 await ledger($, { kind: 'error', error: `auto window: ${errorText(error)}` });
227 }
228}
229
230/** The model a subagent runs on, as routed (the main chat's when the mod did not route it). */
231function subModel(agentId: string): string | undefined {
232 const sub = S.subs.get(agentId);
233 return sub ? S.cfg.models[sub.tier] : S.route.lastModel;
234}
235
236async function saveConfig($: EngineInterface): Promise<void> {
237 const path = `${S.data}/config.json`;
238 await $.fs.write(path, `${JSON.stringify(S.cfg, null, 2)}\n`);
239 S.cfgMtime = (await $.fs.stat(path)).mtimeMs;
240}
241
242async function loadKey($: EngineInterface): Promise<void> {
243 const fromEnv = await $.env.get('OPENROUTER_API_KEY');
244 if (fromEnv && fromEnv.trim()) {
245 S.key = fromEnv.trim();
246 return;
247 }
248 for (const path of [expandHome(S.cfg.jev.keyFile, S.home), `${$.plugin.root}/.openrouter_key`]) {
249 try {
250 const key = String(await $.fs.read(path)).trim();
251 if (key) {
252 S.key = key;
253 return;
254 }
255 } catch {
256 // try the next place
257 }
258 }
259 S.key = undefined;
260}
261
262async function registrySignature($: EngineInterface): Promise<string> {
263 const parts: string[] = [];
264 for (const dir of ['agents', 'skills']) {
265 try {
266 for (const entry of await $.fs.list(`${S.data}/${dir}`)) {
267 if (entry.name.endsWith('.json')) parts.push(`${dir}/${entry.name}:${entry.mtimeMs}`);
268 }
269 } catch {
270 // no such directory yet
271 }
272 }
273 return parts.sort().join('|');
274}
275
276async function loadRegistry($: EngineInterface): Promise<boolean> {
277 const sig = await registrySignature($);
278 if (sig === S.registrySig) return false;
279 const now = nowIso();
280 const agents = new Map<string, AgentRecord>();
281 const skills = new Map<string, SkillRecord>();
282 for (const [dir, into] of [
283 ['skills', skills],
284 ['agents', agents],
285 ] as const) {
286 let entries: { name: string }[] = [];
287 try {
288 entries = await $.fs.list(`${S.data}/${dir}`);
289 } catch {
290 continue;
291 }
292 for (const entry of entries) {
293 if (!entry.name.endsWith('.json')) continue;
294 const raw = await readJson($, `${S.data}/${dir}/${entry.name}`);
295 if (dir === 'skills') {
296 const skill = sanitizeSkill(raw, now);
297 if (skill) skills.set(skill.name, skill);
298 } else {
299 const agent = sanitizeAgent(raw, now);
300 if (agent) (into as Map<string, AgentRecord>).set(agent.name, agent);
301 }
302 }
303 }
304 S.agents = agents;
305 S.skills = skills;
306 S.registrySig = sig;
307 return true;
308}
309
310async function registerAgent($: EngineInterface, agent: AgentRecord): Promise<void> {
311 if (!agent.enabled || !S.cfg.agents.exposeToModel) return;
312 await $.agent.register({
313 name: agent.name,
314 description: agent.description,
315 prompt: composePrompt(agent, S.skills),
316 ...(agent.tools ? { tools: agent.tools } : {}),
317 ...(agent.effort !== 'auto' ? { effort: agent.effort } : {}),
318 });
319}
320
321async function registerAll($: EngineInterface): Promise<void> {
322 for (const agent of S.agents.values()) {
323 try {
324 await registerAgent($, agent);
325 } catch (error) {
326 await ledger($, { kind: 'error', error: `register ${agent.name}: ${errorText(error)}` });
327 }
328 }
329}
330
331async function persistAgent($: EngineInterface, agent: AgentRecord, skills: readonly SkillRecord[]): Promise<void> {
332 for (const skill of skills) {
333 await $.fs.write(`${S.data}/skills/${skill.name}.json`, `${JSON.stringify(skill, null, 2)}\n`);
334 S.skills.set(skill.name, skill);
335 }
336 await $.fs.write(`${S.data}/agents/${agent.name}.json`, `${JSON.stringify(agent, null, 2)}\n`);
337 S.agents.set(agent.name, agent);
338 S.registrySig = await registrySignature($);
339}
340
341/** Appends to ledger/<day>/<session>.jsonl (the file is rewritten whole: $.fs has no append). */
342async function ledger($: EngineInterface, entry: Omit<LedgerEntry, 'ts' | 'session'>): Promise<void> {
343 if (!S.data) return;
344 const ts = nowIso();
345 const line = JSON.stringify({ ts, session: S.session, project: S.project, v: MOD_VERSION, ...entry });
346 const path = `${S.data}/ledger/${ts.slice(0, 10)}/${S.session}.jsonl`;
347 const run = async (): Promise<void> => {
348 let lines = S.ledgerLines.get(path);
349 if (!lines) {
350 lines = [];
351 try {
352 lines = String(await $.fs.read(path)).split('\n').filter(Boolean);
353 } catch {
354 // new file
355 }
356 S.ledgerLines.set(path, lines);
357 }
358 lines.push(line);
359 if (lines.length > LEDGER_MAX_LINES) lines.splice(0, lines.length - LEDGER_MAX_LINES);
360 await $.fs.write(path, `${lines.join('\n')}\n`);
361 };
362 S.ledgerChain = S.ledgerChain.then(run, run).catch(() => undefined);
363 await S.ledgerChain;
364}
365
366async function saveRoute($: EngineInterface): Promise<void> {
367 try {
368 await $.store.set(`route:${S.session}`, { ...S.route, savedAt: Date.now() });
369 } catch {
370 // best effort
371 }
372}
373
374/** The chat's own switches (`/jevg chat`, `/jevg idle`) outlive a restart and a resume of the chat. */
375async function saveChat($: EngineInterface): Promise<void> {
376 try {
377 await $.store.set(`chat:${S.session}`, { ...S.chat, savedAt: Date.now() });
378 } catch {
379 // best effort
380 }
381}
382
383async function restoreChat($: EngineInterface): Promise<void> {
384 try {
385 const saved = (await $.store.get(`chat:${S.session}`)) as { off?: unknown; idle?: unknown } | undefined;
386 S.chat = { off: saved?.off === true, idle: saved?.idle === true };
387 } catch {
388 S.chat = { off: false, idle: false };
389 }
390}
391
392async function restoreRoute($: EngineInterface): Promise<void> {
393 try {
394 const saved = (await $.store.get(`route:${S.session}`)) as Route | undefined;
395 if (saved && typeof saved === 'object') S.route = { ...saved, freeSwitch: Boolean(saved.freeSwitch) };
396 } catch {
397 // fresh route
398 }
399}
400
401// -------------------------------------------------------------------- Jev --
402
403/** A Jev client over the engine's fetch; secrets are replaced in every request and counted in the ledger. */
404function makeAsker($: EngineInterface, key: string, purpose: string): JevAsker {
405 return jevAsker({
406 http: async (url, init) => {
407 const response = await $.http.fetch(url, init);
408 return { status: response.status, ok: response.ok, text: response.text };
409 },
410 endpoint: S.cfg.jev.endpoint,
411 apiKey: key,
412 model: S.cfg.jev.model,
413 onRedacted: (count) => void ledger($, { kind: 'redacted', count, text: purpose }),
414 });
415}
416
417async function askJev(
418 $: EngineInterface,
419 state: object,
420 questions: JevQuestions,
421 timeoutMs = S.cfg.jev.timeoutMs,
422 purpose = 'decision',
423): Promise<{ response: JevResponse; ms: number } | undefined> {
424 const key = jevKey();
425 if (!key) return undefined;
426 const asker = makeAsker($, key, purpose);
427 const started = Date.now();
428 // Routing decisions get one more try after a brief pause: Jev's upstream
429 // drops single requests (503 "no healthy upstream", a timeout) and a turn
430 // without a decision runs on the session's own effort. Not while Jev keeps
431 // failing (each turn would wait twice for nothing), not for trims or
432 // compactions (they keep everything without Jev anyway).
433 const retry = (purpose === 'decision' || purpose === 'subagent') && S.jevFailures === 0;
434 const once = (): Promise<JevResponse> => withTimeout(asker.ask(state, questions), (ms) => $.clock.sleep(ms), timeoutMs);
435 try {
436 let response: JevResponse;
437 try {
438 response = await once();
439 } catch (error) {
440 if (!retry || !isRetriable(error)) throw error;
441 await ledger($, { kind: 'error', error: `jev: ${errorText(error)} (retrying)` });
442 await $.clock.sleep(JEV_RETRY_PAUSE_MS);
443 response = await once();
444 }
445 S.jevFailures = 0;
446 return { response, ms: Date.now() - started };
447 } catch (error) {
448 S.jevFailures++;
449 if (S.jevFailures >= JEV_FAILS_SHOWN && S.cfg.ui.showStatus) {
450 $.ui.status(
451 L(
452 `jev ✕ Jev не отвечает (${S.jevFailures} ${times(S.jevFailures)} подряд): модель и effort — прошлого решения или запасные`,
453 `jev ✕ Jev is not answering (${S.jevFailures} ${times(S.jevFailures)} in a row): model and effort are from the last decision or the fallbacks`,
454 ),
455 );
456 }
457 await ledger($, { kind: 'error', error: `jev: ${errorText(error)}` });
458 return undefined;
459 }
460}
461
462/** A Jev failure worth one more try: a timeout, a network error, 429 or 5xx (not a bad key or request). */
463function isRetriable(error: unknown): boolean {
464 const text = errorText(error);
465 const status = /\((\d{3})\)/.exec(text)?.[1];
466 return status === undefined ? true : status === '429' || status.startsWith('5');
467}
468
469/** "раз" in Russian after a count: 2 раза, 5 раз, 21 раз, 22 раза (English: "times"). */
470function times(n: number): string {
471 if (S.lang !== 'ru') return 'times';
472 const last = n % 10;
473 return last >= 2 && last <= 4 && (n % 100 < 12 || n % 100 > 14) ? 'раза' : 'раз';
474}
475
476async function pressureNow(
477 $: EngineInterface,
478): Promise<{ pressure: Pressure; text: string; contextTokens?: number; rate: Rate }> {
479 try {
480 const usage = await $.session.usage();
481 const paced = budgetPressure(usage.rateLimits, Date.now());
482 const pct = (kind: string): number | undefined => usage.rateLimits.find((w) => w.kind === kind)?.percentUsed;
483 return {
484 pressure: S.cfg.router.budgetAware ? paced.pressure : 0,
485 text: describePressure(paced.windows),
486 contextTokens: usage.context.tokens,
487 rate: { fiveHour: pct('five_hour'), sevenDay: pct('seven_day') },
488 };
489 } catch {
490 return { pressure: 0, text: '', rate: {} };
491 }
492}
493
494// ------------------------------------------------------------- main loop --
495
496async function decideMainTurn($: EngineInterface, text: string): Promise<MainOutcome | undefined> {
497 const cfg = S.cfg;
498 if (!cfg.router.mainModel && !cfg.router.mainEffort) return undefined;
499 const budget = await pressureNow($);
500 const sessionModel = await $.session.model();
501 // A /model change by the person: adopt it, do not fight it this turn.
502 const manual = S.route.lastSessionModel !== undefined && S.route.lastSessionModel !== sessionModel;
503 if (manual) {
504 await ledger($, {
505 kind: 'override',
506 scope: 'main',
507 text: clip(text, 160),
508 model: sessionModel,
509 prevModel: S.route.lastModel,
510 reasons: [`/model ${displayModel(S.route.lastSessionModel ?? '')} → ${displayModel(sessionModel)}`],
511 });
512 }
513 S.route.lastSessionModel = sessionModel;
514 // `lastModel` stays the one the last turn ran on: turn.step sees the /model switch as a switch
515 // (the cache is rewritten), and what follows keeps the person's model.
516 const current = tierOf(manual ? sessionModel : (S.route.lastModel ?? sessionModel), cfg) ?? 'strong';
517 if (manual && S.route.previous) S.route.previous = { ...S.route.previous, tier: current };
518 const idle =
519 S.route.lastRequestAt !== undefined &&
520 Date.now() - S.route.lastRequestAt > cfg.router.cacheTtlMinutes * 60_000;
521 const freeSwitch = S.route.freeSwitch || S.route.lastRequestAt === undefined || idle;
522 // Without Jev's view (no text, a bare "go on", Jev down): the previous decision, else nothing changes.
523 const keep = (why: string, local = false): MainOutcome => ({
524 decision: S.route.previous
525 ? { ...S.route.previous, switched: false, reasons: [`${why}: previous decision kept`] }
526 : { tier: current, effort: cfg.router.defaultEffort, keepEffort: true, switched: false, reasons: [`${why}: nothing changed`] },
527 pressure: budget.pressure,
528 pressureText: budget.text,
529 rate: budget.rate,
530 contextTokens: budget.contextTokens,
531 cacheCold: idle,
532 ...(manual ? { manual: true } : {}),
533 ...(local ? { local: true } : {}),
534 });
535
536 if (!text.trim()) return keep('no prompt text');
537 // "да", "давай", "go on": nothing to judge, so nothing is sent to Jev.
538 if (isShortFollowUp(text)) return keep('short follow-up (no Jev)', true);
539 let messages: HistoryMessage[] = [];
540 try {
541 messages = (await $.session.messages()) as unknown as HistoryMessage[];
542 } catch {
543 messages = [];
544 }
545 // Before the session's first response the engine has no context figure yet.
546 const contextTokens = budget.contextTokens ?? estimateContextTokens(messages);
547 const state = {
548 context: MAIN_STATE_CONTEXT,
549 project: S.project,
550 latest_user_request: clip(text, 3000, 600),
551 recent_conversation: recentHistory(messages, { maxTokens: 3500, maxMessages: 14, latest: text }),
552 current_model: current === 'strong' ? 'Opus 5.5' : 'Sonnet 5.5',
553 };
554 const asked = await askJev($, state, mainQuestions());
555 const signals = asked ? readSignals(asked.response.answers) : undefined;
556 // Jev down: the session's own model would be a switch nobody weighed.
557 if (!signals) {
558 if (S.route.previous || !cfg.router.mainEffort) return keep('Jev unavailable');
559 return { ...keep('Jev unavailable'), decision: fallbackDecision(current, 'Jev unavailable, no earlier decision', cfg) };
560 }
561 const decision = decideMain(
562 signals,
563 {
564 currentTier: current,
565 freeSwitch,
566 contextTokens,
567 pressure: budget.pressure,
568 previous: S.route.previous,
569 turnProfile: S.route.turnProfile,
570 },
571 cfg,
572 );
573 if (manual) {
574 decision.tier = current;
575 decision.switched = false;
576 decision.reasons.push('model set by /model: kept');
577 }
578 if (!cfg.router.mainEffort) decision.reasons.push('effort routing off');
579 return {
580 decision,
581 signals,
582 pressure: budget.pressure,
583 pressureText: budget.text,
584 rate: budget.rate,
585 contextTokens,
586 jevMs: asked?.ms,
587 jevCost: asked?.response.usage?.cost,
588 cacheCold: idle,
589 ...(manual ? { manual: true } : {}),
590 };
591}
592
593// -------------------------------------------------------------- subagents --
594
595async function createAgent(
596 $: EngineInterface,
597 input: { title: string; task: string },
598): Promise<AgentRecord | undefined> {
599 const reply = await $.model.complete({
600 model: S.cfg.agents.draftModel,
601 system: DRAFT_SYSTEM,
602 prompt: draftPrompt({
603 task: input.task,
604 title: input.title,
605 project: S.project,
606 agents: [...S.agents.values()],
607 skills: [...S.skills.values()],
608 }),
609 maxTokens: 1500,
610 effort: 'low',
611 timeoutMs: S.cfg.agents.draftTimeoutMs,
612 });
613 if (!reply.isAnswered) {
614 await ledger($, { kind: 'error', error: `draft agent: ${reply.reason}` });
615 return undefined;
616 }
617 const draft = parseDraft(reply.text, { agents: new Set(S.agents.keys()), skills: S.skills }, nowIso());
618 if (!draft) {
619 await ledger($, { kind: 'error', error: 'draft agent: unusable reply', text: clip(reply.text, 200) });
620 return undefined;
621 }
622 await persistAgent($, draft.agent, draft.skills);
623 await registerAgent($, draft.agent);
624 await ledger($, {
625 kind: 'agent-created',
626 agent: draft.agent.name,
627 text: clip(input.title, 160),
628 reasons: draft.skills.map((s) => `new skill ${s.name}`),
629 });
630 $.ui.toast(`jev-governor: new subagent ${draft.agent.name}`, { timeoutMs: 6000 });
631 return draft.agent;
632}
633
634/**
635 * The session's project root: where it started, or where `/cd` or a worktree
636 * took it. A shell `cd` into a subfolder does not move it (`$.session.cwd()`
637 * would): the project, its capsules and its transcript stay the same.
638 */
639async function projectRoot($: EngineInterface): Promise<string> {
640 try {
641 return await $.session.root();
642 } catch {
643 return await $.session.cwd();
644 }
645}
646
647/** The session may move to another project after it starts: shadow/active follows it. */
648async function refreshCwd($: EngineInterface): Promise<void> {
649 try {
650 S.cwd = await projectRoot($);
651 S.project = basename(S.cwd);
652 } catch {
653 // keep the last known folder
654 }
655}
656
657async function planSpawn(
658 $: EngineInterface,
659 e: { subagentType: string; description: string; prompt: string; parentModel: string; model?: string },
660): Promise<SpawnPlan | undefined> {
661 await loadConfig($, false);
662 await refreshCwd($);
663 const cfg = S.cfg;
664 if (await loadRegistry($)) await registerAll($);
665 const known = new Set(S.agents.keys());
666 const ownName = e.subagentType.startsWith(`${PLUGIN}:`) ? e.subagentType.slice(PLUGIN.length + 1) : undefined;
667 const remap = cfg.agents.enabled && cfg.agents.remapFrom.includes(e.subagentType);
668 const candidates = remap ? rankCandidates([...S.agents.values()], `${e.description} ${e.prompt}`, MAX_CANDIDATES) : [];
669 const questions: JevQuestions = {
670 ...(cfg.router.subagents ? subagentQuestions() : {}),
671 ...(candidates.length > 0 ? { agent: agentQuestion(candidates) } : {}),
672 };
673 const budget = await pressureNow($);
674 const asked =
675 Object.keys(questions).length > 0
676 ? await askJev(
677 $,
678 {
679 context: SUBAGENT_STATE_CONTEXT,
680 project: S.project,
681 parent_request: clip(S.turn?.text ?? '', 600, 200),
682 subagent_type: e.subagentType,
683 task_title: e.description,
684 task: clip(e.prompt, 6000, 1000),
685 },
686 questions,
687 undefined,
688 'subagent',
689 )
690 : undefined;
691 // Jev down: no specialist is matched or drafted (a draft per spawn would duplicate the
692 // registry), but the subagent still gets the fallback effort and the wait note.
693 const jevDown = Object.keys(questions).length > 0 && !asked;
694
695 let agent = ownName ? S.agents.get(ownName) : undefined;
696 let created = false;
697 if (remap && !jevDown) {
698 const pick = asked ? choice(asked.response.answers, 'agent') : undefined;
699 const p = pick ? (pick.probabilities[pick.choice] ?? 0) : 0;
700 if (pick && pick.choice !== NONE && p >= cfg.agents.matchAt && S.agents.get(pick.choice)?.enabled) {
701 agent = S.agents.get(pick.choice);
702 } else if (cfg.agents.autoCreate && !shadow()) {
703 ({ agent, created } = await draftOrReuse($, e, known));
704 }
705 }
706 if (agent && !created && (remap || ownName)) void noteUse($, agent);
707
708 const signals = asked ? readSignals(asked.response.answers) : undefined;
709 const pinnedTier = agent && agent.tier !== 'auto' ? agent.tier : undefined;
710 const pinnedEffort = agent && agent.effort !== 'auto' ? agent.effort : undefined;
711 let decision: Decision | undefined;
712 // The fallback leaves the model alone unless the agent pins it (an Explore agent keeps its own).
713 let keepModel = false;
714 if (cfg.router.subagents && jevDown) {
715 decision = fallbackDecision(pinnedTier ?? tierOf(e.model ?? e.parentModel, cfg) ?? 'strong', 'Jev unavailable', cfg, pinnedEffort);
716 keepModel = pinnedTier === undefined;
717 } else if (cfg.router.subagents && signals) {
718 decision = decideSubagent(
719 signals,
720 {
721 pressure: budget.pressure,
722 subagentType: agent ? agentType(agent.name) : e.subagentType,
723 readOnly:
724 e.subagentType === 'Explore' ||
725 (agent?.tools !== undefined && !agent.tools.some((t) => ['Edit', 'Write', 'NotebookEdit'].includes(t))),
726 pinnedTier,
727 pinnedEffort,
728 },
729 cfg,
730 );
731 } else if (pinnedTier) {
732 decision = { tier: pinnedTier, effort: pinnedEffort ?? cfg.router.defaultEffort, switched: false, reasons: ['pinned by agent'] };
733 }
734
735 let prompt = e.prompt;
736 if (agent && !ownName) prompt = withRolePreamble(agent, S.skills, e.prompt);
737 if (cfg.agents.waitHint) prompt = withWaitNote(prompt);
738 return {
739 rate: budget.rate,
740 prompt,
741 model: decision && !keepModel ? (decision.light && cfg.router.lightSubagents === 'on' ? cfg.models.light : cfg.models[decision.tier]) : undefined,
742 decision,
743 agent,
744 created,
745 signals,
746 pressure: budget.pressure,
747 jevMs: asked?.ms,
748 jevCost: asked?.response.usage?.cost,
749 };
750}
751
752/** Days without a run before a full registry may turn an auto-drafted specialist off. */
753const RETIRE_IDLE_DAYS = 14;
754
755/**
756 * Drafts a specialist for a spawn no existing one fits, one draft at a time. Parallel spawns
757 * of one task otherwise drafted near-copies (`judge-packet-v2-grader` and `-grader-2` four
758 * seconds apart): a spawn that waited first asks Jev whether a specialist drafted meanwhile
759 * fits, and a new draft sees every earlier one. A full registry turns off its longest-idle
760 * auto-drafted specialist to make room.
761 */
762async function draftOrReuse(
763 $: EngineInterface,
764 e: { description: string; prompt: string },
765 known: ReadonlySet<string>,
766): Promise<{ agent?: AgentRecord; created: boolean }> {
767 const prior = S.draftLock;
768 let release = (): void => undefined;
769 S.draftLock = new Promise<void>((resolve) => (release = resolve));
770 try {
771 await prior;
772 const fresh = [...S.agents.values()].filter((a) => a.enabled && !known.has(a.name));
773 if (fresh.length > 0) {
774 const asked = await askJev(
775 $,
776 { context: SUBAGENT_STATE_CONTEXT, project: S.project, task_title: e.description, task: clip(e.prompt, 6000, 1000) },
777 { agent: agentQuestion(fresh) },
778 undefined,
779 'subagent',
780 );
781 const pick = asked ? choice(asked.response.answers, 'agent') : undefined;
782 const p = pick ? (pick.probabilities[pick.choice] ?? 0) : 0;
783 const reuse = pick && pick.choice !== NONE && p >= S.cfg.agents.matchAt ? S.agents.get(pick.choice) : undefined;
784 if (reuse?.enabled) return { agent: reuse, created: false };
785 }
786 const enabled = [...S.agents.values()].filter((a) => a.enabled);
787 if (enabled.length >= S.cfg.agents.maxAgents) {
788 const old = retireCandidate(enabled, Date.now(), RETIRE_IDLE_DAYS);
789 if (!old) return { created: false };
790 await retireAgent($, old);
791 }
792 const agent = await createAgent($, { title: e.description, task: e.prompt }).catch(async (error) => {
793 await ledger($, { kind: 'error', error: `create agent: ${errorText(error)}` });
794 return undefined;
795 });
796 if (agent) await noteUse($, agent);
797 return { agent, created: agent !== undefined };
798 } finally {
799 release();
800 }
801}
802
803async function retireAgent($: EngineInterface, agent: AgentRecord): Promise<void> {
804 const now = nowIso();
805 const off: AgentRecord = { ...agent, enabled: false, retiredAt: now, updatedAt: now };
806 try {
807 await $.fs.write(`${S.data}/agents/${agent.name}.json`, `${JSON.stringify(off, null, 2)}\n`);
808 S.agents.set(agent.name, off);
809 S.registrySig = await registrySignature($);
810 await ledger($, {
811 kind: 'agent-created',
812 agent: agent.name,
813 text: `retired: registry full (${S.cfg.agents.maxAgents})`,
814 reasons: [`uses ${agent.uses ?? 0}, last ${(agent.lastUsedAt ?? agent.createdAt).slice(0, 10)}`],
815 });
816 } catch (error) {
817 await ledger($, { kind: 'error', error: `retire ${agent.name}: ${errorText(error)}` });
818 }
819}
820
821/** Counts a spawn the specialist ran (what a full registry retires by). */
822async function noteUse($: EngineInterface, agent: AgentRecord): Promise<void> {
823 const current = S.agents.get(agent.name) ?? agent;
824 const used: AgentRecord = { ...current, uses: (current.uses ?? 0) + 1, lastUsedAt: nowIso() };
825 try {
826 await $.fs.write(`${S.data}/agents/${agent.name}.json`, `${JSON.stringify(used, null, 2)}\n`);
827 S.agents.set(agent.name, used);
828 S.registrySig = await registrySignature($);
829 } catch {
830 // best effort: a missed count only makes the specialist look idler
831 }
832}
833
834/**
835 * Maps a subagent whose first steps run before its spawn resolved onto that
836 * spawn, by the task text its transcript opens with.
837 */
838async function claimPending($: EngineInterface, agentId: string): Promise<SubState | undefined> {
839 const now = Date.now();
840 S.pending = S.pending.filter((p) => now - p.at < 120_000);
841 if (S.pending.length === 0) return undefined;
842 const messages = await $.session.messages({ agentId });
843 if (!Array.isArray(messages)) return undefined;
844 const first = messages.find((m) => m.role === 'user' && m.text.trim())?.text ?? '';
845 const index = S.pending.findIndex((p) => first.includes(p.promptKey));
846 if (index < 0) return undefined;
847 const [claimed] = S.pending.splice(index, 1);
848 S.subs.set(agentId, claimed!.sub);
849 return claimed!.sub;
850}
851
852/**
853 * Trims one tool result if it qualifies; returns the replacement text and a
854 * ledger record, or undefined to keep it whole. In shadow mode it only
855 * measures (no Jev, no file, nothing replaced).
856 */
857async function trimResult(
858 $: EngineInterface,
859 toolUseId: string,
860 text: string,
861 isError: boolean,
862): Promise<{ text: string; entry: Omit<LedgerEntry, 'ts' | 'session'> } | undefined> {
863 const t = S.cfg.trim;
864 const call = S.calls.get(toolUseId);
865 const tool = call?.tool ?? 'unknown';
866 // Claude Code already replaced a big Bash output with a 2KB preview of its
867 // start: for a test or build run, trim the saved full output instead, so the
868 // outcome and the summary at its end are in the conversation. The file is the
869 // one the engine named for this call; the path in the preview text (which the
870 // command printed, so could forge) only when it is in Claude Code's own folder.
871 const named = tool === 'Bash' && isRunnerCommand(call?.command) ? persistedOutputPath(text) : undefined;
872 const persisted = named === undefined ? undefined : (call?.persisted ?? (isClaudeSavedOutput(named, S.home) ? named : undefined));
873 if (persisted) {
874 let full: string;
875 try {
876 full = String(await $.fs.read(persisted));
877 } catch {
878 return undefined;
879 }
880 // Claude Code saves runs that succeeded (a failing one keeps its tail inline): the outcome,
881 // the summary and any warning lines are what the preview lacked, not 16k of the log.
882 const trimmed = await trimFull($, toolUseId, full, isError, persisted, isError ? undefined : BRIEF_TRIM);
883 if (!trimmed) return undefined;
884 // What entered the history otherwise was the preview: charsAfter above it is a cost, not a saving.
885 trimmed.entry.trim = { ...trimmed.entry.trim!, charsBefore: text.length, persisted: true, fullChars: full.length };
886 return trimmed;
887 }
888 const kind = trimKind(tool, call?.command, text, t, isError);
889 if (!kind) {
890 const logLike = tool === 'Bash' && !isError && t.logs !== 'off' && text.length >= t.logChars && text.length < t.hugeChars && isLogCandidate(call?.command);
891 return logLike ? trimLog($, toolUseId, text) : undefined;
892 }
893 const trimmed = await trimFull($, toolUseId, text, isError, undefined, kind === 'list' ? LIST_TRIM : kind === 'brief' ? BRIEF_TRIM : undefined);
894 if (kind !== 'brief') return trimmed;
895 // A short run that a brief trim barely shortens is not worth a ledger line.
896 if (!trimmed || trimmed.entry.trim?.skipped) return undefined;
897 trimmed.entry.trim = { ...trimmed.entry.trim!, kind: 'brief' };
898 return trimmed;
899}
900
901/** What Jev is told the work is: the subagent's task inside a subagent, else the user's request. */
902function taskFor(call: { agentId?: string } | undefined): string {
903 const sub = call?.agentId ? S.subs.get(call.agentId)?.task : undefined;
904 return sub ?? clip(S.turn?.text ?? '', 1500, 300);
905}
906
907/** Settings for a log: its start, its end, every error and warning; Jev may put middle parts back. */
908const LOG_TRIM = { headLines: 10, tailLines: 30, contextLines: 2, maxChars: 5000, keepLines: LOG_KEY_LINES, collapseSimilar: true } as const;
909/** What the header of a log trim says it was. */
910const LOG_REASON = 'Jev judged this output a log (progress and status lines), not data asked for.';
911
912/**
913 * A command output that may be a log (`isLogCandidate`): Jev is asked in one request
914 * whether it is a log or data the assistant asked for, and which omitted parts still
915 * matter. Trimmed only when it is a log with probability `logAt` or more and the trim
916 * removes MIN_TRIM_GAIN; any doubt, a failed request or nothing to gain keeps it
917 * whole. In `shadow` mode the verdict is logged and nothing is cut.
918 */
919async function trimLog(
920 $: EngineInterface,
921 toolUseId: string,
922 text: string,
923): Promise<{ text: string; entry: Omit<LedgerEntry, 'ts' | 'session'> } | undefined> {
924 const t = { ...S.cfg.trim, ...LOG_TRIM };
925 const call = S.calls.get(toolUseId);
926 if (!jevKey()) return undefined;
927 const plan = planTrim(text, false, t);
928 // The least a trim keeps: when even that saves too little, Jev is not asked.
929 if (!worthTrimming(text.length, renderTrim(plan, { settings: t, originalChars: text.length }).charsAfter)) return undefined;
930 const chunks: Chunk[] = plan.candidates.slice(0, 20);
931 const asked = await askJev(
932 $,
933 {
934 context:
935 'A coding assistant ran a shell command. If its output is only a log, it is trimmed before it enters the conversation: kept_output stays, omitted_chunks are dropped unless needed, and the full output is saved to a file the assistant can search.',
936 task: taskFor(call),
937 command: clip(call?.command ?? 'Bash', 600),
938 output: clip(text, 6000, 3000),
939 kept_output: clip(renderTrim(plan, { settings: t, originalChars: text.length }).text, 5000, 2000),
940 omitted_chunks: Object.fromEntries(chunks.map((c) => [c.id, c.text])),
941 },
942 { ...logQuestion(), ...chunkQuestions(chunks) },
943 undefined,
944 'trim',
945 );
946 if (!asked) return undefined;
947 const logProb = noul(asked.response.answers, 'is_log') ?? 0;
948 const approved = new Set(chunks.filter((c) => (noul(asked.response.answers, `need_${c.id}`) ?? 1) >= t.jevKeepAt).map((c) => c.id));
949 const isLog = logProb >= S.cfg.trim.logAt;
950 const live = S.cfg.trim.logs === 'on' && !shadow();
951 const would = renderTrim(plan, { approved, fullPath: `${outputsDir()}/${toolUseId}.txt`, settings: t, originalChars: text.length, reason: LOG_REASON });
952 const worth = worthTrimming(text.length, would.charsAfter);
953 const skipped = !isLog ? `Jev: data, not a log (${logProb.toFixed(2)})` : !worth ? `removes under ${Math.round(MIN_TRIM_GAIN * 100)}%` : undefined;
954 let applied = live && skipped === undefined;
955 let path: string | undefined;
956 if (applied) {
957 path = `${outputsDir()}/${toolUseId}.txt`;
958 try {
959 await $.fs.write(path, text);
960 } catch {
961 // No saved copy, no trim: what is cut must stay one search away.
962 applied = false;
963 path = undefined;
964 }
965 }
966 const out = applied ? renderTrim(plan, { approved, fullPath: path, settings: t, originalChars: text.length, reason: LOG_REASON }) : would;
967 return {
968 text: out.text,
969 entry: {
970 kind: 'trim',
971 scope: call?.agentId ? 'subagent' : 'main',
972 agentId: call?.agentId,
973 applied,
974 jevMs: asked.ms,
975 jevCost: asked.response.usage?.cost,
976 text: clip(call?.command ?? 'Bash', 160),
977 trim: {
978 tool: 'Bash',
979 command: call?.command ? clip(call.command, 200) : undefined,
980 charsBefore: text.length,
981 charsAfter: out.charsAfter,
982 linesBefore: out.linesBefore,
983 linesAfter: out.linesAfter,
984 outcome: out.outcome.detail,
985 jevChunks: out.jevChunks,
986 path,
987 kind: 'log',
988 logProb: Math.round(logProb * 100) / 100,
989 ...(skipped ? { skipped } : {}),
990 },
991 },
992 };
993}
994
995/** Settings for a short trim: the outcome, summary and signal lines, little else. */
996const BRIEF_TRIM = { headLines: 5, tailLines: 15, contextLines: 1, maxChars: 3000 } as const;
997/** Settings for a long listing: its first and last rows; the rest is one grep away in the saved file. */
998const LIST_TRIM = { headLines: 40, tailLines: 15, contextLines: 0, maxChars: 4000 } as const;
999
1000/**
1001 * Trims `text`; `savedAt` is where the full output already is (else it is saved
1002 * to outputs/). A `preset` (BRIEF_TRIM, LIST_TRIM) replaces the line budgets
1003 * and asks Jev for nothing. A trim that would not remove MIN_TRIM_GAIN of a
1004 * live output keeps it whole (logged as not applied).
1005 */
1006async function trimFull(
1007 $: EngineInterface,
1008 toolUseId: string,
1009 text: string,
1010 isError: boolean,
1011 savedAt?: string,
1012 preset?: typeof BRIEF_TRIM | typeof LIST_TRIM,
1013): Promise<{ text: string; entry: Omit<LedgerEntry, 'ts' | 'session'> } | undefined> {
1014 const t = preset ? { ...S.cfg.trim, ...preset, useJev: false } : S.cfg.trim;
1015 const call = S.calls.get(toolUseId);
1016 const tool = call?.tool ?? 'unknown';
1017 let applied = !shadow();
1018 const plan = planTrim(text, isError, t);
1019 // What the trim keeps without Jev is the least it can keep: if even that saves too little
1020 // (source code is full of "error" and "expected"), Jev is not asked and the output stays whole.
1021 const floor = savedAt ? undefined : renderTrim(plan, { settings: t, originalChars: text.length });
1022 const hopeless = floor !== undefined && !worthTrimming(text.length, floor.charsAfter);
1023 let approved = new Set<string>();
1024 let jevMs: number | undefined;
1025 let jevCost: number | undefined;
1026 if (applied && !hopeless && t.useJev && jevKey() && plan.candidates.length > 0) {
1027 const preview = renderTrim(plan, { settings: t, originalChars: text.length });
1028 const chunks: Chunk[] = plan.candidates.slice(0, 20);
1029 const asked = await askJev(
1030 $,
1031 {
1032 context:
1033 'A coding assistant ran a tool; its long output is being trimmed before it enters the conversation. kept_output is what stays; omitted_chunks would be dropped (the full output stays readable in a file).',
1034 task: taskFor(call),
1035 command: clip(call?.command ?? tool, 400),
1036 kept_output: clip(preview.text, 6000, 2000),
1037 omitted_chunks: Object.fromEntries(chunks.map((c) => [c.id, c.text])),
1038 },
1039 chunkQuestions(chunks),
1040 undefined,
1041 'trim',
1042 );
1043 if (asked) {
1044 jevMs = asked.ms;
1045 jevCost = asked.response.usage?.cost;
1046 approved = new Set(chunks.filter((c) => (noul(asked.response.answers, `need_${c.id}`) ?? 1) >= t.jevKeepAt).map((c) => c.id));
1047 } else {
1048 // Jev unavailable: keep every candidate rather than lose content.
1049 approved = new Set(chunks.map((c) => c.id));
1050 }
1051 }
1052 // A saved Claude Code output replaces a 2KB preview: there, growing is the point.
1053 const skipped = hopeless || (!savedAt && !worthTrimming(text.length, renderTrim(plan, { approved, settings: t, originalChars: text.length }).charsAfter));
1054 if (skipped) applied = false;
1055 let path: string | undefined = savedAt;
1056 if (applied && !savedAt) {
1057 path = `${outputsDir()}/${toolUseId}.txt`;
1058 try {
1059 await $.fs.write(path, text.length > 3_900_000 ? text.slice(0, 3_900_000) : text);
1060 } catch {
1061 path = undefined;
1062 }
1063 }
1064 const out = renderTrim(plan, { approved, fullPath: path, settings: t, originalChars: text.length });
1065 return {
1066 text: out.text,
1067 entry: {
1068 kind: 'trim',
1069 scope: call?.agentId ? 'subagent' : 'main',
1070 agentId: call?.agentId,
1071 applied,
1072 jevMs,
1073 jevCost,
1074 text: clip(call?.command ?? tool, 160),
1075 trim: {
1076 tool,
1077 command: call?.command ? clip(call.command, 200) : undefined,
1078 charsBefore: out.charsBefore,
1079 charsAfter: out.charsAfter,
1080 linesBefore: out.linesBefore,
1081 linesAfter: out.linesAfter,
1082 outcome: out.outcome.detail,
1083 jevChunks: out.jevChunks,
1084 path,
1085 ...(preset === LIST_TRIM ? { kind: 'list' } : {}),
1086 ...(skipped ? { skipped: `removes under ${Math.round(MIN_TRIM_GAIN * 100)}%` } : {}),
1087 },
1088 },
1089 };
1090}
1091
1092/**
1093 * Keeps the saved full outputs bounded (our own directory only): files older
1094 * than `keepDays` go, empty session folders go, and while the total is over
1095 * `maxStorageMb` the oldest session folders go first.
1096 */
1097async function pruneOutputs($: EngineInterface): Promise<void> {
1098 const root = `${S.data}/outputs`;
1099 if (!S.data || !(await $.fs.exists(root))) return;
1100 try {
1101 await $.process.run(['find', root, '-type', 'f', '-name', '*.txt', '-mtime', `+${S.cfg.trim.keepDays}`, '-delete'], {
1102 timeoutMs: 10_000,
1103 });
1104 await $.process.run(['find', root, '-mindepth', '1', '-type', 'd', '-empty', '-delete'], { timeoutMs: 10_000 });
1105 const sessions: { path: string; bytes: number; newest: number }[] = [];
1106 for (const dir of await $.fs.list(root)) {
1107 if (dir.kind !== 'dir' || dir.name === S.session) continue;
1108 const files = await $.fs.list(`${root}/${dir.name}`);
1109 sessions.push({
1110 path: `${root}/${dir.name}`,
1111 bytes: files.reduce((sum, f) => sum + f.size, 0),
1112 newest: files.reduce((max, f) => Math.max(max, f.mtimeMs), 0),
1113 });
1114 }
1115 let total = sessions.reduce((sum, s) => sum + s.bytes, 0);
1116 const cap = S.cfg.trim.maxStorageMb * 1024 * 1024;
1117 for (const old of sessions.sort((a, b) => a.newest - b.newest)) {
1118 if (total <= cap) break;
1119 if (!old.path.startsWith(`${root}/`)) continue;
1120 await $.process.run(['rm', '-rf', old.path], { timeoutMs: 10_000 });
1121 total -= old.bytes;
1122 }
1123 } catch {
1124 // best effort
1125 }
1126}
1127
1128/** Our saved outputs (pruned calls included) and capsules may be read without a prompt; nothing else is decided here. */
1129async function isSavedOutput($: EngineInterface, path: unknown): Promise<boolean> {
1130 if (typeof path !== 'string' || !S.data) return false;
1131 try {
1132 const stat = await $.fs.stat(path, { resolve: true });
1133 for (const dir of ['outputs', 'handoffs']) {
1134 if (!(await $.fs.exists(`${S.data}/${dir}`))) continue;
1135 const root = await $.fs.stat(`${S.data}/${dir}`, { resolve: true });
1136 if (stat.realPath && root.realPath && stat.realPath.startsWith(`${root.realPath}/`)) return true;
1137 }
1138 return false;
1139 } catch {
1140 return false;
1141 }
1142}
1143
1144// ------------------------------------------------------- idle compaction --
1145
1146/**
1147 * Starts a compaction of the main conversation for our own reason. A terminal
1148 * session takes `$.session.compact()`; a headless / SDK-hosted one (the
1149 * desktop app) only compacts through a `/compact` command, which our
1150 * `session.compact` hook then turns into a Jev pruning like any other.
1151 */
1152async function requestCompaction($: EngineInterface, reason: 'threshold' | 'return'): Promise<void> {
1153 S.compacting = true;
1154 S.compactReason = reason;
1155 S.compactVia = 'api';
1156 try {
1157 try {
1158 await $.session.compact();
1159 } catch (error) {
1160 if (!/headless|not available|inside a turn/i.test(errorText(error))) throw error;
1161 S.compactVia = 'command';
1162 await $.command.run({ command: 'compact', args: '' });
1163 }
1164 } finally {
1165 S.compacting = false;
1166 S.compactReason = undefined;
1167 S.compactVia = undefined;
1168 }
1169}
1170
1171/**
1172 * The prompt cache expired while the session sat idle: the next request will
1173 * re-write the whole history anyway, so shrink it now (Jev pruning, no
1174 * summary). Claude Code refuses a compaction from inside a prompt's own
1175 * dispatch, so this runs from a timer (or at session start on resume), before
1176 * the person is back.
1177 */
1178async function compactIdle($: EngineInterface, why: string): Promise<void> {
1179 const k = S.cfg.compaction;
1180 if (!active() || !k.enabled || !idleCompaction() || !jevKey() || shadow() || S.compacting) return;
1181 if (S.route.lastRequestAt === undefined) return;
1182 const idleFor = Date.now() - S.route.lastRequestAt;
1183 if (idleFor <= S.cfg.router.cacheTtlMinutes * 60_000) return;
1184 try {
1185 const { context } = await $.session.usage();
1186 if ((context.tokens ?? 0) < k.onReturnMinTokens) return;
1187 $.ui.log(
1188 `jev-governor: idle ${Math.round(idleFor / 60_000)} min (${why}), the prompt cache expired: compacting ${Math.round((context.tokens ?? 0) / 1000)}k tokens before you are back`,
1189 );
1190 await requestCompaction($, 'return');
1191 } catch (error) {
1192 await ledger($, { kind: 'error', error: `idle compaction: ${errorText(error)}` });
1193 }
1194}
1195
1196/** (Re)arms the idle timer for just after the cache TTL from now. */
1197function scheduleIdle($: EngineInterface): void {
1198 S.idleTimer?.cancel();
1199 S.idleTimer = undefined;
1200 const k = S.cfg.compaction;hooks/lib/compaction/types.ts 220 lines1// Vendored from tamaratran/fast-jev-compaction (MIT, see ./LICENSE) at commit e3f262a.
2// Changes by jev-governor (transport is supplied through JevAsker):
3// - result previews: Jev sees the start and end of each tool output
4// (`resultPreviewChars`, 0 = upstream behaviour), not only its length;
5// - `limitPruning`: puts calls back until at most `maxPruneRatio` is removed.
6
7export type Role = 'user' | 'assistant';
8
9/**
10 * A tool_use block of an assistant message. `text` and `isError` mirror the
11 * outcome once the transcript holds it (Claude Code attaches them).
12 */
13export interface ToolUse {
14 tool_use_id: string;
15 tool: string;
16 input: Record<string, unknown>;
17 text?: string;
18 isError?: boolean;
19}
20
21/** A tool_result block of a user message. */
22export interface ToolResult {
23 tool_use_id: string;
24 text: string;
25 isError?: boolean;
26}
27
28/**
29 * One transcript message. The shape is a subset of Claude Code's
30 * `SessionMessage`, so a session transcript can be passed in as is.
31 */
32export interface Message {
33 role: Role;
34 text: string;
35 toolUses: ToolUse[];
36 toolResults?: ToolResult[];
37}
38
39/** A tool call paired with its result by `tool_use_id`. */
40export interface ToolCall {
41 /** Short id used in the Jev state and question names (`t1`, `t2`, ...). */
42 id: string;
43 tool_use_id: string;
44 tool: string;
45 input: Record<string, unknown>;
46 /** Index of the message holding the tool_use block. */
47 callIndex: number;
48 /** Index of the message holding the tool_result block. */
49 resultIndex: number;
50 resultChars: number;
51 /** The result's text, for the preview Jev sees. */
52 result?: string;
53 isError: boolean;
54 /** In the first or the newest preserved messages; never a candidate. */
55 pinned: boolean;
56}
57
58export interface CallAnswer {
59 /** Jev's probability that the call itself still matters. */
60 keepCall: number;
61 /** Jev's probability that the full result still needs to stay verbatim. */
62 keepResult: number;
63}
64
65export type CallAction = 'keep' | 'drop_result' | 'drop_call';
66
67export interface CallDecision extends CallAnswer {
68 id: string;
69 tool: string;
70 action: CallAction;
71 /** `restored`: Jev dropped it, `limitPruning` put it back. */
72 reason: 'pinned' | 'kept' | 'result_dropped' | 'call_dropped' | 'restored';
73}
74
75export interface HistoryToolCall {
76 id: string;
77 tool: string;
78 input: string;
79 result: string;
80}
81
82export interface HistoryEntry {
83 i: number;
84 role: Role;
85 text: string;
86 /** Structured per call, or one compact line per call once the state has to shrink. */
87 tool_calls?: HistoryToolCall[] | string[];
88}
89
90/** The state sent with every Jev request: the whole history, results omitted. */
91export interface CompactionState {
92 context: string;
93 goal: string;
94 history: HistoryEntry[];
95}
96
97export interface FittedState {
98 state: CompactionState;
99 tokens: number;
100 /** Which fitting stage produced the state, for diagnostics. */
101 stage: string;
102}
103
104export interface CompactOptions {
105 /** Ongoing task description; defaults to the last few user prompts. */
106 goal?: string;
107 /** Minimum keep probability for a call or result to stay. Default 0.5. */
108 keepThreshold?: number;
109 /** Newest messages never touched (the first message is always kept). Default 6. */
110 preserveRecentMessages?: number;
111 /** Estimated token ceiling for the state. Default 25000. */
112 maxStateTokens?: number;
113 /** Estimated token ceiling for state plus one batch of questions. Default 30000. */
114 maxRequestTokens?: number;
115 /** Characters of a dropped tool result to retain. Default 300. */
116 truncateHeadChars?: number;
117 /** Characters of each tool output's start and end shown to Jev. Default 0 (length only). */
118 resultPreviewChars?: number;
119 /** Applied to an output's start and end (with a margin past each cut) before they are shown to Jev. */
120 previewFilter?: (text: string) => string;
121}
122
123export interface ResolvedCompactOptions {
124 goal: string;
125 keepThreshold: number;
126 preserveRecentMessages: number;
127 maxStateTokens: number;
128 maxRequestTokens: number;
129 truncateHeadChars: number;
130 resultPreviewChars: number;
131 previewFilter?: (text: string) => string;
132}
133
134export interface CompactResult {
135 /** The compacted transcript; untouched messages are the input objects. */
136 messages: Message[];
137 decisions: CallDecision[];
138 stats: {
139 messagesBefore: number;
140 messagesAfter: number;
141 charsBefore: number;
142 charsAfter: number;
143 calls: number;
144 kept: number;
145 resultsDropped: number;
146 callsDropped: number;
147 pinned: number;
148 /** Calls put back by `limitPruning`. */
149 restored?: number;
150 stateTokens: number;
151 /** Which fitting stage the state needed, '' when no request was made. */
152 stateStage: string;
153 requests: number;
154 ms: number;
155 };
156}
157
158/** The `state` of a Jev request: a string or any JSON-serialisable object. */
159export type JevState = string | object;
160
161export interface NoulQuestion {
162 type: 'noul';
163 instructions: string;
164 criteria?: {
165 true?: string;
166 false?: string;
167 };
168}
169
170export interface ChoiceQuestion {
171 type: 'choice';
172 instructions: string;
173 criteria: Record<string, string | null>;
174}
175
176export interface ScoreQuestion {
177 type: 'score';
178 instructions: string;
179 criteria: string[];
180}
181
182export type JevQuestion = NoulQuestion | ChoiceQuestion | ScoreQuestion;
183export type JevQuestions = Record<string, JevQuestion>;
184
185export interface NoulAnswer {
186 type?: 'noul';
187 noul: number;
188}
189
190export interface ChoiceAnswer {
191 type?: 'choice';
192 choice: string;
193 confidence: number;
194 probabilities: Record<string, number>;
195}
196
197export interface ScoreAnswer {
198 type?: 'score';
199 score: number;
200 confidence: number;
201 probabilities: Record<string, number>;
202}
203
204export type JevAnswer = NoulAnswer | ChoiceAnswer | ScoreAnswer;
205
206export interface JevResponse {
207 model?: string;
208 answers: Record<string, JevAnswer>;
209 usage?: {
210 input_tokens?: number;
211 output_tokens?: number;
212 };
213 [key: string]: unknown;
214}
215
216/** Anything that can answer Jev questions: `JevClient`, or a host-provided adapter. */
217export interface JevAsker {
218 ask(state: JevState, questions: JevQuestions): Promise<JevResponse>;
219}
220hooks/lib/version.ts 3 lines1/** The mod's version, written into every ledger entry (kept equal to package.json by a test). */
2export const MOD_VERSION = '0.3.4';
3hooks/lib/budget.ts 61 lines1// Budget pressure from the subscription's rate-limit windows, as
2// `$.session.usage().rateLimits` reports them. The idea is pace, not level: a
3// window 60% used with 10% of its time left is fine, 60% used with 80% of its
4// time left means the account is burning faster than it can sustain.
5
6export type RateWindow = { kind: string; percentUsed: number; resetsAt?: string };
7
8/** 0 relaxed, 1 watch, 2 tight, 3 critical. */
9export type Pressure = 0 | 1 | 2 | 3;
10
11const WINDOW_MS: Record<string, number> = {
12 five_hour: 5 * 3600_000,
13 seven_day: 7 * 24 * 3600_000,
14};
15
16export type WindowPace = {
17 kind: string;
18 used: number;
19 /** Fraction of the window's time already gone, when known. */
20 elapsed?: number;
21 /** used − elapsed: positive when ahead of a linear pace. */
22 ahead?: number;
23 pressure: Pressure;
24};
25
26export function windowPace(window: RateWindow, now: number): WindowPace {
27 const used = Math.max(0, window.percentUsed) / 100;
28 const duration = WINDOW_MS[window.kind];
29 let elapsed: number | undefined;
30 if (duration && window.resetsAt) {
31 const resetsAt = Date.parse(window.resetsAt);
32 if (Number.isFinite(resetsAt)) {
33 elapsed = Math.min(1, Math.max(0, 1 - (resetsAt - now) / duration));
34 }
35 }
36 const ahead = elapsed === undefined ? undefined : used - elapsed;
37 let pressure: Pressure = 0;
38 if (used >= 0.95) pressure = 3;
39 else if (used >= 0.85 || (ahead !== undefined && ahead > 0.25)) pressure = 2;
40 else if (used >= 0.7 || (ahead !== undefined && ahead > 0.1)) pressure = 1;
41 // Early in a window a small absolute use is not alarming even when "ahead".
42 if (pressure > 0 && used < 0.3) pressure = 0;
43 return { kind: window.kind, used, elapsed, ahead, pressure };
44}
45
46export function budgetPressure(
47 windows: readonly RateWindow[],
48 now: number,
49): { pressure: Pressure; windows: WindowPace[] } {
50 const paces = windows.map((window) => windowPace(window, now));
51 const pressure = paces.reduce<Pressure>((max, pace) => (pace.pressure > max ? pace.pressure : max), 0);
52 return { pressure, windows: paces };
53}
54
55export function describePressure(windows: readonly WindowPace[]): string {
56 return windows
57 .filter((w) => w.kind in WINDOW_MS)
58 .map((w) => `${w.kind === 'five_hour' ? '5h' : '7d'} ${Math.round(w.used * 100)}%`)
59 .join(' · ');
60}
61hooks/lib/archive.ts 199 lines1// Pruned is not lost: every tool call the Jev compaction drops or truncates is
2// saved to a file, and the compacted history points at it, so the model can
3// read it back instead of re-running the tool or guessing. Pure: the caller
4// writes the files and hands back their paths.
5
6import type { CallDecision, Message, ToolCall } from './compaction/types.ts';
7
8export type PrunedCall = {
9 toolUseId: string;
10 tool: string;
11 input: Record<string, unknown>;
12 result: string;
13 isError: boolean;
14 action: 'drop_result' | 'drop_call';
15};
16
17/** The calls the decisions drop or truncate, with their full input and output. */
18export function prunedCalls(
19 messages: readonly Message[],
20 decisions: readonly CallDecision[],
21 calls: readonly ToolCall[],
22): PrunedCall[] {
23 const results = new Map<string, { text: string; isError: boolean }>();
24 for (const message of messages) {
25 for (const result of message.toolResults ?? []) {
26 results.set(result.tool_use_id, { text: result.text, isError: result.isError ?? false });
27 }
28 }
29 const byId = new Map(calls.map((call) => [call.id, call]));
30 const pruned: PrunedCall[] = [];
31 for (const decision of decisions) {
32 if (decision.action === 'keep') continue;
33 const call = byId.get(decision.id);
34 if (!call) continue;
35 const result = results.get(call.tool_use_id);
36 pruned.push({
37 toolUseId: call.tool_use_id,
38 tool: call.tool,
39 input: call.input,
40 result: result?.text ?? '',
41 isError: result?.isError ?? call.isError,
42 action: decision.action,
43 });
44 }
45 return pruned;
46}
47
48function inputJson(input: Record<string, unknown>): string {
49 try {
50 return JSON.stringify(input, null, 2);
51 } catch {
52 return '[unserializable input]';
53 }
54}
55
56/** The main argument of a call, one line: the command, path, pattern or URL. */
57export function callLabel(tool: string, input: Record<string, unknown>, max = 120): string {
58 const main = input.command ?? input.file_path ?? input.path ?? input.pattern ?? input.url ?? input.description ?? input.prompt;
59 const arg = typeof main === 'string' ? main.replace(/\s+/g, ' ').trim() : '';
60 const clipped = arg.length > max ? `${arg.slice(0, max)}…` : arg;
61 return clipped ? `${tool}(${clipped})` : tool;
62}
63
64/**
65 * What makes a later call a re-run of a pruned one: the same command, file,
66 * search or URL. Undefined for tools whose re-run is not a lookup (edits,
67 * subagents).
68 */
69export function rerunKey(tool: string, input: Record<string, unknown>): string | undefined {
70 const s = (value: unknown): string => (typeof value === 'string' ? value.replace(/\s+/g, ' ').trim() : '');
71 switch (tool) {
72 case 'Bash': {
73 // `cat <file>` reads a file as Read does: the same key.
74 const cat = /^(?:cat|head|tail|less|bat)\s+(?:-\S+\s+)*("[^"]+"|'[^']+'|\S+)$/.exec(s(input.command));
75 if (cat) return `Read:${cat[1]!.replace(/^["']|["']$/g, '')}`;
76 return s(input.command) ? `Bash:${s(input.command)}` : undefined;
77 }
78 case 'Read':
79 return s(input.file_path) ? `Read:${s(input.file_path)}` : undefined;
80 case 'Grep':
81 case 'Glob':
82 return s(input.pattern) ? `${tool}:${s(input.pattern)}|${s(input.path)}` : undefined;
83 case 'WebFetch':
84 return s(input.url) ? `WebFetch:${s(input.url)}` : undefined;
85 case 'WebSearch':
86 return s(input.query) ? `WebSearch:${s(input.query)}` : undefined;
87 default:
88 return undefined;
89 }
90}
91
92/** The file an edit tool changed, if any. */
93export function editedPath(tool: string, input: Record<string, unknown>): string | undefined {
94 if (!['Edit', 'MultiEdit', 'Write', 'NotebookEdit'].includes(tool)) return undefined;
95 const path = input.file_path ?? input.notebook_path;
96 return typeof path === 'string' && path.trim() ? path.trim() : undefined;
97}
98
99/** What a pruned call's archive file holds: the call, then its output verbatim. */
100export function archiveText(call: PrunedCall): string {
101 return [
102 `# ${callLabel(call.tool, call.input, 300)}`,
103 `tool_use_id: ${call.toolUseId}${call.isError ? ' (error)' : ''}`,
104 '',
105 '## input',
106 inputJson(call.input),
107 '',
108 '## output',
109 call.result,
110 '',
111 ].join('\n');
112}
113
114const POINTER_MARK = '[jev-governor pruned ';
115
116/**
117 * A result an earlier compaction already truncated: its full output is in the
118 * archive, and the text is only its head and the pointer. Archiving it again
119 * would overwrite the full output with the stub.
120 */
121export function isPointerText(text: string): boolean {
122 return text.includes(POINTER_MARK) && text.includes('; full output: ');
123}
124
125/**
126 * The text a truncated result keeps: its head and where the rest is. Short
127 * results (within the head plus a little) stay whole, as the library leaves
128 * them, and so does a pointer an earlier compaction left.
129 */
130export function pointerText(text: string, isError: boolean, headChars: number, path: string): string {
131 if (text.length <= headChars + 120 || isPointerText(text)) return text;
132 const head = headChars > 0 ? `${text.slice(0, headChars)}\n` : '';
133 return `${head}${POINTER_MARK}${text.length - headChars} chars of this tool output${isError ? ' (error)' : ''}; full output: ${path} — read it there instead of re-running the tool]`;
134}
135
136/** One line of the pruned-calls index. */
137/**
138 * One line of the session's index: the call, what happened to it and its file. The file is
139 * named without its folder (the index's own): the model reads the index's tail, and the full
140 * path repeated on every line was ~40% of it (an index of 826 lines reached 184k characters).
141 */
142export function indexLine(call: PrunedCall, path: string): string {
143 const what = call.action === 'drop_call' ? 'call removed' : 'output truncated';
144 const size = isPointerText(call.result) ? 'output archived earlier' : `${call.result.length} chars`;
145 const file = path.slice(path.lastIndexOf('/') + 1);
146 return `- ${callLabel(call.tool, call.input)}${call.isError ? ' → error' : ''} · ${what} · ${size} → ${file}`;
147}
148
149/** The note the compacted history carries about calls removed whole. */
150export function indexNote(count: number, indexPath: string): string {
151 return `[jev-governor: ${count} earlier tool call(s) were removed from this conversation to save context; they are listed in ${indexPath} (newest last; search it with grep rather than reading it whole), each with its file in the same folder — read the one you need]`;
152}
153
154/**
155 * Puts the archive into the compacted history: a truncated result's note now
156 * names its file, and the first rebuilt message mentions the index when calls
157 * were removed whole. Messages the library returned untouched stay the same
158 * objects (the engine keeps them whole by their handle).
159 */
160export function withPointers(
161 original: readonly Message[],
162 compacted: readonly Message[],
163 pruned: readonly PrunedCall[],
164 paths: ReadonlyMap<string, string>,
165 headChars: number,
166 index?: { path: string; removed: number },
167): Message[] {
168 const own = new Set<Message>(original);
169 const truncated = new Map<string, PrunedCall>();
170 for (const call of pruned) if (call.action === 'drop_result' && paths.has(call.toolUseId)) truncated.set(call.toolUseId, call);
171 const point = (id: string, text: string, isError: boolean): string => {
172 const call = truncated.get(id);
173 return call ? pointerText(call.result, isError, headChars, paths.get(id)!) : text;
174 };
175 const alreadyNoted = index !== undefined && original.some((m) => m.text.includes(index.path));
176 let noted = alreadyNoted || index === undefined || index.removed === 0;
177 return compacted.map((message) => {
178 if (own.has(message)) return message;
179 const toolUses = message.toolUses.map((tool) =>
180 truncated.has(tool.tool_use_id) && tool.text !== undefined
181 ? { ...tool, text: point(tool.tool_use_id, tool.text, tool.isError ?? false) }
182 : tool,
183 );
184 const rebuilt: Message = { ...message, toolUses };
185 if (message.toolResults) {
186 rebuilt.toolResults = message.toolResults.map((result) =>
187 truncated.has(result.tool_use_id)
188 ? { ...result, text: point(result.tool_use_id, result.text, result.isError ?? false) }
189 : result,
190 );
191 }
192 if (!noted && index) {
193 noted = true;
194 rebuilt.text = rebuilt.text.trim() ? `${rebuilt.text}\n\n${indexNote(index.removed, index.path)}` : indexNote(index.removed, index.path);
195 }
196 return rebuilt;
197 });
198}
199hooks/lib/capsule.ts 420 lines1// Moving work to a new chat: a compact "capsule" of the old one instead of
2// its whole history (the idea: history is not context). A coding
3// session's text is ~1.5% of its history and the rest is tool output, so the
4// capsule is the person's requests and the assistant's final answers turn by
5// turn (newest whole, older condensed, oldest one line each), the files it
6// changed, the last test/build results and, when it was cheap to get, a brief
7// the original chat's model wrote. Pure: the mod reads the session, asks Jev,
8// writes the files.
9
10import { callLabel } from './archive.ts';
11import { estimateTokens } from './compaction/state.ts';
12import type { JevQuestions } from './jev.ts';
13import { detectOutcome } from './trim.ts';
14
15export type CapsuleMessage = {
16 role: 'user' | 'assistant';
17 text: string;
18 toolUses: readonly { tool_use_id?: string; tool: string; input: Record<string, unknown>; text?: string; isError?: boolean }[];
19 toolResults?: readonly { tool_use_id: string; text: string; isError?: boolean }[];
20};
21
22export type Check = { command: string; outcome: string; failed: boolean; turn: number };
23
24export type Turn = {
25 /** 1-based, oldest first. */
26 n: number;
27 user: string;
28 answer: string;
29 calls: number;
30 errors: number;
31 /** Files the turn created or edited. */
32 edited: string[];
33};
34
35export type Skeleton = {
36 turns: Turn[];
37 /** Claude Code's own compaction summary, when the history starts with one. */
38 summary?: string;
39 /** Changed files with their edit counts, most recent last. */
40 files: { path: string; edits: number; lastTurn: number }[];
41 /** The latest result of each test / build / check command, oldest first. */
42 checks: Check[];
43};
44
45export type TurnPlan = 'full' | 'gist' | 'line';
46
47/** Marks the prompt that asks the original chat for its brief (left out of the capsule). */
48export const BRIEF_MARKER = '[jev-governor:getctx-brief]';
49
50/** A capsule reference in a prompt: `jev-ctx:20261005-1912-a3f0`. */
51export const CTX_REF = /\bjev-ctx:(\d{8}-\d{4}-[0-9a-f]{4})\b/;
52
53const COMPACT_SUMMARY = /^This session is being continued from a previous conversation/;
54/** A skill's body the engine inserts as a user message: not something the person typed. */
55const SKILL_BODY = /^Base directory for this skill:/;
56
57const EDIT_TOOLS = new Set(['Edit', 'Write', 'MultiEdit', 'NotebookEdit']);
58
59const CHECK_COMMAND =
60 /\b(test|tests|vitest|jest|pytest|nextest|playwright|cypress|tsc|typecheck|vue-tsc|lint|eslint|clippy|build|check|compile|make)\b/;
61/** Scripts and file edits that merely mention a test word. */
62const NOT_A_CHECK = /^\s*(cat|echo|printf|python3?\s+-|node\s+-e|sed|awk|grep|rg|git|ls|find|head|tail)\b/;
63
64/** Drops what the engine wraps around a prompt (reminders, command echoes, notifications). */
65export function cleanText(text: string): string {
66 return text
67 .replace(/<(system-reminder|command-message|command-name|command-args|local-command-stdout|local-command-stderr|local-command-caveat|task-notification|user-prompt-submit-hook)[^>]*>[\s\S]*?<\/\1>/g, ' ')
68 .replace(/<\/?pasted_content[^>]*>/g, '')
69 .replace(/^Caveat: The messages below were generated by the user while running local commands.*$/gm, '')
70 .replace(/\n{3,}/g, '\n\n')
71 .trim();
72}
73
74/** Head and tail of a long text, cut at whole lines where it can. */
75export function excerpt(text: string, max: number): string {
76 const flat = text.trim();
77 if (flat.length <= max) return flat;
78 const head = Math.round(max * 0.7);
79 const tail = max - head;
80 return `${flat.slice(0, head).trimEnd()}\n[…]\n${flat.slice(-tail).trimStart()}`;
81}
82
83function oneLine(text: string, max: number): string {
84 const flat = text.replace(/\s+/g, ' ').trim();
85 return flat.length > max ? `${flat.slice(0, max)}…` : flat;
86}
87
88function editedPath(input: Record<string, unknown>): string | undefined {
89 const path = input.file_path ?? input.notebook_path ?? input.path;
90 return typeof path === 'string' ? path : undefined;
91}
92
93/**
94 * Splits the conversation into turns (a turn opens with a prompt the person
95 * typed) and collects what the capsule needs from each. Our own brief prompt
96 * and its answer are left out: the brief travels on its own.
97 */
98export function buildSkeleton(messages: readonly CapsuleMessage[], cwd = ''): Skeleton {
99 const relative = (path: string): string => (cwd && path.startsWith(`${cwd}/`) ? path.slice(cwd.length + 1) : path);
100 const results = new Map<string, { text: string; isError: boolean }>();
101 for (const message of messages) {
102 for (const result of message.toolResults ?? []) {
103 results.set(result.tool_use_id, { text: result.text, isError: result.isError ?? false });
104 }
105 }
106 const turns: Turn[] = [];
107 const files = new Map<string, { path: string; edits: number; lastTurn: number }>();
108 const checks = new Map<string, Check>();
109 let summary: string | undefined;
110 let current: Turn | undefined;
111 let skipping = false;
112 for (const message of messages) {
113 if (message.role === 'user') {
114 const text = cleanText(message.text);
115 if (!text) continue;
116 if (SKILL_BODY.test(text)) continue;
117 if (COMPACT_SUMMARY.test(text)) {
118 summary = text;
119 current = undefined;
120 continue;
121 }
122 if ((message.toolResults?.length ?? 0) > 0 && current) continue;
123 skipping = text.includes(BRIEF_MARKER);
124 if (skipping) {
125 current = undefined;
126 continue;
127 }
128 current = { n: turns.length + 1, user: text, answer: '', calls: 0, errors: 0, edited: [] };
129 turns.push(current);
130 continue;
131 }
132 if (skipping || !current) continue;
133 const text = cleanText(message.text);
134 if (text) current.answer = text;
135 for (const tool of message.toolUses) {
136 current.calls++;
137 const result = tool.tool_use_id ? results.get(tool.tool_use_id) : undefined;
138 const isError = tool.isError ?? result?.isError ?? false;
139 if (isError) current.errors++;
140 if (EDIT_TOOLS.has(tool.tool) && !isError) {
141 const found = editedPath(tool.input);
142 const path = found === undefined ? undefined : relative(found);
143 if (path) {
144 if (!current.edited.includes(path)) current.edited.push(path);
145 const file = files.get(path) ?? { path, edits: 0, lastTurn: 0 };
146 file.edits++;
147 file.lastTurn = current.n;
148 files.delete(path);
149 files.set(path, file);
150 }
151 }
152 const command = tool.tool === 'Bash' && typeof tool.input.command === 'string' ? tool.input.command : undefined;
153 // The command proper: its first line without a leading `cd …`, up to the first pipe or separator.
154 const first = (command?.split('\n')[0] ?? '').trim().replace(/^cd\s+\S+\s*&&\s*/, '');
155 const head = first.split(/\s*(?:\||;|&&|\|\|)\s*/)[0] ?? '';
156 if (head && CHECK_COMMAND.test(head) && !NOT_A_CHECK.test(head)) {
157 const output = result?.text ?? tool.text ?? '';
158 const outcome = detectOutcome(output.split('\n'), isError);
159 // Without a test summary or an error there is nothing to report.
160 if (outcome.status !== 'unknown' && !/\(0 passed, 0 failed/.test(outcome.detail)) {
161 const key = oneLine(first, 160);
162 checks.delete(key);
163 checks.set(key, { command: key, outcome: outcome.detail, failed: outcome.status !== 'passed', turn: current.n });
164 }
165 }
166 }
167 }
168 return { turns, summary, files: [...files.values()], checks: [...checks.values()] };
169}
170
171const SIZE = {
172 full: { user: 2_500, answer: 4_000 },
173 gist: { user: 400, answer: 600 },
174 line: { user: 140, answer: 160 },
175};
176
177/** Quoted text must not open sections of its own: its headings become bold lines. */
178function quoted(text: string): string {
179 return text.replace(/^#{1,6}\s+(.+)$/gm, '**$1**');
180}
181
182export function renderTurn(turn: Turn, plan: TurnPlan): string {
183 if (plan === 'line') {
184 const answer = turn.answer ? ` → ${oneLine(turn.answer, SIZE.line.answer)}` : '';
185 return `- Turn ${turn.n}: ${oneLine(turn.user, SIZE.line.user)}${answer}`;
186 }
187 const size = SIZE[plan];
188 const tools = [
189 turn.calls > 0 ? `${turn.calls} tool call(s)${turn.errors > 0 ? `, ${turn.errors} failed` : ''}` : '',
190 turn.edited.length > 0 ? `changed ${turn.edited.slice(0, 8).map((p) => p.split('/').pop()).join(', ')}${turn.edited.length > 8 ? '…' : ''}` : '',
191 ].filter(Boolean);
192 return [
193 `### Turn ${turn.n}${plan === 'gist' ? ' (condensed)' : ''}`,
194 `**User:** ${quoted(excerpt(turn.user, size.user))}`,
195 turn.answer ? `**Assistant:** ${quoted(excerpt(turn.answer, size.answer))}` : '**Assistant:** (no final text)',
196 tools.length > 0 ? `_${tools.join('; ')}_` : '',
197 ]
198 .filter(Boolean)
199 .join('\n\n');
200}
201
202/**
203 * How much of each turn goes in, within `budget` tokens. The newest two go in
204 * full and the first (usually the task itself) at least condensed; the rest
205 * by Jev's `need` (≥ 0.7 full, < 0.3 one line) or, without it, the next four
206 * condensed and older ones one line. Over budget, the least needed turns
207 * (without Jev: the oldest) are reduced first, then left out (the transcript
208 * keeps them); the first and newest two go last.
209 */
210export function planTurns(turns: readonly Turn[], budget: number, need?: ReadonlyMap<number, number>): Map<number, TurnPlan | 'omit'> {
211 const plans = new Map<number, TurnPlan | 'omit'>();
212 const newest = turns.length;
213 const recent = (turn: Turn): boolean => newest - turn.n < 2;
214 for (const turn of turns) {
215 const age = newest - turn.n;
216 const p = need?.get(turn.n);
217 let plan: TurnPlan;
218 if (recent(turn)) plan = 'full';
219 else if (p !== undefined) plan = p >= 0.7 ? 'full' : p >= 0.3 ? 'gist' : 'line';
220 else plan = age < 6 ? 'gist' : 'line';
221 if (turn.n === 1 && plan === 'line') plan = 'gist';
222 plans.set(turn.n, plan);
223 }
224 const cost = (turn: Turn): number => {
225 const plan = plans.get(turn.n)!;
226 return plan === 'omit' ? 0 : estimateTokens(renderTurn(turn, plan));
227 };
228 let total = turns.reduce((sum, turn) => sum + cost(turn), 0);
229 const lower = (turn: Turn, to: TurnPlan | 'omit'): void => {
230 const before = cost(turn);
231 plans.set(turn.n, to);
232 total += cost(turn) - before;
233 };
234 // Least needed first; without Jev's answer a turn counts by its age.
235 const value = (turn: Turn): number => need?.get(turn.n) ?? turn.n / (newest + 1);
236 const middle = turns.filter((t) => !recent(t) && t.n !== 1).sort((a, b) => value(a) - value(b) || a.n - b.n);
237 const order: readonly (TurnPlan | 'omit')[] = ['full', 'gist', 'line', 'omit'];
238 for (let step = 0; step < 3; step++) {
239 for (const turn of middle) {
240 if (total <= budget) return plans;
241 if (plans.get(turn.n) === order[step]) lower(turn, order[step + 1]!);
242 }
243 }
244 // Still too long: the first turn, then the newest two, condensed.
245 for (const turn of [...turns.filter((t) => t.n === 1 && !recent(t)), ...turns.filter(recent)]) {
246 if (total > budget && plans.get(turn.n) === 'full') lower(turn, 'gist');
247 }
248 for (const turn of turns.filter((t) => t.n === 1 && !recent(t))) {
249 if (total > budget && plans.get(turn.n) === 'gist') lower(turn, 'line');
250 }
251 return plans;
252}
253
254export const HANDOFF_STATE_CONTEXT =
255 'A Claude Code coding chat is being continued in a new chat that will not see it. The new chat gets the newest turns in full; for each older turn decide whether it must come along. `turns` lists them (the person\'s request, the assistant\'s final answer), oldest first.';
256
257/** Older turns Jev is asked about (the newest two always go in full). */
258export function candidateTurns(turns: readonly Turn[], max = 60): Turn[] {
259 return turns.slice(0, Math.max(0, turns.length - 2)).slice(-max);
260}
261
262export function turnState(turns: readonly Turn[], candidates: readonly Turn[], focus: string): object {
263 return {
264 context: HANDOFF_STATE_CONTEXT,
265 focus: focus || 'continue the most recent work',
266 newest_turns: turns.slice(-2).map((t) => ({ turn: t.n, user: oneLine(t.user, 600), answer: oneLine(t.answer, 600) })),
267 turns: candidates.map((t) => ({ turn: t.n, user: oneLine(t.user, 300), answer: oneLine(t.answer, 300) })),
268 };
269}
270
271export function turnQuestions(candidates: readonly Turn[]): JevQuestions {
272 return Object.fromEntries(
273 candidates.map((t) => [
274 `need_t${t.n}`,
275 {
276 type: 'noul' as const,
277 instructions: `Turn ${t.n} holds something the assistant needs to continue the work toward focus in the new chat (a decision, constraint, result, open problem or the person's preference) that the newest turns do not already carry`,
278 },
279 ]),
280 );
281}
282
283export type CapsuleMeta = {
284 id: string;
285 project: string;
286 cwd: string;
287 session: string;
288 createdAt: string;
289 focus: string;
290 /** Claude Code's transcript of the old chat, when it was found. */
291 transcript?: string;
292 /** Archive of pruned tool outputs, when the old chat has one. */
293 archive?: string;
294};
295
296export function renderCapsule(input: {
297 meta: CapsuleMeta;
298 skeleton: Skeleton;
299 plans: ReadonlyMap<number, TurnPlan | 'omit'>;
300 brief?: string;
301}): { text: string; tokens: number } {
302 const { meta, skeleton, plans, brief } = input;
303 const where = [
304 meta.transcript ? `full transcript of the previous chat: ${meta.transcript}` : '',
305 meta.archive ? `pruned tool outputs: ${meta.archive}` : '',
306 ].filter(Boolean);
307 const omitted = skeleton.turns.filter((t) => plans.get(t.n) === 'omit').length;
308 const parts: string[] = [
309 `# Context from a previous chat — ${meta.project}`,
310 `jev-ctx:${meta.id} · ${meta.createdAt.slice(0, 16).replace('T', ' ')} · session ${meta.session} · ${skeleton.turns.length} turn(s)`,
311 [
312 'This is a handoff from an earlier Claude Code chat in this project; that chat is not visible here. Continue the work from it. Files on disk are current, so re-read a file before editing it rather than trusting a quote below.',
313 where.length > 0 ? `Details left out here: ${where.join('; ')} — read only what you need.` : '',
314 meta.focus ? `Focus: ${meta.focus}` : '',
315 ]
316 .filter(Boolean)
317 .join('\n'),
318 ];
319 if (brief?.trim()) parts.push(`## Brief from the previous chat\n\n${brief.trim()}`);
320 if (skeleton.summary) parts.push(`## Earlier summary (Claude Code compaction)\n\n${excerpt(skeleton.summary, 6_000)}`);
321 const lines: string[] = [];
322 if (omitted > 0) lines.push(`_${omitted} turn(s) judged least needed are left out (see the transcript)._`);
323 let block: string[] = [];
324 const flush = (): void => {
325 if (block.length > 0) lines.push(block.join('\n'));
326 block = [];
327 };
328 for (const turn of skeleton.turns) {
329 const plan = plans.get(turn.n) ?? 'line';
330 if (plan === 'omit') continue;
331 if (plan === 'line') {
332 block.push(renderTurn(turn, plan));
333 continue;
334 }
335 flush();
336 lines.push(renderTurn(turn, plan));
337 }
338 flush();
339 if (lines.length > 0) parts.push(`## Conversation, oldest first\n\n${lines.join('\n\n')}`);
340 if (skeleton.files.length > 0) {
341 const files = skeleton.files.slice(-40).map((f) => `- ${f.path} (${f.edits} edit${f.edits === 1 ? '' : 's'}, last in turn ${f.lastTurn})`);
342 parts.push(`## Files changed\n\n${files.join('\n')}`);
343 }
344 if (skeleton.checks.length > 0) {
345 const checks = skeleton.checks.slice(-8).map((c) => `- \`${oneLine(c.command, 160)}\` (turn ${c.turn}) → ${c.outcome}`);
346 parts.push(`## Last checks\n\n${checks.join('\n')}`);
347 }
348 const text = `${parts.join('\n\n')}\n`;
349 return { text, tokens: Math.round(estimateTokens(text)) };
350}
351
352/** The sentence the new chat's model follows when the mod did not attach the capsule. */
353export const FALLBACK_HINT = 'если он не подключён к этому сообщению, сначала прочитай этот файл';
354const FALLBACK_HINT_EN = 'if it is not attached to this message, read this file first';
355
356/** What the person pastes into the new chat. */
357export function pastePrompt(meta: Pick<CapsuleMeta, 'id' | 'focus'>, path: string, title: string, lang: 'ru' | 'en' = 'ru'): string {
358 if (lang === 'en') {
359 return [
360 `jev-ctx:${meta.id} — continuing work from the previous chat «${title}».`,
361 `Previous chat context: ${path} (${FALLBACK_HINT_EN}).`,
362 meta.focus ? `Task: ${meta.focus}` : '',
363 ]
364 .filter(Boolean)
365 .join('\n');
366 }
367 return [
368 `jev-ctx:${meta.id} — продолжаем работу из прошлого чата «${title}».`,
369 `Контекст прошлого чата: ${path} (${FALLBACK_HINT}).`,
370 meta.focus ? `Задача: ${meta.focus}` : '',
371 ]
372 .filter(Boolean)
373 .join('\n');
374}
375
376/** The pasted prompt once the mod attached the capsule: no instruction to read the file. */
377export function attachedPrompt(text: string): string {
378 return text
379 .replace(/Контекст прошлого чата: \S+ \([^)]*\)\./, 'Контекст прошлого чата подключён к этому сообщению.')
380 .replace(/Previous chat context: \S+ \([^)]*\)\./, 'Previous chat context is attached to this message.');
381}
382
383export function briefPrompt(focus: string, words: number): string {
384 return [
385 `${BRIEF_MARKER} The person is moving this work to a fresh chat (/jevg getctx or /jevg fresh).`,
386 `Write a handoff brief for the assistant who will continue there. It will not see this chat: only your brief and an abridged list of the person's requests and your final answers.`,
387 `At most ${words} words, Markdown, in the language the person writes in, with these sections: Goal; State (done and verified vs. done but not verified); Decisions (each with its reason, including approaches tried and rejected); Next step and open questions; Key files and commands; Gotchas.`,
388 focus ? `Keep what matters for: ${focus}.` : '',
389 'Facts only. Do not call any tools and do not change anything; reply with the brief alone.',
390 ]
391 .filter(Boolean)
392 .join(' ');
393}
394
395/** `20261005-1912-a3f0`: sortable, short enough to type. */
396export function capsuleId(now: Date, random: number): string {
397 const p = (n: number): string => String(n).padStart(2, '0');
398 const hex = Math.floor(random * 0x10000)
399 .toString(16)
400 .padStart(4, '0');
401 return `${now.getFullYear()}${p(now.getMonth() + 1)}${p(now.getDate())}-${p(now.getHours())}${p(now.getMinutes())}-${hex}`;
402}
403
404/** Claude Code's folder name for a project: every non-alphanumeric character as `-`. */
405export function projectSlug(cwd: string): string {
406 return cwd.replace(/[^A-Za-z0-9]/g, '-');
407}
408
409/** `/jevg getctx --brief focus words` → flags and focus. */
410export function parseGetctxArgs(args: string): { brief?: boolean; focus: string } {
411 let brief: boolean | undefined;
412 const words: string[] = [];
413 for (const word of args.trim().split(/\s+/).filter(Boolean)) {
414 if (word === '--brief') brief = true;
415 else if (word === '--nobrief' || word === '--no-brief') brief = false;
416 else words.push(word);
417 }
418 return { brief, focus: words.join(' ') };
419}
420hooks/lib/compaction/adapter.ts 97 lines1// Session ↔ library mapping for Jev compaction, adapted from
2// tamaratran/fast-jev-compaction hooks/fast-jev.ts (MIT): messages the
3// library returns untouched are handed back as the engine's own objects
4// (handles included), rebuilt ones as fresh messages.
5
6import { compact, limitPruning, reductionRatio, resolveOptions } from './compact.ts';
7import { askFolds, collectFoldCandidates, type FoldDecision, type FoldOptions } from './fold.ts';
8import { collectToolCalls, fitState } from './state.ts';
9import type { CompactOptions, CompactResult, JevAsker, Message, ToolResult, ToolUse } from './types.ts';
10
11/** The engine's SessionMessage, structurally: what the mapping relies on. */
12export type EngineMessage = Message & { handle?: string };
13
14export function toEngineMessages<M extends EngineMessage>(
15 input: readonly M[],
16 output: readonly Message[],
17): EngineMessage[] {
18 const messages = new Map<Message, M>();
19 const uses = new Map<ToolUse, ToolUse>();
20 const results = new Map<ToolResult, ToolResult>();
21 for (const message of input) {
22 messages.set(message, message);
23 for (const tool of message.toolUses) uses.set(tool, tool);
24 for (const result of message.toolResults ?? []) results.set(result, result);
25 }
26 return output.map((message) => {
27 const own = messages.get(message);
28 if (own) return own;
29 const rebuilt: EngineMessage = {
30 role: message.role,
31 text: message.text,
32 toolUses: message.toolUses.map((tool) => {
33 const known = uses.get(tool);
34 if (known) return known;
35 const copy: ToolUse = { tool_use_id: tool.tool_use_id, tool: tool.tool, input: tool.input };
36 if (tool.text !== undefined) copy.text = tool.text;
37 if (tool.isError) copy.isError = true;
38 return copy;
39 }),
40 };
41 if (message.toolResults && message.toolResults.length > 0) {
42 rebuilt.toolResults = message.toolResults.map(
43 (result) =>
44 results.get(result) ?? {
45 tool_use_id: result.tool_use_id,
46 text: result.text,
47 isError: result.isError ?? false,
48 },
49 );
50 }
51 return rebuilt;
52 });
53}
54
55/** One line for the ledger and the toast; `fold`: old messages folded after the pruning, and the reduction of both. */
56export function summarizeCompaction(result: CompactResult, fold?: { folded: number; requests: number; ratio: number }): string {
57 const { stats } = result;
58 const parts = [
59 stats.kept > 0 ? `${stats.kept} kept` : '',
60 stats.resultsDropped > 0 ? `${stats.resultsDropped} results truncated` : '',
61 stats.callsDropped > 0 ? `${stats.callsDropped} calls dropped` : '',
62 stats.pinned > 0 ? `${stats.pinned} pinned` : '',
63 stats.restored ? `${stats.restored} put back (prune cap)` : '',
64 fold && fold.folded > 0 ? `${fold.folded} old message(s) folded` : '',
65 ].filter(Boolean);
66 const ratio = fold ? fold.ratio : reductionRatio(result);
67 return `${Math.round(ratio * 100)}% reduction; ${parts.join(', ') || 'no tool calls'}; ${stats.requests + (fold?.requests ?? 0)} Jev request(s)`;
68}
69
70/** Compacts with Jev, then puts calls back while more than `maxPruneRatio` would go (1 = no cap). */
71export async function runCompaction<M extends EngineMessage>(
72 messages: readonly M[],
73 asker: JevAsker,
74 options: CompactOptions & { maxPruneRatio?: number },
75): Promise<{ result: CompactResult; messages: EngineMessage[]; ratio: number }> {
76 const raw = await compact(messages, asker, options);
77 const result = limitPruning(messages, raw, { ...options, maxPruneRatio: options.maxPruneRatio ?? 1 });
78 return { result, messages: toEngineMessages(messages, result.messages), ratio: reductionRatio(result) };
79}
80
81/**
82 * Asks Jev which old dialog messages to fold (./fold.ts), with the state the
83 * call questions get. Throws when Jev fails; the caller compacts without folds.
84 */
85export async function planFolds(
86 messages: readonly Message[],
87 asker: JevAsker,
88 options: CompactOptions,
89 fold: FoldOptions,
90): Promise<{ decisions: FoldDecision[]; requests: number }> {
91 const candidates = collectFoldCandidates(messages, fold);
92 if (candidates.length === 0) return { decisions: [], requests: 0 };
93 const resolved = resolveOptions(options);
94 const state = fitState(messages, collectToolCalls(messages, resolved.preserveRecentMessages), resolved);
95 return askFolds(asker, state.state, state.tokens, candidates, messages, fold);
96}
97hooks/lib/compaction/compact.ts 409 lines1// Vendored from tamaratran/fast-jev-compaction (MIT, see ./LICENSE) at commit e3f262a.
2// Changes by jev-governor (transport is supplied through JevAsker):
3// - result previews: Jev sees the start and end of each tool output
4// (`resultPreviewChars`, 0 = upstream behaviour), not only its length;
5// - `limitPruning`: puts calls back until at most `maxPruneRatio` is removed.
6
7import { noulAnswer } from './request.ts';
8import { collectToolCalls, estimateTokens, fitState } from './state.ts';
9import type {
10 CallAnswer,
11 CallDecision,
12 CompactOptions,
13 CompactResult,
14 CompactionState,
15 JevAsker,
16 JevQuestions,
17 Message,
18 ResolvedCompactOptions,
19 ToolCall,
20 ToolUse,
21} from './types.ts';
22
23export const DEFAULT_OPTIONS: ResolvedCompactOptions = {
24 goal: '',
25 keepThreshold: 0.5,
26 preserveRecentMessages: 6,
27 maxStateTokens: 25_000,
28 maxRequestTokens: 30_000,
29 truncateHeadChars: 300,
30 resultPreviewChars: 0,
31};
32
33/** Tokens the request envelope (`model`, key names) adds around state and questions. */
34const REQUEST_OVERHEAD_TOKENS = 20;
35
36function finite(value: number | undefined, fallback: number): number {
37 return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
38}
39
40export function resolveOptions(options: CompactOptions = {}): ResolvedCompactOptions {
41 return {
42 goal: options.goal ?? DEFAULT_OPTIONS.goal,
43 keepThreshold: finite(options.keepThreshold, DEFAULT_OPTIONS.keepThreshold),
44 preserveRecentMessages: Math.max(
45 0,
46 Math.floor(
47 finite(options.preserveRecentMessages, DEFAULT_OPTIONS.preserveRecentMessages),
48 ),
49 ),
50 maxStateTokens: Math.max(1, finite(options.maxStateTokens, DEFAULT_OPTIONS.maxStateTokens)),
51 maxRequestTokens: Math.max(
52 1,
53 finite(options.maxRequestTokens, DEFAULT_OPTIONS.maxRequestTokens),
54 ),
55 truncateHeadChars: Math.max(
56 0,
57 Math.floor(finite(options.truncateHeadChars, DEFAULT_OPTIONS.truncateHeadChars)),
58 ),
59 resultPreviewChars: Math.max(
60 0,
61 Math.floor(finite(options.resultPreviewChars, DEFAULT_OPTIONS.resultPreviewChars)),
62 ),
63 ...(options.previewFilter ? { previewFilter: options.previewFilter } : {}),
64 };
65}
66
67/** The two `noul` questions asked about one call: keep the call, keep its result. */
68export function questionsFor(call: ToolCall): JevQuestions {
69 return {
70 [`call_${call.id}`]: {
71 type: 'noul',
72 instructions: `Tool call ${call.id} (${call.tool}) should stay in the history: knowing this call was made, with its input, still matters for what the assistant does next`,
73 },
74 [`result_${call.id}`]: {
75 type: 'noul',
76 instructions: `The full output of tool call ${call.id} (${call.tool}, ${call.resultChars} chars) should stay in the history verbatim: the assistant still needs its contents and re-running the tool would not do`,
77 },
78 };
79}
80
81/**
82 * Splits the candidate calls into batches whose questions, together with the
83 * (always complete) state, fit one request.
84 */
85export function batchCalls(
86 calls: readonly ToolCall[],
87 stateTokens: number,
88 options: Pick<ResolvedCompactOptions, 'maxRequestTokens'>,
89): ToolCall[][] {
90 const budget = options.maxRequestTokens - stateTokens - REQUEST_OVERHEAD_TOKENS;
91 const batches: ToolCall[][] = [];
92 let current: ToolCall[] = [];
93 let currentTokens = 0;
94 for (const call of calls) {
95 const tokens = estimateTokens(JSON.stringify(questionsFor(call)));
96 if (current.length > 0 && currentTokens + tokens > budget) {
97 batches.push(current);
98 current = [];
99 currentTokens = 0;
100 }
101 if (current.length === 0 && tokens > budget) {
102 throw new Error(
103 `state leaves no room for questions (~${stateTokens} of ${options.maxRequestTokens} tokens)`,
104 );
105 }
106 current.push(call);
107 currentTokens += tokens;
108 }
109 if (current.length > 0) batches.push(current);
110 return batches;
111}
112
113export function decideCall(
114 call: Pick<ToolCall, 'id' | 'tool' | 'pinned'>,
115 answer: CallAnswer,
116 options: Pick<ResolvedCompactOptions, 'keepThreshold'>,
117): CallDecision {
118 const base = { id: call.id, tool: call.tool, ...answer };
119 if (call.pinned) return { ...base, action: 'keep', reason: 'pinned' };
120 if (answer.keepResult >= options.keepThreshold) {
121 return { ...base, action: 'keep', reason: 'kept' };
122 }
123 if (answer.keepCall >= options.keepThreshold) {
124 return { ...base, action: 'drop_result', reason: 'result_dropped' };
125 }
126 return { ...base, action: 'drop_call', reason: 'call_dropped' };
127}
128
129async function askBatch(
130 asker: JevAsker,
131 state: CompactionState,
132 batch: readonly ToolCall[],
133): Promise<Map<string, CallAnswer>> {
134 const questions: JevQuestions = Object.assign({}, ...batch.map(questionsFor));
135 const { answers } = await asker.ask(state, questions);
136 return new Map(
137 batch.map((call) => [
138 call.id,
139 {
140 keepCall: noulAnswer(answers, `call_${call.id}`),
141 keepResult: noulAnswer(answers, `result_${call.id}`),
142 },
143 ]),
144 );
145}
146
147function truncatedResultText(text: string, isError: boolean, headChars: number): string {
148 if (text.length <= headChars + 120) return text;
149 const head = headChars > 0 ? `${text.slice(0, headChars)}\n` : '';
150 return `${head}[fast-jev-compaction truncated ${text.length - headChars} chars of this tool result${
151 isError ? ' (error)' : ''
152 }; re-run the tool if needed]`;
153}
154
155/**
156 * Rebuilds the conversation from the decisions. A dropped call disappears
157 * together with its result; a dropped result keeps a bounded head and note.
158 * Messages that lose all their content are removed; untouched messages are
159 * returned as the same objects they came in as.
160 */
161export function applyDecisions(
162 messages: readonly Message[],
163 decisions: readonly CallDecision[],
164 calls: readonly ToolCall[],
165 headChars: number,
166): Message[] {
167 const byId = new Map(calls.map((call) => [call.id, call]));
168 const actions = new Map<string, CallDecision['action']>();
169 for (const decision of decisions) {
170 const call = byId.get(decision.id);
171 if (call && decision.action !== 'keep') actions.set(call.tool_use_id, decision.action);
172 }
173 const kept: Message[] = [];
174 for (const message of messages) {
175 const touched =
176 message.toolUses.some((tool) => actions.has(tool.tool_use_id)) ||
177 (message.toolResults ?? []).some((result) => actions.has(result.tool_use_id));
178 if (!touched) {
179 kept.push(message);
180 continue;
181 }
182 const toolUses = message.toolUses
183 .filter((tool) => actions.get(tool.tool_use_id) !== 'drop_call')
184 .map((tool) => {
185 if (actions.get(tool.tool_use_id) !== 'drop_result') return tool;
186 const text = truncatedResultText(
187 tool.text ?? '',
188 tool.isError ?? false,
189 headChars,
190 );
191 if ((tool.text ?? '') === text) return tool;
192 const copy: ToolUse = {
193 tool_use_id: tool.tool_use_id,
194 tool: tool.tool,
195 input: tool.input,
196 text,
197 };
198 if (tool.isError) copy.isError = true;
199 return copy;
200 });
201 const toolResults = (message.toolResults ?? [])
202 .filter((result) => actions.get(result.tool_use_id) !== 'drop_call')
203 .map((result) => {
204 if (actions.get(result.tool_use_id) !== 'drop_result') return result;
205 const text = truncatedResultText(result.text, result.isError ?? false, headChars);
206 return text === result.text
207 ? result
208 : {
209 tool_use_id: result.tool_use_id,
210 text,
211 isError: result.isError,
212 };
213 });
214 if (
215 !message.toolUses.some(
216 (tool) => actions.get(tool.tool_use_id) === 'drop_call',
217 ) &&
218 !(message.toolResults ?? []).some(
219 (result) => actions.get(result.tool_use_id) === 'drop_call',
220 ) &&
221 toolUses.every((tool, index) => tool === message.toolUses[index]) &&
222 toolResults.every(
223 (result, index) => result === message.toolResults?.[index],
224 )
225 ) {
226 kept.push(message);
227 continue;
228 }
229 if (message.text.trim().length === 0 && toolUses.length === 0 && toolResults.length === 0) {
230 continue;
231 }
232 const rebuilt: Message = { role: message.role, text: message.text, toolUses };
233 if (toolResults.length > 0) rebuilt.toolResults = toolResults;
234 kept.push(rebuilt);
235 }
236 return kept;
237}
238
239/** Characters of text, tool input and tool output a message holds. */
240export function messageChars(message: Message): number {
241 let total = message.text.length;
242 for (const tool of message.toolUses) {
243 try {
244 total += JSON.stringify(tool.input).length;
245 } catch {
246 total += 20;
247 }
248 }
249 for (const result of message.toolResults ?? []) total += result.text.length;
250 return total;
251}
252
253export function reductionRatio(result: Pick<CompactResult, 'stats'>): number {
254 const { charsBefore, charsAfter } = result.stats;
255 return charsBefore === 0 ? 0 : (charsBefore - charsAfter) / charsBefore;
256}
257
258function inputChars(input: Record<string, unknown>): number {
259 try {
260 return JSON.stringify(input).length;
261 } catch {
262 return 20;
263 }
264}
265
266/**
267 * Puts pruned calls back, the ones Jev was least sure about first (highest
268 * `keepResult`, then `keepCall`), until the compaction removes at most
269 * `maxPruneRatio` of the history. One compaction that drops nearly everything
270 * leaves the model guessing what it had; this bounds that. Returns `result`
271 * itself when it is within the limit.
272 */
273export function limitPruning(
274 messages: readonly Message[],
275 result: CompactResult,
276 options: { maxPruneRatio: number } & Partial<Pick<ResolvedCompactOptions, 'preserveRecentMessages' | 'truncateHeadChars'>>,
277): CompactResult {
278 const { charsBefore } = result.stats;
279 const max = options.maxPruneRatio;
280 if (!(max < 1) || charsBefore === 0 || reductionRatio(result) <= max) return result;
281 const resolved = resolveOptions(options);
282 const calls = collectToolCalls(messages, resolved.preserveRecentMessages);
283 const byId = new Map(calls.map((call) => [call.id, call]));
284 const resultText = new Map<string, string>();
285 for (const message of messages) {
286 for (const r of message.toolResults ?? []) resultText.set(r.tool_use_id, r.text);
287 }
288 const saved = (decision: CallDecision): number => {
289 const call = byId.get(decision.id);
290 if (!call) return 0;
291 const text = resultText.get(call.tool_use_id) ?? '';
292 if (decision.action === 'drop_call') return inputChars(call.input) + text.length;
293 if (decision.action === 'drop_result') {
294 return text.length - truncatedResultText(text, call.isError, resolved.truncateHeadChars).length;
295 }
296 return 0;
297 };
298 const decisions = result.decisions.map((decision) => ({ ...decision }));
299 const position = new Map(decisions.map((decision, i) => [decision, i]));
300 // Least sure first; on a tie the newer call (decisions run oldest first).
301 const order = decisions
302 .filter((decision) => decision.action !== 'keep' && saved(decision) > 0)
303 .sort((a, b) => b.keepResult - a.keepResult || b.keepCall - a.keepCall || position.get(b)! - position.get(a)!);
304 let after = result.stats.charsAfter;
305 let restored = 0;
306 const excess = (): number => charsBefore - after - max * charsBefore;
307 const restore = (decision: CallDecision): void => {
308 after += saved(decision);
309 decision.action = 'keep';
310 decision.reason = 'restored';
311 restored++;
312 };
313 // First pass, in that order: calls that do not overshoot the cap by more than
314 // a tenth of the history (one huge output put back would undo the whole
315 // compaction). Second, what is still needed with as little put back as
316 // possible: the smallest call that covers it alone, else the largest first.
317 const slack = 0.1 * charsBefore;
318 for (const decision of order) {
319 if (excess() <= 0) break;
320 if (saved(decision) <= excess() + slack) restore(decision);
321 }
322 if (excess() > 0) {
323 const left = order.filter((decision) => decision.action !== 'keep');
324 const alone = left.filter((decision) => saved(decision) >= excess()).sort((a, b) => saved(a) - saved(b))[0];
325 for (const decision of alone ? [alone] : left.sort((a, b) => saved(b) - saved(a))) {
326 if (excess() <= 0) break;
327 restore(decision);
328 }
329 }
330 const kept = applyDecisions(messages, decisions, calls, resolved.truncateHeadChars);
331 return {
332 messages: kept,
333 decisions,
334 stats: {
335 ...result.stats,
336 messagesAfter: kept.length,
337 charsAfter: kept.reduce((sum, message) => sum + messageChars(message), 0),
338 resultsDropped: count(decisions, 'result_dropped'),
339 callsDropped: count(decisions, 'call_dropped'),
340 restored,
341 },
342 };
343}
344
345function count(decisions: readonly CallDecision[], reason: CallDecision['reason']): number {
346 return decisions.filter((decision) => decision.reason === reason).length;
347}
348
349/**
350 * Compacts a transcript by asking Jev, for every tool call outside the pinned
351 * first and newest messages, whether the call and whether its result must
352 * stay. The whole history (results omitted, fitted into `maxStateTokens`) is
353 * sent as state with every batch of questions. Throws when Jev fails or the
354 * history cannot be fitted; the caller decides whether to fall back.
355 */
356export async function compact(
357 messages: readonly Message[],
358 asker: JevAsker,
359 options: CompactOptions = {},
360): Promise<CompactResult> {
361 const started = Date.now();
362 const resolved = resolveOptions(options);
363 const calls = collectToolCalls(messages, resolved.preserveRecentMessages);
364 const candidates = calls.filter((call) => !call.pinned);
365 const charsBefore = messages.reduce((sum, message) => sum + messageChars(message), 0);
366
367 let fitted: { tokens: number; stage: string } = { tokens: 0, stage: '' };
368 let batches: ToolCall[][] = [];
369 const answers = new Map<string, CallAnswer>();
370 if (candidates.length > 0) {
371 const state = fitState(messages, calls, resolved);
372 fitted = state;
373 batches = batchCalls(candidates, state.tokens, resolved);
374 const answered = await Promise.all(
375 batches.map((batch) => askBatch(asker, state.state, batch)),
376 );
377 for (const map of answered) for (const [id, answer] of map) answers.set(id, answer);
378 }
379
380 const decisions = calls.map((call) =>
381 decideCall(call, answers.get(call.id) ?? { keepCall: 1, keepResult: 1 }, resolved),
382 );
383 const kept = applyDecisions(
384 messages,
385 decisions,
386 calls,
387 resolved.truncateHeadChars,
388 );
389 return {
390 messages: kept,
391 decisions,
392 stats: {
393 messagesBefore: messages.length,
394 messagesAfter: kept.length,
395 charsBefore,
396 charsAfter: kept.reduce((sum, message) => sum + messageChars(message), 0),
397 calls: calls.length,
398 kept: count(decisions, 'kept'),
399 resultsDropped: count(decisions, 'result_dropped'),
400 callsDropped: count(decisions, 'call_dropped'),
401 pinned: count(decisions, 'pinned'),
402 stateTokens: fitted.tokens,
403 stateStage: fitted.stage,
404 requests: batches.length,
405 ms: Date.now() - started,
406 },
407 };
408}
409hooks/lib/compaction/fold.ts 361 lines1// jev-governor addition to the vendored compaction (not upstream): Jev folds
2// old dialog text. compact.ts only ever touches tool calls, so in a long chat
3// earlier answers, background-task notifications and long pastes pile up and
4// are carried through every compaction (one real chat kept ~96k tokens of
5// them, written into the cache again at each of its 67 compactions). A folded
6// message keeps its first lines and a pointer to the archived full text, which
7// the model reads back when it needs it.
8
9import { noulAnswer } from './request.ts';
10import { estimateTokens, isPinned } from './state.ts';
11import type { CompactionState, JevAsker, JevQuestions, Message, Role } from './types.ts';
12
13/**
14 * `answer`: the assistant's text; `paste`: a long user message; `agent`,
15 * `monitor`, `task`: a `<task-notification>` with an agent's result, a
16 * monitor's event, or anything else (a background command finishing).
17 */
18export type FoldKind = 'answer' | 'paste' | 'agent' | 'monitor' | 'task';
19
20export interface FoldOptions {
21 /** The newest user prompts whose turns are never folded (superseded notifications aside). */
22 keepTurns: number;
23 /** Newest messages never touched; the first message is always kept. */
24 preserveRecentMessages: number;
25 /** Shortest text of each kind worth folding. */
26 minChars: Record<FoldKind, number>;
27 /** Keep probability at or above which a message of each kind stays verbatim. */
28 keepAt: Record<FoldKind, number>;
29 /** Characters of a folded text (a notification: of its body) left in place. */
30 headChars: Record<FoldKind, number>;
31 /** Characters of a candidate's start quoted in its question (and half that of its end). */
32 previewChars: number;
33 /** Estimated token ceiling for the state plus one batch of questions. */
34 maxRequestTokens: number;
35 /** Applied to the quoted start and end before Jev sees them (secrets). */
36 previewFilter?: (text: string) => string;
37}
38
39export const DEFAULT_FOLD_OPTIONS: FoldOptions = {
40 keepTurns: 2,
41 preserveRecentMessages: 6,
42 minChars: { answer: 800, paste: 3000, agent: 800, monitor: 200, task: 400 },
43 // Jev's keep probabilities for old dialog sit between ~0.15 and ~0.35. On 340
44 // messages of a long chat graded in hindsight (did the work after the
45 // compaction need more than the first 300 chars?) none was needed beyond its
46 // head; Jev ranked the ones related to the ongoing work highest (AUC 0.84).
47 // These cut-offs fold ~80% of the text and keep about half of the related
48 // ones. A user's own words are folded only when Jev is quite sure.
49 keepAt: { answer: 0.3, paste: 0.2, agent: 0.25, monitor: 0.35, task: 0.35 },
50 // A progress event's gist is in its first line; a newer one usually repeats it.
51 headChars: { answer: 300, paste: 300, agent: 300, monitor: 120, task: 160 },
52 previewChars: 240,
53 maxRequestTokens: 30_000,
54};
55
56export interface FoldCandidate {
57 /** Question key suffix, `m<index>`. */
58 id: string;
59 /** Index in the messages the candidates were collected from. */
60 index: number;
61 role: Role;
62 kind: FoldKind;
63 chars: number;
64 /** User prompts after this message. */
65 turnsAgo: number;
66 /** A notification's `<summary>`. */
67 label?: string;
68 /** A newer notification of the same task id. */
69 supersededBy?: number;
70}
71
72export interface FoldDecision extends FoldCandidate {
73 /** Jev's probability that the message must stay verbatim. */
74 keep: number;
75 fold: boolean;
76}
77
78/** Starts every folded text, so a later compaction leaves it alone. */
79export const FOLD_MARK = '[jev-governor folded';
80
81/** What a fold must take away beyond the head it leaves, or the pointer is not worth it. */
82const MIN_SAVING = 250;
83
84const NOTICE = '<task-notification>';
85/** User-side text the engine writes, not a prompt the person typed. */
86const ENGINE_TEXT = /^<(task-notification|system-reminder|local-command-|command-|bash-)/;
87
88/** Text the person typed (not a notification, reminder or command echo the engine wrote). */
89export function isTypedText(text: string): boolean {
90 const trimmed = text.trimStart();
91 return trimmed.length > 0 && !ENGINE_TEXT.test(trimmed);
92}
93
94/** A prompt the person wrote: starts a turn. */
95export function isPrompt(message: Message): boolean {
96 if (message.role !== 'user' || (message.toolResults ?? []).length > 0) return false;
97 return isTypedText(message.text);
98}
99
100/** A file name for a folded text: the same text always gets the same name (FNV-1a and length). */
101export function foldFileName(text: string): string {
102 let hash = 0x811c9dc5;
103 for (let i = 0; i < text.length; i++) {
104 hash ^= text.charCodeAt(i);
105 hash = Math.imul(hash, 0x01000193) >>> 0;
106 }
107 return `${hash.toString(16).padStart(8, '0')}-${text.length}.md`;
108}
109
110function noticeKind(text: string): FoldKind | undefined {
111 if (!text.trimStart().startsWith(NOTICE)) return undefined;
112 if (text.includes('<result>')) return 'agent';
113 if (text.includes('<event>')) return 'monitor';
114 return 'task';
115}
116
117function tag(text: string, name: string): string | undefined {
118 return new RegExp(`<${name}>([\\s\\S]*?)</${name}>`).exec(text)?.[1]?.trim();
119}
120
121/** The part a fold shortens: a notification's result or event, any other message's whole text. */
122function foldableChars(text: string, kind: FoldKind): number {
123 if (kind === 'answer' || kind === 'paste') return text.length;
124 return (tag(text, 'result') ?? tag(text, 'event') ?? text).length;
125}
126
127/**
128 * The messages Jev is asked about: older than the newest `keepTurns` prompts
129 * (a notification a newer one of the same task supersedes: anywhere), outside
130 * the pinned first and newest messages, long enough to matter, not folded yet.
131 */
132export function collectFoldCandidates(
133 messages: readonly Message[],
134 options: Pick<FoldOptions, 'keepTurns' | 'preserveRecentMessages' | 'minChars' | 'headChars'>,
135): FoldCandidate[] {
136 const prompts: number[] = [];
137 const lastOfTask = new Map<string, number>();
138 messages.forEach((message, index) => {
139 if (isPrompt(message)) prompts.push(index);
140 else if (message.role === 'user' && noticeKind(message.text)) {
141 const id = tag(message.text, 'task-id');
142 if (id) lastOfTask.set(id, index);
143 }
144 });
145 const keepTurns = Math.max(0, Math.floor(options.keepTurns));
146 const cutoff = keepTurns === 0 ? messages.length : prompts.length > keepTurns ? prompts[prompts.length - keepTurns]! : 0;
147 const candidates: FoldCandidate[] = [];
148 let seen = 0;
149 messages.forEach((message, index) => {
150 while (seen < prompts.length && prompts[seen]! <= index) seen++;
151 if (isPinned(index, messages.length, options.preserveRecentMessages)) return;
152 const text = message.text;
153 if (!text || text.includes(FOLD_MARK)) return;
154 let kind: FoldKind | undefined;
155 let label: string | undefined;
156 let supersededBy: number | undefined;
157 if (message.role === 'assistant') kind = 'answer';
158 else {
159 kind = noticeKind(text);
160 if (kind) {
161 label = tag(text, 'summary');
162 const id = tag(text, 'task-id');
163 const last = id === undefined ? undefined : lastOfTask.get(id);
164 if (last !== undefined && last > index) supersededBy = last;
165 } else if (isPrompt(message)) kind = 'paste';
166 }
167 if (!kind) return;
168 const foldable = foldableChars(text, kind);
169 if (foldable < options.minChars[kind] || foldable < options.headChars[kind] + MIN_SAVING) return;
170 if (index >= cutoff && supersededBy === undefined) return;
171 candidates.push({
172 id: `m${index}`,
173 index,
174 role: message.role,
175 kind,
176 chars: text.length,
177 turnsAgo: prompts.length - seen,
178 ...(label ? { label: label.slice(0, 160) } : {}),
179 ...(supersededBy !== undefined ? { supersededBy } : {}),
180 });
181 });
182 return candidates;
183}
184
185const KIND_TEXT: Record<FoldKind, string> = {
186 answer: 'an earlier message of the assistant',
187 paste: 'a long earlier message of the user',
188 agent: 'the result of a background agent',
189 monitor: 'an event of a background monitor',
190 task: 'a background task notification',
191};
192
193function preview(text: string, chars: number, filter?: (text: string) => string): string {
194 const flat = (part: string): string => (filter ? filter(part) : part).replace(/\s+/g, ' ').trim();
195 const tail = Math.floor(chars / 2);
196 if (text.length <= chars + tail + 40) return `«${flat(text)}»`;
197 // The filter sees a margin past each cut: a secret across the cut is whole when it is replaced.
198 return `«${flat(text.slice(0, chars + 80)).slice(0, chars)} […] ${flat(text.slice(-(tail + 80))).slice(-tail)}»`;
199}
200
201/** The one `noul` question asked about a candidate: must it stay verbatim. */
202export function foldQuestion(
203 candidate: FoldCandidate,
204 text: string,
205 options: Pick<FoldOptions, 'previewChars' | 'previewFilter'>,
206): JevQuestions {
207 const what = [
208 KIND_TEXT[candidate.kind],
209 candidate.label ? `"${candidate.label}"` : '',
210 `${candidate.chars} chars`,
211 candidate.turnsAgo > 0 ? `${candidate.turnsAgo} user prompt(s) ago` : 'in the current turn',
212 ]
213 .filter(Boolean)
214 .join(', ');
215 const later =
216 candidate.supersededBy === undefined ? '' : ` A newer notification of the same task follows at i=${candidate.supersededBy}.`;
217 return {
218 [`keep_${candidate.id}`]: {
219 type: 'noul',
220 instructions: `Message i=${candidate.index} (${what}) should stay in the history verbatim: what the assistant does next still depends on its exact wording or details. Otherwise it is folded to its first lines and a pointer to a file with the full text, which the assistant can read back at any time.${later} It reads: ${preview(text, options.previewChars, options.previewFilter)}`,
221 },
222 };
223}
224
225/** Splits candidates into batches whose questions fit one request next to the state. */
226export function batchFolds(
227 candidates: readonly FoldCandidate[],
228 messages: readonly Message[],
229 stateTokens: number,
230 options: Pick<FoldOptions, 'previewChars' | 'previewFilter' | 'maxRequestTokens'>,
231): FoldCandidate[][] {
232 const budget = options.maxRequestTokens - stateTokens - 20;
233 const batches: FoldCandidate[][] = [];
234 let current: FoldCandidate[] = [];
235 let used = 0;
236 for (const candidate of candidates) {
237 const tokens = estimateTokens(JSON.stringify(foldQuestion(candidate, messages[candidate.index]?.text ?? '', options)));
238 if (tokens > budget) throw new Error(`state leaves no room for fold questions (~${stateTokens} of ${options.maxRequestTokens} tokens)`);
239 if (current.length > 0 && used + tokens > budget) {
240 batches.push(current);
241 current = [];
242 used = 0;
243 }
244 current.push(candidate);
245 used += tokens;
246 }
247 if (current.length > 0) batches.push(current);
248 return batches;
249}
250
251/** Asks Jev about every candidate (batches in parallel). Throws when Jev fails. */
252export async function askFolds(
253 asker: JevAsker,
254 state: CompactionState,
255 stateTokens: number,
256 candidates: readonly FoldCandidate[],
257 messages: readonly Message[],
258 options: FoldOptions,
259): Promise<{ decisions: FoldDecision[]; requests: number }> {
260 if (candidates.length === 0) return { decisions: [], requests: 0 };
261 const batches = batchFolds(candidates, messages, stateTokens, options);
262 const answered = await Promise.all(
263 batches.map(async (batch) => {
264 const questions: JevQuestions = Object.assign(
265 {},
266 ...batch.map((candidate) => foldQuestion(candidate, messages[candidate.index]?.text ?? '', options)),
267 );
268 const { answers } = await asker.ask(state, questions);
269 return batch.map((candidate): FoldDecision => {
270 const keep = noulAnswer(answers, `keep_${candidate.id}`);
271 return { ...candidate, keep, fold: keep < options.keepAt[candidate.kind] };
272 });
273 }),
274 );
275 return { decisions: answered.flat(), requests: batches.length };
276}
277
278/** The first `chars` of a text, cut at a line or sentence end where one is near. */
279export function headOf(text: string, chars: number): string {
280 const trimmed = text.trim();
281 if (chars <= 0) return '';
282 if (trimmed.length <= chars) return trimmed;
283 const cut = trimmed.slice(0, chars);
284 const floor = Math.floor(chars * 0.6);
285 const line = cut.lastIndexOf('\n');
286 if (line >= floor) return `${cut.slice(0, line).trimEnd()}\n…`;
287 const sentence = Math.max(cut.lastIndexOf('. '), cut.lastIndexOf('! '), cut.lastIndexOf('? '));
288 if (sentence >= floor) return `${cut.slice(0, sentence + 1)} …`;
289 return `${cut.trimEnd()}…`;
290}
291
292/** What a folded message reads: the pointer, then its first lines (a notification keeps its envelope). */
293export function foldedText(text: string, candidate: Pick<FoldCandidate, 'kind' | 'chars'>, path: string, headChars: number): string {
294 const note = `${FOLD_MARK} ${candidate.chars} chars to save context; the full text is in ${path} — read it there if you need it]`;
295 if (candidate.kind !== 'answer' && candidate.kind !== 'paste') {
296 // Task id, output file, status and summary stay; the boilerplate note goes; the body folds.
297 const envelope = text.replace(/\n?<note>[\s\S]*?<\/note>/, '');
298 for (const name of ['result', 'event']) {
299 const body = tag(envelope, name);
300 if (body === undefined) continue;
301 const head = headOf(body, headChars);
302 return envelope.replace(new RegExp(`<${name}>[\\s\\S]*?</${name}>`), () => `<${name}>${note}${head ? `\n${head}` : ''}</${name}>`);
303 }
304 }
305 const head = headOf(text, headChars);
306 return `${note}${head ? `\n${head}` : ''}`;
307}
308
309/** What the archive file of a folded message holds. */
310export function foldArchiveText(text: string, candidate: Pick<FoldCandidate, 'kind' | 'role' | 'label'>): string {
311 return `# Folded ${candidate.role} message (${KIND_TEXT[candidate.kind]}${candidate.label ? `: ${candidate.label}` : ''})\n\n${text}\n`;
312}
313
314/**
315 * The messages with every folded candidate that has an archive path replaced
316 * by a copy holding its folded text (when that is shorter by a margin); all
317 * other messages are the same objects.
318 */
319export function applyFolds(
320 messages: readonly Message[],
321 decisions: readonly FoldDecision[],
322 paths: ReadonlyMap<number, string>,
323 headChars: Record<FoldKind, number>,
324): Message[] {
325 const byIndex = new Map(decisions.filter((d) => d.fold && paths.has(d.index)).map((d) => [d.index, d]));
326 return messages.map((message, index) => {
327 const decision = byIndex.get(index);
328 if (!decision) return message;
329 const text = foldedText(message.text, decision, paths.get(index)!, headChars[decision.kind]);
330 if (text.length > message.text.length - 100) return message;
331 const folded: Message = { role: message.role, text, toolUses: message.toolUses };
332 if (message.toolResults) folded.toolResults = message.toolResults;
333 return folded;
334 });
335}
336
337/** Jev's answer on each candidate, compact for the ledger: `kind:keep:chars:turnsAgo`, at most 40. */
338export function foldKeeps(decisions: readonly FoldDecision[]): string[] {
339 return decisions.slice(0, 40).map((d) => `${d.kind}:${d.keep.toFixed(2)}:${d.chars}:${d.turnsAgo}`);
340}
341
342/** Counts for the ledger: candidates, folded, and characters folded away, per kind. */
343export function foldStats(
344 decisions: readonly FoldDecision[],
345 before: readonly Message[],
346 after: readonly Message[],
347): { candidates: number; folded: number; charsSaved: number; byKind: Partial<Record<FoldKind, number>> } {
348 const byKind: Partial<Record<FoldKind, number>> = {};
349 let folded = 0;
350 let charsSaved = 0;
351 for (const decision of decisions) {
352 const was = before[decision.index];
353 const now = after[decision.index];
354 if (!was || !now || was === now) continue;
355 folded++;
356 byKind[decision.kind] = (byKind[decision.kind] ?? 0) + 1;
357 charsSaved += was.text.length - now.text.length;
358 }
359 return { candidates: decisions.length, folded, charsSaved, byKind };
360}
361hooks/lib/lang.ts 32 lines1// The language of the mod's own messages (chat notices, toasts, command replies):
2// the setting `ui.language` is 'en', 'ru' or 'auto'; 'auto' follows Claude Code's
3// `language` setting, then the system locale.
4
5export type Lang = 'en' | 'ru';
6
7const isRussian = (s: unknown): boolean => typeof s === 'string' && /^(ru|rus|russian|рус)/i.test(s.trim());
8
9/**
10 * `claudeLanguage` is `language` from Claude Code's settings.json (a name or code),
11 * `locales` the environment's (LC_ALL, LANG, the system locale), most specific first.
12 */
13export function resolveLang(setting: string | undefined, claudeLanguage?: unknown, locales: readonly (string | undefined)[] = []): Lang {
14 if (setting === 'ru') return 'ru';
15 if (setting === 'en') return 'en';
16 if (typeof claudeLanguage === 'string' && claudeLanguage.trim()) return isRussian(claudeLanguage) ? 'ru' : 'en';
17 const locale = locales.find((l) => l && l !== 'C' && l !== 'POSIX');
18 return isRussian(locale) ? 'ru' : 'en';
19}
20
21/** The locales of this process: LC_ALL, LANG, and Intl's (the system's on macOS and Windows). */
22export function systemLocales(): (string | undefined)[] {
23 let intl: string | undefined;
24 try {
25 intl = Intl.DateTimeFormat().resolvedOptions().locale;
26 } catch {
27 // none
28 }
29 const env = (globalThis as { process?: { env?: Record<string, string | undefined> } }).process?.env;
30 return [env?.LC_ALL, env?.LANG, intl];
31}
32hooks/lib/config.ts 337 lines1import { EFFORTS, type Effort, type GovernorConfig } from './types.ts';
2
3export const DEFAULT_CONFIG: GovernorConfig = {
4 version: 1,
5 enabled: true,
6 mode: 'active',
7 shadowPaths: [],
8 activePaths: [],
9 excludeProjects: [],
10 jev: {
11 model: 'jev-latest',
12 endpoint: 'https://openrouter.ai/api/v1/systemone',
13 keyFile: '~/.claude/jev-governor/openrouter.key',
14 timeoutMs: 2500,
15 },
16 models: { standard: 'claude-sonnet-5-5', strong: 'claude-opus-5-5', light: 'claude-haiku-4-5-20251001' },
17 router: {
18 mainModel: true,
19 mainEffort: true,
20 subagents: true,
21 upgradeAt: 0.55,
22 forceUpgradeAt: 0.8,
23 downgradeAt: 0.65,
24 subagentStrongAt: 0.5,
25 lightSubagents: 'shadow',
26 lightBelow: 0.15,
27 lightMaxSteps: 30,
28 cheapSwitchTokens: 30_000,
29 expectedTurns: 4,
30 standardMaxContextTokens: 180_000,
31 cacheTtlMinutes: 60,
32 minEffort: 'medium',
33 maxEffort: 'max',
34 defaultEffort: 'medium',
35 fallbackEffort: 'high',
36 effortConfidenceAt: 0.5,
37 escalateAfterErrors: 2,
38 errorWindow: 6,
39 riskyAt: 0.7,
40 continuationAt: 0.6,
41 budgetAware: true,
42 },
43 agents: {
44 enabled: true,
45 autoCreate: true,
46 remapFrom: ['general-purpose'],
47 matchAt: 0.6,
48 maxAgents: 40,
49 exposeToModel: false,
50 draftModel: 'claude-sonnet-5-5',
51 draftTimeoutMs: 30_000,
52 waitHint: true,
53 },
54 compaction: {
55 enabled: true,
56 compactAtTokens: 200_000,
57 compactAtPercent: 0,
58 recompactAfterTokens: 60_000,
59 thresholdMinReduction: 0.4,
60 minReductionRatio: 0.25,
61 keepThreshold: 0.5,
62 preserveRecentMessages: 6,
63 truncateHeadChars: 300,
64 maxStateTokens: 25_000,
65 maxRequestTokens: 30_000,
66 onReturn: false,
67 onReturnMinTokens: 60_000,
68 onReturnMinReduction: 0.15,
69 archive: true,
70 maxPruneRatio: 0.8,
71 resultPreviewChars: 150,
72 autoWindowTokens: 250_000,
73 fold: true,
74 foldKeepTurns: 2,
75 },
76 handoff: {
77 enabled: true,
78 maxTokens: 14_000,
79 brief: 'auto',
80 briefWords: 350,
81 useJev: true,
82 suggestAtTokens: 150_000,
83 suggestAt: 0.8,
84 // Off: nobody asked for it (and compaction after a pause is off by default, compaction.onReturn); on 2026-10-06 none of the hints was taken.
85 suggestColdAtTokens: 0,
86 keepDays: 30,
87 },
88 trim: {
89 enabled: true,
90 minChars: 6_000,
91 hugeChars: 40_000,
92 listChars: 6_000,
93 briefChars: 2_000,
94 logs: 'shadow',
95 logChars: 3_000,
96 // Measured 2026-10-07: 29 real script/report outputs ≤ 0.29, realistic logs (push, deploy, compose, dev server) 0.55–0.73.
97 logAt: 0.5,
98 headLines: 30,
99 tailLines: 100,
100 contextLines: 3,
101 maxChars: 16_000,
102 useJev: true,
103 jevKeepAt: 0.35,
104 keepDays: 7,
105 maxStorageMb: 200,
106 },
107 projects: {
108 enabled: true,
109 descriptionLanguage: 'en',
110 learnFromClaude: true,
111 minSuccesses: 2,
112 describeWithClaude: true,
113 showWorktrees: false,
114 },
115 ui: { port: 4777, language: 'auto', showStatus: true, nodePath: 'node', terminal: 'Terminal' },
116 savings: { effortFactor: 0.3, defaultBaseModel: 'claude-opus-5-5', defaultBaseEffort: 'xhigh' },
117 codex: { enabled: false },
118};
119
120type Json = Record<string, unknown>;
121
122function isObject(value: unknown): value is Json {
123 return typeof value === 'object' && value !== null && !Array.isArray(value);
124}
125
126function num(value: unknown, fallback: number, min: number, max: number): number {
127 return typeof value === 'number' && Number.isFinite(value)
128 ? Math.min(max, Math.max(min, value))
129 : fallback;
130}
131
132function bool(value: unknown, fallback: boolean): boolean {
133 return typeof value === 'boolean' ? value : fallback;
134}
135
136function str(value: unknown, fallback: string): string {
137 return typeof value === 'string' && value.trim().length > 0 ? value.trim() : fallback;
138}
139
140function effort(value: unknown, fallback: Effort): Effort {
141 return typeof value === 'string' && (EFFORTS as readonly string[]).includes(value)
142 ? (value as Effort)
143 : fallback;
144}
145
146function strings(value: unknown, fallback: string[]): string[] {
147 return Array.isArray(value) && value.every((v) => typeof v === 'string')
148 ? (value as string[])
149 : fallback;
150}
151
152/**
153 * Merges a stored (possibly partial or hand-edited) config over the defaults,
154 * clamping every number to a sane range; unknown keys are dropped.
155 */
156export function resolveConfig(raw: unknown): GovernorConfig {
157 const d = DEFAULT_CONFIG;
158 const c = isObject(raw) ? raw : {};
159 const jev = isObject(c.jev) ? c.jev : {};
160 const models = isObject(c.models) ? c.models : {};
161 const r = isObject(c.router) ? c.router : {};
162 const a = isObject(c.agents) ? c.agents : {};
163 const k = isObject(c.compaction) ? c.compaction : {};
164 const t = isObject(c.trim) ? c.trim : {};
165 const h = isObject(c.handoff) ? c.handoff : {};
166 const ui = isObject(c.ui) ? c.ui : {};
167 const pr = isObject(c.projects) ? c.projects : {};
168 const sv = isObject(c.savings) ? c.savings : {};
169 const codex = isObject(c.codex) ? c.codex : {};
170 const minEffort = effort(r.minEffort, d.router.minEffort);
171 let maxEffort = effort(r.maxEffort, d.router.maxEffort);
172 if (EFFORTS.indexOf(maxEffort) < EFFORTS.indexOf(minEffort)) maxEffort = minEffort;
173 return {
174 version: 1,
175 enabled: bool(c.enabled, d.enabled),
176 mode: c.mode === 'shadow' ? 'shadow' : 'active',
177 shadowPaths: strings(c.shadowPaths, d.shadowPaths).map((p) => p.trim()).filter(Boolean),
178 activePaths: strings(c.activePaths, d.activePaths).map((p) => p.trim()).filter(Boolean),
179 excludeProjects: strings(c.excludeProjects, d.excludeProjects).map((p) => p.trim()).filter(Boolean),
180 jev: {
181 model: str(jev.model, d.jev.model),
182 endpoint: str(jev.endpoint, d.jev.endpoint),
183 keyFile: str(jev.keyFile, d.jev.keyFile),
184 timeoutMs: num(jev.timeoutMs, d.jev.timeoutMs, 300, 20_000),
185 },
186 models: {
187 standard: str(models.standard, d.models.standard),
188 strong: str(models.strong, d.models.strong),
189 light: str(models.light, d.models.light),
190 },
191 router: {
192 mainModel: bool(r.mainModel, d.router.mainModel),
193 mainEffort: bool(r.mainEffort, d.router.mainEffort),
194 subagents: bool(r.subagents, d.router.subagents),
195 upgradeAt: num(r.upgradeAt, d.router.upgradeAt, 0, 1),
196 forceUpgradeAt: num(r.forceUpgradeAt, d.router.forceUpgradeAt, 0, 1),
197 downgradeAt: num(r.downgradeAt, d.router.downgradeAt, 0, 1),
198 subagentStrongAt: num(r.subagentStrongAt, d.router.subagentStrongAt, 0, 1),
199 lightSubagents: r.lightSubagents === 'on' || r.lightSubagents === 'off' ? r.lightSubagents : 'shadow',
200 lightBelow: num(r.lightBelow, d.router.lightBelow, 0, 1),
201 lightMaxSteps: num(r.lightMaxSteps, d.router.lightMaxSteps, 1, 500),
202 cheapSwitchTokens: num(r.cheapSwitchTokens, d.router.cheapSwitchTokens, 0, 2_000_000),
203 expectedTurns: num(r.expectedTurns, d.router.expectedTurns, 1, 100),
204 standardMaxContextTokens: num(
205 r.standardMaxContextTokens,
206 d.router.standardMaxContextTokens,
207 10_000,
208 2_000_000,
209 ),
210 cacheTtlMinutes: num(r.cacheTtlMinutes, d.router.cacheTtlMinutes, 1, 24 * 60),
211 minEffort,
212 maxEffort,
213 defaultEffort: effort(r.defaultEffort, d.router.defaultEffort),
214 fallbackEffort: effort(r.fallbackEffort, d.router.fallbackEffort),
215 effortConfidenceAt: num(r.effortConfidenceAt, d.router.effortConfidenceAt, 0, 1),
216 escalateAfterErrors: num(r.escalateAfterErrors, d.router.escalateAfterErrors, 0, 20),
217 errorWindow: Math.round(num(r.errorWindow, d.router.errorWindow, 0, 100)),
218 riskyAt: num(r.riskyAt, d.router.riskyAt, 0, 1),
219 continuationAt: num(r.continuationAt, d.router.continuationAt, 0, 1),
220 budgetAware: bool(r.budgetAware, d.router.budgetAware),
221 },
222 agents: {
223 enabled: bool(a.enabled, d.agents.enabled),
224 autoCreate: bool(a.autoCreate, d.agents.autoCreate),
225 remapFrom: strings(a.remapFrom, d.agents.remapFrom),
226 matchAt: num(a.matchAt, d.agents.matchAt, 0, 1),
227 maxAgents: num(a.maxAgents, d.agents.maxAgents, 0, 500),
228 exposeToModel: bool(a.exposeToModel, d.agents.exposeToModel),
229 draftModel: str(a.draftModel, d.agents.draftModel),
230 draftTimeoutMs: num(a.draftTimeoutMs, d.agents.draftTimeoutMs, 5_000, 120_000),
231 waitHint: bool(a.waitHint, d.agents.waitHint),
232 },
233 compaction: {
234 enabled: bool(k.enabled, d.compaction.enabled),
235 compactAtTokens: num(k.compactAtTokens, d.compaction.compactAtTokens, 0, 2_000_000),
236 compactAtPercent: num(k.compactAtPercent, d.compaction.compactAtPercent, 0, 95),
237 recompactAfterTokens: num(k.recompactAfterTokens, d.compaction.recompactAfterTokens, 0, 2_000_000),
238 thresholdMinReduction: num(k.thresholdMinReduction, d.compaction.thresholdMinReduction, 0, 1),
239 minReductionRatio: num(k.minReductionRatio, d.compaction.minReductionRatio, 0, 1),
240 keepThreshold: num(k.keepThreshold, d.compaction.keepThreshold, 0, 1),
241 preserveRecentMessages: num(
242 k.preserveRecentMessages,
243 d.compaction.preserveRecentMessages,
244 0,
245 100,
246 ),
247 truncateHeadChars: num(k.truncateHeadChars, d.compaction.truncateHeadChars, 0, 5_000),
248 maxStateTokens: num(k.maxStateTokens, d.compaction.maxStateTokens, 2_000, 30_000),
249 maxRequestTokens: num(k.maxRequestTokens, d.compaction.maxRequestTokens, 3_000, 31_000),
250 onReturn: bool(k.onReturn, d.compaction.onReturn),
251 onReturnMinTokens: num(k.onReturnMinTokens, d.compaction.onReturnMinTokens, 0, 2_000_000),
252 onReturnMinReduction: num(k.onReturnMinReduction, d.compaction.onReturnMinReduction, 0, 1),
253 archive: bool(k.archive, d.compaction.archive),
254 // Below ~0.5 a capped compaction falls under the minimum reductions and is skipped every time.
255 maxPruneRatio: num(k.maxPruneRatio, d.compaction.maxPruneRatio, 0.5, 1),
256 resultPreviewChars: num(k.resultPreviewChars, d.compaction.resultPreviewChars, 0, 1_000),
257 // Under ~100k the engine would compact every few steps.
258 autoWindowTokens: Math.round(num(k.autoWindowTokens, d.compaction.autoWindowTokens, 0, 2_000_000)),
259 fold: bool(k.fold, d.compaction.fold),
260 foldKeepTurns: Math.round(num(k.foldKeepTurns, d.compaction.foldKeepTurns, 1, 50)),
261 },
262 handoff: {
263 enabled: bool(h.enabled, d.handoff.enabled),
264 maxTokens: num(h.maxTokens, d.handoff.maxTokens, 2_000, 60_000),
265 brief: h.brief === 'always' || h.brief === 'never' ? h.brief : 'auto',
266 briefWords: num(h.briefWords, d.handoff.briefWords, 100, 1_500),
267 useJev: bool(h.useJev, d.handoff.useJev),
268 suggestAtTokens: num(h.suggestAtTokens, d.handoff.suggestAtTokens, 0, 2_000_000),
269 suggestAt: num(h.suggestAt, d.handoff.suggestAt, 0, 1),
270 suggestColdAtTokens: num(h.suggestColdAtTokens, d.handoff.suggestColdAtTokens, 0, 2_000_000),
271 keepDays: num(h.keepDays, d.handoff.keepDays, 1, 365),
272 },
273 trim: {
274 enabled: bool(t.enabled, d.trim.enabled),
275 minChars: num(t.minChars, d.trim.minChars, 1_000, 1_000_000),
276 hugeChars: num(t.hugeChars, d.trim.hugeChars, 5_000, 4_000_000),
277 listChars: t.listChars === 0 ? 0 : num(t.listChars, d.trim.listChars, 1_000, 1_000_000),
278 briefChars: t.briefChars === 0 ? 0 : num(t.briefChars, d.trim.briefChars, 1_000, 1_000_000),
279 logs: t.logs === 'off' || t.logs === 'shadow' || t.logs === 'on' ? t.logs : d.trim.logs,
280 logChars: num(t.logChars, d.trim.logChars, 1_000, 1_000_000),
281 logAt: num(t.logAt, d.trim.logAt, 0.5, 1),
282 headLines: num(t.headLines, d.trim.headLines, 0, 1_000),
283 tailLines: num(t.tailLines, d.trim.tailLines, 0, 2_000),
284 contextLines: num(t.contextLines, d.trim.contextLines, 0, 50),
285 maxChars: num(t.maxChars, d.trim.maxChars, 2_000, 200_000),
286 useJev: bool(t.useJev, d.trim.useJev),
287 jevKeepAt: num(t.jevKeepAt, d.trim.jevKeepAt, 0, 1),
288 keepDays: num(t.keepDays, d.trim.keepDays, 1, 90),
289 maxStorageMb: num(t.maxStorageMb, d.trim.maxStorageMb, 10, 10_000),
290 },
291 projects: {
292 enabled: bool(pr.enabled, d.projects.enabled),
293 descriptionLanguage: pr.descriptionLanguage === 'ru' ? 'ru' : 'en',
294 learnFromClaude: bool(pr.learnFromClaude, d.projects.learnFromClaude),
295 minSuccesses: num(pr.minSuccesses, d.projects.minSuccesses, 1, 50),
296 describeWithClaude: bool(pr.describeWithClaude, d.projects.describeWithClaude),
297 showWorktrees: bool(pr.showWorktrees, d.projects.showWorktrees),
298 },
299 ui: {
300 terminal: 'Terminal',
301 port: num(ui.port, d.ui.port, 1024, 65_535),
302 language: ui.language === 'ru' || ui.language === 'en' ? ui.language : 'auto',
303 showStatus: bool(ui.showStatus, d.ui.showStatus),
304 nodePath: str(ui.nodePath, d.ui.nodePath),
305 },
306 savings: {
307 effortFactor: num(sv.effortFactor, d.savings.effortFactor, 0, 0.6),
308 defaultBaseModel: str(sv.defaultBaseModel, d.savings.defaultBaseModel),
309 defaultBaseEffort: effort(sv.defaultBaseEffort, d.savings.defaultBaseEffort),
310 },
311 codex: { enabled: bool(codex.enabled, d.codex.enabled) },
312 };
313}
314
315/** `cwd` is `path` or inside it (`~` expanded). */
316function under(path: string, cwd: string, home: string): boolean {
317 const root = expandHome(path, home).replace(/\/+$/, '');
318 return root.length > 0 && (cwd === root || cwd.startsWith(`${root}/`));
319}
320
321/** Whether a session in `cwd` only observes (shadow mode). */
322export function isShadow(config: GovernorConfig, cwd: string, home: string): boolean {
323 if (config.activePaths.some((p) => under(p, cwd, home))) return false;
324 return config.mode === 'shadow' || config.shadowPaths.some((p) => under(p, cwd, home));
325}
326
327/** Whether a session in `cwd` must not send anything to Jev (`excludeProjects`). */
328export function isExcluded(config: GovernorConfig, cwd: string, home: string): boolean {
329 return config.excludeProjects.some((p) => under(p, cwd, home));
330}
331
332/** Expands a leading `~` against the home directory. */
333export function expandHome(path: string, home: string): string {
334 if (path === '~') return home;
335 return path.startsWith('~/') ? `${home}/${path.slice(2)}` : path;
336}
337hooks/lib/history.ts 97 lines1// A small, recent view of the conversation for routing decisions: enough for
2// Jev to read a terse follow-up ("yes, do it") in context, small enough to
3// keep a decision well under a second.
4
5import { estimateTokens } from './compaction/state.ts';
6
7export type HistoryMessage = {
8 role: 'user' | 'assistant';
9 text: string;
10 toolUses: readonly {
11 tool: string;
12 input: Record<string, unknown>;
13 isError?: boolean;
14 }[];
15 toolResults?: readonly { isError?: boolean }[];
16};
17
18/** Drops the wrapper Claude Code puts around pasted text, keeping the text. */
19export function unwrapPasted(text: string): string {
20 return text.replace(/<\/?pasted_content[^>]*>/g, ' ');
21}
22
23export function clip(text: string, head: number, tail = 0): string {
24 const flat = unwrapPasted(text).replace(/\s+/g, ' ').trim();
25 if (flat.length <= head + tail + 20) return flat;
26 return tail > 0
27 ? `${flat.slice(0, head)} […] ${flat.slice(-tail)}`
28 : `${flat.slice(0, head)}…`;
29}
30
31function toolLine(tool: HistoryMessage['toolUses'][number]): string {
32 const input = tool.input ?? {};
33 const main =
34 input.command ?? input.file_path ?? input.pattern ?? input.description ?? input.url ?? input.prompt;
35 const arg = typeof main === 'string' ? clip(main, 80) : '';
36 return `${tool.tool}${arg ? `(${arg})` : ''}${tool.isError ? ' → error' : ''}`;
37}
38
39/** System prompt, tools and memory files a session carries before any message. */
40const BASE_CONTEXT_TOKENS = 30_000;
41
42/**
43 * A rough context size from the transcript, for when the engine has no figure
44 * yet (before the session's first response): base overhead plus the text,
45 * tool inputs and tool outputs at ~3.5 characters a token.
46 */
47export function estimateContextTokens(
48 messages: readonly (HistoryMessage & { toolUses: readonly { text?: string }[] })[],
49): number {
50 let chars = 0;
51 for (const message of messages) {
52 chars += message.text.length;
53 for (const tool of message.toolUses) {
54 chars += JSON.stringify(tool.input ?? {}).length + (tool.text?.length ?? 0);
55 }
56 }
57 return BASE_CONTEXT_TOKENS + Math.round(chars / 3.5);
58}
59
60/**
61 * The newest messages as one line each (text abridged, tool calls as
62 * `Tool(arg)`), oldest first, within `maxTokens`. The latest user message is
63 * left out when it equals `latest` (it is sent on its own).
64 */
65export function recentHistory(
66 messages: readonly HistoryMessage[],
67 options: { maxTokens: number; maxMessages: number; latest?: string },
68): string[] {
69 const lines: string[] = [];
70 let tokens = 0;
71 let skippedLatest = false;
72 for (let i = messages.length - 1; i >= 0 && lines.length < options.maxMessages; i--) {
73 const message = messages[i]!;
74 if (
75 !skippedLatest &&
76 options.latest !== undefined &&
77 message.role === 'user' &&
78 message.text.trim() === options.latest.trim()
79 ) {
80 skippedLatest = true;
81 continue;
82 }
83 const parts: string[] = [];
84 if (message.text.trim()) parts.push(clip(message.text, message.role === 'user' ? 500 : 300, 150));
85 if (message.toolUses.length > 0) parts.push(message.toolUses.map(toolLine).join('; '));
86 const errors = (message.toolResults ?? []).filter((r) => r.isError).length;
87 if (errors > 0) parts.push(`${errors} tool error(s)`);
88 if (parts.length === 0) continue;
89 const line = `${message.role}: ${parts.join(' | ')}`;
90 const cost = estimateTokens(line);
91 if (tokens + cost > options.maxTokens) break;
92 tokens += cost;
93 lines.push(line);
94 }
95 return lines.reverse();
96}
97