Jev-powered verbatim compaction: once context reaches 60%, prunes stale tool calls with zero generation. Skips (never falls back to the LLM summarizer) when…

<img src="docs/repo-hero.png" alt="VexJoy Agent" width="100%">
Essays and writing behind this toolkit live at vexjoy.com.
VexJoy Agent connects plain-English requests to specialist agents, skills, and workflows. /do selects the knowledge and tools needed for your task. Hooks enforce specific checks, and scripts handle repeatable work.
The aim is to give capable models useful domain knowledge without making you learn the toolkit's catalog.
<!-- Counts here must match the Four Layers table (~line 143). Verify both: python3 scripts/validate-doc-counts.py --> 43 agents, 62 skills, 78 hooks, 147 scripts. Agents carry domain knowledge, skills provide reusable methods, hooks enforce selected checks, and scripts handle repeatable plumbing.
Works across Claude Code (/do), Codex ($do), Factory (/do), Reasonix (/do).
$ claude
> /do debug this Go test
Routing: go-engineer + systematic-debugging
Phase 1/4: Reproduce: running test, capturing failure...
Phase 2/4: Hypothesize: 3 candidates from stack trace...
Phase 3/4: Verify: isolated root cause in connection pool timeout
Phase 4/4: Fix: patch applied, test passing, PR opened
✓ Delivered: PR #847, fix connection pool timeout in health check
The router pairs a Go agent with a debugging skill, then follows the task through verification and delivery.
ROUTE PLAN EXECUTE VERIFY DELIVER RECORD
┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐
│ /do │───▶│ Task │───▶│Agent │───▶│Tests │───▶│ PR │───▶│Route │
│Router│ │ Plan │ │+Skill│ │Gates │ │Branch│ │Result│
└──────┘ └──────┘ └──────┘ └──────┘ └──────┘ └──────┘
/d requires TypeSafe's Jev. Jev classifies the request, checks that the selected route preserves the requested outcome, and then dispatches the agent, skill, and pipeline.
Intent preservation is production behavior: /d restates the requested outcome and gates dispatch on a Jev receipt for that exact proposed intent, preventing an agent from quietly expanding an apple into an orchard.
Choose either transport:
# Preferred: Jev through Vercel AI Gateway
export JEV_TRANSPORT=vercel
export AI_GATEWAY_API_KEY=...
# Alternative: Jev's direct API
export JEV_TRANSPORT=direct
export TYPESAFE_API_KEY=...
JEV_TRANSPORT=auto is the default. It prefers Vercel when AI_GATEWAY_API_KEY is set, then uses the direct API when only TYPESAFE_API_KEY is set. An explicit transport never silently switches to the other one. The TypeSafe MCP plugin is not required for /d.
Intent-alignment judgments are recorded in learning.db without raw request text. Receipts keep the Jev judge model separate from the executing agent's configured model and effort level. View proposed-intent difference rates by agent profile, judge model, and transport:
python3 scripts/jev-intent-stats.py --days 30
Example:
> /d fix the flaky test in the payments module
Intent alignment (/d):
-> Restated outcome: Fix the flaky payments test without changing unrelated behavior.
-> Jev: aligned
ROUTING (/d): testing-automation-engineer + testing-preferred-patterns
Source: jev (confidence: medium)
Invoking...
Checks require evidence rather than confidence.
| Agent Says | What Happens |
|---|---|
| "Code looks correct, skip tests" | Exit gate requires test output. Blocked. |
| "Trivial change, no verification" | Hook blocks completion without evidence. |
| "Similar to before" | Skill demands case-specific proof. |
| "User is in a hurry" | Protocol overrides time pressure. |
| "I'm confident" | Gate demands exit code, not assertion. |
Hooks run at configured events. Skills state what to verify; blocking hooks enforce the checks they cover. Coverage depends on the runtime and tool path.
The content engine researches, drafts in a calibrated voice, checks 397 writing patterns, and adapts finished pieces for each platform. /html produces a self-contained report, deck, prototype, chart, or diagram.
Toolkit changes use direct review and relevant checks. Model comparisons can settle specific uncertainties; they are not required for every edit. PHILOSOPHY.md explains the validation policy. what-didnt-work.md records failed experiments, routing reversals, unvalidated A/B citations, disabled lint rules, and program refutations.
The automated nightly evolution loop (/evolve, writes to evolution-reports/) ran regularly through mid-May 2026. It is currently dormant; recent evidence has come from manual PRs instead.
git clone https://github.com/notque/vexjoy-agent.git ~/vexjoy-agent
cd ~/vexjoy-agent
./install.sh
Installs into ~/.claude/ and into ~/.codex/, ~/.factory/, ~/.hermes/, and ~/.reasonix/ when the runtime command is on PATH or its home directory exists. install.sh wraps the vexinstall engine (scripts/vexinstall/). Each SessionStart in the repo re-syncs through the same engine. Symlinks (default in a git checkout) follow git pull; --copy gives a stable snapshot.
| Flag | Effect |
|---|---|
--dry-run | Print the plan; write nothing |
--target <t> | Only claude, codex, factory, hermes, reasonix, or all |
--no-takeover | Leave old unowned copies in place (default: move them to trash and replace) |
--uninstall / --rollback | Remove owned entries / restore the newest trash and settings backup |
Layout and safety rules: docs/installer-layout.md.
Want only part of the toolkit? Run ./install.sh --configure, or copy .local.example/profile.yaml to .local/profile.yaml and edit it. Without a profile, the full toolkit installs. Details: .local.example/README.md.
| CLI | Entry Point |
|---|---|
| Claude Code | /do |
| Codex | $do |
| Factory | /do |
| Reasonix | /do |
Jev Auto-Compact plugin (optional, requires TYPESAFE_API_KEY):
claude plugin marketplace add ./plugins/jev-auto-compact
claude plugin install jev-auto-compact@jev-auto-compact -y
Replaces generated compaction summaries with Jev-judged verbatim pruning after context reaches 60%. Inspect recorded before/after tokens and duration with python3 scripts/jev-compact-evidence.py.
Full setup: docs/start-here.md
Mirrors agents, skills, and supported hooks into ~/.codex/. Codex v0.144.1+ supports 59 of the 70 unique Claude hook registrations: 29 native, 30 adapter-backed, and 11 unsupported. These are registrations, not unique hook files.
The adapter converts apply_patch operations into the Write/Edit payload expected by existing guards. It cannot intercept writes through unified_exec, unmatched MCP tools, WebSearch, or other unsupported paths. PreCompact and Stop receive less telemetry than in Claude Code. This is expanded compatibility, not full parity.
After install or any hook-definition change, run /hooks in Codex and review the new definitions before trusting them. Codex hash-trusts hook commands and skips changed, unreviewed definitions.
Mirrors agents (as "droids"), skills, and hooks into ~/.factory/. Hook config merges into ~/.factory/settings.json with paths rewritten.
Mirrors skills, 147 scripts, and 10 allowlisted hook registrations into ~/.reasonix/. Reasonix has no agent or custom-command surface; /do arrives as a skill. It exposes four events: PreToolUse, PostToolUse, UserPromptSubmit, and Stop. MCP, model, and permissions in ~/.reasonix/config.json remain user-owned.
The toolkit supplies its own routing, domain knowledge, methodology, and enforcement. The default system prompt duplicates most of that.
claude --system-prompt "."
Strips built-in tool-use instructions. The toolkit's agents, skills, hooks, and CLAUDE.md provide the project-specific guidance.
<!-- Counts here must match the intro line (~line 13). Verify both: python3 scripts/validate-doc-counts.py -->
| Layer | Count | Does |
|---|---|---|
| Agents | 43 | Domain knowledge: idiom tables, failure mode catalogs, error-to-fix mappings |
| Skills | 62 | Reusable guidance and methodology for recurring work. |
| Hooks | 78 | Lifecycle checks, context injection, and telemetry. |
| Scripts | 147 | Repeatable validation, orchestration, and plumbing. |
Full skill catalog: docs/skills.md.
┌─────────────────────────────────────────────────┐
│ SKILL.md │
│ ┌─ Frontmatter ─────────────────────────────┐ │
│ │ triggers, pairs_with, success-criteria │ │
│ └────────────────────────────────────────────┘ │
│ Reference Loading Table (conditional imports) │
│ Phased Instructions (numbered, with gates) │
│ Verification (evidence requirements) │
└─────────────────────────────────────────────────┘
A game built entirely by Claude Code using these agents, skills, and pipelines:
<video src="https://github.com/user-attachments/assets/0e74abeb-dc7e-42ba-8239-a7a98cb1ab09" width="100%" autoplay loop muted playsinline></video>
I just want to use it Install, learn /do, done.
I do knowledge work Writing, research, data analysis, moderation, HTML artifacts. No code.
I'm a developer Architecture, extension points, adding agents and skills.
I'm an AI power user Routing tables, pipelines, hooks, telemetry DB.
I'm an AI agent Machine-dense inventory. Tables, paths, schemas.
Full design philosophy: PHILOSOPHY.md
One report-only script surfaces upkeep work; it prints a digest and never edits, deletes, or blocks.
python3 scripts/stale-skill-scan.py --top 20 ranks stale skills and agents as pruning candidates. Run it quarterly; see docs/deprecation-template.md.Scheduled work follows the same boundary as everything else: judgment uses models; repeatable plumbing uses 147 scripts.
| Need | Use |
|---|---|
| Run a deterministic command on a schedule | scripts/agent-scheduler.py with runner: "command" |
| Run an agent judgment on a schedule, webhook, or file change | scripts/agent-scheduler.py with the default runner: "claude" |
| Install or remove a user crontab entry safely | scripts/crontab-manager.py |
| Audit shell cron reliability | cron-automation |
| Keep one interactive objective moving until criteria verify | objective-loop |
See CONTRIBUTING.md.
MIT. See LICENSE.
hooks/jev-auto-compact.mjs 860 lines1/**
2 * Jev Auto-Compact: constant verbatim compaction via function hooks.
3 *
4 * Two hooks:
5 * turn.complete — triggers $.session.compact() once context >= 60%
6 * session.compact — Jev-judges each tool call, returns pruned {messages}
7 *
8 * Three tiers (programs -> Jev -> no LLM):
9 * Tier 1: pair tool calls with results, pin recent, pre-filter obvious drops
10 * Tier 2: Jev judges ambiguous calls (3 criteria-rich Nouls per call)
11 * Tier 3: not needed -- zero generation, verbatim pruning only
12 *
13 * Zero dependencies. Uses $.http.fetch for Jev API, nothing else external.
14 *
15 * Nothing leaves the host with a secret in it: every message text and tool
16 * input is passed through ./redact.mjs before it enters the Jev state.
17 */
18
19import { redactText, redactToolInput } from './redact.mjs';
20
21// ---------------------------------------------------------------------------
22// Constants
23// ---------------------------------------------------------------------------
24
25const MIN_REDUCTION_RATIO = 0.15;
26/** Context-window percent at which turn.complete triggers a compaction. */
27const COMPACT_AT_PERCENT = 60;
28const PRESERVE_RECENT = 6;
29const KEEP_THRESHOLD = 0.5;
30const TRUNCATE_HEAD_CHARS = 300;
31const MAX_STATE_TOKENS = 12000;
32/**
33 * Jev accepts 64k tokens per request, and `state` plus the longest question
34 * must stay under 32k. The two limits below leave margin for the rough token
35 * estimate. State is billed once per request, so every request should carry
36 * as many questions as fit.
37 */
38const MAX_REQUEST_TOKENS = 56000;
39const STATE_HARD_LIMIT_TOKENS = 28000;
40/** A compaction that needs more requests than this goes to the fallback path. */
41const MAX_REQUESTS_PER_COMPACTION = 4;
42/** Auth and billing failures never succeed on retry; stop sending after one. */
43const STOP_ON_STATUS = new Set([401, 402]);
44const MAX_GOAL_CHARS = 500;
45const ABRIDGE_TEXT_THRESHOLD = 550;
46const ABRIDGE_HEAD = 400;
47const ABRIDGE_TAIL = 150;
48const JEV_URL = 'https://api.typesafe.ai/v1/systemone';
49const JEV_MODEL = 'jev-latest';
50/**
51 * After a `{skip}` (Jev found nothing to prune) the transcript is unchanged,
52 * so asking again next turn only repeats the same cold-cache Jev pass. Retry
53 * once the conversation has grown by this many messages ...
54 */
55const SKIP_RETRY_MIN_NEW_MESSAGES = 8;
56/** ... or context has risen by this many percentage points. */
57const SKIP_RETRY_MIN_PERCENT = 5;
58
59/** Per-tool threshold overrides (lower = keep more aggressively). */
60const TOOL_THRESHOLDS = { Edit: 0.35, Write: 0.35, NotebookEdit: 0.35 };
61
62// ---------------------------------------------------------------------------
63// Token estimation
64// ---------------------------------------------------------------------------
65
66function estimateTokens(text) {
67 if (!text) return 0;
68 let tokens = 0;
69 for (const ch of text) {
70 if (/[a-zA-Z]/.test(ch)) tokens += 1 / 6;
71 else if (/[0-9]/.test(ch)) tokens += 0.5;
72 else tokens += 0.9;
73 }
74 return Math.floor(tokens) + 1;
75}
76
77// ---------------------------------------------------------------------------
78// Tool call collection from SessionMessage[]
79// ---------------------------------------------------------------------------
80
81function collectCalls(messages) {
82 const uses = new Map();
83 const results = new Map();
84
85 for (let i = 0; i < messages.length; i++) {
86 const msg = messages[i];
87 for (const tu of msg.toolUses || []) {
88 uses.set(tu.tool_use_id, { use: tu, msgIdx: i });
89 }
90 for (const tr of msg.toolResults || []) {
91 results.set(tr.tool_use_id, { result: tr, msgIdx: i });
92 }
93 }
94
95 const total = messages.length;
96 const calls = [];
97 let seq = 0;
98
99 for (const [uid, { use, msgIdx: useIdx }] of uses) {
100 const res = results.get(uid);
101 if (!res) continue;
102 const { result, msgIdx: resIdx } = res;
103
104 seq++;
105 const pinned =
106 useIdx === 0 ||
107 resIdx === 0 ||
108 useIdx >= total - PRESERVE_RECENT ||
109 resIdx >= total - PRESERVE_RECENT;
110
111 const resultText = result.text || '';
112 const inputStr =
113 typeof use.input === 'string' ? use.input : JSON.stringify(use.input || {});
114
115 calls.push({
116 uid,
117 seqId: `t${seq}`,
118 tool: use.tool || 'unknown',
119 input: inputStr,
120 resultText,
121 resultChars: resultText.length,
122 useIdx,
123 resIdx,
124 pinned,
125 isError: result.isError || false,
126 useRef: use,
127 resRef: result,
128 });
129 }
130 return calls;
131}
132
133// ---------------------------------------------------------------------------
134// Tier 1: programmatic pre-filters
135// ---------------------------------------------------------------------------
136
137function prefilter(call) {
138 const { tool, resultChars, isError, input } = call;
139
140 // Obvious keeps: long error results (diagnostic info)
141 if (isError && resultChars > 200) return false;
142
143 // Obvious drops
144 if (tool === 'Glob' || tool === 'ListDirectory') return true;
145
146 if (tool === 'Bash') {
147 let cmd = input;
148 try {
149 const parsed = JSON.parse(input);
150 cmd = parsed.command || input;
151 } catch { /* not JSON, use raw */ }
152 const stripped = cmd.trim();
153 if (['ls', 'pwd', 'ls -la', 'ls -l', 'ls -a'].includes(stripped)) return true;
154 if (stripped.startsWith('git status')) return true;
155 }
156
157 return null; // Jev decides
158}
159
160function prefilterCalls(calls) {
161 const candidates = [];
162 const decisions = {};
163
164 for (const call of calls) {
165 if (call.pinned) {
166 decisions[call.seqId] = { action: 'keep', reason: 'pinned' };
167 continue;
168 }
169 const verdict = prefilter(call);
170 if (verdict === true) {
171 decisions[call.seqId] = { action: 'drop_call', reason: 'prefilter_drop' };
172 } else if (verdict === false) {
173 decisions[call.seqId] = { action: 'keep', reason: 'prefilter_keep' };
174 } else {
175 candidates.push(call);
176 }
177 }
178 return { candidates, decisions };
179}
180
181// ---------------------------------------------------------------------------
182// Multi-stage state fitting (ported from Python jev-compact.py)
183// ---------------------------------------------------------------------------
184
185/** Last three user prompts, redacted. Returns { goal, redactions }. */
186function extractGoal(messages) {
187 const texts = [];
188 let redactions = 0;
189 for (let i = messages.length - 1; i >= 0 && texts.length < 3; i--) {
190 if (messages[i].role === 'user' && messages[i].text) {
191 const r = redactText(messages[i].text);
192 redactions += r.count;
193 texts.unshift(r.text.substring(0, MAX_GOAL_CHARS));
194 }
195 }
196 return { goal: texts.join('\n---\n'), redactions };
197}
198
199// `msg.text` and each call's `input` arrive here already redacted (fitState
200// does that once, before the fitting stages, so a redaction counts once).
201function buildHistoryEntry(msg, idx, callsByMsg, inputLimit, abridge, collapse, pinnedIndices) {
202 const role = msg.role || 'unknown';
203 const text = msg.text || '';
204
205 // Collapse non-pinned messages without tool calls
206 if (collapse && !pinnedIndices.has(idx) && !callsByMsg.has(idx)) {
207 if (!text.trim()) return null;
208 return { i: idx, role, text: `[... ${text.length} chars omitted ...]` };
209 }
210
211 // Abridge long non-pinned text
212 let displayText = text;
213 if (abridge && !pinnedIndices.has(idx) && text.length > ABRIDGE_TEXT_THRESHOLD) {
214 displayText =
215 text.substring(0, ABRIDGE_HEAD) +
216 `\n[... ${text.length - ABRIDGE_HEAD - ABRIDGE_TAIL} chars omitted ...]\n` +
217 text.substring(text.length - ABRIDGE_TAIL);
218 }
219
220 const entry = { i: idx, role, text: displayText };
221
222 const mc = callsByMsg.get(idx);
223 if (mc) {
224 entry.tool_calls = mc.map((c) => ({
225 id: c.seqId,
226 tool: c.tool,
227 input: c.input.substring(0, inputLimit),
228 result: `${c.isError ? 'error' : 'ok'}, ${c.resultChars} chars (omitted)`,
229 }));
230 }
231
232 return entry;
233}
234
235/**
236 * Build the Jev state, shrinking it stage by stage until it fits.
237 * Redacts every message text and tool input first. Returns { state, redactions }.
238 */
239function fitState(rawMessages, rawCalls, goal) {
240 let redactions = 0;
241 const messages = rawMessages.map((m) => {
242 if (!m.text) return m;
243 const r = redactText(m.text);
244 redactions += r.count;
245 return r.count ? { ...m, text: r.text } : m;
246 });
247 const calls = rawCalls.map((c) => {
248 const r = redactToolInput(c.input);
249 redactions += r.count;
250 return r.count ? { ...c, input: r.text } : c;
251 });
252
253 const callsByMsg = new Map();
254 const pinnedIndices = new Set();
255
256 for (const c of calls) {
257 if (!callsByMsg.has(c.useIdx)) callsByMsg.set(c.useIdx, []);
258 callsByMsg.get(c.useIdx).push(c);
259 if (c.pinned) {
260 pinnedIndices.add(c.useIdx);
261 pinnedIndices.add(c.resIdx);
262 }
263 }
264
265 // Pin recent messages and first message
266 const total = messages.length;
267 for (let i = Math.max(0, total - PRESERVE_RECENT); i < total; i++) {
268 pinnedIndices.add(i);
269 }
270 pinnedIndices.add(0);
271
272 const contextText =
273 'A coding assistant conversation is being compacted to free context. ' +
274 '`history` is the whole conversation; tool outputs are replaced by notes. ' +
275 'Each question asks whether one tool call or result still needs to stay verbatim. ' +
276 'Whatever is not kept is deleted permanently, but the assistant can re-run any tool.';
277
278 // Progressive stages: [inputLimit, abridge, collapse, dropTextOnly]
279 const stages = [
280 [1000, false, false, false],
281 [200, false, false, false],
282 [60, false, false, false],
283 [60, true, false, false],
284 [60, true, true, false],
285 [60, true, true, true],
286 ];
287
288 let state;
289 for (const [inputLimit, abridge, collapse, dropTextOnly] of stages) {
290 const history = [];
291 for (let i = 0; i < messages.length; i++) {
292 const entry = buildHistoryEntry(
293 messages[i], i, callsByMsg, inputLimit, abridge, collapse, pinnedIndices
294 );
295 if (entry === null) continue;
296 if (dropTextOnly && !pinnedIndices.has(i) && !callsByMsg.has(i)) continue;
297 history.push(entry);
298 }
299
300 state = { context: contextText, goal, history };
301 const stateStr = JSON.stringify(state);
302 if (estimateTokens(stateStr) <= MAX_STATE_TOKENS) return { state, redactions };
303 }
304
305 return { state, redactions }; // smallest version
306}
307
308// ---------------------------------------------------------------------------
309// Build Jev questions
310// ---------------------------------------------------------------------------
311
312function questionsFor(call) {
313 const { seqId, tool, resultChars } = call;
314 return {
315 [`call_${seqId}`]: {
316 type: 'noul',
317 instructions: {
318 question: `Tool call ${seqId} (${tool}) should stay in the history: knowing this call was made, with its input, still matters for what the assistant does next.`,
319 criteria: {
320 true: 'The call\'s input or the fact it was made is referenced, constrains, or informs later actions.',
321 false: 'The call was exploratory, superseded by a later call to the same tool on the same target, or its input is fully captured in later context.',
322 },
323 },
324 },
325 [`result_${seqId}`]: {
326 type: 'noul',
327 instructions: {
328 question: `The full output of tool call ${seqId} (${tool}, ${resultChars} chars) should stay verbatim: the assistant still needs its contents and re-running the tool would not reproduce it.`,
329 criteria: {
330 true: 'The result contains information not available elsewhere: error details, file contents being modified, test output. Re-running would not reproduce it.',
331 false: 'The result is a simple acknowledgment, a file listing available from later reads, or output reproducible by re-running.',
332 },
333 },
334 },
335 [`referenced_${seqId}`]: {
336 type: 'noul',
337 instructions: {
338 question: `The result of tool call ${seqId} (${tool}) was referenced or used in a later assistant message or tool call input.`,
339 },
340 },
341 };
342}
343
344// ---------------------------------------------------------------------------
345// Batching
346// ---------------------------------------------------------------------------
347
348function batchCalls(candidates, stateTokens) {
349 if (!candidates.length) return [];
350 // State over the hard limit cannot be judged in any request.
351 if (stateTokens > STATE_HARD_LIMIT_TOKENS) return null;
352
353 const available = MAX_REQUEST_TOKENS - stateTokens;
354 const batches = [];
355 let current = [];
356 let used = 0;
357 for (const c of candidates) {
358 const cost = estimateTokens(JSON.stringify(questionsFor(c)));
359 if (current.length > 0 && used + cost > available) {
360 batches.push(current);
361 current = [];
362 used = 0;
363 }
364 current.push(c);
365 used += cost;
366 }
367 if (current.length > 0) batches.push(current);
368
369 // Each request re-sends the whole state. Past the cap the run costs more
370 // than it saves, so the caller takes its fallback path.
371 if (batches.length > MAX_REQUESTS_PER_COMPACTION) return null;
372 return batches;
373}
374
375// ---------------------------------------------------------------------------
376// Jev API call
377// ---------------------------------------------------------------------------
378
379async function callJev(fetchFn, apiKey, state, questions) {
380 const payload = { model: JEV_MODEL, state, questions };
381 const response = await fetchFn(JEV_URL, {
382 method: 'POST',
383 headers: {
384 Authorization: `Bearer ${apiKey}`,
385 'Content-Type': 'application/json',
386 },
387 body: JSON.stringify(payload),
388 });
389
390 if (!response.ok) {
391 const err = new Error(`Jev HTTP ${response.status}`);
392 err.status = response.status; // status only: the body is never logged
393 throw err;
394 }
395
396 const data = JSON.parse(response.text);
397 if (!data.answers) throw new Error('Jev response missing answers');
398 return data;
399}
400
401// ---------------------------------------------------------------------------
402// Decision logic
403// ---------------------------------------------------------------------------
404
405function noulValue(answers, key) {
406 const a = answers[key];
407 if (!a || typeof a.noul !== 'number') return null;
408 return a.noul;
409}
410
411function decideCall(call, keepCall, keepResult, referenced) {
412 const threshold = TOOL_THRESHOLDS[call.tool] || KEEP_THRESHOLD;
413
414 let effectiveResult = keepResult;
415 if (referenced >= 0.6) {
416 effectiveResult = Math.max(keepResult, 0.5 + (referenced - 0.6) * 0.5);
417 }
418
419 if (effectiveResult >= threshold) {
420 return { action: 'keep', reason: 'kept', keepCall, keepResult, referenced };
421 } else if (keepCall >= threshold) {
422 return { action: 'drop_result', reason: 'result_dropped', keepCall, keepResult, referenced };
423 } else {
424 return { action: 'drop_call', reason: 'call_dropped', keepCall, keepResult, referenced };
425 }
426}
427
428// ---------------------------------------------------------------------------
429// Apply decisions to SessionMessage[]
430// ---------------------------------------------------------------------------
431
432function applyDecisions(messages, decisions, calls) {
433 const uidDecision = new Map();
434 for (const call of calls) {
435 const d = decisions[call.seqId];
436 if (d) uidDecision.set(call.uid, d);
437 }
438
439 const result = [];
440 for (const msg of messages) {
441 let changed = false;
442 let newToolUses = msg.toolUses || [];
443 let newToolResults = msg.toolResults || [];
444
445 if (newToolUses.length > 0) {
446 const filtered = newToolUses.filter((tu) => {
447 const d = uidDecision.get(tu.tool_use_id);
448 if (d && d.action === 'drop_call') { changed = true; return false; }
449 return true;
450 });
451 if (filtered.length !== newToolUses.length) {
452 newToolUses = filtered;
453 changed = true;
454 }
455 }
456
457 if (newToolResults.length > 0) {
458 const filtered = [];
459 for (const tr of newToolResults) {
460 const d = uidDecision.get(tr.tool_use_id);
461 if (d && d.action === 'drop_call') { changed = true; continue; }
462 if (d && d.action === 'drop_result') {
463 const text = tr.text || '';
464 if (text.length > TRUNCATE_HEAD_CHARS) {
465 const truncated = text.length - TRUNCATE_HEAD_CHARS;
466 filtered.push({
467 ...tr,
468 text: text.substring(0, TRUNCATE_HEAD_CHARS) +
469 `\n[jev-compact truncated ${truncated} chars; re-run the tool if needed]`,
470 });
471 changed = true;
472 } else {
473 filtered.push(tr);
474 }
475 } else {
476 filtered.push(tr);
477 }
478 }
479 if (changed || filtered.length !== newToolResults.length) {
480 newToolResults = filtered;
481 changed = true;
482 }
483 }
484
485 if (!msg.text && newToolUses.length === 0 && newToolResults.length === 0) continue;
486
487 if (!changed) {
488 result.push(msg);
489 } else {
490 const rebuilt = { ...msg, text: msg.text || '', toolUses: newToolUses };
491 if (newToolResults.length > 0) rebuilt.toolResults = newToolResults;
492 else delete rebuilt.toolResults;
493 result.push(rebuilt);
494 }
495 }
496 return result;
497}
498
499// ---------------------------------------------------------------------------
500// Full compaction pipeline
501// ---------------------------------------------------------------------------
502
503/**
504 * @param {(url, init) => Promise<{status, ok, text}>} fetchFn
505 * @param {string} apiKey
506 * @param {SessionMessage[]} messages
507 * @param {(line: string) => void} [log] status-only diagnostics (never a payload)
508 */
509export async function compact(fetchFn, apiKey, messages, log = () => {}) {
510 const calls = collectCalls(messages);
511 if (calls.length === 0) return { messages, stats: { totalCalls: 0, ratio: 0, errorBatches: 0, redactions: 0 } };
512
513 // Tier 1: pre-filter
514 const { candidates, decisions } = prefilterCalls(calls);
515 const pinned = Object.values(decisions).filter((d) => d.reason === 'pinned').length;
516 const prefiltered = Object.values(decisions).filter((d) => d.reason === 'prefilter_drop').length;
517
518 if (candidates.length === 0) {
519 const result = applyDecisions(messages, decisions, calls);
520 return {
521 messages: result,
522 decisions,
523 stats: { totalCalls: calls.length, pinned, prefiltered, jevJudged: 0, ratio: 0, errorBatches: 0, redactions: 0 },
524 };
525 }
526
527 // Build state with multi-stage fitting. Nothing below this line may read
528 // `messages` text or `calls[].input` for the Jev payload: only `state`.
529 const { goal, redactions: goalRedactions } = extractGoal(messages);
530 const { state, redactions: stateRedactions } = fitState(messages, calls, goal);
531 const redactions = goalRedactions + stateRedactions;
532 const stateTokens = estimateTokens(JSON.stringify(state));
533
534 // Batch questions to fit Jev limits
535 const batches = batchCalls(candidates, stateTokens);
536 if (batches === null) {
537 log(`[jev-auto-compact] ${candidates.length} tool calls with a ${stateTokens}-token state do not fit in ${MAX_REQUESTS_PER_COMPACTION} Jev requests`);
538 return {
539 messages,
540 decisions,
541 stats: { totalCalls: calls.length, pinned, prefiltered, jevJudged: 0, jevCalls: 0, ratio: 0, errorBatches: 0, redactions, overBudget: true, callLog: [] },
542 };
543 }
544
545 const start = Date.now();
546 const answers = {};
547 let jevCalls = 0;
548 let errorBatches = 0;
549 const callLog = []; // one entry per Jev request, for the call log
550
551 // Process batches sequentially ($.http.fetch may not support concurrent)
552 for (const batch of batches) {
553 const batchQuestions = {};
554 for (const c of batch) {
555 Object.assign(batchQuestions, questionsFor(c));
556 }
557
558 const callStart = Date.now();
559 const callEntry = {
560 ts: new Date(callStart).toISOString(),
561 n_questions: Object.keys(batchQuestions).length,
562 };
563 try {
564 const jevData = await callJev(fetchFn, apiKey, state, batchQuestions);
565 Object.assign(answers, jevData.answers || {});
566 jevCalls++;
567 const u = jevData.usage || {};
568 callLog.push({
569 ...callEntry,
570 ok: true,
571 latency_ms: Date.now() - callStart,
572 input_tokens: typeof u.input_tokens === 'number' ? u.input_tokens : null,
573 output_tokens: typeof u.output_tokens === 'number' ? u.output_tokens : null,
574 });
575 } catch (err) {
576 callLog.push({
577 ...callEntry,
578 ok: false,
579 latency_ms: Date.now() - callStart,
580 error: typeof err?.status === 'number' ? `HTTP ${err.status}` : err?.name || 'error',
581 });
582 // On batch failure, keep all calls in this batch (fail-safe). Say so
583 // once per compaction, status only: never the request or response body.
584 errorBatches++;
585 if (errorBatches === 1) {
586 const status = typeof err?.status === 'number' ? `HTTP ${err.status}` : (err?.name || 'error');
587 log(`[jev-auto-compact] Jev batch failed (${status}); keeping its ${batch.length} calls`);
588 }
589 for (const c of batch) {
590 decisions[c.seqId] = { action: 'keep', reason: 'jev_error' };
591 }
592 if (STOP_ON_STATUS.has(err?.status)) {
593 // Keep everything not yet judged and send nothing more.
594 for (const c of candidates) {
595 if (!decisions[c.seqId]) decisions[c.seqId] = { action: 'keep', reason: 'jev_error' };
596 }
597 break;
598 }
599 }
600 }
601
602 const latencyMs = Date.now() - start;
603
604 // Decide per candidate
605 for (const call of candidates) {
606 if (decisions[call.seqId]) continue; // already decided (error fallback)
607 const kc = noulValue(answers, `call_${call.seqId}`);
608 const kr = noulValue(answers, `result_${call.seqId}`);
609 const ref = noulValue(answers, `referenced_${call.seqId}`);
610
611 if (kc === null) {
612 decisions[call.seqId] = { action: 'keep', reason: 'missing_answer' };
613 } else {
614 decisions[call.seqId] = decideCall(call, kc, kr ?? 0.5, ref ?? 0.5);
615 }
616 }
617
618 // Apply
619 const result = applyDecisions(messages, decisions, calls);
620 const kept = Object.values(decisions).filter((d) => d.action === 'keep').length;
621 const droppedResult = Object.values(decisions).filter((d) => d.action === 'drop_result').length;
622 const droppedCall = Object.values(decisions).filter((d) => d.action === 'drop_call').length;
623
624 const origLen = JSON.stringify(messages).length;
625 const resultLen = JSON.stringify(result).length;
626 const ratio = origLen > 0 ? 1 - resultLen / origLen : 0;
627
628 return {
629 messages: result,
630 decisions,
631 stats: {
632 totalCalls: calls.length,
633 pinned,
634 prefiltered,
635 jevJudged: candidates.length,
636 kept,
637 droppedResult,
638 droppedCall,
639 ratio: Math.round(ratio * 10000) / 10000,
640 latencyMs,
641 jevCalls,
642 errorBatches,
643 redactions,
644 callLog,
645 },
646 };
647}
648
649// ---------------------------------------------------------------------------
650// Evidence: bounded ring buffer in the plugin's own $.store (a JSON file the
651// engine keeps under the user's Claude Code config dir). $.fs.write to a
652// home-dir path is refused by the plugin sandbox, so the store is the one
653// durable place. scripts/jev-compact-evidence.py ingests it into learning.db.
654// ---------------------------------------------------------------------------
655
656const EVIDENCE_KEY = 'events';
657const EVIDENCE_MAX = 500;
658
659async function appendEvidence($, record) {
660 try {
661 const prior = await $.store.get(EVIDENCE_KEY);
662 let events = Array.isArray(prior) ? prior : [];
663 events.push(...(Array.isArray(record) ? record : [record]));
664 if (events.length > EVIDENCE_MAX) events = events.slice(-EVIDENCE_MAX);
665 await $.store.set(EVIDENCE_KEY, events);
666 } catch (error) {
667 // Evidence is best-effort; never let it break compaction. Say why once.
668 $.ui.log(`[jev-auto-compact] evidence not recorded (${error instanceof Error ? error.message : String(error)})`);
669 }
670}
671
672async function safeUsage($) {
673 try {
674 const u = await $.session.usage();
675 return {
676 tokens: u?.context?.tokens ?? null,
677 window: u?.context?.window ?? null,
678 percent: u?.context?.percent ?? null,
679 cost_usd: u?.cost?.usd ?? null,
680 };
681 } catch {
682 return { tokens: null, window: null, percent: null, cost_usd: null };
683 }
684}
685
686async function safeSessionId($) {
687 try {
688 return await $.session.id();
689 } catch {
690 return null;
691 }
692}
693
694function pct(ratio) {
695 return `${(ratio * 100).toFixed(0)}%`;
696}
697
698// ---------------------------------------------------------------------------
699// Plugin registration
700// ---------------------------------------------------------------------------
701
702/** @type {import('claude-code').Register} */
703export const register = (on, options) => {
704 let compacting = false;
705 /** Set when a plugin-triggered compaction answered {skip}: the transcript
706 * size it saw. turn.complete does not re-trigger until it has grown. */
707 let skippedAt = null; // { messages: number, percent: number|null }
708
709 on('session.compact', async ($, event, next) => {
710 const trigger = event.trigger || 'unknown';
711
712 // precompute installs nothing; a subagent's own transcript is covered by
713 // the PreCompact Python hook (function hooks only see the main loop).
714 if (trigger === 'precompute' || event.agentId) return next(event);
715
716 const startedAt = Date.now();
717 const ts = new Date(startedAt).toISOString();
718 const sessionId = await safeSessionId($);
719 const usage = await safeUsage($);
720 const base = {
721 kind: 'compaction',
722 session_id: sessionId,
723 ts,
724 source: 'plugin',
725 trigger,
726 messages_before: event.messages.length,
727 tokens_before: usage.tokens,
728 };
729
730 try {
731 const apiKey =
732 (await $.env.get('TYPESAFE_API_KEY')) ||
733 ((await $.settings.read())?.env || {})['TYPESAFE_API_KEY'];
734
735 if (!apiKey) {
736 $.ui.log('[jev-auto-compact] no TYPESAFE_API_KEY, falling back to built-in');
737 await appendEvidence($, { ...base, engine: 'builtin', note: 'no api key' });
738 return next(event);
739 }
740
741 const fetchFn = async (url, init) => {
742 const resp = await $.http.fetch(url, init);
743 return { status: resp.status, ok: resp.ok, text: resp.text };
744 };
745
746 const result = await compact(fetchFn, apiKey, event.messages, (line) => $.ui.log(line));
747 const { stats } = result;
748 if (Array.isArray(stats.callLog) && stats.callLog.length > 0) {
749 await appendEvidence(
750 $,
751 stats.callLog.map((c) => ({ kind: 'jev_call', session_id: sessionId, ...c }))
752 );
753 }
754 const statsRecord = {
755 messages_after: result.messages.length,
756 reduction_ratio: stats.ratio,
757 dropped_calls: stats.droppedCall ?? 0,
758 truncated_results: stats.droppedResult ?? 0,
759 pinned: stats.pinned ?? 0,
760 prefiltered: stats.prefiltered ?? 0,
761 jev_judged: stats.jevJudged ?? 0,
762 jev_api_calls: stats.jevCalls ?? 0,
763 jev_error_batches: stats.errorBatches ?? 0,
764 redactions: stats.redactions ?? 0,
765 latency_ms: stats.latencyMs ?? 0,
766 duration_ms: Date.now() - startedAt,
767 };
768
769 if (stats.ratio < MIN_REDUCTION_RATIO) {
770 if (trigger === 'plugin') {
771 // We asked for this compaction and Jev found nothing worth pruning.
772 // Leave the conversation as it is: never hand a self-triggered
773 // compaction to the built-in summarizer (minutes of LLM time).
774 const reason = stats.overBudget
775 ? 'too large for one Jev pass'
776 : `nothing to prune (${pct(stats.ratio)} < ${pct(MIN_REDUCTION_RATIO)} min)`;
777 $.ui.log(`[jev-auto-compact] skipped: ${reason}`);
778 await appendEvidence($, { ...base, ...statsRecord, engine: 'skipped', note: reason });
779 skippedAt = { messages: event.messages.length, percent: usage.percent };
780 return { skip: `jev-auto-compact: ${reason}` };
781 }
782 // The engine or the user needs this compaction; Jev alone is not
783 // enough, so the built-in summarizer runs.
784 const reason = stats.overBudget
785 ? 'too large for one Jev pass'
786 : `${pct(stats.ratio)} < ${pct(MIN_REDUCTION_RATIO)} min`;
787 $.ui.log(`[jev-auto-compact] fallback to built-in (${reason})`);
788 await appendEvidence($, { ...base, ...statsRecord, engine: 'builtin', note: reason });
789 return next(event);
790 }
791
792 $.ui.log(
793 `[jev-auto-compact] ${result.messages.length}/${event.messages.length} msgs, ` +
794 `${pct(stats.ratio)} reduction, ` +
795 `${stats.droppedCall} dropped, ${stats.droppedResult} truncated, ` +
796 `${stats.pinned} pinned, ${stats.prefiltered} prefiltered ` +
797 `(${stats.latencyMs}ms, ${stats.jevJudged} Jev-judged, ${stats.jevCalls} API calls` +
798 (stats.errorBatches ? `, ${stats.errorBatches} failed` : '') +
799 `, redactions: ${stats.redactions ?? 0}` +
800 (usage.tokens ? `, ctx ${usage.tokens} tokens before)` : ')')
801 );
802 await appendEvidence($, { ...base, ...statsRecord, engine: 'jev' });
803 skippedAt = null;
804
805 const compacted = { messages: result.messages };
806 if (typeof usage.tokens === 'number') compacted.tokensBefore = usage.tokens;
807 return compacted;
808 } catch (error) {
809 const message = error instanceof Error ? error.message : String(error);
810 $.ui.log(`[jev-auto-compact] fallback to built-in (${message})`);
811 await appendEvidence($, { ...base, engine: 'error', note: message.substring(0, 200) });
812 return next(event);
813 }
814 });
815
816 on('turn.complete', async ($, event, next) => {
817 // Main-loop turns only; a subagent's turn.complete carries agentId.
818 if (event.agentId || compacting) return next(event);
819 try {
820 compacting = true;
821 const usage = await safeUsage($);
822 const sessionId = await safeSessionId($);
823 await appendEvidence($, {
824 kind: 'usage',
825 session_id: sessionId,
826 ts: new Date().toISOString(),
827 phase: 'turn_complete',
828 context_tokens: usage.tokens,
829 context_window: usage.window,
830 context_percent: usage.percent,
831 cost_usd: usage.cost_usd,
832 });
833 // Threshold-gated. Every compaction is a cold KV-cache rewrite of the
834 // whole prefix (measured: ~$2 vs ~$0.44 for a normal turn), so firing
835 // every turn multiplies cost with no benefit. Fire only when context is
836 // genuinely filling; if usage is unavailable, do nothing.
837 if (typeof usage.percent !== 'number' || usage.percent < COMPACT_AT_PERCENT) {
838 return next(event);
839 }
840 if (skippedAt) {
841 // The last self-triggered compaction found nothing to prune. Re-ask
842 // only once the conversation has moved on enough for a new answer.
843 const messageCount = (await $.session.messages()).length;
844 const grewBy = messageCount - skippedAt.messages;
845 const roseBy = typeof skippedAt.percent === 'number' ? usage.percent - skippedAt.percent : Infinity;
846 if (grewBy < SKIP_RETRY_MIN_NEW_MESSAGES && roseBy < SKIP_RETRY_MIN_PERCENT) return next(event);
847 }
848 $.ui.log(`[jev-auto-compact] context at ${usage.percent.toFixed(0)}% >= ${COMPACT_AT_PERCENT}%, compacting`);
849 await $.session.compact();
850 } catch (error) {
851 $.ui.log(
852 `[jev-auto-compact] auto-compact skipped (${error instanceof Error ? error.message : String(error)})`
853 );
854 } finally {
855 compacting = false;
856 }
857 return next(event);
858 });
859};
860hooks/redact.mjs 119 lines1/**
2 * Secret redaction for text that leaves the host to the Jev API.
3 *
4 * Owner rule: nothing leaves the host to Jev with a secret in it. Every tool
5 * input and message text passes through `redactText` before it enters the
6 * compaction state; tool inputs that name a protected path are withheld whole.
7 *
8 * TYPE names are shared with the Python twin (`scripts/jev_redact.py`):
9 * bearer, cookie, jwt, aws_key, aws_secret, github, openai, slack, google,
10 * pem, url_credentials, kv_secret
11 *
12 * Replacement: `<redacted:TYPE:last4>` when the value is >= 12 chars,
13 * `<redacted:TYPE>` otherwise. Zero dependencies.
14 */
15
16const MIN_LAST4_LEN = 12;
17const MARK = '<redacted:';
18
19export function mark(type, value) {
20 const v = String(value ?? '');
21 return v.length >= MIN_LAST4_LEN ? `${MARK}${type}:${v.slice(-4)}>` : `${MARK}${type}>`;
22}
23
24// PEM armor, assembled from pieces so a secret scanner never sees a
25// key-shaped literal in this source file.
26const ARMOR = '-'.repeat(5);
27const PEM_BLOCK = new RegExp(
28 `${ARMOR}BEGIN [A-Z ]*PRIVATE KEY${ARMOR}[\\s\\S]*?${ARMOR}END [A-Z ]*PRIVATE KEY${ARMOR}`,
29 'g'
30);
31
32// Each rule: [type, regex, replacer(match, ...groups) -> string]. Order matters:
33// multi-line and structured forms first, the generic KEY=VALUE rule last, and
34// a value already marked `<redacted:` is never redacted twice.
35const RULES = [
36 ['pem', PEM_BLOCK, (m) => mark('pem', m)],
37 ['url_credentials', /\b([a-z][a-z0-9+.-]*:\/\/)(?!<redacted:)([^\s/:@<]+):([^\s/@<]+)@/gi,
38 (m, scheme, user, pass) => `${scheme}${mark('url_credentials', `${user}:${pass}`)}@`],
39 // JWT before bearer: `Bearer eyJ...` is reported as a jwt, the more specific type.
40 ['jwt', /\beyJ[A-Za-z0-9_-]{5,}\.[A-Za-z0-9_-]{5,}\.[A-Za-z0-9_-]{5,}\b/g,
41 (m) => mark('jwt', m)],
42 // Header with a scheme (Bearer/Basic/Token) or a bare `Bearer x`/`Basic x`.
43 ['bearer', /\b(authorization\s*[:=]\s*["']?)(bearer|basic|token)(\s+)(?!<redacted:)([A-Za-z0-9\-._~+/]+=*)/gi,
44 (m, hdr, scheme, sp, tok) => `${hdr}${scheme}${sp}${mark('bearer', tok)}`],
45 ['bearer', /(?<!redacted:)\b(bearer|basic)(\s+)(?!<redacted:)([A-Za-z0-9\-._~+/]{8,}=*)/gi,
46 (m, scheme, sp, tok) => `${scheme}${sp}${mark('bearer', tok)}`],
47 // Header with no scheme: the whole value is the credential.
48 ['bearer', /\b(authorization\s*[:=]\s*["']?)(?!<redacted:)(?!bearer\b|basic\b|token\b)([^\s"',;<>]{8,})/gi,
49 (m, hdr, tok) => `${hdr}${mark('bearer', tok)}`],
50 // Not preceded by `:` or a word char: the `cookie:xxxx>` inside a mark is not a header.
51 ['cookie', /(?<![:\w])((?:set-)?cookie\s*[:=]\s*["']?)([^\s"'<][^\r\n"']*)/gi,
52 (m, hdr, val) => `${hdr}${mark('cookie', val)}`],
53 ['aws_secret', /\b(aws_secret_access_key|aws[_-]?secret(?:[_-]?key)?)(\s*[:=]\s*["']?)(?!<redacted:)([A-Za-z0-9/+=]{40})\b/gi,
54 (m, key, sep, val) => `${key}${sep}${mark('aws_secret', val)}`],
55 ['aws_key', /\b(?:AKIA|ASIA)[A-Z0-9]{16}\b/g, (m) => mark('aws_key', m)],
56 ['github', /\bgh[pousr]_[A-Za-z0-9]{20,}\b/g, (m) => mark('github', m)],
57 ['openai', /\bsk-(?:proj-|live-|test-)?[A-Za-z0-9_-]{16,}\b|\bsk_(?:live|test)_[A-Za-z0-9]{8,}\b/g,
58 (m) => mark('openai', m)],
59 ['slack', /\bxox[abp]-[A-Za-z0-9-]{10,}\b/g, (m) => mark('slack', m)],
60 ['google', /\bAIza[0-9A-Za-z_-]{35}\b/g, (m) => mark('google', m)],
61 // A key inside an existing mark (`<redacted:kv_secret:…`) is never a key,
62 // and a value never spans into a mark's `<`/`>`.
63 ['kv_secret',
64 /(?<!redacted:)\b([A-Za-z0-9_.-]*(?:secret|token|passw|api[_-]?key|private|credential|auth)[A-Za-z0-9_.-]*)("?\s*[:=]\s*["']?)(?!<redacted:)([^\s"',;<>]{8,})/gi,
65 (m, key, sep, val) => `${key}${sep}${mark('kv_secret', val)}`],
66];
67
68/**
69 * Redact secrets in `text`.
70 * @param {string} text
71 * @returns {{ text: string, count: number }}
72 */
73export function redactText(text) {
74 let out = String(text ?? '');
75 if (!out) return { text: out, count: 0 };
76 let count = 0;
77 for (const [, re, fn] of RULES) {
78 out = out.replace(re, (...args) => {
79 count++;
80 return fn(...args);
81 });
82 }
83 return { text: out, count };
84}
85
86// A tool input that names one of these is withheld whole: its content is a
87// credential file by construction, and a path alone says enough to Jev.
88const PROTECTED_PATH = new RegExp(
89 [
90 String.raw`(?:^|[\s"'=/\\])\.env(?:\.[A-Za-z0-9_-]+)?\b`,
91 String.raw`\.pem\b`,
92 String.raw`\.key\b`,
93 String.raw`(?:^|[\s"'=/\\])id_[A-Za-z0-9]+\b`,
94 String.raw`\.ssh[/\\]`,
95 String.raw`\.aws[/\\]`,
96 String.raw`\.gnupg[/\\]`,
97 String.raw`\btoken\.json\b`,
98 String.raw`\bcredentials[^/\\\s"']*`,
99 ].join('|'),
100 'i'
101);
102
103export const WITHHELD_INPUT = '[input withheld: protected path]';
104
105/** True when a tool input string names a protected path. */
106export function namesProtectedPath(input) {
107 return PROTECTED_PATH.test(String(input ?? ''));
108}
109
110/**
111 * Redact a tool input: withheld whole for a protected path, else redactText.
112 * A withheld input counts as one redaction.
113 * @returns {{ text: string, count: number }}
114 */
115export function redactToolInput(input) {
116 if (namesProtectedPath(input)) return { text: WITHHELD_INPUT, count: 1 };
117 return redactText(input);
118}
119