Route work through the smallest sufficient loop: handle bounded tasks directly, fan out clear branches, or coordinate durable app-rooted Claude and Codex…

English | 繁體中文
You call the shots. Give boss-say one task or a backlog and it chooses the smallest sufficient loop: carry bounded work here, fan out clear branches, or coordinate app-rooted Claude Code, Codex CLI, and Antigravity workrooms when separate ownership or continuity is useful. It works in a single app out of the box and coordinates across a monorepo when needed.
Named after the ranch foreman who works the ground alongside the crew, not from an office.
Bounded work should stay bounded. When a task benefits from its own workroom, straw-boss roots that worker in the app it owns instead of copying the app's context into a summary that can drift. Claude Code workers load that app's .claude/skills/ and .claude/settings.json hooks there; Claude Code, Codex CLI, and Antigravity workers all operate from the correct app directory and local instructions. The same routing applies to implementation, audits, research, and diagnosis. Cross-app routing through resolving-app is available for monorepos, not required for a single app. Full rationale: docs/architecture.md.
boss-say — hand over the work; it selects the owner, execution tier, coordination graph, and reality anchor./loop, started by boss-say itself.boss-say; the original window leaves that scope.pi-herdr-agents.From a source checkout, install or update every supported CLI available on this machine from the GitHub marketplace source and verify the installed version:
bash scripts/install.sh
For development against this checkout instead, opt into its machine-local path explicitly:
bash scripts/install.sh --local
Restart active agent sessions afterward. The equivalent manual commands are below.
/plugin marketplace add https://github.com/wayne930242/straw-boss
/plugin install straw-boss@straw-boss
Then run once per project:
/straw-boss:init
codex plugin marketplace add wayne930242/straw-boss --ref main
codex plugin add straw-boss@straw-boss
Start a new Codex session so it loads the installed skills and hooks, review and trust the bundled hooks when prompted, then run once per project:
$straw-boss:init
You can also browse or manage the installed plugin interactively by starting codex and entering /plugins. Plugins are not available in the Codex IDE extension.
agy plugin install wayne930242/straw-boss
Then run once per project:
/straw-boss:init
init confirms the managed apps and work routes, writes .straw-boss/apps.json, syncs the root AGENTS.md and CLAUDE.md, offers to fill in each app's missing instruction files, and checks the Herdr dispatch requirement.
pi install git:github.com/wayne930242/straw-boss
Pi loads boss-say, resolving-app, choosing-graph, shipping-task, reporting-to-user, and dispatching-work, plus the dispatch_control extension. A Pi main agent keeps the same workflow and dispatches Pi workers through pi-herdr-agents; it runs none of the bundled Python scripts. Worker models come from the role models in your pi-herdr-agents configuration. Its skills are a separate Pi set under pi/skills/, so the Claude Code, Codex, and Antigravity skills stay as they are. See Pi dispatching-work.
For a single app, init is a bonus — boss-say works the moment the plugin's installed. Run it to check Herdr readiness, configure per-app options like forbidDirectCommit/localFiles, or a monorepo's apps configured.
Once init's run, hand everything to the main agent:
boss-say fix the login redirect
boss-say audit the payments module against our rules
boss-say work through docs/backlog.md
boss-say decides the rest: the owning skill, whether the current agent can carry the work or a separate workroom is useful, the coordination graph and reality anchor, one task or a batch, and /loop when a backlog needs its own pacing. It states what it picked, and you can override it in one sentence.
Call a skill by name when the situation fits; anything unlisted goes to boss-say.
| When you want to… | Use |
|---|---|
| Fix, build, audit, research, or diagnose; work through a backlog; ask what is running; close out a dispatch | boss-say |
| Set up managed apps and work routes, or check Herdr readiness | init |
| Find which app owns a request | resolving-app |
| See what a dispatch is doing without joining or interrupting it | peeking-work |
| Move a scope and its in-progress dispatches to a new orchestrator tab | handoff-orchestrator |
| Have this window take over dispatches another main agent coordinates | boss-say; it moves them only when you ask |
| Resolve friction in Straw Boss's own coordination | boss-assistant |
Give an app without AGENTS.md or CLAUDE.md a minimal agent system | create-great-harness |
| From inside a dispatched worker, bring in a coworker for review or pairing | bringing-coworker |
The main agent and workers run the other skills on their own: i-am-orchestrator, choosing-graph, dispatching-work, shipping-task, contacting-orchestrators, reporting-to-user, notifying-main-agent, asking-peer-agents, and agent-feedback. docs/architecture.md describes each one.
init writes the managed apps and each app's lifecycle configuration to .straw-boss/apps.json. The format is the apps config schema. The app summary and the project's work routes are synced into the root AGENTS.md and CLAUDE.md.
An app's agentKind names its default agent. The project's work routes name the provider profile, model, effort, and an optional Claude advisor, which init syncs into both instruction files; a Codex route records advisor: none.
Configuration is read from .straw-boss/apps.json first, falling back to .claude/straw-boss/apps.json when the new path does not exist. init writes the confirmed old configuration to the new path, leaves the old file in place, and reports that it has been superseded.
Jev pruning is opt-in and leaves ordinary context renewal unchanged. Enable it for one newly launched Claude session (with TYPESAFE_API_KEY in the environment):
STRAW_BOSS_JEV=1 CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude
Start a session with the experiment explicitly off:
STRAW_BOSS_JEV=0 claude
Restart or resume with the desired switch; keep pruning off globally during the renewal acceptance window through 2026-09-25. See benchmark, recovery, and implementation limits.
For Codex in Herdr, use STRAW_BOSS_JEV=1 codex with the same nonblank key. At the 200k renewal point, Jev attempts a pruned codex resume in the same pane; gate misses and failures use ordinary continuity renewal. See the Codex renewal behavior and accounting costs.
hooks/jev-pruning.ts 162 lines1import type { CoreEngineInterface, Register } from 'claude-code';
2import { compact } from '../vendor/fast-jev-compaction/src/compact.js';
3import { buildJevRequest, parseJevResponse } from '../vendor/fast-jev-compaction/src/request.js';
4import { toSessionMessages } from './jev/messages.js';
5
6type Engine = CoreEngineInterface;
7
8export async function active($: Engine): Promise<boolean> {
9 return (await $.env.get('STRAW_BOSS_JEV')) === '1' &&
10 !!(await $.env.get('TYPESAFE_API_KEY'))?.trim();
11}
12
13export async function host($: Engine, command: string, data: unknown): Promise<any> {
14 const result = await $.process.run(['python3', `${$.plugin.root}/scripts/jev-runtime.py`, command], {
15 stdin: JSON.stringify(data), timeoutMs: 90_000,
16 });
17 if (result.exitCode !== 0) {
18 let reason = 'host-operation-failed';
19 try { reason = JSON.parse(result.stdout).error ?? reason; } catch { /* Use operation code. */ }
20 throw new Error(`${command}:${reason}`);
21 }
22 if (!result.stdout.trim()) throw new Error(`${command}:inactive`);
23 return JSON.parse(result.stdout);
24}
25
26export const register: Register = (on) => {
27 let pendingRun: string | undefined;
28 // `$.session.model()` answers the /model label (an alias such as `opus`),
29 // which the token-count API rejects; count with the id the API reported.
30 let apiModel: string | undefined;
31
32 on('session.start', async ($, event, next) => {
33 // Clear inherited readiness so an unloaded or child session cannot opt out
34 // of the normal renewal behavior through its parent's readiness marker.
35 await $.env.set('STRAW_BOSS_JEV_READY_SESSION', undefined);
36 await $.env.set('STRAW_BOSS_JEV_WAITING_USAGE', undefined);
37 if (await active($)) {
38 try {
39 const config = await host($, 'config', { session_id: await $.session.id() });
40 pendingRun = config.pending?.run_id;
41 await $.env.set('STRAW_BOSS_JEV_READY_SESSION', await $.session.id());
42 } catch { $.ui.log('Jev unavailable; existing renewal remains active.'); }
43 }
44 return next(event);
45 });
46
47 on('session.compact', async ($, event, next) => {
48 if (!(await active($))) return next(event);
49 // Precompute installs no history. Its returned candidate may be used much
50 // later, so delegate until the engine invokes an actual compaction.
51 if (event.trigger === 'precompute' || event.agentId) return next(event);
52 const started = Date.now();
53 const session = await $.session.id();
54 const runId = `${session}-${started}-${Math.random().toString(36).slice(2, 8)}`;
55 const metrics: any[] = [];
56 const record: any = {
57 schema_version: 2, run_id: runId, session_id: session, trigger: event.trigger,
58 started_at: new Date(started).toISOString(), tokens_before: null, tokens_after: null,
59 context_window: null, criteria_version: null, decisions: [], reduction_pct: null,
60 jev: { model: null, requests: 0, input_tokens: 0, latency_ms: 0, request_metrics: metrics },
61 outcome: 'fallback', fallback_reason: null,
62 };
63 let candidate: any;
64 let apply = false;
65 try {
66 const config = await host($, 'config', {});
67 record.criteria_version = config.criteria_version;
68 record.policy = config.policy;
69 record.policy_version = config.policy_version;
70 const usage = await $.session.usage();
71 record.context_window = usage.context.window;
72 const policy = config.policy;
73 const key = (await $.env.get('TYPESAFE_API_KEY'))!;
74 const result = await compact(event.messages, {
75 async ask(state, questions) {
76 const start = Date.now();
77 record.jev.requests++;
78 const metric: any = { request: record.jev.requests, input_tokens: null, latency_ms: null };
79 metrics.push(metric);
80 try {
81 const request = buildJevRequest({ apiKey: key }, state, questions);
82 const response = await $.http.fetch(request.url, {
83 method: request.method, headers: request.headers, body: request.body,
84 });
85 // Keep server error bodies out of logs and benchmark records.
86 if (!response.ok) throw new Error(`jev-http-${response.status}`);
87 const parsed = parseJevResponse(response.status, response.ok, response.text);
88 if (!parsed.model || !Number.isFinite(parsed.usage?.input_tokens)) {
89 throw new Error('jev-missing-model-or-usage');
90 }
91 metric.model = parsed.model;
92 metric.input_tokens = parsed.usage!.input_tokens!;
93 record.jev.model = parsed.model;
94 record.jev.input_tokens += metric.input_tokens;
95 return parsed;
96 } finally {
97 metric.latency_ms = Date.now() - start;
98 record.jev.latency_ms += metric.latency_ms;
99 }
100 },
101 }, {
102 criteria: config.criteria, protectedInputPatterns: policy.protected_input_patterns, keepThreshold: policy.keep_threshold,
103 preserveRecentMessages: policy.preserve_recent_messages,
104 maxStateTokens: policy.max_state_tokens, maxRequestTokens: policy.max_request_tokens,
105 truncateHeadChars: policy.truncate_head_chars,
106 });
107 candidate = toSessionMessages(event.messages, result.messages);
108 record.decisions = result.decisions;
109 record.state = { estimated_tokens: result.stats.stateTokens, stage: result.stats.stateStage };
110 if (result.decisions.every((d) => d.action === 'keep')) {
111 record.fallback_reason = 'no-removable-records';
112 } else {
113 Object.assign(record, await host($, 'measure', {
114 model: apiModel ?? await $.session.model(), before: event.messages, after: candidate,
115 live_input_tokens: usage.context.tokens,
116 }));
117 apply = record.gate_reduction_pct >= policy.min_reduction_ratio * 100;
118 record.fallback_reason = apply ? null : 'under-reduction-gate';
119 }
120 } catch (error) {
121 record.fallback_reason = error instanceof Error ? error.message : 'jev-failure';
122 }
123 record.decision = apply ? 'apply' : 'fallback';
124 record.outcome = apply ? null : 'fallback';
125 record.application = { status: 'awaiting-backend-usage' };
126 record.elapsed_ms = Date.now() - started;
127 try {
128 await host($, 'persist', { record, original_messages: event.messages, candidate_messages: candidate });
129 } catch {
130 $.ui.log('Jev recovery record unavailable; using built-in compaction.');
131 return next(event);
132 }
133 pendingRun = runId;
134 // A manual compaction can end before another backend request. Its old
135 // transcript usage cannot trigger renewal while the new window is unmeasured.
136 await $.env.set('STRAW_BOSS_JEV_WAITING_USAGE', session);
137 if (apply) {
138 $.ui.log(`Jev candidate selected: ${record.gate_reduction_pct.toFixed(1)}% backend-count reduction; recovery ${runId}.`);
139 return { messages: candidate };
140 }
141 $.ui.log(`Jev fallback: ${record.fallback_reason}.`);
142 return next(event);
143 });
144
145 on('turn.complete', async ($, event, next) => {
146 if (!event.agentId && event.usage?.model) apiModel = event.usage.model;
147 if (await active($)) {
148 const usage = await $.session.usage();
149 if (usage.context.tokens !== undefined) {
150 await $.env.set('STRAW_BOSS_JEV_WAITING_USAGE', undefined);
151 if (pendingRun) {
152 await host($, 'observe', { session_id: await $.session.id(), run_id: pendingRun,
153 actual_session_tokens_after: usage.context.tokens, context_window: usage.context.window,
154 source: 'first-observed-post-compaction-backend-input', observed_at: new Date().toISOString() });
155 pendingRun = undefined;
156 }
157 }
158 }
159 return next(event);
160 });
161};
162vendor/fast-jev-compaction/src/compact.ts 325 lines1import { noulAnswer } from './request.js';
2import { collectToolCalls, estimateTokens, fitState } from './state.js';
3import type {
4 CallAnswer,
5 CallDecision,
6 CompactOptions,
7 CompactResult,
8 CompactionState,
9 JevAsker,
10 JevQuestions,
11 Message,
12 ResolvedCompactOptions,
13 ToolCall,
14 ToolUse,
15 RetentionCriteria,
16} from './types.js';
17
18export const DEFAULT_OPTIONS: ResolvedCompactOptions = {
19 goal: '',
20 keepThreshold: 0.5,
21 preserveRecentMessages: 6,
22 maxStateTokens: 25_000,
23 maxRequestTokens: 30_000,
24 truncateHeadChars: 300,
25};
26
27/** Tokens the request envelope (`model`, key names) adds around state and questions. */
28const REQUEST_OVERHEAD_TOKENS = 20;
29
30function finite(value: number | undefined, fallback: number): number {
31 return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
32}
33
34export function resolveOptions(options: CompactOptions = {}): ResolvedCompactOptions {
35 return {
36 goal: options.goal ?? DEFAULT_OPTIONS.goal,
37 ...(options.criteria ? { criteria: options.criteria } : {}),
38 ...(options.protectedInputPatterns ? { protectedInputPatterns: options.protectedInputPatterns } : {}),
39 keepThreshold: finite(options.keepThreshold, DEFAULT_OPTIONS.keepThreshold),
40 preserveRecentMessages: Math.max(
41 0,
42 Math.floor(
43 finite(options.preserveRecentMessages, DEFAULT_OPTIONS.preserveRecentMessages),
44 ),
45 ),
46 maxStateTokens: Math.max(1, finite(options.maxStateTokens, DEFAULT_OPTIONS.maxStateTokens)),
47 maxRequestTokens: Math.max(
48 1,
49 finite(options.maxRequestTokens, DEFAULT_OPTIONS.maxRequestTokens),
50 ),
51 truncateHeadChars: Math.max(
52 0,
53 Math.floor(finite(options.truncateHeadChars, DEFAULT_OPTIONS.truncateHeadChars)),
54 ),
55 };
56}
57
58/** The two `noul` questions asked about one call: keep the call, keep its result. */
59export function questionsFor(call: ToolCall, criteria?: RetentionCriteria): JevQuestions {
60 if (criteria) {
61 return Object.fromEntries((["keepCall", "keepResult"] as const).map((kind) => [
62 `${kind === "keepCall" ? "call" : "result"}_${call.id}`,
63 { type: "noul" as const, instructions: `Target: tool call ${call.id} (${call.tool}, ${call.resultChars} result chars). ${criteria[kind]}`,
64 criteria: { true: criteria[kind], false: "The proposition is false; this material can be pruned according to the stated retention contract." } },
65 ]));
66 }
67 return {
68 [`call_${call.id}`]: {
69 type: 'noul',
70 instructions: `Tool call ${call.id} (${call.tool}) should stay in the history: knowing this call was made, with its input, still matters for what the assistant does next`,
71 },
72 [`result_${call.id}`]: {
73 type: 'noul',
74 instructions: `The full output of tool call ${call.id} (${call.tool}, ${call.resultChars} chars) should stay in the history verbatim: the assistant still needs its contents and re-running the tool would not do`,
75 },
76 };
77}
78
79/**
80 * Splits the candidate calls into batches whose questions, together with the
81 * (always complete) state, fit one request.
82 */
83export function batchCalls(
84 calls: readonly ToolCall[],
85 stateTokens: number,
86 options: Pick<ResolvedCompactOptions, 'maxRequestTokens' | 'criteria'>,
87): ToolCall[][] {
88 const budget = options.maxRequestTokens - stateTokens - REQUEST_OVERHEAD_TOKENS;
89 const batches: ToolCall[][] = [];
90 let current: ToolCall[] = [];
91 let currentTokens = 0;
92 for (const call of calls) {
93 const tokens = estimateTokens(JSON.stringify(questionsFor(call, options.criteria)));
94 if (current.length > 0 && currentTokens + tokens > budget) {
95 batches.push(current);
96 current = [];
97 currentTokens = 0;
98 }
99 if (current.length === 0 && tokens > budget) {
100 throw new Error(
101 `state leaves no room for questions (~${stateTokens} of ${options.maxRequestTokens} tokens)`,
102 );
103 }
104 current.push(call);
105 currentTokens += tokens;
106 }
107 if (current.length > 0) batches.push(current);
108 return batches;
109}
110
111export function decideCall(
112 call: Pick<ToolCall, 'id' | 'tool' | 'pinned' | 'governingSource'>,
113 answer: CallAnswer,
114 options: Pick<ResolvedCompactOptions, 'keepThreshold'>,
115): CallDecision {
116 const base = { id: call.id, tool: call.tool, ...answer };
117 if (call.pinned) return { ...base, action: 'keep', reason: call.governingSource ? 'governing_source' : 'pinned' };
118 if (answer.keepResult >= options.keepThreshold) {
119 return { ...base, action: 'keep', reason: 'kept' };
120 }
121 if (answer.keepCall >= options.keepThreshold) {
122 return { ...base, action: 'drop_result', reason: 'result_dropped' };
123 }
124 return { ...base, action: 'drop_call', reason: 'call_dropped' };
125}
126
127async function askBatch(
128 asker: JevAsker,
129 state: CompactionState,
130 batch: readonly ToolCall[],
131 criteria?: RetentionCriteria,
132): Promise<Map<string, CallAnswer>> {
133 const questions: JevQuestions = Object.assign({}, ...batch.map((call) => questionsFor(call, criteria)));
134 const { answers } = await asker.ask(state, questions);
135 return new Map(
136 batch.map((call) => [
137 call.id,
138 {
139 keepCall: noulAnswer(answers, `call_${call.id}`),
140 keepResult: noulAnswer(answers, `result_${call.id}`),
141 },
142 ]),
143 );
144}
145
146function truncatedResultText(text: string, isError: boolean, headChars: number): string {
147 if (text.length <= headChars) return text;
148 const head = headChars > 0 ? `${text.slice(0, headChars)}\n` : '';
149 return `${head}[fast-jev-compaction truncated ${text.length - headChars} chars of this tool result${
150 isError ? ' (error)' : ''
151 }; original saved in the recovery record]`;
152}
153
154/**
155 * Rebuilds the conversation from the decisions. A dropped call disappears
156 * together with its result; a dropped result keeps a bounded head and note.
157 * Messages that lose all their content are removed; untouched messages are
158 * returned as the same objects they came in as.
159 */
160export function applyDecisions(
161 messages: readonly Message[],
162 decisions: readonly CallDecision[],
163 calls: readonly ToolCall[],
164 headChars: number,
165): Message[] {
166 const byId = new Map(calls.map((call) => [call.id, call]));
167 const actions = new Map<string, CallDecision['action']>();
168 for (const decision of decisions) {
169 const call = byId.get(decision.id);
170 if (call && decision.action !== 'keep') actions.set(call.tool_use_id, decision.action);
171 }
172 const kept: Message[] = [];
173 for (const message of messages) {
174 const touched =
175 message.toolUses.some((tool) => actions.has(tool.tool_use_id)) ||
176 (message.toolResults ?? []).some((result) => actions.has(result.tool_use_id));
177 if (!touched) {
178 kept.push(message);
179 continue;
180 }
181 const toolUses = message.toolUses
182 .filter((tool) => actions.get(tool.tool_use_id) !== 'drop_call')
183 .map((tool) => {
184 if (actions.get(tool.tool_use_id) !== 'drop_result') return tool;
185 const text = truncatedResultText(
186 tool.text ?? '',
187 tool.isError ?? false,
188 headChars,
189 );
190 if ((tool.text ?? '') === text) return tool;
191 const copy: ToolUse = {
192 tool_use_id: tool.tool_use_id,
193 tool: tool.tool,
194 input: tool.input,
195 text,
196 };
197 if (tool.isError) copy.isError = true;
198 return copy;
199 });
200 const toolResults = (message.toolResults ?? [])
201 .filter((result) => actions.get(result.tool_use_id) !== 'drop_call')
202 .map((result) => {
203 if (actions.get(result.tool_use_id) !== 'drop_result') return result;
204 const text = truncatedResultText(result.text, result.isError ?? false, headChars);
205 return text === result.text
206 ? result
207 : {
208 tool_use_id: result.tool_use_id,
209 text,
210 isError: result.isError,
211 };
212 });
213 if (
214 !message.toolUses.some(
215 (tool) => actions.get(tool.tool_use_id) === 'drop_call',
216 ) &&
217 !(message.toolResults ?? []).some(
218 (result) => actions.get(result.tool_use_id) === 'drop_call',
219 ) &&
220 toolUses.every((tool, index) => tool === message.toolUses[index]) &&
221 toolResults.every(
222 (result, index) => result === message.toolResults?.[index],
223 )
224 ) {
225 kept.push(message);
226 continue;
227 }
228 if (message.text.length === 0 && toolUses.length === 0 && toolResults.length === 0) {
229 continue;
230 }
231 const rebuilt: Message = { role: message.role, text: message.text, toolUses };
232 if (toolResults.length > 0) rebuilt.toolResults = toolResults;
233 kept.push(rebuilt);
234 }
235 return kept;
236}
237
238/** Characters of text, tool input and tool output a message holds. */
239export function messageChars(message: Message): number {
240 let total = message.text.length;
241 for (const tool of message.toolUses) {
242 try {
243 total += JSON.stringify(tool.input).length;
244 } catch {
245 total += 20;
246 }
247 }
248 for (const result of message.toolResults ?? []) total += result.text.length;
249 return total;
250}
251
252export function reductionRatio(result: Pick<CompactResult, 'stats'>): number {
253 const { charsBefore, charsAfter } = result.stats;
254 return charsBefore === 0 ? 0 : (charsBefore - charsAfter) / charsBefore;
255}
256
257function count(decisions: readonly CallDecision[], reason: CallDecision['reason']): number {
258 return decisions.filter((decision) => decision.reason === reason).length;
259}
260
261/**
262 * Compacts a transcript by asking Jev, for every tool call outside the pinned
263 * first and newest messages, whether the call and whether its result must
264 * stay. The whole history (results omitted, fitted into `maxStateTokens`) is
265 * sent as state with every batch of questions. Throws when Jev fails or the
266 * history cannot be fitted; the caller decides whether to fall back.
267 */
268export async function compact(
269 messages: readonly Message[],
270 asker: JevAsker,
271 options: CompactOptions = {},
272): Promise<CompactResult> {
273 const started = Date.now();
274 const resolved = resolveOptions(options);
275 const calls = collectToolCalls(messages, resolved.preserveRecentMessages, resolved.protectedInputPatterns);
276 const candidates = calls.filter((call) => !call.pinned);
277 const charsBefore = messages.reduce((sum, message) => sum + messageChars(message), 0);
278
279 let fitted: { tokens: number; stage: string } = { tokens: 0, stage: '' };
280 let batches: ToolCall[][] = [];
281 const answers = new Map<string, CallAnswer>();
282 if (candidates.length > 0) {
283 const state = fitState(messages, calls, resolved);
284 fitted = state;
285 batches = batchCalls(candidates, state.tokens, resolved);
286 const settled = await Promise.allSettled(
287 batches.map((batch) => askBatch(asker, state.state, batch, resolved.criteria)),
288 );
289 const failed = settled.find((item) => item.status === "rejected");
290 if (failed?.status === "rejected") throw failed.reason;
291 for (const item of settled) if (item.status === "fulfilled") {
292 for (const [id, answer] of item.value) answers.set(id, answer);
293 }
294 }
295
296 const decisions = calls.map((call) =>
297 decideCall(call, answers.get(call.id) ?? { keepCall: 1, keepResult: 1 }, resolved),
298 );
299 const kept = applyDecisions(
300 messages,
301 decisions,
302 calls,
303 resolved.truncateHeadChars,
304 );
305 return {
306 messages: kept,
307 decisions,
308 stats: {
309 messagesBefore: messages.length,
310 messagesAfter: kept.length,
311 charsBefore,
312 charsAfter: kept.reduce((sum, message) => sum + messageChars(message), 0),
313 calls: calls.length,
314 kept: count(decisions, 'kept'),
315 resultsDropped: count(decisions, 'result_dropped'),
316 callsDropped: count(decisions, 'call_dropped'),
317 pinned: count(decisions, 'pinned') + count(decisions, 'governing_source'),
318 stateTokens: fitted.tokens,
319 stateStage: fitted.stage,
320 requests: batches.length,
321 ms: Date.now() - started,
322 },
323 };
324}
325vendor/fast-jev-compaction/src/request.ts 82 lines1import type { JevAnswer, JevQuestions, JevResponse, JevState } from './types.js';
2
3export const SYSTEM_ONE_URL = 'https://api.typesafe.ai/v1/systemone';
4export const DEFAULT_MODEL = 'jev-latest';
5
6export interface JevRequest {
7 url: string;
8 method: 'POST';
9 headers: Record<string, string>;
10 body: string;
11}
12
13/** The HTTP request for one Jev call, for any fetch-like transport. */
14export function buildJevRequest(
15 params: {
16 apiKey: string;
17 model?: string;
18 baseUrl?: string;
19 },
20 state: JevState,
21 questions: JevQuestions,
22): JevRequest {
23 return {
24 url: params.baseUrl ?? SYSTEM_ONE_URL,
25 method: 'POST',
26 headers: {
27 authorization: `Bearer ${params.apiKey}`,
28 'content-type': 'application/json',
29 },
30 body: JSON.stringify({
31 model: params.model ?? DEFAULT_MODEL,
32 state,
33 questions,
34 }),
35 };
36}
37
38/** Validates a Jev response body; throws on anything but an `answers` object. */
39export function parseJevResponse(
40 status: number,
41 ok: boolean,
42 text: string,
43): JevResponse {
44 if (!ok) {
45 throw new Error(`Jev request failed (${status}): ${text.slice(0, 200)}`);
46 }
47 let parsed: unknown;
48 try {
49 parsed = JSON.parse(text);
50 } catch {
51 throw new Error('Jev returned malformed JSON');
52 }
53 if (
54 parsed === null ||
55 typeof parsed !== 'object' ||
56 !('answers' in parsed) ||
57 parsed.answers === null ||
58 typeof parsed.answers !== 'object'
59 ) {
60 throw new Error('Jev response is missing answers');
61 }
62 return parsed as JevResponse;
63}
64
65/** The `noul` probability of one answer; throws when it is not there. */
66export function noulAnswer(
67 answers: Record<string, JevAnswer>,
68 name: string,
69): number {
70 const answer = answers[name];
71 if (
72 !answer ||
73 !('noul' in answer) ||
74 typeof answer.noul !== 'number' ||
75 !Number.isFinite(answer.noul) ||
76 answer.noul < 0 || answer.noul > 1
77 ) {
78 throw new Error(`Invalid Jev answer for ${name}`);
79 }
80 return answer.noul;
81}
82hooks/jev/messages.ts 60 lines1import type { SessionMessage, ToolUseSummary, ToolResultSummary } from 'claude-code';
2import type { Message, ToolUse, ToolResult } from '../../vendor/fast-jev-compaction/src/types.js';
3
4function toolUseSummary(tool: ToolUse): ToolUseSummary {
5 const summary: ToolUseSummary = {
6 tool_use_id: tool.tool_use_id,
7 tool: tool.tool,
8 input: tool.input,
9 };
10 if (tool.text !== undefined) summary.text = tool.text;
11 if (tool.isError) summary.isError = true;
12 return summary;
13}
14
15function toolResultSummary(result: ToolResult): ToolResultSummary {
16 return {
17 tool_use_id: result.tool_use_id,
18 text: result.text,
19 isError: result.isError ?? false,
20 };
21}
22
23/**
24 * Maps the library's output back onto session messages. Whatever came back
25 * Every returned message is fresh so Claude persists a new parent chain across
26 * restart. Reusing message handles lets 2.1.278 reconnect dropped ancestors
27 * when resuming. Text and unmodified tool objects remain verbatim.
28 */
29export function toSessionMessages(
30 input: readonly SessionMessage[],
31 output: readonly Message[],
32): SessionMessage[] {
33 const messages = new Map<Message, SessionMessage>();
34 const uses = new Map<ToolUse, ToolUseSummary>();
35 const results = new Map<ToolResult, ToolResultSummary>();
36 for (const message of input) {
37 messages.set(message, message);
38 for (const tool of message.toolUses) uses.set(tool, tool);
39 for (const result of message.toolResults ?? []) results.set(result, result);
40 }
41 return output.map((message) => {
42 const own = messages.get(message);
43 if (own) {
44 const { handle: _handle, ...fresh } = own;
45 return fresh;
46 }
47 const rebuilt: SessionMessage = {
48 role: message.role,
49 text: message.text,
50 toolUses: message.toolUses.map((tool) => uses.get(tool) ?? toolUseSummary(tool)),
51 };
52 if (message.toolResults && message.toolResults.length > 0) {
53 rebuilt.toolResults = message.toolResults.map(
54 (result) => results.get(result) ?? toolResultSummary(result),
55 );
56 }
57 return rebuilt;
58 });
59}
60vendor/fast-jev-compaction/src/state.ts 308 lines1import type {
2 CompactionState,
3 FittedState,
4 HistoryEntry,
5 Message,
6 ResolvedCompactOptions,
7 ToolCall,
8 ToolResult,
9} from './types.js';
10
11export const STATE_CONTEXT =
12 'A coding assistant history is being compacted. `history` runs oldest first, with tool results omitted and long text abridged. Score the named call and its full result for the unfinished task. Keep governing instructions and required evidence; archived originals are recoverable, but some tool actions cannot be reproduced.';
13
14/** Successive caps on the serialised tool input included per call. */
15const INPUT_CHARS = [1000, 200, 60] as const;
16const TEXT_HEAD = 400;
17const TEXT_TAIL = 150;
18
19const TOKEN_PIECES = /[A-Za-z]+|\d+|[^\sA-Za-z\d]/g;
20
21/**
22 * Estimates tokens without a tokenizer: a word costs one token per six
23 * letters, a digit half a token, any other symbol nine tenths. Calibrated
24 * against the usage Jev reports for real transcripts, where it lands 2–18%
25 * above the true count; a plain characters-per-token ratio undercounts the
26 * JSON-heavy states by up to 40%.
27 */
28export function estimateTokens(text: string): number {
29 let tokens = 0;
30 for (const [piece] of text.matchAll(TOKEN_PIECES)) {
31 const first = piece.charCodeAt(0);
32 if (first >= 48 && first <= 57) tokens += piece.length / 2;
33 else if ((first >= 65 && first <= 90) || (first >= 97 && first <= 122)) {
34 tokens += 1 + Math.floor((piece.length - 1) / 6);
35 } else tokens += 0.9;
36 }
37 return Math.ceil(tokens);
38}
39
40export function truncate(text: string, limit: number): string {
41 return text.length <= limit ? text : `${text.slice(0, Math.max(0, limit - 1))}…`;
42}
43
44function abridge(text: string, head: number, tail: number): string {
45 if (text.length <= head + tail + 40) return text;
46 const omitted = text.length - head - tail;
47 return `${text.slice(0, head)}\n[… ${omitted} chars omitted …]\n${text.slice(-tail)}`;
48}
49
50export function isPinned(
51 index: number,
52 total: number,
53 preserveRecentMessages: number,
54): boolean {
55 return index === 0 || index >= total - preserveRecentMessages;
56}
57
58/**
59 * Pairs every tool_use with its tool_result by `tool_use_id`. Calls without a
60 * result are not candidates (there is nothing to drop yet).
61 */
62export function collectToolCalls(
63 messages: readonly Message[],
64 preserveRecentMessages: number,
65 protectedInputPatterns: string[] = [],
66): ToolCall[] {
67 const results = new Map<string, { index: number; result: ToolResult }>();
68 messages.forEach((message, index) => {
69 for (const result of message.toolResults ?? []) {
70 results.set(result.tool_use_id, { index, result });
71 }
72 });
73 const calls: ToolCall[] = [];
74 messages.forEach((message, callIndex) => {
75 for (const tool of message.toolUses) {
76 const found = results.get(tool.tool_use_id);
77 if (!found) continue;
78 const governingSource = protectedInputPatterns.some((pattern) => new RegExp(pattern).test(JSON.stringify(tool.input)));
79 calls.push({
80 id: `t${calls.length + 1}`,
81 tool_use_id: tool.tool_use_id,
82 tool: tool.tool,
83 input: tool.input,
84 callIndex,
85 resultIndex: found.index,
86 resultChars: found.result.text.length,
87 isError: found.result.isError ?? false,
88 ...(governingSource ? { governingSource: true } : {}),
89 pinned:
90 governingSource || isPinned(callIndex, messages.length, preserveRecentMessages) ||
91 isPinned(found.index, messages.length, preserveRecentMessages),
92 });
93 }
94 });
95 return calls;
96}
97
98function inputText(input: Record<string, unknown>, limit: number): string {
99 let json = '';
100 try {
101 json = JSON.stringify(input);
102 } catch {
103 json = '[unserializable input]';
104 }
105 return truncate(json, limit);
106}
107
108function resultNote(call: ToolCall): string {
109 return `${call.isError ? 'error' : 'ok'}, ${call.resultChars} chars (omitted)`;
110}
111
112/** One call as a single line, for when the structured form is too costly. */
113function compactCall(call: ToolCall): string {
114 const input = Object.entries(call.input)
115 .map(([key, value]) => {
116 const text = typeof value === 'string' ? value : inputText({ [key]: value }, 200);
117 return `${key}=${text.replace(/\s+/g, ' ')}`;
118 })
119 .join(' ');
120 return `${call.id} ${call.tool} ${truncate(input, INPUT_CHARS[2])} → ${
121 call.isError ? 'error' : 'ok'
122 } ${call.resultChars}ch`;
123}
124
125/**
126 * Folds runs of adjacent call-only entries into one entry each, so the
127 * per-entry envelope is paid once per run; the call lines keep their ids.
128 */
129function mergeCallRuns(history: readonly HistoryEntry[], pinned: (e: HistoryEntry) => boolean): HistoryEntry[] {
130 const merged: HistoryEntry[] = [];
131 for (const entry of history) {
132 const previous = merged[merged.length - 1];
133 const foldable = (e: HistoryEntry): boolean =>
134 !pinned(e) && e.text.length === 0 && typeof e.tool_calls?.[0] === 'string';
135 if (previous && foldable(previous) && foldable(entry) && previous.role === entry.role) {
136 previous.tool_calls = [...(previous.tool_calls as string[]), ...(entry.tool_calls as string[])];
137 continue;
138 }
139 merged.push({ ...entry });
140 }
141 return merged;
142}
143
144function callsByMessage(calls: readonly ToolCall[]): Map<number, ToolCall[]> {
145 const byMessage = new Map<number, ToolCall[]>();
146 for (const call of calls) {
147 const list = byMessage.get(call.callIndex) ?? [];
148 list.push(call);
149 byMessage.set(call.callIndex, list);
150 }
151 return byMessage;
152}
153
154function historyEntries(
155 messages: readonly Message[],
156 calls: readonly ToolCall[],
157 inputChars: number,
158): HistoryEntry[] {
159 const byMessage = callsByMessage(calls);
160 const entries: HistoryEntry[] = [];
161 messages.forEach((message, i) => {
162 const toolCalls = (byMessage.get(i) ?? []).map((call) => ({
163 id: call.id,
164 tool: call.tool,
165 input: inputText(call.input, inputChars),
166 result: resultNote(call),
167 }));
168 if (message.text.trim().length === 0 && toolCalls.length === 0) return;
169 const entry: HistoryEntry = { i, role: message.role, text: message.text };
170 if (toolCalls.length > 0) entry.tool_calls = toolCalls;
171 entries.push(entry);
172 });
173 return entries;
174}
175
176/** The last three user prompts, as the default `goal`. */
177export function goalFromMessages(messages: readonly Message[]): string {
178 return messages
179 .filter(
180 (message) =>
181 message.role === 'user' &&
182 message.text.trim().length > 0 &&
183 (message.toolResults ?? []).length === 0,
184 )
185 .slice(-3)
186 .map((message) => truncate(message.text, 500))
187 .join('\n');
188}
189
190/**
191 * Builds the Jev state from the whole conversation and shrinks it in stages
192 * until it fits `maxStateTokens`: tool inputs are truncated, then long texts
193 * are abridged oldest-first (pinned messages last), then old messages collapse
194 * to a one-line note, then old tool calls shrink to one line each, then old
195 * messages that carry no call are left out, then runs of old call-only
196 * messages are folded into one entry. Throws when even that is too big.
197 */
198export function fitState(
199 messages: readonly Message[],
200 calls: readonly ToolCall[],
201 options: Pick<ResolvedCompactOptions, 'maxStateTokens' | 'preserveRecentMessages' | 'goal'>,
202): FittedState {
203 const goal = options.goal || goalFromMessages(messages);
204 const stateOf = (history: HistoryEntry[]): CompactionState => ({
205 context: STATE_CONTEXT,
206 goal,
207 history,
208 });
209 const entryTokens = (entry: HistoryEntry): number => estimateTokens(JSON.stringify(entry)) + 1;
210 const baseTokens = estimateTokens(JSON.stringify(stateOf([])));
211 const fitted = (history: HistoryEntry[], tokens: number, stage: string): FittedState => ({
212 state: stateOf(history),
213 tokens,
214 stage,
215 });
216
217 let history: HistoryEntry[] = [];
218 let perEntry: number[] = [];
219 let tokens = 0;
220 const rebuild = (inputChars: number): void => {
221 history = historyEntries(messages, calls, inputChars);
222 perEntry = history.map(entryTokens);
223 tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
224 };
225 const fits = (): boolean => tokens <= options.maxStateTokens;
226 const shrink = (index: number, change: (entry: HistoryEntry) => void): void => {
227 const entry = history[index];
228 if (!entry) return;
229 change(entry);
230 const now = entryTokens(entry);
231 tokens += now - (perEntry[index] ?? 0);
232 perEntry[index] = now;
233 };
234
235 rebuild(INPUT_CHARS[0]);
236 if (fits()) return fitted(history, tokens, 'full');
237
238 for (const limit of INPUT_CHARS.slice(1)) {
239 rebuild(limit);
240 if (fits()) return fitted(history, tokens, `inputs<=${limit}`);
241 }
242
243 const pinned = (entry: HistoryEntry): boolean =>
244 isPinned(entry.i, messages.length, options.preserveRecentMessages);
245 const indices = history.map((_, index) => index);
246 const order = [
247 ...indices.filter((index) => !pinned(history[index]!)),
248 ...indices.filter((index) => pinned(history[index]!)),
249 ];
250
251 for (const index of order) {
252 const entry = history[index]!;
253 if (entry.text.length <= TEXT_HEAD + TEXT_TAIL + 40) continue;
254 shrink(index, (e) => {
255 e.text = abridge(e.text, TEXT_HEAD, TEXT_TAIL);
256 });
257 if (fits()) return fitted(history, tokens, 'texts abridged');
258 }
259
260 for (const index of order) {
261 const entry = history[index]!;
262 if (pinned(entry) || entry.text.length === 0) continue;
263 const original = messages[entry.i]?.text.length ?? entry.text.length;
264 shrink(index, (e) => {
265 e.text = `[… ${original} chars omitted …]`;
266 });
267 if (fits()) return fitted(history, tokens, 'old messages collapsed');
268 }
269
270 const byMessage = callsByMessage(calls);
271 for (const index of order) {
272 const entry = history[index]!;
273 const own = byMessage.get(entry.i);
274 if (pinned(entry) || !own) continue;
275 shrink(index, (e) => {
276 e.tool_calls = own.map(compactCall);
277 });
278 if (fits()) return fitted(history, tokens, 'old calls compacted');
279 }
280
281 const left = new Set<number>();
282 for (const index of order) {
283 const entry = history[index]!;
284 if (pinned(entry) || entry.tool_calls) continue;
285 left.add(index);
286 tokens -= perEntry[index] ?? 0;
287 if (fits()) {
288 return fitted(
289 history.filter((_, i) => !left.has(i)),
290 tokens,
291 'old messages left out',
292 );
293 }
294 }
295
296 history = mergeCallRuns(
297 history.filter((_, i) => !left.has(i)),
298 pinned,
299 );
300 perEntry = history.map(entryTokens);
301 tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
302 if (fits()) return fitted(history, tokens, 'old calls merged');
303
304 throw new Error(
305 `history too large for Jev (~${tokens} tokens after truncation, limit ${options.maxStateTokens})`,
306 );
307}
308vendor/fast-jev-compaction/src/types.ts 213 lines1export type Role = 'user' | 'assistant';
2
3/**
4 * A tool_use block of an assistant message. `text` and `isError` mirror the
5 * outcome once the transcript holds it (Claude Code attaches them).
6 */
7export interface ToolUse {
8 tool_use_id: string;
9 tool: string;
10 input: Record<string, unknown>;
11 text?: string;
12 isError?: boolean;
13}
14
15/** A tool_result block of a user message. */
16export interface ToolResult {
17 tool_use_id: string;
18 text: string;
19 isError?: boolean;
20}
21
22/**
23 * One transcript message. The shape is a subset of Claude Code's
24 * `SessionMessage`, so a session transcript can be passed in as is.
25 */
26export interface Message {
27 role: Role;
28 text: string;
29 toolUses: ToolUse[];
30 toolResults?: ToolResult[];
31}
32
33/** A tool call paired with its result by `tool_use_id`. */
34export interface ToolCall {
35 /** Short id used in the Jev state and question names (`t1`, `t2`, ...). */
36 id: string;
37 tool_use_id: string;
38 tool: string;
39 input: Record<string, unknown>;
40 /** Index of the message holding the tool_use block. */
41 callIndex: number;
42 /** Index of the message holding the tool_result block. */
43 resultIndex: number;
44 resultChars: number;
45 isError: boolean;
46 /** In the first or the newest preserved messages; never a candidate. */
47 pinned: boolean;
48 governingSource?: boolean;
49}
50
51export interface CallAnswer {
52 /** Jev's probability that the call itself still matters. */
53 keepCall: number;
54 /** Jev's probability that the full result still needs to stay verbatim. */
55 keepResult: number;
56}
57
58export type CallAction = 'keep' | 'drop_result' | 'drop_call';
59
60export interface CallDecision extends CallAnswer {
61 id: string;
62 tool: string;
63 action: CallAction;
64 reason: 'pinned' | 'governing_source' | 'kept' | 'result_dropped' | 'call_dropped';
65}
66
67export interface HistoryToolCall {
68 id: string;
69 tool: string;
70 input: string;
71 result: string;
72}
73
74export interface HistoryEntry {
75 i: number;
76 role: Role;
77 text: string;
78 /** Structured per call, or one compact line per call once the state has to shrink. */
79 tool_calls?: HistoryToolCall[] | string[];
80}
81
82/** The state sent with every Jev request: the whole history, results omitted. */
83export interface CompactionState {
84 context: string;
85 goal: string;
86 history: HistoryEntry[];
87}
88
89export interface FittedState {
90 state: CompactionState;
91 tokens: number;
92 /** Which fitting stage produced the state, for diagnostics. */
93 stage: string;
94}
95
96export interface RetentionCriteria {
97 keepCall: string;
98 keepResult: string;
99}
100
101export interface CompactOptions {
102 criteria?: RetentionCriteria;
103 protectedInputPatterns?: string[];
104 /** Ongoing task description; defaults to the last few user prompts. */
105 goal?: string;
106 /** Minimum keep probability for a call or result to stay. Default 0.5. */
107 keepThreshold?: number;
108 /** Newest messages never touched (the first message is always kept). Default 6. */
109 preserveRecentMessages?: number;
110 /** Estimated token ceiling for the state. Default 25000. */
111 maxStateTokens?: number;
112 /** Estimated token ceiling for state plus one batch of questions. Default 30000. */
113 maxRequestTokens?: number;
114 /** Characters of a dropped tool result to retain. Default 300. */
115 truncateHeadChars?: number;
116}
117
118export interface ResolvedCompactOptions {
119 criteria?: RetentionCriteria;
120 protectedInputPatterns?: string[];
121 goal: string;
122 keepThreshold: number;
123 preserveRecentMessages: number;
124 maxStateTokens: number;
125 maxRequestTokens: number;
126 truncateHeadChars: number;
127}
128
129export interface CompactResult {
130 /** The compacted transcript; untouched messages are the input objects. */
131 messages: Message[];
132 decisions: CallDecision[];
133 stats: {
134 messagesBefore: number;
135 messagesAfter: number;
136 charsBefore: number;
137 charsAfter: number;
138 calls: number;
139 kept: number;
140 resultsDropped: number;
141 callsDropped: number;
142 pinned: number;
143 stateTokens: number;
144 /** Which fitting stage the state needed, '' when no request was made. */
145 stateStage: string;
146 requests: number;
147 ms: number;
148 };
149}
150
151/** The `state` of a Jev request: a string or any JSON-serialisable object. */
152export type JevState = string | object;
153
154export interface NoulQuestion {
155 type: 'noul';
156 instructions: string;
157 criteria?: {
158 true?: string;
159 false?: string;
160 };
161}
162
163export interface ChoiceQuestion {
164 type: 'choice';
165 instructions: string;
166 criteria: Record<string, string | null>;
167}
168
169export interface ScoreQuestion {
170 type: 'score';
171 instructions: string;
172 criteria: string[];
173}
174
175export type JevQuestion = NoulQuestion | ChoiceQuestion | ScoreQuestion;
176export type JevQuestions = Record<string, JevQuestion>;
177
178export interface NoulAnswer {
179 type?: 'noul';
180 noul: number;
181}
182
183export interface ChoiceAnswer {
184 type?: 'choice';
185 choice: string;
186 confidence: number;
187 probabilities: Record<string, number>;
188}
189
190export interface ScoreAnswer {
191 type?: 'score';
192 score: number;
193 confidence: number;
194 probabilities: Record<string, number>;
195}
196
197export type JevAnswer = NoulAnswer | ChoiceAnswer | ScoreAnswer;
198
199export interface JevResponse {
200 model?: string;
201 answers: Record<string, JevAnswer>;
202 usage?: {
203 input_tokens?: number;
204 output_tokens?: number;
205 };
206 [key: string]: unknown;
207}
208
209/** Anything that can answer Jev questions: `JevClient`, or a host-provided adapter. */
210export interface JevAsker {
211 ask(state: JevState, questions: JevQuestions): Promise<JevResponse>;
212}
213