SLOPSHOPPER

straw-boss

Route work through the smallest sufficient loop: handle bounded tasks directly, fan out clear branches, or coordinate durable app-rooted Claude and Codex…

newprocessnetwork
★ 4v0.31.7MITupdated 2026-10-05wayne930242/straw-boss
A shopper browsing a rack in a slop shop
README

straw-boss

English | 繁體中文

You call the shots. Give boss-say one task or a backlog and it chooses the smallest sufficient loop: carry bounded work here, fan out clear branches, or coordinate app-rooted Claude Code, Codex CLI, and Antigravity workrooms when separate ownership or continuity is useful. It works in a single app out of the box and coordinates across a monorepo when needed.

Named after the ranch foreman who works the ground alongside the crew, not from an office.

Why

Bounded work should stay bounded. When a task benefits from its own workroom, straw-boss roots that worker in the app it owns instead of copying the app's context into a summary that can drift. Claude Code workers load that app's .claude/skills/ and .claude/settings.json hooks there; Claude Code, Codex CLI, and Antigravity workers all operate from the correct app directory and local instructions. The same routing applies to implementation, audits, research, and diagnosis. Cross-app routing through resolving-app is available for monorepos, not required for a single app. Full rationale: docs/architecture.md.

Highlights

  • One door: boss-say — hand over the work; it selects the owner, execution tier, coordination graph, and reality anchor.
  • Smallest sufficient loop — bounded work stays with the current agent; clear branches fan out; durable app-rooted work gets its own workroom.
  • Claude Code, Codex CLI, and Antigravity workers — choose provider, profile, model, and effort per work route; Claude routes can also use a native advisor.
  • Event-driven coordination — persisted checkpoints and terminal status drive scheduling, handoffs, and cleanup.
  • Worktree isolation — team-mode tasks can run side by side on their own feature branches.
  • Cross-main-agent resource lock — a file lock for ports and shared-DB migrations worktrees can't isolate.
  • Self-paced batches — a backlog too big for one turn gets its own /loop, started by boss-say itself.
  • Independent orchestrator handoff — with your approval, move one scope and its in-progress dispatches into a named Herdr tab whose orchestrator takes over through boss-say; the original window leaves that scope.
  • herdr for human-in-the-loop — watch or join a dispatched Claude Code, Codex CLI, or Antigravity workroom and answer questions there.

Requirements

  • Claude Code with plugins enabled, Codex CLI with plugin support, Google Antigravity (AGY CLI), or Pi with pi-herdr-agents.
  • Python 3 for the bundled lifecycle and installation scripts.
  • Herdr, required for dispatch. Claude Code, Codex CLI, and Antigravity workers all run in a Herdr pane you can watch and join. Installing, reading configuration, and tidying persisted state all work on their own; starting a dispatch needs a running Herdr service and a current pane.

Install

From a source checkout, install or update every supported CLI available on this machine from the GitHub marketplace source and verify the installed version:

bash scripts/install.sh

For development against this checkout instead, opt into its machine-local path explicitly:

bash scripts/install.sh --local

Restart active agent sessions afterward. The equivalent manual commands are below.

Claude Code

/plugin marketplace add https://github.com/wayne930242/straw-boss
/plugin install straw-boss@straw-boss

Then run once per project:

/straw-boss:init

Codex CLI

codex plugin marketplace add wayne930242/straw-boss --ref main
codex plugin add straw-boss@straw-boss

Start a new Codex session so it loads the installed skills and hooks, review and trust the bundled hooks when prompted, then run once per project:

$straw-boss:init

You can also browse or manage the installed plugin interactively by starting codex and entering /plugins. Plugins are not available in the Codex IDE extension.

Antigravity (AGY CLI)

agy plugin install wayne930242/straw-boss

Then run once per project:

/straw-boss:init

init confirms the managed apps and work routes, writes .straw-boss/apps.json, syncs the root AGENTS.md and CLAUDE.md, offers to fill in each app's missing instruction files, and checks the Herdr dispatch requirement.

Pi

pi install git:github.com/wayne930242/straw-boss

Pi loads boss-say, resolving-app, choosing-graph, shipping-task, reporting-to-user, and dispatching-work, plus the dispatch_control extension. A Pi main agent keeps the same workflow and dispatches Pi workers through pi-herdr-agents; it runs none of the bundled Python scripts. Worker models come from the role models in your pi-herdr-agents configuration. Its skills are a separate Pi set under pi/skills/, so the Claude Code, Codex, and Antigravity skills stay as they are. See Pi dispatching-work.

For a single app, init is a bonus — boss-say works the moment the plugin's installed. Run it to check Herdr readiness, configure per-app options like forbidDirectCommit/localFiles, or a monorepo's apps configured.

Usage

Once init's run, hand everything to the main agent:

boss-say fix the login redirect
boss-say audit the payments module against our rules
boss-say work through docs/backlog.md

boss-say decides the rest: the owning skill, whether the current agent can carry the work or a separate workroom is useful, the coordination graph and reality anchor, one task or a batch, and /loop when a backlog needs its own pacing. It states what it picked, and you can override it in one sentence.

Skills

Call a skill by name when the situation fits; anything unlisted goes to boss-say.

When you want to…Use
Fix, build, audit, research, or diagnose; work through a backlog; ask what is running; close out a dispatchboss-say
Set up managed apps and work routes, or check Herdr readinessinit
Find which app owns a requestresolving-app
See what a dispatch is doing without joining or interrupting itpeeking-work
Move a scope and its in-progress dispatches to a new orchestrator tabhandoff-orchestrator
Have this window take over dispatches another main agent coordinatesboss-say; it moves them only when you ask
Resolve friction in Straw Boss's own coordinationboss-assistant
Give an app without AGENTS.md or CLAUDE.md a minimal agent systemcreate-great-harness
From inside a dispatched worker, bring in a coworker for review or pairingbringing-coworker

The main agent and workers run the other skills on their own: i-am-orchestrator, choosing-graph, dispatching-work, shipping-task, contacting-orchestrators, reporting-to-user, notifying-main-agent, asking-peer-agents, and agent-feedback. docs/architecture.md describes each one.

Configuration

init writes the managed apps and each app's lifecycle configuration to .straw-boss/apps.json. The format is the apps config schema. The app summary and the project's work routes are synced into the root AGENTS.md and CLAUDE.md.

An app's agentKind names its default agent. The project's work routes name the provider profile, model, effort, and an optional Claude advisor, which init syncs into both instruction files; a Codex route records advisor: none.

Configuration is read from .straw-boss/apps.json first, falling back to .claude/straw-boss/apps.json when the new path does not exist. init writes the confirmed old configuration to the new path, leaves the old file in place, and reports that it has been superseded.

License

MIT

Experimental Jev pruning

Jev pruning is opt-in and leaves ordinary context renewal unchanged. Enable it for one newly launched Claude session (with TYPESAFE_API_KEY in the environment):

STRAW_BOSS_JEV=1 CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude

Start a session with the experiment explicitly off:

STRAW_BOSS_JEV=0 claude

Restart or resume with the desired switch; keep pruning off globally during the renewal acceptance window through 2026-09-25. See benchmark, recovery, and implementation limits.

For Codex in Herdr, use STRAW_BOSS_JEV=1 codex with the same nonblank key. At the 200k renewal point, Jev attempts a pruned codex resume in the same pane; gate misses and failures use ordinary continuity renewal. See the Codex renewal behavior and accounting costs.

Source 6 files
hooks/jev-pruning.ts 162 lines
1import type { CoreEngineInterface, Register } from 'claude-code';
2import { compact } from '../vendor/fast-jev-compaction/src/compact.js';
3import { buildJevRequest, parseJevResponse } from '../vendor/fast-jev-compaction/src/request.js';
4import { toSessionMessages } from './jev/messages.js';
5
6type Engine = CoreEngineInterface;
7
8export async function active($: Engine): Promise<boolean> {
9  return (await $.env.get('STRAW_BOSS_JEV')) === '1' &&
10    !!(await $.env.get('TYPESAFE_API_KEY'))?.trim();
11}
12
13export async function host($: Engine, command: string, data: unknown): Promise<any> {
14  const result = await $.process.run(['python3', `${$.plugin.root}/scripts/jev-runtime.py`, command], {
15    stdin: JSON.stringify(data), timeoutMs: 90_000,
16  });
17  if (result.exitCode !== 0) {
18    let reason = 'host-operation-failed';
19    try { reason = JSON.parse(result.stdout).error ?? reason; } catch { /* Use operation code. */ }
20    throw new Error(`${command}:${reason}`);
21  }
22  if (!result.stdout.trim()) throw new Error(`${command}:inactive`);
23  return JSON.parse(result.stdout);
24}
25
26export const register: Register = (on) => {
27  let pendingRun: string | undefined;
28  // `$.session.model()` answers the /model label (an alias such as `opus`),
29  // which the token-count API rejects; count with the id the API reported.
30  let apiModel: string | undefined;
31
32  on('session.start', async ($, event, next) => {
33    // Clear inherited readiness so an unloaded or child session cannot opt out
34    // of the normal renewal behavior through its parent's readiness marker.
35    await $.env.set('STRAW_BOSS_JEV_READY_SESSION', undefined);
36    await $.env.set('STRAW_BOSS_JEV_WAITING_USAGE', undefined);
37    if (await active($)) {
38      try {
39        const config = await host($, 'config', { session_id: await $.session.id() });
40        pendingRun = config.pending?.run_id;
41        await $.env.set('STRAW_BOSS_JEV_READY_SESSION', await $.session.id());
42      } catch { $.ui.log('Jev unavailable; existing renewal remains active.'); }
43    }
44    return next(event);
45  });
46
47  on('session.compact', async ($, event, next) => {
48    if (!(await active($))) return next(event);
49    // Precompute installs no history. Its returned candidate may be used much
50    // later, so delegate until the engine invokes an actual compaction.
51    if (event.trigger === 'precompute' || event.agentId) return next(event);
52    const started = Date.now();
53    const session = await $.session.id();
54    const runId = `${session}-${started}-${Math.random().toString(36).slice(2, 8)}`;
55    const metrics: any[] = [];
56    const record: any = {
57      schema_version: 2, run_id: runId, session_id: session, trigger: event.trigger,
58      started_at: new Date(started).toISOString(), tokens_before: null, tokens_after: null,
59      context_window: null, criteria_version: null, decisions: [], reduction_pct: null,
60      jev: { model: null, requests: 0, input_tokens: 0, latency_ms: 0, request_metrics: metrics },
61      outcome: 'fallback', fallback_reason: null,
62    };
63    let candidate: any;
64    let apply = false;
65    try {
66      const config = await host($, 'config', {});
67      record.criteria_version = config.criteria_version;
68      record.policy = config.policy;
69      record.policy_version = config.policy_version;
70      const usage = await $.session.usage();
71      record.context_window = usage.context.window;
72      const policy = config.policy;
73      const key = (await $.env.get('TYPESAFE_API_KEY'))!;
74      const result = await compact(event.messages, {
75        async ask(state, questions) {
76          const start = Date.now();
77          record.jev.requests++;
78          const metric: any = { request: record.jev.requests, input_tokens: null, latency_ms: null };
79          metrics.push(metric);
80          try {
81            const request = buildJevRequest({ apiKey: key }, state, questions);
82            const response = await $.http.fetch(request.url, {
83              method: request.method, headers: request.headers, body: request.body,
84            });
85            // Keep server error bodies out of logs and benchmark records.
86            if (!response.ok) throw new Error(`jev-http-${response.status}`);
87            const parsed = parseJevResponse(response.status, response.ok, response.text);
88            if (!parsed.model || !Number.isFinite(parsed.usage?.input_tokens)) {
89              throw new Error('jev-missing-model-or-usage');
90            }
91            metric.model = parsed.model;
92            metric.input_tokens = parsed.usage!.input_tokens!;
93            record.jev.model = parsed.model;
94            record.jev.input_tokens += metric.input_tokens;
95            return parsed;
96          } finally {
97            metric.latency_ms = Date.now() - start;
98            record.jev.latency_ms += metric.latency_ms;
99          }
100        },
101      }, {
102        criteria: config.criteria, protectedInputPatterns: policy.protected_input_patterns, keepThreshold: policy.keep_threshold,
103        preserveRecentMessages: policy.preserve_recent_messages,
104        maxStateTokens: policy.max_state_tokens, maxRequestTokens: policy.max_request_tokens,
105        truncateHeadChars: policy.truncate_head_chars,
106      });
107      candidate = toSessionMessages(event.messages, result.messages);
108      record.decisions = result.decisions;
109      record.state = { estimated_tokens: result.stats.stateTokens, stage: result.stats.stateStage };
110      if (result.decisions.every((d) => d.action === 'keep')) {
111        record.fallback_reason = 'no-removable-records';
112      } else {
113        Object.assign(record, await host($, 'measure', {
114          model: apiModel ?? await $.session.model(), before: event.messages, after: candidate,
115          live_input_tokens: usage.context.tokens,
116        }));
117        apply = record.gate_reduction_pct >= policy.min_reduction_ratio * 100;
118        record.fallback_reason = apply ? null : 'under-reduction-gate';
119      }
120    } catch (error) {
121      record.fallback_reason = error instanceof Error ? error.message : 'jev-failure';
122    }
123    record.decision = apply ? 'apply' : 'fallback';
124    record.outcome = apply ? null : 'fallback';
125    record.application = { status: 'awaiting-backend-usage' };
126    record.elapsed_ms = Date.now() - started;
127    try {
128      await host($, 'persist', { record, original_messages: event.messages, candidate_messages: candidate });
129    } catch {
130      $.ui.log('Jev recovery record unavailable; using built-in compaction.');
131      return next(event);
132    }
133    pendingRun = runId;
134    // A manual compaction can end before another backend request. Its old
135    // transcript usage cannot trigger renewal while the new window is unmeasured.
136    await $.env.set('STRAW_BOSS_JEV_WAITING_USAGE', session);
137    if (apply) {
138      $.ui.log(`Jev candidate selected: ${record.gate_reduction_pct.toFixed(1)}% backend-count reduction; recovery ${runId}.`);
139      return { messages: candidate };
140    }
141    $.ui.log(`Jev fallback: ${record.fallback_reason}.`);
142    return next(event);
143  });
144
145  on('turn.complete', async ($, event, next) => {
146    if (!event.agentId && event.usage?.model) apiModel = event.usage.model;
147    if (await active($)) {
148      const usage = await $.session.usage();
149      if (usage.context.tokens !== undefined) {
150        await $.env.set('STRAW_BOSS_JEV_WAITING_USAGE', undefined);
151        if (pendingRun) {
152          await host($, 'observe', { session_id: await $.session.id(), run_id: pendingRun,
153            actual_session_tokens_after: usage.context.tokens, context_window: usage.context.window,
154            source: 'first-observed-post-compaction-backend-input', observed_at: new Date().toISOString() });
155          pendingRun = undefined;
156        }
157      }
158    }
159    return next(event);
160  });
161};
162
vendor/fast-jev-compaction/src/compact.ts 325 lines
1import { noulAnswer } from './request.js';
2import { collectToolCalls, estimateTokens, fitState } from './state.js';
3import type {
4  CallAnswer,
5  CallDecision,
6  CompactOptions,
7  CompactResult,
8  CompactionState,
9  JevAsker,
10  JevQuestions,
11  Message,
12  ResolvedCompactOptions,
13  ToolCall,
14  ToolUse,
15  RetentionCriteria,
16} from './types.js';
17
18export const DEFAULT_OPTIONS: ResolvedCompactOptions = {
19  goal: '',
20  keepThreshold: 0.5,
21  preserveRecentMessages: 6,
22  maxStateTokens: 25_000,
23  maxRequestTokens: 30_000,
24  truncateHeadChars: 300,
25};
26
27/** Tokens the request envelope (`model`, key names) adds around state and questions. */
28const REQUEST_OVERHEAD_TOKENS = 20;
29
30function finite(value: number | undefined, fallback: number): number {
31  return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
32}
33
34export function resolveOptions(options: CompactOptions = {}): ResolvedCompactOptions {
35  return {
36    goal: options.goal ?? DEFAULT_OPTIONS.goal,
37    ...(options.criteria ? { criteria: options.criteria } : {}),
38    ...(options.protectedInputPatterns ? { protectedInputPatterns: options.protectedInputPatterns } : {}),
39    keepThreshold: finite(options.keepThreshold, DEFAULT_OPTIONS.keepThreshold),
40    preserveRecentMessages: Math.max(
41      0,
42      Math.floor(
43        finite(options.preserveRecentMessages, DEFAULT_OPTIONS.preserveRecentMessages),
44      ),
45    ),
46    maxStateTokens: Math.max(1, finite(options.maxStateTokens, DEFAULT_OPTIONS.maxStateTokens)),
47    maxRequestTokens: Math.max(
48      1,
49      finite(options.maxRequestTokens, DEFAULT_OPTIONS.maxRequestTokens),
50    ),
51    truncateHeadChars: Math.max(
52      0,
53      Math.floor(finite(options.truncateHeadChars, DEFAULT_OPTIONS.truncateHeadChars)),
54    ),
55  };
56}
57
58/** The two `noul` questions asked about one call: keep the call, keep its result. */
59export function questionsFor(call: ToolCall, criteria?: RetentionCriteria): JevQuestions {
60  if (criteria) {
61    return Object.fromEntries((["keepCall", "keepResult"] as const).map((kind) => [
62      `${kind === "keepCall" ? "call" : "result"}_${call.id}`,
63      { type: "noul" as const, instructions: `Target: tool call ${call.id} (${call.tool}, ${call.resultChars} result chars). ${criteria[kind]}`,
64        criteria: { true: criteria[kind], false: "The proposition is false; this material can be pruned according to the stated retention contract." } },
65    ]));
66  }
67  return {
68    [`call_${call.id}`]: {
69      type: 'noul',
70      instructions: `Tool call ${call.id} (${call.tool}) should stay in the history: knowing this call was made, with its input, still matters for what the assistant does next`,
71    },
72    [`result_${call.id}`]: {
73      type: 'noul',
74      instructions: `The full output of tool call ${call.id} (${call.tool}, ${call.resultChars} chars) should stay in the history verbatim: the assistant still needs its contents and re-running the tool would not do`,
75    },
76  };
77}
78
79/**
80 * Splits the candidate calls into batches whose questions, together with the
81 * (always complete) state, fit one request.
82 */
83export function batchCalls(
84  calls: readonly ToolCall[],
85  stateTokens: number,
86  options: Pick<ResolvedCompactOptions, 'maxRequestTokens' | 'criteria'>,
87): ToolCall[][] {
88  const budget = options.maxRequestTokens - stateTokens - REQUEST_OVERHEAD_TOKENS;
89  const batches: ToolCall[][] = [];
90  let current: ToolCall[] = [];
91  let currentTokens = 0;
92  for (const call of calls) {
93    const tokens = estimateTokens(JSON.stringify(questionsFor(call, options.criteria)));
94    if (current.length > 0 && currentTokens + tokens > budget) {
95      batches.push(current);
96      current = [];
97      currentTokens = 0;
98    }
99    if (current.length === 0 && tokens > budget) {
100      throw new Error(
101        `state leaves no room for questions (~${stateTokens} of ${options.maxRequestTokens} tokens)`,
102      );
103    }
104    current.push(call);
105    currentTokens += tokens;
106  }
107  if (current.length > 0) batches.push(current);
108  return batches;
109}
110
111export function decideCall(
112  call: Pick<ToolCall, 'id' | 'tool' | 'pinned' | 'governingSource'>,
113  answer: CallAnswer,
114  options: Pick<ResolvedCompactOptions, 'keepThreshold'>,
115): CallDecision {
116  const base = { id: call.id, tool: call.tool, ...answer };
117  if (call.pinned) return { ...base, action: 'keep', reason: call.governingSource ? 'governing_source' : 'pinned' };
118  if (answer.keepResult >= options.keepThreshold) {
119    return { ...base, action: 'keep', reason: 'kept' };
120  }
121  if (answer.keepCall >= options.keepThreshold) {
122    return { ...base, action: 'drop_result', reason: 'result_dropped' };
123  }
124  return { ...base, action: 'drop_call', reason: 'call_dropped' };
125}
126
127async function askBatch(
128  asker: JevAsker,
129  state: CompactionState,
130  batch: readonly ToolCall[],
131  criteria?: RetentionCriteria,
132): Promise<Map<string, CallAnswer>> {
133  const questions: JevQuestions = Object.assign({}, ...batch.map((call) => questionsFor(call, criteria)));
134  const { answers } = await asker.ask(state, questions);
135  return new Map(
136    batch.map((call) => [
137      call.id,
138      {
139        keepCall: noulAnswer(answers, `call_${call.id}`),
140        keepResult: noulAnswer(answers, `result_${call.id}`),
141      },
142    ]),
143  );
144}
145
146function truncatedResultText(text: string, isError: boolean, headChars: number): string {
147  if (text.length <= headChars) return text;
148  const head = headChars > 0 ? `${text.slice(0, headChars)}\n` : '';
149  return `${head}[fast-jev-compaction truncated ${text.length - headChars} chars of this tool result${
150    isError ? ' (error)' : ''
151  }; original saved in the recovery record]`;
152}
153
154/**
155 * Rebuilds the conversation from the decisions. A dropped call disappears
156 * together with its result; a dropped result keeps a bounded head and note.
157 * Messages that lose all their content are removed; untouched messages are
158 * returned as the same objects they came in as.
159 */
160export function applyDecisions(
161  messages: readonly Message[],
162  decisions: readonly CallDecision[],
163  calls: readonly ToolCall[],
164  headChars: number,
165): Message[] {
166  const byId = new Map(calls.map((call) => [call.id, call]));
167  const actions = new Map<string, CallDecision['action']>();
168  for (const decision of decisions) {
169    const call = byId.get(decision.id);
170    if (call && decision.action !== 'keep') actions.set(call.tool_use_id, decision.action);
171  }
172  const kept: Message[] = [];
173  for (const message of messages) {
174    const touched =
175      message.toolUses.some((tool) => actions.has(tool.tool_use_id)) ||
176      (message.toolResults ?? []).some((result) => actions.has(result.tool_use_id));
177    if (!touched) {
178      kept.push(message);
179      continue;
180    }
181    const toolUses = message.toolUses
182      .filter((tool) => actions.get(tool.tool_use_id) !== 'drop_call')
183      .map((tool) => {
184        if (actions.get(tool.tool_use_id) !== 'drop_result') return tool;
185        const text = truncatedResultText(
186          tool.text ?? '',
187          tool.isError ?? false,
188          headChars,
189        );
190        if ((tool.text ?? '') === text) return tool;
191        const copy: ToolUse = {
192          tool_use_id: tool.tool_use_id,
193          tool: tool.tool,
194          input: tool.input,
195          text,
196        };
197        if (tool.isError) copy.isError = true;
198        return copy;
199      });
200    const toolResults = (message.toolResults ?? [])
201      .filter((result) => actions.get(result.tool_use_id) !== 'drop_call')
202      .map((result) => {
203        if (actions.get(result.tool_use_id) !== 'drop_result') return result;
204        const text = truncatedResultText(result.text, result.isError ?? false, headChars);
205        return text === result.text
206          ? result
207          : {
208              tool_use_id: result.tool_use_id,
209              text,
210              isError: result.isError,
211            };
212      });
213    if (
214      !message.toolUses.some(
215        (tool) => actions.get(tool.tool_use_id) === 'drop_call',
216      ) &&
217      !(message.toolResults ?? []).some(
218        (result) => actions.get(result.tool_use_id) === 'drop_call',
219      ) &&
220      toolUses.every((tool, index) => tool === message.toolUses[index]) &&
221      toolResults.every(
222        (result, index) => result === message.toolResults?.[index],
223      )
224    ) {
225      kept.push(message);
226      continue;
227    }
228    if (message.text.length === 0 && toolUses.length === 0 && toolResults.length === 0) {
229      continue;
230    }
231    const rebuilt: Message = { role: message.role, text: message.text, toolUses };
232    if (toolResults.length > 0) rebuilt.toolResults = toolResults;
233    kept.push(rebuilt);
234  }
235  return kept;
236}
237
238/** Characters of text, tool input and tool output a message holds. */
239export function messageChars(message: Message): number {
240  let total = message.text.length;
241  for (const tool of message.toolUses) {
242    try {
243      total += JSON.stringify(tool.input).length;
244    } catch {
245      total += 20;
246    }
247  }
248  for (const result of message.toolResults ?? []) total += result.text.length;
249  return total;
250}
251
252export function reductionRatio(result: Pick<CompactResult, 'stats'>): number {
253  const { charsBefore, charsAfter } = result.stats;
254  return charsBefore === 0 ? 0 : (charsBefore - charsAfter) / charsBefore;
255}
256
257function count(decisions: readonly CallDecision[], reason: CallDecision['reason']): number {
258  return decisions.filter((decision) => decision.reason === reason).length;
259}
260
261/**
262 * Compacts a transcript by asking Jev, for every tool call outside the pinned
263 * first and newest messages, whether the call and whether its result must
264 * stay. The whole history (results omitted, fitted into `maxStateTokens`) is
265 * sent as state with every batch of questions. Throws when Jev fails or the
266 * history cannot be fitted; the caller decides whether to fall back.
267 */
268export async function compact(
269  messages: readonly Message[],
270  asker: JevAsker,
271  options: CompactOptions = {},
272): Promise<CompactResult> {
273  const started = Date.now();
274  const resolved = resolveOptions(options);
275  const calls = collectToolCalls(messages, resolved.preserveRecentMessages, resolved.protectedInputPatterns);
276  const candidates = calls.filter((call) => !call.pinned);
277  const charsBefore = messages.reduce((sum, message) => sum + messageChars(message), 0);
278
279  let fitted: { tokens: number; stage: string } = { tokens: 0, stage: '' };
280  let batches: ToolCall[][] = [];
281  const answers = new Map<string, CallAnswer>();
282  if (candidates.length > 0) {
283    const state = fitState(messages, calls, resolved);
284    fitted = state;
285    batches = batchCalls(candidates, state.tokens, resolved);
286    const settled = await Promise.allSettled(
287      batches.map((batch) => askBatch(asker, state.state, batch, resolved.criteria)),
288    );
289    const failed = settled.find((item) => item.status === "rejected");
290    if (failed?.status === "rejected") throw failed.reason;
291    for (const item of settled) if (item.status === "fulfilled") {
292      for (const [id, answer] of item.value) answers.set(id, answer);
293    }
294  }
295
296  const decisions = calls.map((call) =>
297    decideCall(call, answers.get(call.id) ?? { keepCall: 1, keepResult: 1 }, resolved),
298  );
299  const kept = applyDecisions(
300    messages,
301    decisions,
302    calls,
303    resolved.truncateHeadChars,
304  );
305  return {
306    messages: kept,
307    decisions,
308    stats: {
309      messagesBefore: messages.length,
310      messagesAfter: kept.length,
311      charsBefore,
312      charsAfter: kept.reduce((sum, message) => sum + messageChars(message), 0),
313      calls: calls.length,
314      kept: count(decisions, 'kept'),
315      resultsDropped: count(decisions, 'result_dropped'),
316      callsDropped: count(decisions, 'call_dropped'),
317      pinned: count(decisions, 'pinned') + count(decisions, 'governing_source'),
318      stateTokens: fitted.tokens,
319      stateStage: fitted.stage,
320      requests: batches.length,
321      ms: Date.now() - started,
322    },
323  };
324}
325
vendor/fast-jev-compaction/src/request.ts 82 lines
1import type { JevAnswer, JevQuestions, JevResponse, JevState } from './types.js';
2
3export const SYSTEM_ONE_URL = 'https://api.typesafe.ai/v1/systemone';
4export const DEFAULT_MODEL = 'jev-latest';
5
6export interface JevRequest {
7  url: string;
8  method: 'POST';
9  headers: Record<string, string>;
10  body: string;
11}
12
13/** The HTTP request for one Jev call, for any fetch-like transport. */
14export function buildJevRequest(
15  params: {
16    apiKey: string;
17    model?: string;
18    baseUrl?: string;
19  },
20  state: JevState,
21  questions: JevQuestions,
22): JevRequest {
23  return {
24    url: params.baseUrl ?? SYSTEM_ONE_URL,
25    method: 'POST',
26    headers: {
27      authorization: `Bearer ${params.apiKey}`,
28      'content-type': 'application/json',
29    },
30    body: JSON.stringify({
31      model: params.model ?? DEFAULT_MODEL,
32      state,
33      questions,
34    }),
35  };
36}
37
38/** Validates a Jev response body; throws on anything but an `answers` object. */
39export function parseJevResponse(
40  status: number,
41  ok: boolean,
42  text: string,
43): JevResponse {
44  if (!ok) {
45    throw new Error(`Jev request failed (${status}): ${text.slice(0, 200)}`);
46  }
47  let parsed: unknown;
48  try {
49    parsed = JSON.parse(text);
50  } catch {
51    throw new Error('Jev returned malformed JSON');
52  }
53  if (
54    parsed === null ||
55    typeof parsed !== 'object' ||
56    !('answers' in parsed) ||
57    parsed.answers === null ||
58    typeof parsed.answers !== 'object'
59  ) {
60    throw new Error('Jev response is missing answers');
61  }
62  return parsed as JevResponse;
63}
64
65/** The `noul` probability of one answer; throws when it is not there. */
66export function noulAnswer(
67  answers: Record<string, JevAnswer>,
68  name: string,
69): number {
70  const answer = answers[name];
71  if (
72    !answer ||
73    !('noul' in answer) ||
74    typeof answer.noul !== 'number' ||
75    !Number.isFinite(answer.noul) ||
76    answer.noul < 0 || answer.noul > 1
77  ) {
78    throw new Error(`Invalid Jev answer for ${name}`);
79  }
80  return answer.noul;
81}
82
hooks/jev/messages.ts 60 lines
1import type { SessionMessage, ToolUseSummary, ToolResultSummary } from 'claude-code';
2import type { Message, ToolUse, ToolResult } from '../../vendor/fast-jev-compaction/src/types.js';
3
4function toolUseSummary(tool: ToolUse): ToolUseSummary {
5  const summary: ToolUseSummary = {
6    tool_use_id: tool.tool_use_id,
7    tool: tool.tool,
8    input: tool.input,
9  };
10  if (tool.text !== undefined) summary.text = tool.text;
11  if (tool.isError) summary.isError = true;
12  return summary;
13}
14
15function toolResultSummary(result: ToolResult): ToolResultSummary {
16  return {
17    tool_use_id: result.tool_use_id,
18    text: result.text,
19    isError: result.isError ?? false,
20  };
21}
22
23/**
24 * Maps the library's output back onto session messages. Whatever came back
25 * Every returned message is fresh so Claude persists a new parent chain across
26 * restart. Reusing message handles lets 2.1.278 reconnect dropped ancestors
27 * when resuming. Text and unmodified tool objects remain verbatim.
28 */
29export function toSessionMessages(
30  input: readonly SessionMessage[],
31  output: readonly Message[],
32): SessionMessage[] {
33  const messages = new Map<Message, SessionMessage>();
34  const uses = new Map<ToolUse, ToolUseSummary>();
35  const results = new Map<ToolResult, ToolResultSummary>();
36  for (const message of input) {
37    messages.set(message, message);
38    for (const tool of message.toolUses) uses.set(tool, tool);
39    for (const result of message.toolResults ?? []) results.set(result, result);
40  }
41  return output.map((message) => {
42    const own = messages.get(message);
43    if (own) {
44      const { handle: _handle, ...fresh } = own;
45      return fresh;
46    }
47    const rebuilt: SessionMessage = {
48      role: message.role,
49      text: message.text,
50      toolUses: message.toolUses.map((tool) => uses.get(tool) ?? toolUseSummary(tool)),
51    };
52    if (message.toolResults && message.toolResults.length > 0) {
53      rebuilt.toolResults = message.toolResults.map(
54        (result) => results.get(result) ?? toolResultSummary(result),
55      );
56    }
57    return rebuilt;
58  });
59}
60
vendor/fast-jev-compaction/src/state.ts 308 lines
1import type {
2  CompactionState,
3  FittedState,
4  HistoryEntry,
5  Message,
6  ResolvedCompactOptions,
7  ToolCall,
8  ToolResult,
9} from './types.js';
10
11export const STATE_CONTEXT =
12  'A coding assistant history is being compacted. `history` runs oldest first, with tool results omitted and long text abridged. Score the named call and its full result for the unfinished task. Keep governing instructions and required evidence; archived originals are recoverable, but some tool actions cannot be reproduced.';
13
14/** Successive caps on the serialised tool input included per call. */
15const INPUT_CHARS = [1000, 200, 60] as const;
16const TEXT_HEAD = 400;
17const TEXT_TAIL = 150;
18
19const TOKEN_PIECES = /[A-Za-z]+|\d+|[^\sA-Za-z\d]/g;
20
21/**
22 * Estimates tokens without a tokenizer: a word costs one token per six
23 * letters, a digit half a token, any other symbol nine tenths. Calibrated
24 * against the usage Jev reports for real transcripts, where it lands 2–18%
25 * above the true count; a plain characters-per-token ratio undercounts the
26 * JSON-heavy states by up to 40%.
27 */
28export function estimateTokens(text: string): number {
29  let tokens = 0;
30  for (const [piece] of text.matchAll(TOKEN_PIECES)) {
31    const first = piece.charCodeAt(0);
32    if (first >= 48 && first <= 57) tokens += piece.length / 2;
33    else if ((first >= 65 && first <= 90) || (first >= 97 && first <= 122)) {
34      tokens += 1 + Math.floor((piece.length - 1) / 6);
35    } else tokens += 0.9;
36  }
37  return Math.ceil(tokens);
38}
39
40export function truncate(text: string, limit: number): string {
41  return text.length <= limit ? text : `${text.slice(0, Math.max(0, limit - 1))}…`;
42}
43
44function abridge(text: string, head: number, tail: number): string {
45  if (text.length <= head + tail + 40) return text;
46  const omitted = text.length - head - tail;
47  return `${text.slice(0, head)}\n[… ${omitted} chars omitted …]\n${text.slice(-tail)}`;
48}
49
50export function isPinned(
51  index: number,
52  total: number,
53  preserveRecentMessages: number,
54): boolean {
55  return index === 0 || index >= total - preserveRecentMessages;
56}
57
58/**
59 * Pairs every tool_use with its tool_result by `tool_use_id`. Calls without a
60 * result are not candidates (there is nothing to drop yet).
61 */
62export function collectToolCalls(
63  messages: readonly Message[],
64  preserveRecentMessages: number,
65  protectedInputPatterns: string[] = [],
66): ToolCall[] {
67  const results = new Map<string, { index: number; result: ToolResult }>();
68  messages.forEach((message, index) => {
69    for (const result of message.toolResults ?? []) {
70      results.set(result.tool_use_id, { index, result });
71    }
72  });
73  const calls: ToolCall[] = [];
74  messages.forEach((message, callIndex) => {
75    for (const tool of message.toolUses) {
76      const found = results.get(tool.tool_use_id);
77      if (!found) continue;
78      const governingSource = protectedInputPatterns.some((pattern) => new RegExp(pattern).test(JSON.stringify(tool.input)));
79      calls.push({
80        id: `t${calls.length + 1}`,
81        tool_use_id: tool.tool_use_id,
82        tool: tool.tool,
83        input: tool.input,
84        callIndex,
85        resultIndex: found.index,
86        resultChars: found.result.text.length,
87        isError: found.result.isError ?? false,
88        ...(governingSource ? { governingSource: true } : {}),
89        pinned:
90          governingSource || isPinned(callIndex, messages.length, preserveRecentMessages) ||
91          isPinned(found.index, messages.length, preserveRecentMessages),
92      });
93    }
94  });
95  return calls;
96}
97
98function inputText(input: Record<string, unknown>, limit: number): string {
99  let json = '';
100  try {
101    json = JSON.stringify(input);
102  } catch {
103    json = '[unserializable input]';
104  }
105  return truncate(json, limit);
106}
107
108function resultNote(call: ToolCall): string {
109  return `${call.isError ? 'error' : 'ok'}, ${call.resultChars} chars (omitted)`;
110}
111
112/** One call as a single line, for when the structured form is too costly. */
113function compactCall(call: ToolCall): string {
114  const input = Object.entries(call.input)
115    .map(([key, value]) => {
116      const text = typeof value === 'string' ? value : inputText({ [key]: value }, 200);
117      return `${key}=${text.replace(/\s+/g, ' ')}`;
118    })
119    .join(' ');
120  return `${call.id} ${call.tool} ${truncate(input, INPUT_CHARS[2])} → ${
121    call.isError ? 'error' : 'ok'
122  } ${call.resultChars}ch`;
123}
124
125/**
126 * Folds runs of adjacent call-only entries into one entry each, so the
127 * per-entry envelope is paid once per run; the call lines keep their ids.
128 */
129function mergeCallRuns(history: readonly HistoryEntry[], pinned: (e: HistoryEntry) => boolean): HistoryEntry[] {
130  const merged: HistoryEntry[] = [];
131  for (const entry of history) {
132    const previous = merged[merged.length - 1];
133    const foldable = (e: HistoryEntry): boolean =>
134      !pinned(e) && e.text.length === 0 && typeof e.tool_calls?.[0] === 'string';
135    if (previous && foldable(previous) && foldable(entry) && previous.role === entry.role) {
136      previous.tool_calls = [...(previous.tool_calls as string[]), ...(entry.tool_calls as string[])];
137      continue;
138    }
139    merged.push({ ...entry });
140  }
141  return merged;
142}
143
144function callsByMessage(calls: readonly ToolCall[]): Map<number, ToolCall[]> {
145  const byMessage = new Map<number, ToolCall[]>();
146  for (const call of calls) {
147    const list = byMessage.get(call.callIndex) ?? [];
148    list.push(call);
149    byMessage.set(call.callIndex, list);
150  }
151  return byMessage;
152}
153
154function historyEntries(
155  messages: readonly Message[],
156  calls: readonly ToolCall[],
157  inputChars: number,
158): HistoryEntry[] {
159  const byMessage = callsByMessage(calls);
160  const entries: HistoryEntry[] = [];
161  messages.forEach((message, i) => {
162    const toolCalls = (byMessage.get(i) ?? []).map((call) => ({
163      id: call.id,
164      tool: call.tool,
165      input: inputText(call.input, inputChars),
166      result: resultNote(call),
167    }));
168    if (message.text.trim().length === 0 && toolCalls.length === 0) return;
169    const entry: HistoryEntry = { i, role: message.role, text: message.text };
170    if (toolCalls.length > 0) entry.tool_calls = toolCalls;
171    entries.push(entry);
172  });
173  return entries;
174}
175
176/** The last three user prompts, as the default `goal`. */
177export function goalFromMessages(messages: readonly Message[]): string {
178  return messages
179    .filter(
180      (message) =>
181        message.role === 'user' &&
182        message.text.trim().length > 0 &&
183        (message.toolResults ?? []).length === 0,
184    )
185    .slice(-3)
186    .map((message) => truncate(message.text, 500))
187    .join('\n');
188}
189
190/**
191 * Builds the Jev state from the whole conversation and shrinks it in stages
192 * until it fits `maxStateTokens`: tool inputs are truncated, then long texts
193 * are abridged oldest-first (pinned messages last), then old messages collapse
194 * to a one-line note, then old tool calls shrink to one line each, then old
195 * messages that carry no call are left out, then runs of old call-only
196 * messages are folded into one entry. Throws when even that is too big.
197 */
198export function fitState(
199  messages: readonly Message[],
200  calls: readonly ToolCall[],
201  options: Pick<ResolvedCompactOptions, 'maxStateTokens' | 'preserveRecentMessages' | 'goal'>,
202): FittedState {
203  const goal = options.goal || goalFromMessages(messages);
204  const stateOf = (history: HistoryEntry[]): CompactionState => ({
205    context: STATE_CONTEXT,
206    goal,
207    history,
208  });
209  const entryTokens = (entry: HistoryEntry): number => estimateTokens(JSON.stringify(entry)) + 1;
210  const baseTokens = estimateTokens(JSON.stringify(stateOf([])));
211  const fitted = (history: HistoryEntry[], tokens: number, stage: string): FittedState => ({
212    state: stateOf(history),
213    tokens,
214    stage,
215  });
216
217  let history: HistoryEntry[] = [];
218  let perEntry: number[] = [];
219  let tokens = 0;
220  const rebuild = (inputChars: number): void => {
221    history = historyEntries(messages, calls, inputChars);
222    perEntry = history.map(entryTokens);
223    tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
224  };
225  const fits = (): boolean => tokens <= options.maxStateTokens;
226  const shrink = (index: number, change: (entry: HistoryEntry) => void): void => {
227    const entry = history[index];
228    if (!entry) return;
229    change(entry);
230    const now = entryTokens(entry);
231    tokens += now - (perEntry[index] ?? 0);
232    perEntry[index] = now;
233  };
234
235  rebuild(INPUT_CHARS[0]);
236  if (fits()) return fitted(history, tokens, 'full');
237
238  for (const limit of INPUT_CHARS.slice(1)) {
239    rebuild(limit);
240    if (fits()) return fitted(history, tokens, `inputs<=${limit}`);
241  }
242
243  const pinned = (entry: HistoryEntry): boolean =>
244    isPinned(entry.i, messages.length, options.preserveRecentMessages);
245  const indices = history.map((_, index) => index);
246  const order = [
247    ...indices.filter((index) => !pinned(history[index]!)),
248    ...indices.filter((index) => pinned(history[index]!)),
249  ];
250
251  for (const index of order) {
252    const entry = history[index]!;
253    if (entry.text.length <= TEXT_HEAD + TEXT_TAIL + 40) continue;
254    shrink(index, (e) => {
255      e.text = abridge(e.text, TEXT_HEAD, TEXT_TAIL);
256    });
257    if (fits()) return fitted(history, tokens, 'texts abridged');
258  }
259
260  for (const index of order) {
261    const entry = history[index]!;
262    if (pinned(entry) || entry.text.length === 0) continue;
263    const original = messages[entry.i]?.text.length ?? entry.text.length;
264    shrink(index, (e) => {
265      e.text = `[… ${original} chars omitted …]`;
266    });
267    if (fits()) return fitted(history, tokens, 'old messages collapsed');
268  }
269
270  const byMessage = callsByMessage(calls);
271  for (const index of order) {
272    const entry = history[index]!;
273    const own = byMessage.get(entry.i);
274    if (pinned(entry) || !own) continue;
275    shrink(index, (e) => {
276      e.tool_calls = own.map(compactCall);
277    });
278    if (fits()) return fitted(history, tokens, 'old calls compacted');
279  }
280
281  const left = new Set<number>();
282  for (const index of order) {
283    const entry = history[index]!;
284    if (pinned(entry) || entry.tool_calls) continue;
285    left.add(index);
286    tokens -= perEntry[index] ?? 0;
287    if (fits()) {
288      return fitted(
289        history.filter((_, i) => !left.has(i)),
290        tokens,
291        'old messages left out',
292      );
293    }
294  }
295
296  history = mergeCallRuns(
297    history.filter((_, i) => !left.has(i)),
298    pinned,
299  );
300  perEntry = history.map(entryTokens);
301  tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
302  if (fits()) return fitted(history, tokens, 'old calls merged');
303
304  throw new Error(
305    `history too large for Jev (~${tokens} tokens after truncation, limit ${options.maxStateTokens})`,
306  );
307}
308
vendor/fast-jev-compaction/src/types.ts 213 lines
1export type Role = 'user' | 'assistant';
2
3/**
4 * A tool_use block of an assistant message. `text` and `isError` mirror the
5 * outcome once the transcript holds it (Claude Code attaches them).
6 */
7export interface ToolUse {
8  tool_use_id: string;
9  tool: string;
10  input: Record<string, unknown>;
11  text?: string;
12  isError?: boolean;
13}
14
15/** A tool_result block of a user message. */
16export interface ToolResult {
17  tool_use_id: string;
18  text: string;
19  isError?: boolean;
20}
21
22/**
23 * One transcript message. The shape is a subset of Claude Code's
24 * `SessionMessage`, so a session transcript can be passed in as is.
25 */
26export interface Message {
27  role: Role;
28  text: string;
29  toolUses: ToolUse[];
30  toolResults?: ToolResult[];
31}
32
33/** A tool call paired with its result by `tool_use_id`. */
34export interface ToolCall {
35  /** Short id used in the Jev state and question names (`t1`, `t2`, ...). */
36  id: string;
37  tool_use_id: string;
38  tool: string;
39  input: Record<string, unknown>;
40  /** Index of the message holding the tool_use block. */
41  callIndex: number;
42  /** Index of the message holding the tool_result block. */
43  resultIndex: number;
44  resultChars: number;
45  isError: boolean;
46  /** In the first or the newest preserved messages; never a candidate. */
47  pinned: boolean;
48  governingSource?: boolean;
49}
50
51export interface CallAnswer {
52  /** Jev's probability that the call itself still matters. */
53  keepCall: number;
54  /** Jev's probability that the full result still needs to stay verbatim. */
55  keepResult: number;
56}
57
58export type CallAction = 'keep' | 'drop_result' | 'drop_call';
59
60export interface CallDecision extends CallAnswer {
61  id: string;
62  tool: string;
63  action: CallAction;
64  reason: 'pinned' | 'governing_source' | 'kept' | 'result_dropped' | 'call_dropped';
65}
66
67export interface HistoryToolCall {
68  id: string;
69  tool: string;
70  input: string;
71  result: string;
72}
73
74export interface HistoryEntry {
75  i: number;
76  role: Role;
77  text: string;
78  /** Structured per call, or one compact line per call once the state has to shrink. */
79  tool_calls?: HistoryToolCall[] | string[];
80}
81
82/** The state sent with every Jev request: the whole history, results omitted. */
83export interface CompactionState {
84  context: string;
85  goal: string;
86  history: HistoryEntry[];
87}
88
89export interface FittedState {
90  state: CompactionState;
91  tokens: number;
92  /** Which fitting stage produced the state, for diagnostics. */
93  stage: string;
94}
95
96export interface RetentionCriteria {
97  keepCall: string;
98  keepResult: string;
99}
100
101export interface CompactOptions {
102  criteria?: RetentionCriteria;
103  protectedInputPatterns?: string[];
104  /** Ongoing task description; defaults to the last few user prompts. */
105  goal?: string;
106  /** Minimum keep probability for a call or result to stay. Default 0.5. */
107  keepThreshold?: number;
108  /** Newest messages never touched (the first message is always kept). Default 6. */
109  preserveRecentMessages?: number;
110  /** Estimated token ceiling for the state. Default 25000. */
111  maxStateTokens?: number;
112  /** Estimated token ceiling for state plus one batch of questions. Default 30000. */
113  maxRequestTokens?: number;
114  /** Characters of a dropped tool result to retain. Default 300. */
115  truncateHeadChars?: number;
116}
117
118export interface ResolvedCompactOptions {
119  criteria?: RetentionCriteria;
120  protectedInputPatterns?: string[];
121  goal: string;
122  keepThreshold: number;
123  preserveRecentMessages: number;
124  maxStateTokens: number;
125  maxRequestTokens: number;
126  truncateHeadChars: number;
127}
128
129export interface CompactResult {
130  /** The compacted transcript; untouched messages are the input objects. */
131  messages: Message[];
132  decisions: CallDecision[];
133  stats: {
134    messagesBefore: number;
135    messagesAfter: number;
136    charsBefore: number;
137    charsAfter: number;
138    calls: number;
139    kept: number;
140    resultsDropped: number;
141    callsDropped: number;
142    pinned: number;
143    stateTokens: number;
144    /** Which fitting stage the state needed, '' when no request was made. */
145    stateStage: string;
146    requests: number;
147    ms: number;
148  };
149}
150
151/** The `state` of a Jev request: a string or any JSON-serialisable object. */
152export type JevState = string | object;
153
154export interface NoulQuestion {
155  type: 'noul';
156  instructions: string;
157  criteria?: {
158    true?: string;
159    false?: string;
160  };
161}
162
163export interface ChoiceQuestion {
164  type: 'choice';
165  instructions: string;
166  criteria: Record<string, string | null>;
167}
168
169export interface ScoreQuestion {
170  type: 'score';
171  instructions: string;
172  criteria: string[];
173}
174
175export type JevQuestion = NoulQuestion | ChoiceQuestion | ScoreQuestion;
176export type JevQuestions = Record<string, JevQuestion>;
177
178export interface NoulAnswer {
179  type?: 'noul';
180  noul: number;
181}
182
183export interface ChoiceAnswer {
184  type?: 'choice';
185  choice: string;
186  confidence: number;
187  probabilities: Record<string, number>;
188}
189
190export interface ScoreAnswer {
191  type?: 'score';
192  score: number;
193  confidence: number;
194  probabilities: Record<string, number>;
195}
196
197export type JevAnswer = NoulAnswer | ChoiceAnswer | ScoreAnswer;
198
199export interface JevResponse {
200  model?: string;
201  answers: Record<string, JevAnswer>;
202  usage?: {
203    input_tokens?: number;
204    output_tokens?: number;
205  };
206  [key: string]: unknown;
207}
208
209/** Anything that can answer Jev questions: `JevClient`, or a host-provided adapter. */
210export interface JevAsker {
211  ask(state: JevState, questions: JevQuestions): Promise<JevResponse>;
212}
213