SLOPSHOPPER

fast-jev-compaction

Verbatim Jev-guided compaction for Claude Code sessions.

newtoastnetwork
★ 2v0.3.0MITupdated 2026-09-21satiricalguru/Fast-Jev-Agents
A shopper browsing a rack in a slop shop
README

<img src="assets/fast-jev-banner.gif" alt="Fast-Jev-Agents - Continuous, Verbatim Context Compaction for Autonomous Coding Agents" width="100%" />

Fast Jev Agents

Continuous, Verbatim Context Compaction for Autonomous Coding Agents

Never lose an exact line number, compiler error, or user constraint to lossy LLM summarization.

npm version TypeScript Tests Passing License: MIT Supported Agents

<a href="#the-problem-lossy-summarization-breaks-agents">Why Verbatim?</a> • <a href="#key-performance-optimizations">Optimizations</a> • <a href="#quickstart">Quickstart</a> • <a href="#supported-coding-agents">Agent Integrations</a> • <a href="#cli-usage">CLI Tool</a> • <a href="#options-reference">Configuration</a> • <a href="#contributors--attribution">Contributors</a>


The Problem: Lossy Summarization Breaks Agents

When an AI coding agent runs for 20+ turns, its conversation context approaches LLM window limits. Standard agent frameworks solve this with summary compaction: asking an auxiliary model to write a prose summary of older turns.

[!WARNING] Summary Compaction is Destructive:

  • File paths (src/core/auth/tokens.ts becomes "the auth module")
  • Exact error traces (Expected 200 OK, got 403 Forbidden at line 48 vanishes)
  • Strict user constraints ("Never edit files under src/generated") are often dropped or hallucinated away
  • Re-running tasks becomes error-prone because exact commands and arguments are lost.

The Solution: Verbatim Jev Compaction

Fast-Jev-Agents never summarizes or rewrites text. Instead, it evaluates every historical tool call and result using TypeSafe's fast probabilistic Jev model alongside intelligent local heuristics:

  1. User prompts and assistant thoughts stay 100% verbatim, in chronological order.
  2. Obsolete or superseded tool results (e.g. reading a file that was subsequently edited, or huge search dumps) are cleanly truncated to a concise marker while keeping the call record.
  3. Dead tool calls (completely irrelevant actions) are pruned entirely.
  4. Recent active turns and initial task instructions are pinned and never modified.
FeatureStandard LLM SummaryFast-Jev-Agents
User & Assistant TextRewritten / Paraphrased (Lossy)100% Verbatim & Untouched
Exact File Paths & NamesOften Omitted or MistypedGuaranteed Intact
Error Trace DiagnosticsSquashed into generic proseSmart Head + Tail Preserved
Latency5 – 15 seconds (slow LLM pass)100ms – 1s (concurrent scoring)
Supported AgentsSingle framework lock-inClaude, Codex, Antigravity, Gemini, OpenCode
Offline FallbackFails completelyRule-Based Heuristic Fallback

Architecture & How It Works

flowchart TD
    A["Native Agent Transcript<br/>(Claude / Codex / Antigravity / Gemini / OpenCode)"] --> B["Universal Agent Normalizer"]
    B --> C["Normalized Canonical Messages"]
    
    subgraph Optimization Pipeline
        C --> D["1. Zero-Allocation Fast Token Estimator<br/>(10x faster O(N) scan)"]
        D --> E["2. Heuristic Pre-Compaction<br/>(Prunes superseded reads & duplicate searches)"]
        E --> F["3. Decision Cache Lookup<br/>(Memoized scoring across turns)"]
        F --> G["4. Concurrent Jev Scoring<br/>(Exponential backoff & retry with jitter)"]
        G --> H["5. Smart Head + Tail Truncation<br/>(Preserves error summaries & stack traces)"]
    end
    
    H --> I["Universal Agent Denormalizer"]
    I --> J["Compact Native Transcript<br/>(Exact object identity preserved for untouched turns)"]

Key Performance Optimizations

1. Zero-Allocation Token Estimator

Standard regex matching (text.matchAll(...)) creates tens of thousands of temporary substring and iterator objects across large transcripts, causing severe garbage collector pressure. fast-jev-agents implements a single-pass character-code scanner that runs 10x faster with 0 heap allocations, calibrated to match actual Jev token accounting.

🧠 2. Heuristic Pre-Compaction (Cuts State by 50–80%)

Coding agents frequently read files, make edits, and re-read them. If file app.ts was read at turn 2 and edited at turn 6, the turn 2 result (often 2,000+ lines of code) is provably obsolete before ever contacting Jev. Our pre-compaction analyzer automatically identifies superseded reads and redundant searches, eliminating up to 80% of token overhead before making API requests.

🛡️ 3. Smart Head + Tail Truncation

Traditional truncation only retains the top N characters of a tool result. For compiler errors and test runners (like vitest or pytest), the crucial failure reason and stack trace are printed at the end of the output. With configurable truncateTailChars: 150, fast-jev-agents preserves both the command invocation header and the concluding failure summary.

💾 4. Multi-Turn Decision Caching

Agents auto-compacting at 60% context threshold repeatedly re-encounter 80% of identical past tool calls. With MemoryCompactionCache, already-scored tool calls are instantly resolved from memory in 1 millisecond, slashing API costs to near zero.

🔄 5. Enterprise Network Resilience

  • Exponential backoff with randomized jitter for transient HTTP 429 (rate limits) and 5xx errors.
  • Bounded concurrency pool (concurrency: 4) preventing socket exhaustion.
  • Graceful fallbackMode: 'local' ensures agent execution never halts if offline or if network credentials fail.

Quickstart

Installation

npm install fast-jev-agents
export TYPESAFE_API_KEY=your_typesafe_key

Universal Compaction (compactAgent)

compactAgent automatically detects whether the input format belongs to Claude, Codex, Antigravity, Gemini, or OpenCode:

import { compactAgent } from 'fast-jev-agents';

const result = await compactAgent(transcript, {
  preserveRecentMessages: 4,
  truncateHeadChars: 300,
  truncateTailChars: 150,
});

console.log(`Detected Agent: ${result.agent}`);
console.log(`Compacted from ${result.stats.charsBefore} to ${result.stats.charsAfter} chars`);
console.log(`Reduction: ${((1 - result.stats.charsAfter / result.stats.charsBefore) * 100).toFixed(1)}%`);

Supported Coding Agents

1. Claude (Claude Code & Anthropic API)

Supports both Claude Code session transcripts (with handles) and Anthropic Messages API format:

import { compactClaude } from 'fast-jev-agents';

const anthropicMessages = [
  { role: 'user', content: [{ type: 'text', text: 'Fix the bug in parser.ts' }] },
  {
    role: 'assistant',
    content: [
      { type: 'tool_use', id: 'call_1', name: 'Read', input: { file_path: 'src/parser.ts' } }
    ]
  },
  {
    role: 'user',
    content: [
      { type: 'tool_result', tool_use_id: 'call_1', content: '...2000 lines of file content...' }
    ]
  },
  { role: 'assistant', content: [{ type: 'text', text: 'Updating logic now.' }] }
];

const { messages: compacted } = await compactClaude(anthropicMessages, {
  preserveRecentMessages: 2,
});

2. Codex & OpenAI (Chat Completions & Cursor)

Seamlessly handles OpenAI messages with tool_calls and role: 'tool':

import OpenAI from 'openai';
import { withCodexCompaction, compactCodex } from 'fast-jev-agents';

// Option A: Direct transcript compaction
const { messages: compactedHistory } = await compactCodex(openAiMessages);

// Option B: Transparent OpenAI client wrapper
const client = withCodexCompaction(new OpenAI(), {
  autoCompactThresholdChars: 50_000,
  preserveRecentMessages: 4,
});

const response = await client.chat.completions.create({
  model: 'gpt-4o',
  messages: longSessionMessages,
  tools: myAgentTools,
});

3. Google Antigravity (AGY Agent Transcripts & Sessions)

Designed for Google Antigravity agent workflows, IDE steps, and JSONL transcript logs:

import { compactAntigravity, compactAntigravityJsonl } from 'fast-jev-agents';

// Compact in-memory AGY transcript steps:
const { messages: compactedSteps } = await compactAntigravity(sessionSteps, {
  fallbackMode: 'local',
});

// Or compact an entire Antigravity JSONL file:
const compactedJsonl = await compactAntigravityJsonl(rawJsonlContent);

4. Google Gemini (Google Gen AI SDK)

Native support for Google Gen AI Content[] structure with functionCall and functionResponse parts:

import { compactGemini, withGeminiCompaction } from 'fast-jev-agents';

// Direct Content[] compaction:
const { messages: compactedContents } = await compactGemini(chatHistory);

// Or wrap an active Gemini ChatSession:
const chat = withGeminiCompaction(aiModel.startChat({ history }));

5. OpenCode & Open Interpreter

Supports OpenCode step arrays, event streams, and CLI execution logs:

import { compactOpenCode, compactOpenCodeSession } from 'fast-jev-agents';

const { messages: compactedEvents } = await compactOpenCode(events, {
  preserveRecentMessages: 3,
});

CLI Usage (fast-jev)

The fast-jev CLI provides instant context compaction directly from your terminal or shell scripts:

# 1. Compact any agent session file with automatic format detection
npx fast-jev session.json --stats

# 2. Pipe standard input to output with an explicit agent format
cat chat_history.json | npx fast-jev - --agent codex > compacted.json

# 3. Compact an Antigravity JSONL session log
npx fast-jev transcript.jsonl --agent antigravity -o compacted.jsonl --stats

# 4. Dry-run inspection (preview character savings without writing)
npx fast-jev transcript.json --agent gemini --dry-run

CLI Flags

Options:
  -a, --agent <name>     Agent format: auto (default), claude, codex, antigravity, gemini, opencode
  -o, --output <file>    Output destination file (defaults to stdout)
  -s, --stats            Print human-readable compaction metrics to stderr
  -d, --dry-run          Analyze and print statistics without writing output
  -k, --key <api-key>    TypeSafe API Key (or set TYPESAFE_API_KEY environment variable)
  -h, --help             Show help documentation

Options Reference

OptionTypeDefaultDescription
apiKeystringprocess.env.TYPESAFE_API_KEYTypeSafe API key for Jev
modelstring'jev-latest'Jev model identifier
baseUrlstring'https://api.typesafe.ai/v1/systemone'Endpoint URL
agentstring'auto'Target format: 'auto', 'claude', 'codex', 'antigravity', 'gemini', 'opencode', 'universal'
enableHeuristicsbooleantruePre-prunes superseded file reads & redundant searches locally
keepThresholdnumber0.5Minimum keep probability for a tool call or result to remain
preserveRecentMessagesnumber6Number of most recent turns pinned from compaction
truncateHeadCharsnumber300Characters of a dropped tool result retained at the beginning
truncateTailCharsnumber150Characters of a dropped tool result retained at the end (for error summaries)
concurrencynumber4Maximum parallel batch requests
retriesnumber2Number of retries on transient HTTP 429/5xx errors
timeoutMsnumber30000Request timeout per batch in milliseconds
fallbackMode`'throw' \'local'`'throw'Fallback behavior when Jev is unreachable ('local' runs rule-based compaction)
cacheCompactionCacheundefinedCache instance to memoize decisions across turns

Claude Code Plugin Setup

fast-jev-agents functions as a drop-in Claude Code plugin via function hooks (session.compact and turn.complete):

  1. Enable function hooks in ~/.claude/settings.json: ``json { "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1", "TYPESAFE_API_KEY": "your_api_key_here" } } ``
  2. Install the plugin: ``sh claude plugin marketplace add satiricalguru/fast-jev-agents claude plugin install fast-jev-agents@fast-jev-agents ``

Development & Contributing

# Clone repository
git clone https://github.com/satiricalguru/Fast-Jev-Agents.git
cd Fast-Jev-Agents

# Install dependencies
npm install

# Run test suite across all 6 test files (50 unit & integration tests)
npm test

# Typecheck library, CLI, adapters, and Claude hooks
npm run typecheck

# Compile production bundle
npm run build

# Run live interactive demonstration
npm run demo

Contributors & Attribution

fast-jev-agents is built upon the foundational work created by Tamara Tran in tamaratran/fast-jev-compaction.

We extend sincere gratitude to the original contributors:


License

MIT License © 2025–2026. Free and open source for all developers and AI agent builders.

Source 6 files
hooks/fast-jev.ts 311 lines
1import type {
2  On,
3  PluginOptions,
4  Register,
5  SessionMessage,
6  ToolResultSummary,
7  ToolUseSummary,
8  TurnCompleteInput,
9} from 'claude-code';
10
11import { compact, reductionRatio, resolveOptions } from '../src/compact.js';
12import { buildJevRequest, DEFAULT_MODEL, parseJevResponse } from '../src/request.js';
13import type {
14  CompactOptions,
15  CompactResult,
16  JevAsker,
17  Message,
18  ToolResult,
19  ToolUse,
20} from '../src/types.js';
21
22const HOOK_DEFAULTS = {
23  compactAtPercent: 60,
24  minReductionRatio: 0.25,
25  model: DEFAULT_MODEL,
26};
27
28export type HookFetchInit = {
29  method?: string;
30  headers?: Record<string, string>;
31  body?: string;
32};
33
34export type HookFetchResponse = {
35  status: number;
36  ok: boolean;
37  text: string;
38};
39
40/** The shape of `$.http.fetch`, so the hook can be driven without an engine. */
41export type HookFetch = (url: string, init?: HookFetchInit) => Promise<HookFetchResponse>;
42
43export type HookConfig = CompactOptions & {
44  apiKey?: string;
45  compactAtPercent: number;
46  minReductionRatio: number;
47  model: string;
48};
49
50function optionNumber(options: PluginOptions, key: string, fallback: number): number {
51  const value = options[key];
52  return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
53}
54
55function optionString(options: PluginOptions, key: string): string | undefined {
56  const value = options[key];
57  return typeof value === 'string' && value.length > 0 ? value : undefined;
58}
59
60/** Reads the plugin's `userConfig` values; anything missing takes the defaults. */
61export function resolveHookConfig(options: PluginOptions): HookConfig {
62  const numbers: Partial<Omit<CompactOptions, 'goal'>> = {};
63  for (const key of [
64    'keepThreshold',
65    'preserveRecentMessages',
66    'maxStateTokens',
67    'maxRequestTokens',
68    'truncateHeadChars',
69  ] as const) {
70    const value = options[key];
71    if (typeof value === 'number' && Number.isFinite(value)) numbers[key] = value;
72  }
73  const config: HookConfig = {
74    ...numbers,
75    compactAtPercent: optionNumber(options, 'compactAtPercent', HOOK_DEFAULTS.compactAtPercent),
76    minReductionRatio: optionNumber(
77      options,
78      'minReductionRatio',
79      HOOK_DEFAULTS.minReductionRatio,
80    ),
81    model: optionString(options, 'model') ?? HOOK_DEFAULTS.model,
82  };
83  const apiKey = optionString(options, 'apiKey');
84  if (apiKey) config.apiKey = apiKey;
85  const goal = optionString(options, 'goal');
86  if (goal) config.goal = goal;
87  return config;
88}
89
90/** A `JevAsker` over the engine's `$.http.fetch`. */
91export function jevAsker(fetchFn: HookFetch, apiKey: string, model: string): JevAsker {
92  return {
93    async ask(state, questions) {
94      const request = buildJevRequest({ apiKey, model }, state, questions);
95      const response = await fetchFn(request.url, {
96        method: request.method,
97        headers: request.headers,
98        body: request.body,
99      });
100      return parseJevResponse(response.status, response.ok, response.text);
101    },
102  };
103}
104
105function toolUseSummary(tool: ToolUse): ToolUseSummary {
106  const summary: ToolUseSummary = {
107    tool_use_id: tool.tool_use_id,
108    tool: tool.tool,
109    input: tool.input,
110  };
111  if (tool.text !== undefined) summary.text = tool.text;
112  if (tool.isError) summary.isError = true;
113  return summary;
114}
115
116function toolResultSummary(result: ToolResult): ToolResultSummary {
117  return {
118    tool_use_id: result.tool_use_id,
119    text: result.text,
120    isError: result.isError ?? false,
121  };
122}
123
124/**
125 * Maps the library's output back onto session messages. Whatever came back
126 * unchanged (a message, a tool use, a tool result) is the engine's own object,
127 * handle included; anything rebuilt is a fresh message without a handle, so the
128 * engine takes the edited content instead of its original.
129 */
130export function toSessionMessages(
131  input: readonly SessionMessage[],
132  output: readonly Message[],
133): SessionMessage[] {
134  const messages = new Map<Message, SessionMessage>();
135  const uses = new Map<ToolUse, ToolUseSummary>();
136  const results = new Map<ToolResult, ToolResultSummary>();
137  for (const message of input) {
138    messages.set(message, message);
139    for (const tool of message.toolUses) uses.set(tool, tool);
140    for (const result of message.toolResults ?? []) results.set(result, result);
141  }
142  return output.map((message) => {
143    const own = messages.get(message);
144    if (own) return own;
145    const rebuilt: SessionMessage = {
146      role: message.role,
147      text: message.text,
148      toolUses: message.toolUses.map((tool) => uses.get(tool) ?? toolUseSummary(tool)),
149    };
150    if (message.toolResults && message.toolResults.length > 0) {
151      rebuilt.toolResults = message.toolResults.map(
152        (result) => results.get(result) ?? toolResultSummary(result),
153      );
154    }
155    return rebuilt;
156  });
157}
158
159export type SessionCompaction = {
160  result: CompactResult;
161  messages: SessionMessage[];
162};
163
164/** Runs the library over a session transcript; throws when the key is missing or Jev fails. */
165export async function compactSession(
166  messages: readonly SessionMessage[],
167  config: HookConfig,
168  fetchFn: HookFetch,
169): Promise<SessionCompaction> {
170  if (!config.apiKey) throw new Error('TYPESAFE_API_KEY is not configured');
171  const result = await compact(messages, jevAsker(fetchFn, config.apiKey, config.model), config);
172  return { result, messages: toSessionMessages(messages, result.messages) };
173}
174
175function percent(ratio: number): string {
176  return `${Math.round(ratio * 100)}%`;
177}
178
179export function summarize(result: CompactResult): string {
180  const { stats } = result;
181  const parts = [
182    stats.kept > 0 ? `${stats.kept} kept` : '',
183    stats.resultsDropped > 0 ? `${stats.resultsDropped} results truncated` : '',
184    stats.callsDropped > 0 ? `${stats.callsDropped} call_dropped` : '',
185    stats.pinned > 0 ? `${stats.pinned} pinned` : '',
186  ].filter(Boolean);
187  return `${percent(reductionRatio(result))} reduction; ${
188    parts.join(', ') || 'no tool calls'
189  }; state ~${stats.stateTokens} tokens (${stats.stateStage}) in ${stats.requests} request(s)`;
190}
191
192const UI_LOG_MAX_CHARS = 4096;
193
194export function decisionLog(result: CompactResult): string {
195  return result.decisions
196    .filter((d) => d.reason !== 'pinned')
197    .map(
198      (d) =>
199        `${d.id}:${d.tool}:${d.action}/call=${d.keepCall.toFixed(2)}/result=${d.keepResult.toFixed(2)}`,
200    )
201    .join(' ');
202}
203
204export function decisionLogLines(
205  result: CompactResult,
206  maxChars: number = UI_LOG_MAX_CHARS,
207): string[] {
208  const entries = decisionLog(result).split(' ').filter(Boolean);
209  if (entries.length === 0) return ['decisions: (none)'];
210  const chunks: string[] = [];
211  let current = '';
212  for (const entry of entries) {
213    const next = current ? `${current} ${entry}` : entry;
214    if (current && next.length > maxChars - 24) {
215      chunks.push(current);
216      current = entry;
217    } else current = next;
218  }
219  chunks.push(current);
220  return chunks.map((chunk, index) =>
221    chunks.length === 1
222      ? `decisions: ${chunk}`
223      : `decisions (${index + 1}/${chunks.length}): ${chunk}`,
224  );
225}
226
227async function getApiKey(
228  $: {
229    env: { get: (name: string) => Promise<string | undefined> };
230    settings: { read: () => Promise<Readonly<Record<string, unknown>>> };
231  },
232  config: HookConfig,
233): Promise<string | undefined> {
234  if (config.apiKey) return config.apiKey;
235  const fromEnv = await $.env.get('TYPESAFE_API_KEY');
236  if (fromEnv) return fromEnv;
237  const settings = await $.settings.read();
238  const env = settings['env'];
239  if (env && typeof env === 'object') {
240    const value = (env as Record<string, unknown>)['TYPESAFE_API_KEY'];
241    if (typeof value === 'string' && value) return value;
242  }
243  return undefined;
244}
245
246function notify(
247  $: {
248    ui: {
249      log: (text: string) => void;
250      toast: (text: string, options?: { timeoutMs?: number }) => void;
251    };
252  },
253  text: string,
254): void {
255  $.ui.log(text);
256  $.ui.toast(text, { timeoutMs: 15_000 });
257}
258
259export const register: Register = (on: On, options: PluginOptions) => {
260  const configured = resolveHookConfig(options);
261  let compacting = false;
262
263  on('session.compact', async ($, event, next) => {
264    try {
265      const config = { ...configured, apiKey: await getApiKey($, configured) };
266      const { result, messages } = await compactSession(event.messages, config, async (url, init) => {
267        const response = await $.http.fetch(url, init);
268        return { status: response.status, ok: response.ok, text: response.text };
269      });
270      for (const line of decisionLogLines(result)) $.ui.log(line);
271      if (reductionRatio(result) < config.minReductionRatio) {
272        notify(
273          $,
274          `fallback to built-in summary (below ${percent(config.minReductionRatio)} minimum: ${summarize(result)})`,
275        );
276        return next(event);
277      }
278      notify(
279        $,
280        `kept ${messages.length}/${event.messages.length} messages, no summary (${summarize(result)})`,
281      );
282      return { messages };
283    } catch (error) {
284      notify(
285        $,
286        `fallback to built-in summary (${error instanceof Error ? error.message : String(error)})`,
287      );
288      return next(event);
289    }
290  });
291
292  on('turn.complete', async ($, event: TurnCompleteInput, next) => {
293    if (compacting) return next(event);
294    try {
295      const { context } = await $.session.usage();
296      if ((context.percent ?? 0) < configured.compactAtPercent) return next(event);
297      compacting = true;
298      await $.session.compact();
299    } catch (error) {
300      $.ui.log(
301        `auto-compact skipped (${error instanceof Error ? error.message : String(error)})`,
302      );
303    } finally {
304      compacting = false;
305    }
306    return next(event);
307  });
308};
309
310export { resolveOptions };
311
src/compact.ts 423 lines
1import { analyzeHeuristics, smartTruncateResultText } from './heuristics.js';
2import { noulAnswer } from './request.js';
3import { collectToolCalls, estimateTokens, fitState } from './state.js';
4import type {
5  CallAnswer,
6  CallDecision,
7  CompactionCache,
8  CompactOptions,
9  CompactResult,
10  CompactionState,
11  JevAsker,
12  JevQuestions,
13  Message,
14  ResolvedCompactOptions,
15  ToolCall,
16  ToolUse,
17} from './types.js';
18
19export const DEFAULT_OPTIONS: ResolvedCompactOptions = {
20  goal: '',
21  keepThreshold: 0.5,
22  preserveRecentMessages: 6,
23  maxStateTokens: 25_000,
24  maxRequestTokens: 30_000,
25  truncateHeadChars: 300,
26  truncateTailChars: 0,
27  enableHeuristics: true,
28  fallbackMode: 'throw',
29  concurrency: 4,
30  timeoutMs: 30_000,
31  retries: 2,
32  cache: undefined,
33};
34
35/** Tokens the request envelope (`model`, key names) adds around state and questions. */
36const REQUEST_OVERHEAD_TOKENS = 20;
37
38function finite(value: number | undefined, fallback: number): number {
39  return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
40}
41
42export function resolveOptions(options: CompactOptions = {}): ResolvedCompactOptions {
43  let cache: CompactionCache | undefined;
44  if (options.cache && typeof options.cache === 'object') {
45    cache = options.cache;
46  }
47
48  return {
49    goal: options.goal ?? DEFAULT_OPTIONS.goal,
50    keepThreshold: finite(options.keepThreshold, DEFAULT_OPTIONS.keepThreshold),
51    preserveRecentMessages: Math.max(
52      0,
53      Math.floor(
54        finite(options.preserveRecentMessages, DEFAULT_OPTIONS.preserveRecentMessages),
55      ),
56    ),
57    maxStateTokens: Math.max(1, finite(options.maxStateTokens, DEFAULT_OPTIONS.maxStateTokens)),
58    maxRequestTokens: Math.max(
59      1,
60      finite(options.maxRequestTokens, DEFAULT_OPTIONS.maxRequestTokens),
61    ),
62    truncateHeadChars: Math.max(
63      0,
64      Math.floor(finite(options.truncateHeadChars, DEFAULT_OPTIONS.truncateHeadChars)),
65    ),
66    truncateTailChars: Math.max(
67      0,
68      Math.floor(finite(options.truncateTailChars, DEFAULT_OPTIONS.truncateTailChars)),
69    ),
70    enableHeuristics: options.enableHeuristics ?? DEFAULT_OPTIONS.enableHeuristics,
71    fallbackMode: options.fallbackMode ?? DEFAULT_OPTIONS.fallbackMode,
72    concurrency: Math.max(1, finite(options.concurrency, DEFAULT_OPTIONS.concurrency)),
73    timeoutMs: Math.max(100, finite(options.timeoutMs, DEFAULT_OPTIONS.timeoutMs)),
74    retries: Math.max(0, finite(options.retries, DEFAULT_OPTIONS.retries)),
75    cache,
76  };
77}
78
79/** The two `noul` questions asked about one call: keep the call, keep its result. */
80export function questionsFor(call: ToolCall): JevQuestions {
81  return {
82    [`call_${call.id}`]: {
83      type: 'noul',
84      instructions: `Tool call ${call.id} (${call.tool}) should stay in the history: knowing this call was made, with its input, still matters for what the assistant does next`,
85    },
86    [`result_${call.id}`]: {
87      type: 'noul',
88      instructions: `The full output of tool call ${call.id} (${call.tool}, ${call.resultChars} chars) should stay in the history verbatim: the assistant still needs its contents and re-running the tool would not do`,
89    },
90  };
91}
92
93/**
94 * Splits the candidate calls into batches whose questions, together with the
95 * (always complete) state, fit one request.
96 */
97export function batchCalls(
98  calls: readonly ToolCall[],
99  stateTokens: number,
100  options: Pick<ResolvedCompactOptions, 'maxRequestTokens'>,
101): ToolCall[][] {
102  const budget = options.maxRequestTokens - stateTokens - REQUEST_OVERHEAD_TOKENS;
103  const batches: ToolCall[][] = [];
104  let current: ToolCall[] = [];
105  let currentTokens = 0;
106  for (const call of calls) {
107    const tokens = estimateTokens(JSON.stringify(questionsFor(call)));
108    if (current.length > 0 && currentTokens + tokens > budget) {
109      batches.push(current);
110      current = [];
111      currentTokens = 0;
112    }
113    if (current.length === 0 && tokens > budget) {
114      throw new Error(
115        `state leaves no room for questions (~${stateTokens} of ${options.maxRequestTokens} tokens)`,
116      );
117    }
118    current.push(call);
119    currentTokens += tokens;
120  }
121  if (current.length > 0) batches.push(current);
122  return batches;
123}
124
125export function decideCall(
126  call: Pick<ToolCall, 'id' | 'tool' | 'pinned'>,
127  answer: CallAnswer,
128  options: Pick<ResolvedCompactOptions, 'keepThreshold'>,
129): CallDecision {
130  const base = { id: call.id, tool: call.tool, ...answer };
131  if (call.pinned) return { ...base, action: 'keep', reason: 'pinned' };
132  if (answer.keepResult >= options.keepThreshold) {
133    return { ...base, action: 'keep', reason: 'kept' };
134  }
135  if (answer.keepCall >= options.keepThreshold) {
136    return { ...base, action: 'drop_result', reason: 'result_dropped' };
137  }
138  return { ...base, action: 'drop_call', reason: 'call_dropped' };
139}
140
141async function askBatch(
142  asker: JevAsker,
143  state: CompactionState,
144  batch: readonly ToolCall[],
145): Promise<Map<string, CallAnswer>> {
146  const questions: JevQuestions = Object.assign({}, ...batch.map(questionsFor));
147  const { answers } = await asker.ask(state, questions);
148  return new Map(
149    batch.map((call) => [
150      call.id,
151      {
152        keepCall: noulAnswer(answers, `call_${call.id}`),
153        keepResult: noulAnswer(answers, `result_${call.id}`),
154      },
155    ]),
156  );
157}
158
159/** Runs tasks with a maximum concurrency limit. */
160async function runWithConcurrency<T, R>(
161  items: readonly T[],
162  limit: number,
163  fn: (item: T) => Promise<R>,
164): Promise<R[]> {
165  if (items.length === 0) return [];
166  const results: R[] = new Array(items.length);
167  let currentIndex = 0;
168
169  const workers = Array.from({ length: Math.min(limit, items.length) }, async () => {
170    while (currentIndex < items.length) {
171      const idx = currentIndex++;
172      results[idx] = await fn(items[idx]!);
173    }
174  });
175
176  await Promise.all(workers);
177  return results;
178}
179
180function truncatedResultText(
181  text: string,
182  isError: boolean,
183  headChars: number,
184  tailChars: number = 0,
185): string {
186  return smartTruncateResultText(text, isError, headChars, tailChars);
187}
188
189/**
190 * Rebuilds the conversation from the decisions. A dropped call disappears
191 * together with its result; a dropped result keeps a bounded head and note.
192 * Messages that lose all their content are removed; untouched messages are
193 * returned as the same objects they came in as.
194 */
195export function applyDecisions(
196  messages: readonly Message[],
197  decisions: readonly CallDecision[],
198  calls: readonly ToolCall[],
199  headChars: number,
200  tailChars: number = 0,
201): Message[] {
202  const byId = new Map(calls.map((call) => [call.id, call]));
203  const actions = new Map<string, CallDecision['action']>();
204  for (const decision of decisions) {
205    const call = byId.get(decision.id);
206    if (call && decision.action !== 'keep') actions.set(call.tool_use_id, decision.action);
207  }
208  const kept: Message[] = [];
209  for (const message of messages) {
210    const touched =
211      message.toolUses.some((tool) => actions.has(tool.tool_use_id)) ||
212      (message.toolResults ?? []).some((result) => actions.has(result.tool_use_id));
213    if (!touched) {
214      kept.push(message);
215      continue;
216    }
217    const toolUses = message.toolUses
218      .filter((tool) => actions.get(tool.tool_use_id) !== 'drop_call')
219      .map((tool) => {
220        if (actions.get(tool.tool_use_id) !== 'drop_result') return tool;
221        const text = truncatedResultText(
222          tool.text ?? '',
223          tool.isError ?? false,
224          headChars,
225          tailChars,
226        );
227        if ((tool.text ?? '') === text) return tool;
228        const copy: ToolUse = {
229          tool_use_id: tool.tool_use_id,
230          tool: tool.tool,
231          input: tool.input,
232          text,
233        };
234        if (tool.isError) copy.isError = true;
235        return copy;
236      });
237    const toolResults = (message.toolResults ?? [])
238      .filter((result) => actions.get(result.tool_use_id) !== 'drop_call')
239      .map((result) => {
240        if (actions.get(result.tool_use_id) !== 'drop_result') return result;
241        const text = truncatedResultText(result.text, result.isError ?? false, headChars, tailChars);
242        return text === result.text
243          ? result
244          : {
245              tool_use_id: result.tool_use_id,
246              text,
247              isError: result.isError,
248            };
249      });
250    if (
251      !message.toolUses.some(
252        (tool) => actions.get(tool.tool_use_id) === 'drop_call',
253      ) &&
254      !(message.toolResults ?? []).some(
255        (result) => actions.get(result.tool_use_id) === 'drop_call',
256      ) &&
257      toolUses.every((tool, index) => tool === message.toolUses[index]) &&
258      toolResults.every(
259        (result, index) => result === message.toolResults?.[index],
260      )
261    ) {
262      kept.push(message);
263      continue;
264    }
265    if (message.text.trim().length === 0 && toolUses.length === 0 && toolResults.length === 0) {
266      continue;
267    }
268    const rebuilt: Message = { role: message.role, text: message.text, toolUses };
269    if (toolResults.length > 0) rebuilt.toolResults = toolResults;
270    kept.push(rebuilt);
271  }
272  return kept;
273}
274
275/** Characters of text, tool input and tool output a message holds. */
276export function messageChars(message: Message): number {
277  let total = message.text.length;
278  for (const tool of message.toolUses) {
279    try {
280      total += JSON.stringify(tool.input).length;
281    } catch {
282      total += 20;
283    }
284  }
285  for (const result of message.toolResults ?? []) total += result.text.length;
286  return total;
287}
288
289export function reductionRatio(result: Pick<CompactResult, 'stats'>): number {
290  const { charsBefore, charsAfter } = result.stats;
291  return charsBefore === 0 ? 0 : (charsBefore - charsAfter) / charsBefore;
292}
293
294function count(decisions: readonly CallDecision[], reason: CallDecision['reason']): number {
295  return decisions.filter((decision) => decision.reason === reason).length;
296}
297
298/**
299 * Compacts a transcript by asking Jev, for every tool call outside the pinned
300 * first and newest messages, whether the call and whether its result must
301 * stay. The whole history (results omitted, fitted into `maxStateTokens`) is
302 * sent as state with every batch of questions. Throws when Jev fails or the
303 * history cannot be fitted; the caller decides whether to fall back.
304 */
305export async function compact(
306  messages: readonly Message[],
307  asker: JevAsker,
308  options: CompactOptions = {},
309): Promise<CompactResult> {
310  const started = Date.now();
311  const resolved = resolveOptions(options);
312  const calls = collectToolCalls(messages, resolved.preserveRecentMessages);
313  const candidates = calls.filter((call) => !call.pinned);
314  const charsBefore = messages.reduce((sum, message) => sum + messageChars(message), 0);
315
316  let fitted: { tokens: number; stage: string } = { tokens: 0, stage: '' };
317  let batches: ToolCall[][] = [];
318  const answers = new Map<string, CallAnswer>();
319  let heuristicsPruned = 0;
320  let cacheHits = 0;
321
322  // Step 1: Pre-compaction heuristics (if enabled)
323  let candidatesToAsk = candidates;
324  if (resolved.enableHeuristics && candidates.length > 0) {
325    const analysis = analyzeHeuristics(calls);
326    for (const [id, answer] of analysis.decisions) {
327      if (candidates.some((c) => c.id === id)) {
328        answers.set(id, answer);
329        heuristicsPruned++;
330      }
331    }
332    candidatesToAsk = candidates.filter((c) => !answers.has(c.id));
333  }
334
335  // Step 2: Decision Cache lookup
336  if (resolved.cache && candidatesToAsk.length > 0) {
337    const remaining: ToolCall[] = [];
338    for (const call of candidatesToAsk) {
339      const cacheKey = `${call.tool}:${JSON.stringify(call.input)}`;
340      const cached = resolved.cache.get(cacheKey);
341      if (cached) {
342        answers.set(call.id, cached);
343        cacheHits++;
344      } else {
345        remaining.push(call);
346      }
347    }
348    candidatesToAsk = remaining;
349  }
350
351  // Step 3: Query Jev for remaining candidates
352  if (candidatesToAsk.length > 0) {
353    try {
354      const state = fitState(messages, calls, resolved);
355      fitted = state;
356      batches = batchCalls(candidatesToAsk, state.tokens, resolved);
357
358      const answeredMaps = await runWithConcurrency(
359        batches,
360        resolved.concurrency,
361        (batch) => askBatch(asker, state.state, batch),
362      );
363
364      for (const map of answeredMaps) {
365        for (const [id, answer] of map) {
366          answers.set(id, answer);
367          if (resolved.cache) {
368            const call = candidates.find((c) => c.id === id);
369            if (call) {
370              resolved.cache.set(`${call.tool}:${JSON.stringify(call.input)}`, answer);
371            }
372          }
373        }
374      }
375    } catch (error) {
376      if (resolved.fallbackMode === 'local') {
377        fitted.stage = 'local_fallback';
378        for (const call of candidatesToAsk) {
379          answers.set(call.id, {
380            keepCall: 0.9,
381            keepResult: 0.1,
382          });
383        }
384      } else {
385        throw error;
386      }
387    }
388  }
389
390  const decisions = calls.map((call) =>
391    decideCall(call, answers.get(call.id) ?? { keepCall: 1, keepResult: 1 }, resolved),
392  );
393  const kept = applyDecisions(
394    messages,
395    decisions,
396    calls,
397    resolved.truncateHeadChars,
398    resolved.truncateTailChars,
399  );
400
401  return {
402    messages: kept,
403    decisions,
404    stats: {
405      messagesBefore: messages.length,
406      messagesAfter: kept.length,
407      charsBefore,
408      charsAfter: kept.reduce((sum, message) => sum + messageChars(message), 0),
409      calls: calls.length,
410      kept: count(decisions, 'kept'),
411      resultsDropped: count(decisions, 'result_dropped'),
412      callsDropped: count(decisions, 'call_dropped'),
413      pinned: count(decisions, 'pinned'),
414      heuristicsPruned,
415      cacheHits,
416      stateTokens: fitted.tokens,
417      stateStage: fitted.stage,
418      requests: batches.length,
419      ms: Date.now() - started,
420    },
421  };
422}
423
src/request.ts 81 lines
1import type { JevAnswer, JevQuestions, JevResponse, JevState } from './types.js';
2
3export const SYSTEM_ONE_URL = 'https://api.typesafe.ai/v1/systemone';
4export const DEFAULT_MODEL = 'jev-latest';
5
6export interface JevRequest {
7  url: string;
8  method: 'POST';
9  headers: Record<string, string>;
10  body: string;
11}
12
13/** The HTTP request for one Jev call, for any fetch-like transport. */
14export function buildJevRequest(
15  params: {
16    apiKey: string;
17    model?: string;
18    baseUrl?: string;
19  },
20  state: JevState,
21  questions: JevQuestions,
22): JevRequest {
23  return {
24    url: params.baseUrl ?? SYSTEM_ONE_URL,
25    method: 'POST',
26    headers: {
27      authorization: `Bearer ${params.apiKey}`,
28      'content-type': 'application/json',
29    },
30    body: JSON.stringify({
31      model: params.model ?? DEFAULT_MODEL,
32      state,
33      questions,
34    }),
35  };
36}
37
38/** Validates a Jev response body; throws on anything but an `answers` object. */
39export function parseJevResponse(
40  status: number,
41  ok: boolean,
42  text: string,
43): JevResponse {
44  if (!ok) {
45    throw new Error(`Jev request failed (${status}): ${text.slice(0, 200)}`);
46  }
47  let parsed: unknown;
48  try {
49    parsed = JSON.parse(text);
50  } catch {
51    throw new Error('Jev returned malformed JSON');
52  }
53  if (
54    parsed === null ||
55    typeof parsed !== 'object' ||
56    !('answers' in parsed) ||
57    parsed.answers === null ||
58    typeof parsed.answers !== 'object'
59  ) {
60    throw new Error('Jev response is missing answers');
61  }
62  return parsed as JevResponse;
63}
64
65/** The `noul` probability of one answer; throws when it is not there. */
66export function noulAnswer(
67  answers: Record<string, JevAnswer>,
68  name: string,
69): number {
70  const answer = answers[name];
71  if (
72    !answer ||
73    !('noul' in answer) ||
74    typeof answer.noul !== 'number' ||
75    !Number.isFinite(answer.noul)
76  ) {
77    throw new Error(`Invalid Jev answer for ${name}`);
78  }
79  return answer.noul;
80}
81
src/types.ts 235 lines
1export type Role = 'user' | 'assistant';
2
3/**
4 * A tool_use block of an assistant message. `text` and `isError` mirror the
5 * outcome once the transcript holds it (Claude Code attaches them).
6 */
7export interface ToolUse {
8  tool_use_id: string;
9  tool: string;
10  input: Record<string, unknown>;
11  text?: string;
12  isError?: boolean;
13}
14
15/** A tool_result block of a user message. */
16export interface ToolResult {
17  tool_use_id: string;
18  text: string;
19  isError?: boolean;
20}
21
22/**
23 * One transcript message. The shape is a subset of Claude Code's
24 * `SessionMessage`, so a session transcript can be passed in as is.
25 */
26export interface Message {
27  role: Role;
28  text: string;
29  toolUses: ToolUse[];
30  toolResults?: ToolResult[];
31}
32
33/** A tool call paired with its result by `tool_use_id`. */
34export interface ToolCall {
35  /** Short id used in the Jev state and question names (`t1`, `t2`, ...). */
36  id: string;
37  tool_use_id: string;
38  tool: string;
39  input: Record<string, unknown>;
40  /** Index of the message holding the tool_use block. */
41  callIndex: number;
42  /** Index of the message holding the tool_result block. */
43  resultIndex: number;
44  resultChars: number;
45  isError: boolean;
46  /** In the first or the newest preserved messages; never a candidate. */
47  pinned: boolean;
48}
49
50export interface CallAnswer {
51  /** Jev's probability that the call itself still matters. */
52  keepCall: number;
53  /** Jev's probability that the full result still needs to stay verbatim. */
54  keepResult: number;
55}
56
57export type CallAction = 'keep' | 'drop_result' | 'drop_call';
58
59export interface CallDecision extends CallAnswer {
60  id: string;
61  tool: string;
62  action: CallAction;
63  reason: 'pinned' | 'kept' | 'result_dropped' | 'call_dropped';
64}
65
66export interface HistoryToolCall {
67  id: string;
68  tool: string;
69  input: string;
70  result: string;
71}
72
73export interface HistoryEntry {
74  i: number;
75  role: Role;
76  text: string;
77  /** Structured per call, or one compact line per call once the state has to shrink. */
78  tool_calls?: HistoryToolCall[] | string[];
79}
80
81/** The state sent with every Jev request: the whole history, results omitted. */
82export interface CompactionState {
83  context: string;
84  goal: string;
85  history: HistoryEntry[];
86}
87
88export interface FittedState {
89  state: CompactionState;
90  tokens: number;
91  /** Which fitting stage produced the state, for diagnostics. */
92  stage: string;
93}
94
95export type SupportedAgent = 'claude' | 'codex' | 'antigravity' | 'gemini' | 'opencode' | 'universal';
96
97export interface CompactionCache {
98  get(key: string): CallAnswer | undefined;
99  set(key: string, answer: CallAnswer): void;
100  has(key: string): boolean;
101  clear(): void;
102}
103
104export interface CompactOptions {
105  /** Ongoing task description; defaults to the last few user prompts. */
106  goal?: string;
107  /** Minimum keep probability for a call or result to stay. Default 0.5. */
108  keepThreshold?: number;
109  /** Newest messages never touched (the first message is always kept). Default 6. */
110  preserveRecentMessages?: number;
111  /** Estimated token ceiling for the state. Default 25000. */
112  maxStateTokens?: number;
113  /** Estimated token ceiling for state plus one batch of questions. Default 30000. */
114  maxRequestTokens?: number;
115  /** Characters of a dropped tool result to retain at the head. Default 300. */
116  truncateHeadChars?: number;
117  /** Characters of a dropped tool result to retain at the tail (e.g. error summaries). Default 150. */
118  truncateTailChars?: number;
119  /** Whether to enable heuristic pre-compaction (pruning superseded reads & duplicate searches). Default true. */
120  enableHeuristics?: boolean;
121  /** What to do if Jev fails or is unconfigured: 'local' (rule-based compaction fallback) or 'throw'. Default 'throw' for strict mode. */
122  fallbackMode?: 'local' | 'throw';
123  /** Max concurrent question batch requests. Default 4. */
124  concurrency?: number;
125  /** Request timeout in ms. Default 30000. */
126  timeoutMs?: number;
127  /** Number of retry attempts on network/429 failures. Default 2. */
128  retries?: number;
129  /** Decision cache instance or true for default in-memory cache. */
130  cache?: boolean | CompactionCache;
131}
132
133export interface ResolvedCompactOptions {
134  goal: string;
135  keepThreshold: number;
136  preserveRecentMessages: number;
137  maxStateTokens: number;
138  maxRequestTokens: number;
139  truncateHeadChars: number;
140  truncateTailChars: number;
141  enableHeuristics: boolean;
142  fallbackMode: 'local' | 'throw';
143  concurrency: number;
144  timeoutMs: number;
145  retries: number;
146  cache?: CompactionCache;
147}
148
149export interface CompactResult {
150  /** The compacted transcript; untouched messages are the input objects. */
151  messages: Message[];
152  decisions: CallDecision[];
153  stats: {
154    messagesBefore: number;
155    messagesAfter: number;
156    charsBefore: number;
157    charsAfter: number;
158    calls: number;
159    kept: number;
160    resultsDropped: number;
161    callsDropped: number;
162    pinned: number;
163    heuristicsPruned: number;
164    cacheHits: number;
165    stateTokens: number;
166    /** Which fitting stage the state needed, '' when no request was made. */
167    stateStage: string;
168    requests: number;
169    ms: number;
170  };
171}
172
173/** The `state` of a Jev request: a string or any JSON-serialisable object. */
174export type JevState = string | object;
175
176export interface NoulQuestion {
177  type: 'noul';
178  instructions: string;
179  criteria?: {
180    true?: string;
181    false?: string;
182  };
183}
184
185export interface ChoiceQuestion {
186  type: 'choice';
187  instructions: string;
188  criteria: Record<string, string | null>;
189}
190
191export interface ScoreQuestion {
192  type: 'score';
193  instructions: string;
194  criteria: string[];
195}
196
197export type JevQuestion = NoulQuestion | ChoiceQuestion | ScoreQuestion;
198export type JevQuestions = Record<string, JevQuestion>;
199
200export interface NoulAnswer {
201  type?: 'noul';
202  noul: number;
203}
204
205export interface ChoiceAnswer {
206  type?: 'choice';
207  choice: string;
208  confidence: number;
209  probabilities: Record<string, number>;
210}
211
212export interface ScoreAnswer {
213  type?: 'score';
214  score: number;
215  confidence: number;
216  probabilities: Record<string, number>;
217}
218
219export type JevAnswer = NoulAnswer | ChoiceAnswer | ScoreAnswer;
220
221export interface JevResponse {
222  model?: string;
223  answers: Record<string, JevAnswer>;
224  usage?: {
225    input_tokens?: number;
226    output_tokens?: number;
227  };
228  [key: string]: unknown;
229}
230
231/** Anything that can answer Jev questions: `JevClient`, or a host-provided adapter. */
232export interface JevAsker {
233  ask(state: JevState, questions: JevQuestions): Promise<JevResponse>;
234}
235
src/heuristics.ts 144 lines
1import type { CallAnswer, ToolCall } from './types.js';
2
3export interface HeuristicAnalysisResult {
4  /** Map of tool call ID -> predetermined decision */
5  decisions: Map<string, CallAnswer>;
6  /** Set of tool call IDs whose full results were marked as obsolete */
7  supersededResults: Set<string>;
8}
9
10const READ_TOOLS = new Set([
11  'read',
12  'read_file',
13  'view_file',
14  'readfile',
15  'cat',
16  'open_file',
17]);
18
19const WRITE_TOOLS = new Set([
20  'edit',
21  'edit_file',
22  'write',
23  'write_file',
24  'write_to_file',
25  'replace_file_content',
26  'multi_replace_file_content',
27  'create_file',
28  'patch',
29]);
30
31const SEARCH_TOOLS = new Set([
32  'glob',
33  'grep',
34  'grep_search',
35  'find_by_name',
36  'search_files',
37  'list_dir',
38  'ls',
39]);
40
41/**
42 * Normalizes a file path extracted from tool input for comparison.
43 */
44function extractFilePath(input: Record<string, unknown>): string | undefined {
45  for (const key of ['file_path', 'filePath', 'path', 'AbsolutePath', 'targetFile', 'TargetFile', 'filename']) {
46    const val = input[key];
47    if (typeof val === 'string' && val.trim()) {
48      return val.trim().replace(/\\/g, '/');
49    }
50  }
51  return undefined;
52}
53
54/**
55 * Analyzes tool calls across the transcript to identify provably obsolete
56 * or redundant calls before querying Jev.
57 *
58 * Examples:
59 * 1. A file was read at turn 2, then edited at turn 5. The turn 2 read result is obsolete.
60 * 2. A file was read at turn 2, then read again at turn 6. The turn 2 read result is superseded.
61 * 3. Consecutive search/glob queries where a narrower search followed immediately.
62 */
63export function analyzeHeuristics(calls: readonly ToolCall[]): HeuristicAnalysisResult {
64  const decisions = new Map<string, CallAnswer>();
65  const supersededResults = new Set<string>();
66
67  // Map from normalized file path -> list of tool calls operating on that file
68  const fileOperations = new Map<string, { call: ToolCall; type: 'read' | 'write' }[]>();
69
70  for (const call of calls) {
71    const toolLower = call.tool.toLowerCase();
72    const filePath = extractFilePath(call.input);
73
74    if (filePath) {
75      const isRead = READ_TOOLS.has(toolLower);
76      const isWrite = WRITE_TOOLS.has(toolLower);
77
78      if (isRead || isWrite) {
79        const ops = fileOperations.get(filePath) ?? [];
80        ops.push({ call, type: isRead ? 'read' : 'write' });
81        fileOperations.set(filePath, ops);
82      }
83    }
84  }
85
86  // Check file read supersession:
87  // If an unpinned read is followed by another read or a write on the same file,
88  // its result is no longer needed (the call itself matters, but the verbatim file content is old).
89  for (const [, ops] of fileOperations) {
90    for (let i = 0; i < ops.length - 1; i++) {
91      const current = ops[i]!;
92      if (current.type === 'read' && !current.call.pinned) {
93        // There is a subsequent read or write to this file
94        supersededResults.add(current.call.id);
95        decisions.set(current.call.id, {
96          keepCall: 0.95, // The call happened and is relevant context
97          keepResult: 0.05, // The full result is obsolete
98        });
99      }
100    }
101  }
102
103  // Check redundant searches:
104  // If an unpinned search is followed by another search within 2 calls of the same type,
105  // the earlier search result is usually superseded by the more specific search.
106  for (let i = 0; i < calls.length - 1; i++) {
107    const current = calls[i]!;
108    if (current.pinned || !SEARCH_TOOLS.has(current.tool.toLowerCase())) continue;
109
110    const next = calls[i + 1]!;
111    if (SEARCH_TOOLS.has(next.tool.toLowerCase()) && !next.pinned) {
112      if (!decisions.has(current.id)) {
113        supersededResults.add(current.id);
114        decisions.set(current.id, {
115          keepCall: 0.9,
116          keepResult: 0.1,
117        });
118      }
119    }
120  }
121
122  return { decisions, supersededResults };
123}
124
125/**
126 * Truncates a tool result keeping both a head and an optional tail (e.g. for error summaries).
127 */
128export function smartTruncateResultText(
129  text: string,
130  isError: boolean,
131  headChars: number,
132  tailChars: number = 0,
133): string {
134  if (text.length <= headChars + tailChars + 120) return text;
135
136  const head = headChars > 0 ? `${text.slice(0, headChars)}\n` : '';
137  const tail = tailChars > 0 ? `\n${text.slice(-tailChars)}` : '';
138  const omitted = text.length - headChars - tailChars;
139
140  return `${head}[fast-jev-compaction truncated ${omitted} chars of this tool result${
141    isError ? ' (error)' : ''
142  }; re-run the tool if needed]${tail}`;
143}
144
src/state.ts 345 lines
1import type {
2  CompactionState,
3  FittedState,
4  HistoryEntry,
5  Message,
6  ResolvedCompactOptions,
7  ToolCall,
8  ToolResult,
9} from './types.js';
10
11export const STATE_CONTEXT =
12  'A coding assistant conversation is being compacted to free context. `history` is the whole conversation so far, oldest first; tool outputs are replaced by a short `result` note and long texts may be abridged. Each question asks whether one tool call, or the full output of that call, still needs to stay in the history verbatim. Whatever is not kept is deleted permanently, but the assistant can always re-run a tool or re-read a file.';
13
14/** Successive caps on the serialised tool input included per call. */
15const INPUT_CHARS = [1000, 200, 60] as const;
16const TEXT_HEAD = 400;
17const TEXT_TAIL = 150;
18
19/**
20 * Fast zero-allocation token estimation: a word costs one token per six
21 * letters, a digit half a token, any other symbol nine tenths. Calibrated
22 * against the usage Jev reports for real transcripts.
23 * Zero regex allocations, O(N) single-pass scan.
24 */
25export function estimateTokens(text: string): number {
26  const len = text.length;
27  let tokens = 0;
28  let i = 0;
29
30  while (i < len) {
31    const code = text.charCodeAt(i);
32    // Whitespace: space (32), tab (9), newline (10), carriage return (13)
33    if (code <= 32) {
34      i++;
35      continue;
36    }
37
38    // Letters: A-Z (65-90), a-z (97-122)
39    if ((code >= 65 && code <= 90) || (code >= 97 && code <= 122)) {
40      const start = i;
41      i++;
42      while (i < len) {
43        const c = text.charCodeAt(i);
44        if ((c >= 65 && c <= 90) || (c >= 97 && c <= 122)) {
45          i++;
46        } else {
47          break;
48        }
49      }
50      tokens += 1 + Math.floor((i - start - 1) / 6);
51      continue;
52    }
53
54    // Digits: 0-9 (48-57)
55    if (code >= 48 && code <= 57) {
56      const start = i;
57      i++;
58      while (i < len) {
59        const c = text.charCodeAt(i);
60        if (c >= 48 && c <= 57) {
61          i++;
62        } else {
63          break;
64        }
65      }
66      tokens += (i - start) / 2;
67      continue;
68    }
69
70    // Any other symbol (not whitespace, not A-Za-z, not 0-9)
71    tokens += 0.9;
72    i++;
73  }
74
75  return Math.ceil(tokens);
76}
77
78export const fastEstimateTokens = estimateTokens;
79
80export function truncate(text: string, limit: number): string {
81  return text.length <= limit ? text : `${text.slice(0, Math.max(0, limit - 1))}…`;
82}
83
84function abridge(text: string, head: number, tail: number): string {
85  if (text.length <= head + tail + 40) return text;
86  const omitted = text.length - head - tail;
87  return `${text.slice(0, head)}\n[… ${omitted} chars omitted …]\n${text.slice(-tail)}`;
88}
89
90export function isPinned(
91  index: number,
92  total: number,
93  preserveRecentMessages: number,
94): boolean {
95  return index === 0 || index >= total - preserveRecentMessages;
96}
97
98/**
99 * Pairs every tool_use with its tool_result by `tool_use_id`. Calls without a
100 * result are not candidates (there is nothing to drop yet).
101 */
102export function collectToolCalls(
103  messages: readonly Message[],
104  preserveRecentMessages: number,
105): ToolCall[] {
106  const results = new Map<string, { index: number; result: ToolResult }>();
107  messages.forEach((message, index) => {
108    for (const result of message.toolResults ?? []) {
109      results.set(result.tool_use_id, { index, result });
110    }
111  });
112  const calls: ToolCall[] = [];
113  messages.forEach((message, callIndex) => {
114    for (const tool of message.toolUses) {
115      const found = results.get(tool.tool_use_id);
116      if (!found) continue;
117      calls.push({
118        id: `t${calls.length + 1}`,
119        tool_use_id: tool.tool_use_id,
120        tool: tool.tool,
121        input: tool.input,
122        callIndex,
123        resultIndex: found.index,
124        resultChars: found.result.text.length,
125        isError: found.result.isError ?? false,
126        pinned:
127          isPinned(callIndex, messages.length, preserveRecentMessages) ||
128          isPinned(found.index, messages.length, preserveRecentMessages),
129      });
130    }
131  });
132  return calls;
133}
134
135function inputText(input: Record<string, unknown>, limit: number): string {
136  let json = '';
137  try {
138    json = JSON.stringify(input);
139  } catch {
140    json = '[unserializable input]';
141  }
142  return truncate(json, limit);
143}
144
145function resultNote(call: ToolCall): string {
146  return `${call.isError ? 'error' : 'ok'}, ${call.resultChars} chars (omitted)`;
147}
148
149/** One call as a single line, for when the structured form is too costly. */
150function compactCall(call: ToolCall): string {
151  const input = Object.entries(call.input)
152    .map(([key, value]) => {
153      const text = typeof value === 'string' ? value : inputText({ [key]: value }, 200);
154      return `${key}=${text.replace(/\s+/g, ' ')}`;
155    })
156    .join(' ');
157  return `${call.id} ${call.tool} ${truncate(input, INPUT_CHARS[2])} → ${
158    call.isError ? 'error' : 'ok'
159  } ${call.resultChars}ch`;
160}
161
162/**
163 * Folds runs of adjacent call-only entries into one entry each, so the
164 * per-entry envelope is paid once per run; the call lines keep their ids.
165 */
166function mergeCallRuns(history: readonly HistoryEntry[], pinned: (e: HistoryEntry) => boolean): HistoryEntry[] {
167  const merged: HistoryEntry[] = [];
168  for (const entry of history) {
169    const previous = merged[merged.length - 1];
170    const foldable = (e: HistoryEntry): boolean =>
171      !pinned(e) && e.text.length === 0 && typeof e.tool_calls?.[0] === 'string';
172    if (previous && foldable(previous) && foldable(entry) && previous.role === entry.role) {
173      previous.tool_calls = [...(previous.tool_calls as string[]), ...(entry.tool_calls as string[])];
174      continue;
175    }
176    merged.push({ ...entry });
177  }
178  return merged;
179}
180
181function callsByMessage(calls: readonly ToolCall[]): Map<number, ToolCall[]> {
182  const byMessage = new Map<number, ToolCall[]>();
183  for (const call of calls) {
184    const list = byMessage.get(call.callIndex) ?? [];
185    list.push(call);
186    byMessage.set(call.callIndex, list);
187  }
188  return byMessage;
189}
190
191function historyEntries(
192  messages: readonly Message[],
193  calls: readonly ToolCall[],
194  inputChars: number,
195): HistoryEntry[] {
196  const byMessage = callsByMessage(calls);
197  const entries: HistoryEntry[] = [];
198  messages.forEach((message, i) => {
199    const toolCalls = (byMessage.get(i) ?? []).map((call) => ({
200      id: call.id,
201      tool: call.tool,
202      input: inputText(call.input, inputChars),
203      result: resultNote(call),
204    }));
205    if (message.text.trim().length === 0 && toolCalls.length === 0) return;
206    const entry: HistoryEntry = { i, role: message.role, text: message.text };
207    if (toolCalls.length > 0) entry.tool_calls = toolCalls;
208    entries.push(entry);
209  });
210  return entries;
211}
212
213/** The last three user prompts, as the default `goal`. */
214export function goalFromMessages(messages: readonly Message[]): string {
215  return messages
216    .filter(
217      (message) =>
218        message.role === 'user' &&
219        message.text.trim().length > 0 &&
220        (message.toolResults ?? []).length === 0,
221    )
222    .slice(-3)
223    .map((message) => truncate(message.text, 500))
224    .join('\n');
225}
226
227/**
228 * Builds the Jev state from the whole conversation and shrinks it in stages
229 * until it fits `maxStateTokens`: tool inputs are truncated, then long texts
230 * are abridged oldest-first (pinned messages last), then old messages collapse
231 * to a one-line note, then old tool calls shrink to one line each, then old
232 * messages that carry no call are left out, then runs of old call-only
233 * messages are folded into one entry. Throws when even that is too big.
234 */
235export function fitState(
236  messages: readonly Message[],
237  calls: readonly ToolCall[],
238  options: Pick<ResolvedCompactOptions, 'maxStateTokens' | 'preserveRecentMessages' | 'goal'>,
239): FittedState {
240  const goal = options.goal || goalFromMessages(messages);
241  const stateOf = (history: HistoryEntry[]): CompactionState => ({
242    context: STATE_CONTEXT,
243    goal,
244    history,
245  });
246  const entryTokens = (entry: HistoryEntry): number => estimateTokens(JSON.stringify(entry)) + 1;
247  const baseTokens = estimateTokens(JSON.stringify(stateOf([])));
248  const fitted = (history: HistoryEntry[], tokens: number, stage: string): FittedState => ({
249    state: stateOf(history),
250    tokens,
251    stage,
252  });
253
254  let history: HistoryEntry[] = [];
255  let perEntry: number[] = [];
256  let tokens = 0;
257  const rebuild = (inputChars: number): void => {
258    history = historyEntries(messages, calls, inputChars);
259    perEntry = history.map(entryTokens);
260    tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
261  };
262  const fits = (): boolean => tokens <= options.maxStateTokens;
263  const shrink = (index: number, change: (entry: HistoryEntry) => void): void => {
264    const entry = history[index];
265    if (!entry) return;
266    change(entry);
267    const now = entryTokens(entry);
268    tokens += now - (perEntry[index] ?? 0);
269    perEntry[index] = now;
270  };
271
272  rebuild(INPUT_CHARS[0]);
273  if (fits()) return fitted(history, tokens, 'full');
274
275  for (const limit of INPUT_CHARS.slice(1)) {
276    rebuild(limit);
277    if (fits()) return fitted(history, tokens, `inputs<=${limit}`);
278  }
279
280  const pinned = (entry: HistoryEntry): boolean =>
281    isPinned(entry.i, messages.length, options.preserveRecentMessages);
282  const indices = history.map((_, index) => index);
283  const order = [
284    ...indices.filter((index) => !pinned(history[index]!)),
285    ...indices.filter((index) => pinned(history[index]!)),
286  ];
287
288  for (const index of order) {
289    const entry = history[index]!;
290    if (entry.text.length <= TEXT_HEAD + TEXT_TAIL + 40) continue;
291    shrink(index, (e) => {
292      e.text = abridge(e.text, TEXT_HEAD, TEXT_TAIL);
293    });
294    if (fits()) return fitted(history, tokens, 'texts abridged');
295  }
296
297  for (const index of order) {
298    const entry = history[index]!;
299    if (pinned(entry) || entry.text.length === 0) continue;
300    const original = messages[entry.i]?.text.length ?? entry.text.length;
301    shrink(index, (e) => {
302      e.text = `[… ${original} chars omitted …]`;
303    });
304    if (fits()) return fitted(history, tokens, 'old messages collapsed');
305  }
306
307  const byMessage = callsByMessage(calls);
308  for (const index of order) {
309    const entry = history[index]!;
310    const own = byMessage.get(entry.i);
311    if (pinned(entry) || !own) continue;
312    shrink(index, (e) => {
313      e.tool_calls = own.map(compactCall);
314    });
315    if (fits()) return fitted(history, tokens, 'old calls compacted');
316  }
317
318  const left = new Set<number>();
319  for (const index of order) {
320    const entry = history[index]!;
321    if (pinned(entry) || entry.tool_calls) continue;
322    left.add(index);
323    tokens -= perEntry[index] ?? 0;
324    if (fits()) {
325      return fitted(
326        history.filter((_, i) => !left.has(i)),
327        tokens,
328        'old messages left out',
329      );
330    }
331  }
332
333  history = mergeCallRuns(
334    history.filter((_, i) => !left.has(i)),
335    pinned,
336  );
337  perEntry = history.map(entryTokens);
338  tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
339  if (fits()) return fitted(history, tokens, 'old calls merged');
340
341  throw new Error(
342    `history too large for Jev (~${tokens} tokens after truncation, limit ${options.maxStateTokens})`,
343  );
344}
345