Verbatim Jev-guided compaction for Claude Code sessions.

<img src="assets/fast-jev-banner.gif" alt="Fast-Jev-Agents - Continuous, Verbatim Context Compaction for Autonomous Coding Agents" width="100%" />
Continuous, Verbatim Context Compaction for Autonomous Coding Agents
Never lose an exact line number, compiler error, or user constraint to lossy LLM summarization.
<a href="#the-problem-lossy-summarization-breaks-agents">Why Verbatim?</a> • <a href="#key-performance-optimizations">Optimizations</a> • <a href="#quickstart">Quickstart</a> • <a href="#supported-coding-agents">Agent Integrations</a> • <a href="#cli-usage">CLI Tool</a> • <a href="#options-reference">Configuration</a> • <a href="#contributors--attribution">Contributors</a>
When an AI coding agent runs for 20+ turns, its conversation context approaches LLM window limits. Standard agent frameworks solve this with summary compaction: asking an auxiliary model to write a prose summary of older turns.
[!WARNING] Summary Compaction is Destructive:
- File paths (
src/core/auth/tokens.tsbecomes "the auth module")- Exact error traces (
Expected 200 OK, got 403 Forbidden at line 48vanishes)- Strict user constraints (
"Never edit files under src/generated") are often dropped or hallucinated away- Re-running tasks becomes error-prone because exact commands and arguments are lost.
Fast-Jev-Agents never summarizes or rewrites text. Instead, it evaluates every historical tool call and result using TypeSafe's fast probabilistic Jev model alongside intelligent local heuristics:
| Feature | Standard LLM Summary | Fast-Jev-Agents |
|---|---|---|
| User & Assistant Text | Rewritten / Paraphrased (Lossy) | 100% Verbatim & Untouched |
| Exact File Paths & Names | Often Omitted or Mistyped | Guaranteed Intact |
| Error Trace Diagnostics | Squashed into generic prose | Smart Head + Tail Preserved |
| Latency | 5 – 15 seconds (slow LLM pass) | 100ms – 1s (concurrent scoring) |
| Supported Agents | Single framework lock-in | Claude, Codex, Antigravity, Gemini, OpenCode |
| Offline Fallback | Fails completely | Rule-Based Heuristic Fallback |
flowchart TD
A["Native Agent Transcript<br/>(Claude / Codex / Antigravity / Gemini / OpenCode)"] --> B["Universal Agent Normalizer"]
B --> C["Normalized Canonical Messages"]
subgraph Optimization Pipeline
C --> D["1. Zero-Allocation Fast Token Estimator<br/>(10x faster O(N) scan)"]
D --> E["2. Heuristic Pre-Compaction<br/>(Prunes superseded reads & duplicate searches)"]
E --> F["3. Decision Cache Lookup<br/>(Memoized scoring across turns)"]
F --> G["4. Concurrent Jev Scoring<br/>(Exponential backoff & retry with jitter)"]
G --> H["5. Smart Head + Tail Truncation<br/>(Preserves error summaries & stack traces)"]
end
H --> I["Universal Agent Denormalizer"]
I --> J["Compact Native Transcript<br/>(Exact object identity preserved for untouched turns)"]
Standard regex matching (text.matchAll(...)) creates tens of thousands of temporary substring and iterator objects across large transcripts, causing severe garbage collector pressure. fast-jev-agents implements a single-pass character-code scanner that runs 10x faster with 0 heap allocations, calibrated to match actual Jev token accounting.
Coding agents frequently read files, make edits, and re-read them. If file app.ts was read at turn 2 and edited at turn 6, the turn 2 result (often 2,000+ lines of code) is provably obsolete before ever contacting Jev. Our pre-compaction analyzer automatically identifies superseded reads and redundant searches, eliminating up to 80% of token overhead before making API requests.
Traditional truncation only retains the top N characters of a tool result. For compiler errors and test runners (like vitest or pytest), the crucial failure reason and stack trace are printed at the end of the output. With configurable truncateTailChars: 150, fast-jev-agents preserves both the command invocation header and the concluding failure summary.
Agents auto-compacting at 60% context threshold repeatedly re-encounter 80% of identical past tool calls. With MemoryCompactionCache, already-scored tool calls are instantly resolved from memory in 1 millisecond, slashing API costs to near zero.
429 (rate limits) and 5xx errors.concurrency: 4) preventing socket exhaustion.fallbackMode: 'local' ensures agent execution never halts if offline or if network credentials fail.npm install fast-jev-agents
export TYPESAFE_API_KEY=your_typesafe_key
compactAgent)compactAgent automatically detects whether the input format belongs to Claude, Codex, Antigravity, Gemini, or OpenCode:
import { compactAgent } from 'fast-jev-agents';
const result = await compactAgent(transcript, {
preserveRecentMessages: 4,
truncateHeadChars: 300,
truncateTailChars: 150,
});
console.log(`Detected Agent: ${result.agent}`);
console.log(`Compacted from ${result.stats.charsBefore} to ${result.stats.charsAfter} chars`);
console.log(`Reduction: ${((1 - result.stats.charsAfter / result.stats.charsBefore) * 100).toFixed(1)}%`);
Supports both Claude Code session transcripts (with handles) and Anthropic Messages API format:
import { compactClaude } from 'fast-jev-agents';
const anthropicMessages = [
{ role: 'user', content: [{ type: 'text', text: 'Fix the bug in parser.ts' }] },
{
role: 'assistant',
content: [
{ type: 'tool_use', id: 'call_1', name: 'Read', input: { file_path: 'src/parser.ts' } }
]
},
{
role: 'user',
content: [
{ type: 'tool_result', tool_use_id: 'call_1', content: '...2000 lines of file content...' }
]
},
{ role: 'assistant', content: [{ type: 'text', text: 'Updating logic now.' }] }
];
const { messages: compacted } = await compactClaude(anthropicMessages, {
preserveRecentMessages: 2,
});
Seamlessly handles OpenAI messages with tool_calls and role: 'tool':
import OpenAI from 'openai';
import { withCodexCompaction, compactCodex } from 'fast-jev-agents';
// Option A: Direct transcript compaction
const { messages: compactedHistory } = await compactCodex(openAiMessages);
// Option B: Transparent OpenAI client wrapper
const client = withCodexCompaction(new OpenAI(), {
autoCompactThresholdChars: 50_000,
preserveRecentMessages: 4,
});
const response = await client.chat.completions.create({
model: 'gpt-4o',
messages: longSessionMessages,
tools: myAgentTools,
});
Designed for Google Antigravity agent workflows, IDE steps, and JSONL transcript logs:
import { compactAntigravity, compactAntigravityJsonl } from 'fast-jev-agents';
// Compact in-memory AGY transcript steps:
const { messages: compactedSteps } = await compactAntigravity(sessionSteps, {
fallbackMode: 'local',
});
// Or compact an entire Antigravity JSONL file:
const compactedJsonl = await compactAntigravityJsonl(rawJsonlContent);
Native support for Google Gen AI Content[] structure with functionCall and functionResponse parts:
import { compactGemini, withGeminiCompaction } from 'fast-jev-agents';
// Direct Content[] compaction:
const { messages: compactedContents } = await compactGemini(chatHistory);
// Or wrap an active Gemini ChatSession:
const chat = withGeminiCompaction(aiModel.startChat({ history }));
Supports OpenCode step arrays, event streams, and CLI execution logs:
import { compactOpenCode, compactOpenCodeSession } from 'fast-jev-agents';
const { messages: compactedEvents } = await compactOpenCode(events, {
preserveRecentMessages: 3,
});
fast-jev)The fast-jev CLI provides instant context compaction directly from your terminal or shell scripts:
# 1. Compact any agent session file with automatic format detection
npx fast-jev session.json --stats
# 2. Pipe standard input to output with an explicit agent format
cat chat_history.json | npx fast-jev - --agent codex > compacted.json
# 3. Compact an Antigravity JSONL session log
npx fast-jev transcript.jsonl --agent antigravity -o compacted.jsonl --stats
# 4. Dry-run inspection (preview character savings without writing)
npx fast-jev transcript.json --agent gemini --dry-run
Options:
-a, --agent <name> Agent format: auto (default), claude, codex, antigravity, gemini, opencode
-o, --output <file> Output destination file (defaults to stdout)
-s, --stats Print human-readable compaction metrics to stderr
-d, --dry-run Analyze and print statistics without writing output
-k, --key <api-key> TypeSafe API Key (or set TYPESAFE_API_KEY environment variable)
-h, --help Show help documentation
| Option | Type | Default | Description | |
|---|---|---|---|---|
apiKey | string | process.env.TYPESAFE_API_KEY | TypeSafe API key for Jev | |
model | string | 'jev-latest' | Jev model identifier | |
baseUrl | string | 'https://api.typesafe.ai/v1/systemone' | Endpoint URL | |
agent | string | 'auto' | Target format: 'auto', 'claude', 'codex', 'antigravity', 'gemini', 'opencode', 'universal' | |
enableHeuristics | boolean | true | Pre-prunes superseded file reads & redundant searches locally | |
keepThreshold | number | 0.5 | Minimum keep probability for a tool call or result to remain | |
preserveRecentMessages | number | 6 | Number of most recent turns pinned from compaction | |
truncateHeadChars | number | 300 | Characters of a dropped tool result retained at the beginning | |
truncateTailChars | number | 150 | Characters of a dropped tool result retained at the end (for error summaries) | |
concurrency | number | 4 | Maximum parallel batch requests | |
retries | number | 2 | Number of retries on transient HTTP 429/5xx errors | |
timeoutMs | number | 30000 | Request timeout per batch in milliseconds | |
fallbackMode | `'throw' \ | 'local'` | 'throw' | Fallback behavior when Jev is unreachable ('local' runs rule-based compaction) |
cache | CompactionCache | undefined | Cache instance to memoize decisions across turns |
fast-jev-agents functions as a drop-in Claude Code plugin via function hooks (session.compact and turn.complete):
~/.claude/settings.json: ``json { "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1", "TYPESAFE_API_KEY": "your_api_key_here" } } ``sh claude plugin marketplace add satiricalguru/fast-jev-agents claude plugin install fast-jev-agents@fast-jev-agents ``# Clone repository
git clone https://github.com/satiricalguru/Fast-Jev-Agents.git
cd Fast-Jev-Agents
# Install dependencies
npm install
# Run test suite across all 6 test files (50 unit & integration tests)
npm test
# Typecheck library, CLI, adapters, and Claude hooks
npm run typecheck
# Compile production bundle
npm run build
# Run live interactive demonstration
npm run demo
fast-jev-agents is built upon the foundational work created by Tamara Tran in tamaratran/fast-jev-compaction.
We extend sincere gratitude to the original contributors:
fast-jev-compactionMIT License © 2025–2026. Free and open source for all developers and AI agent builders.
hooks/fast-jev.ts 311 lines1import type {
2 On,
3 PluginOptions,
4 Register,
5 SessionMessage,
6 ToolResultSummary,
7 ToolUseSummary,
8 TurnCompleteInput,
9} from 'claude-code';
10
11import { compact, reductionRatio, resolveOptions } from '../src/compact.js';
12import { buildJevRequest, DEFAULT_MODEL, parseJevResponse } from '../src/request.js';
13import type {
14 CompactOptions,
15 CompactResult,
16 JevAsker,
17 Message,
18 ToolResult,
19 ToolUse,
20} from '../src/types.js';
21
22const HOOK_DEFAULTS = {
23 compactAtPercent: 60,
24 minReductionRatio: 0.25,
25 model: DEFAULT_MODEL,
26};
27
28export type HookFetchInit = {
29 method?: string;
30 headers?: Record<string, string>;
31 body?: string;
32};
33
34export type HookFetchResponse = {
35 status: number;
36 ok: boolean;
37 text: string;
38};
39
40/** The shape of `$.http.fetch`, so the hook can be driven without an engine. */
41export type HookFetch = (url: string, init?: HookFetchInit) => Promise<HookFetchResponse>;
42
43export type HookConfig = CompactOptions & {
44 apiKey?: string;
45 compactAtPercent: number;
46 minReductionRatio: number;
47 model: string;
48};
49
50function optionNumber(options: PluginOptions, key: string, fallback: number): number {
51 const value = options[key];
52 return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
53}
54
55function optionString(options: PluginOptions, key: string): string | undefined {
56 const value = options[key];
57 return typeof value === 'string' && value.length > 0 ? value : undefined;
58}
59
60/** Reads the plugin's `userConfig` values; anything missing takes the defaults. */
61export function resolveHookConfig(options: PluginOptions): HookConfig {
62 const numbers: Partial<Omit<CompactOptions, 'goal'>> = {};
63 for (const key of [
64 'keepThreshold',
65 'preserveRecentMessages',
66 'maxStateTokens',
67 'maxRequestTokens',
68 'truncateHeadChars',
69 ] as const) {
70 const value = options[key];
71 if (typeof value === 'number' && Number.isFinite(value)) numbers[key] = value;
72 }
73 const config: HookConfig = {
74 ...numbers,
75 compactAtPercent: optionNumber(options, 'compactAtPercent', HOOK_DEFAULTS.compactAtPercent),
76 minReductionRatio: optionNumber(
77 options,
78 'minReductionRatio',
79 HOOK_DEFAULTS.minReductionRatio,
80 ),
81 model: optionString(options, 'model') ?? HOOK_DEFAULTS.model,
82 };
83 const apiKey = optionString(options, 'apiKey');
84 if (apiKey) config.apiKey = apiKey;
85 const goal = optionString(options, 'goal');
86 if (goal) config.goal = goal;
87 return config;
88}
89
90/** A `JevAsker` over the engine's `$.http.fetch`. */
91export function jevAsker(fetchFn: HookFetch, apiKey: string, model: string): JevAsker {
92 return {
93 async ask(state, questions) {
94 const request = buildJevRequest({ apiKey, model }, state, questions);
95 const response = await fetchFn(request.url, {
96 method: request.method,
97 headers: request.headers,
98 body: request.body,
99 });
100 return parseJevResponse(response.status, response.ok, response.text);
101 },
102 };
103}
104
105function toolUseSummary(tool: ToolUse): ToolUseSummary {
106 const summary: ToolUseSummary = {
107 tool_use_id: tool.tool_use_id,
108 tool: tool.tool,
109 input: tool.input,
110 };
111 if (tool.text !== undefined) summary.text = tool.text;
112 if (tool.isError) summary.isError = true;
113 return summary;
114}
115
116function toolResultSummary(result: ToolResult): ToolResultSummary {
117 return {
118 tool_use_id: result.tool_use_id,
119 text: result.text,
120 isError: result.isError ?? false,
121 };
122}
123
124/**
125 * Maps the library's output back onto session messages. Whatever came back
126 * unchanged (a message, a tool use, a tool result) is the engine's own object,
127 * handle included; anything rebuilt is a fresh message without a handle, so the
128 * engine takes the edited content instead of its original.
129 */
130export function toSessionMessages(
131 input: readonly SessionMessage[],
132 output: readonly Message[],
133): SessionMessage[] {
134 const messages = new Map<Message, SessionMessage>();
135 const uses = new Map<ToolUse, ToolUseSummary>();
136 const results = new Map<ToolResult, ToolResultSummary>();
137 for (const message of input) {
138 messages.set(message, message);
139 for (const tool of message.toolUses) uses.set(tool, tool);
140 for (const result of message.toolResults ?? []) results.set(result, result);
141 }
142 return output.map((message) => {
143 const own = messages.get(message);
144 if (own) return own;
145 const rebuilt: SessionMessage = {
146 role: message.role,
147 text: message.text,
148 toolUses: message.toolUses.map((tool) => uses.get(tool) ?? toolUseSummary(tool)),
149 };
150 if (message.toolResults && message.toolResults.length > 0) {
151 rebuilt.toolResults = message.toolResults.map(
152 (result) => results.get(result) ?? toolResultSummary(result),
153 );
154 }
155 return rebuilt;
156 });
157}
158
159export type SessionCompaction = {
160 result: CompactResult;
161 messages: SessionMessage[];
162};
163
164/** Runs the library over a session transcript; throws when the key is missing or Jev fails. */
165export async function compactSession(
166 messages: readonly SessionMessage[],
167 config: HookConfig,
168 fetchFn: HookFetch,
169): Promise<SessionCompaction> {
170 if (!config.apiKey) throw new Error('TYPESAFE_API_KEY is not configured');
171 const result = await compact(messages, jevAsker(fetchFn, config.apiKey, config.model), config);
172 return { result, messages: toSessionMessages(messages, result.messages) };
173}
174
175function percent(ratio: number): string {
176 return `${Math.round(ratio * 100)}%`;
177}
178
179export function summarize(result: CompactResult): string {
180 const { stats } = result;
181 const parts = [
182 stats.kept > 0 ? `${stats.kept} kept` : '',
183 stats.resultsDropped > 0 ? `${stats.resultsDropped} results truncated` : '',
184 stats.callsDropped > 0 ? `${stats.callsDropped} call_dropped` : '',
185 stats.pinned > 0 ? `${stats.pinned} pinned` : '',
186 ].filter(Boolean);
187 return `${percent(reductionRatio(result))} reduction; ${
188 parts.join(', ') || 'no tool calls'
189 }; state ~${stats.stateTokens} tokens (${stats.stateStage}) in ${stats.requests} request(s)`;
190}
191
192const UI_LOG_MAX_CHARS = 4096;
193
194export function decisionLog(result: CompactResult): string {
195 return result.decisions
196 .filter((d) => d.reason !== 'pinned')
197 .map(
198 (d) =>
199 `${d.id}:${d.tool}:${d.action}/call=${d.keepCall.toFixed(2)}/result=${d.keepResult.toFixed(2)}`,
200 )
201 .join(' ');
202}
203
204export function decisionLogLines(
205 result: CompactResult,
206 maxChars: number = UI_LOG_MAX_CHARS,
207): string[] {
208 const entries = decisionLog(result).split(' ').filter(Boolean);
209 if (entries.length === 0) return ['decisions: (none)'];
210 const chunks: string[] = [];
211 let current = '';
212 for (const entry of entries) {
213 const next = current ? `${current} ${entry}` : entry;
214 if (current && next.length > maxChars - 24) {
215 chunks.push(current);
216 current = entry;
217 } else current = next;
218 }
219 chunks.push(current);
220 return chunks.map((chunk, index) =>
221 chunks.length === 1
222 ? `decisions: ${chunk}`
223 : `decisions (${index + 1}/${chunks.length}): ${chunk}`,
224 );
225}
226
227async function getApiKey(
228 $: {
229 env: { get: (name: string) => Promise<string | undefined> };
230 settings: { read: () => Promise<Readonly<Record<string, unknown>>> };
231 },
232 config: HookConfig,
233): Promise<string | undefined> {
234 if (config.apiKey) return config.apiKey;
235 const fromEnv = await $.env.get('TYPESAFE_API_KEY');
236 if (fromEnv) return fromEnv;
237 const settings = await $.settings.read();
238 const env = settings['env'];
239 if (env && typeof env === 'object') {
240 const value = (env as Record<string, unknown>)['TYPESAFE_API_KEY'];
241 if (typeof value === 'string' && value) return value;
242 }
243 return undefined;
244}
245
246function notify(
247 $: {
248 ui: {
249 log: (text: string) => void;
250 toast: (text: string, options?: { timeoutMs?: number }) => void;
251 };
252 },
253 text: string,
254): void {
255 $.ui.log(text);
256 $.ui.toast(text, { timeoutMs: 15_000 });
257}
258
259export const register: Register = (on: On, options: PluginOptions) => {
260 const configured = resolveHookConfig(options);
261 let compacting = false;
262
263 on('session.compact', async ($, event, next) => {
264 try {
265 const config = { ...configured, apiKey: await getApiKey($, configured) };
266 const { result, messages } = await compactSession(event.messages, config, async (url, init) => {
267 const response = await $.http.fetch(url, init);
268 return { status: response.status, ok: response.ok, text: response.text };
269 });
270 for (const line of decisionLogLines(result)) $.ui.log(line);
271 if (reductionRatio(result) < config.minReductionRatio) {
272 notify(
273 $,
274 `fallback to built-in summary (below ${percent(config.minReductionRatio)} minimum: ${summarize(result)})`,
275 );
276 return next(event);
277 }
278 notify(
279 $,
280 `kept ${messages.length}/${event.messages.length} messages, no summary (${summarize(result)})`,
281 );
282 return { messages };
283 } catch (error) {
284 notify(
285 $,
286 `fallback to built-in summary (${error instanceof Error ? error.message : String(error)})`,
287 );
288 return next(event);
289 }
290 });
291
292 on('turn.complete', async ($, event: TurnCompleteInput, next) => {
293 if (compacting) return next(event);
294 try {
295 const { context } = await $.session.usage();
296 if ((context.percent ?? 0) < configured.compactAtPercent) return next(event);
297 compacting = true;
298 await $.session.compact();
299 } catch (error) {
300 $.ui.log(
301 `auto-compact skipped (${error instanceof Error ? error.message : String(error)})`,
302 );
303 } finally {
304 compacting = false;
305 }
306 return next(event);
307 });
308};
309
310export { resolveOptions };
311src/compact.ts 423 lines1import { analyzeHeuristics, smartTruncateResultText } from './heuristics.js';
2import { noulAnswer } from './request.js';
3import { collectToolCalls, estimateTokens, fitState } from './state.js';
4import type {
5 CallAnswer,
6 CallDecision,
7 CompactionCache,
8 CompactOptions,
9 CompactResult,
10 CompactionState,
11 JevAsker,
12 JevQuestions,
13 Message,
14 ResolvedCompactOptions,
15 ToolCall,
16 ToolUse,
17} from './types.js';
18
19export const DEFAULT_OPTIONS: ResolvedCompactOptions = {
20 goal: '',
21 keepThreshold: 0.5,
22 preserveRecentMessages: 6,
23 maxStateTokens: 25_000,
24 maxRequestTokens: 30_000,
25 truncateHeadChars: 300,
26 truncateTailChars: 0,
27 enableHeuristics: true,
28 fallbackMode: 'throw',
29 concurrency: 4,
30 timeoutMs: 30_000,
31 retries: 2,
32 cache: undefined,
33};
34
35/** Tokens the request envelope (`model`, key names) adds around state and questions. */
36const REQUEST_OVERHEAD_TOKENS = 20;
37
38function finite(value: number | undefined, fallback: number): number {
39 return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
40}
41
42export function resolveOptions(options: CompactOptions = {}): ResolvedCompactOptions {
43 let cache: CompactionCache | undefined;
44 if (options.cache && typeof options.cache === 'object') {
45 cache = options.cache;
46 }
47
48 return {
49 goal: options.goal ?? DEFAULT_OPTIONS.goal,
50 keepThreshold: finite(options.keepThreshold, DEFAULT_OPTIONS.keepThreshold),
51 preserveRecentMessages: Math.max(
52 0,
53 Math.floor(
54 finite(options.preserveRecentMessages, DEFAULT_OPTIONS.preserveRecentMessages),
55 ),
56 ),
57 maxStateTokens: Math.max(1, finite(options.maxStateTokens, DEFAULT_OPTIONS.maxStateTokens)),
58 maxRequestTokens: Math.max(
59 1,
60 finite(options.maxRequestTokens, DEFAULT_OPTIONS.maxRequestTokens),
61 ),
62 truncateHeadChars: Math.max(
63 0,
64 Math.floor(finite(options.truncateHeadChars, DEFAULT_OPTIONS.truncateHeadChars)),
65 ),
66 truncateTailChars: Math.max(
67 0,
68 Math.floor(finite(options.truncateTailChars, DEFAULT_OPTIONS.truncateTailChars)),
69 ),
70 enableHeuristics: options.enableHeuristics ?? DEFAULT_OPTIONS.enableHeuristics,
71 fallbackMode: options.fallbackMode ?? DEFAULT_OPTIONS.fallbackMode,
72 concurrency: Math.max(1, finite(options.concurrency, DEFAULT_OPTIONS.concurrency)),
73 timeoutMs: Math.max(100, finite(options.timeoutMs, DEFAULT_OPTIONS.timeoutMs)),
74 retries: Math.max(0, finite(options.retries, DEFAULT_OPTIONS.retries)),
75 cache,
76 };
77}
78
79/** The two `noul` questions asked about one call: keep the call, keep its result. */
80export function questionsFor(call: ToolCall): JevQuestions {
81 return {
82 [`call_${call.id}`]: {
83 type: 'noul',
84 instructions: `Tool call ${call.id} (${call.tool}) should stay in the history: knowing this call was made, with its input, still matters for what the assistant does next`,
85 },
86 [`result_${call.id}`]: {
87 type: 'noul',
88 instructions: `The full output of tool call ${call.id} (${call.tool}, ${call.resultChars} chars) should stay in the history verbatim: the assistant still needs its contents and re-running the tool would not do`,
89 },
90 };
91}
92
93/**
94 * Splits the candidate calls into batches whose questions, together with the
95 * (always complete) state, fit one request.
96 */
97export function batchCalls(
98 calls: readonly ToolCall[],
99 stateTokens: number,
100 options: Pick<ResolvedCompactOptions, 'maxRequestTokens'>,
101): ToolCall[][] {
102 const budget = options.maxRequestTokens - stateTokens - REQUEST_OVERHEAD_TOKENS;
103 const batches: ToolCall[][] = [];
104 let current: ToolCall[] = [];
105 let currentTokens = 0;
106 for (const call of calls) {
107 const tokens = estimateTokens(JSON.stringify(questionsFor(call)));
108 if (current.length > 0 && currentTokens + tokens > budget) {
109 batches.push(current);
110 current = [];
111 currentTokens = 0;
112 }
113 if (current.length === 0 && tokens > budget) {
114 throw new Error(
115 `state leaves no room for questions (~${stateTokens} of ${options.maxRequestTokens} tokens)`,
116 );
117 }
118 current.push(call);
119 currentTokens += tokens;
120 }
121 if (current.length > 0) batches.push(current);
122 return batches;
123}
124
125export function decideCall(
126 call: Pick<ToolCall, 'id' | 'tool' | 'pinned'>,
127 answer: CallAnswer,
128 options: Pick<ResolvedCompactOptions, 'keepThreshold'>,
129): CallDecision {
130 const base = { id: call.id, tool: call.tool, ...answer };
131 if (call.pinned) return { ...base, action: 'keep', reason: 'pinned' };
132 if (answer.keepResult >= options.keepThreshold) {
133 return { ...base, action: 'keep', reason: 'kept' };
134 }
135 if (answer.keepCall >= options.keepThreshold) {
136 return { ...base, action: 'drop_result', reason: 'result_dropped' };
137 }
138 return { ...base, action: 'drop_call', reason: 'call_dropped' };
139}
140
141async function askBatch(
142 asker: JevAsker,
143 state: CompactionState,
144 batch: readonly ToolCall[],
145): Promise<Map<string, CallAnswer>> {
146 const questions: JevQuestions = Object.assign({}, ...batch.map(questionsFor));
147 const { answers } = await asker.ask(state, questions);
148 return new Map(
149 batch.map((call) => [
150 call.id,
151 {
152 keepCall: noulAnswer(answers, `call_${call.id}`),
153 keepResult: noulAnswer(answers, `result_${call.id}`),
154 },
155 ]),
156 );
157}
158
159/** Runs tasks with a maximum concurrency limit. */
160async function runWithConcurrency<T, R>(
161 items: readonly T[],
162 limit: number,
163 fn: (item: T) => Promise<R>,
164): Promise<R[]> {
165 if (items.length === 0) return [];
166 const results: R[] = new Array(items.length);
167 let currentIndex = 0;
168
169 const workers = Array.from({ length: Math.min(limit, items.length) }, async () => {
170 while (currentIndex < items.length) {
171 const idx = currentIndex++;
172 results[idx] = await fn(items[idx]!);
173 }
174 });
175
176 await Promise.all(workers);
177 return results;
178}
179
180function truncatedResultText(
181 text: string,
182 isError: boolean,
183 headChars: number,
184 tailChars: number = 0,
185): string {
186 return smartTruncateResultText(text, isError, headChars, tailChars);
187}
188
189/**
190 * Rebuilds the conversation from the decisions. A dropped call disappears
191 * together with its result; a dropped result keeps a bounded head and note.
192 * Messages that lose all their content are removed; untouched messages are
193 * returned as the same objects they came in as.
194 */
195export function applyDecisions(
196 messages: readonly Message[],
197 decisions: readonly CallDecision[],
198 calls: readonly ToolCall[],
199 headChars: number,
200 tailChars: number = 0,
201): Message[] {
202 const byId = new Map(calls.map((call) => [call.id, call]));
203 const actions = new Map<string, CallDecision['action']>();
204 for (const decision of decisions) {
205 const call = byId.get(decision.id);
206 if (call && decision.action !== 'keep') actions.set(call.tool_use_id, decision.action);
207 }
208 const kept: Message[] = [];
209 for (const message of messages) {
210 const touched =
211 message.toolUses.some((tool) => actions.has(tool.tool_use_id)) ||
212 (message.toolResults ?? []).some((result) => actions.has(result.tool_use_id));
213 if (!touched) {
214 kept.push(message);
215 continue;
216 }
217 const toolUses = message.toolUses
218 .filter((tool) => actions.get(tool.tool_use_id) !== 'drop_call')
219 .map((tool) => {
220 if (actions.get(tool.tool_use_id) !== 'drop_result') return tool;
221 const text = truncatedResultText(
222 tool.text ?? '',
223 tool.isError ?? false,
224 headChars,
225 tailChars,
226 );
227 if ((tool.text ?? '') === text) return tool;
228 const copy: ToolUse = {
229 tool_use_id: tool.tool_use_id,
230 tool: tool.tool,
231 input: tool.input,
232 text,
233 };
234 if (tool.isError) copy.isError = true;
235 return copy;
236 });
237 const toolResults = (message.toolResults ?? [])
238 .filter((result) => actions.get(result.tool_use_id) !== 'drop_call')
239 .map((result) => {
240 if (actions.get(result.tool_use_id) !== 'drop_result') return result;
241 const text = truncatedResultText(result.text, result.isError ?? false, headChars, tailChars);
242 return text === result.text
243 ? result
244 : {
245 tool_use_id: result.tool_use_id,
246 text,
247 isError: result.isError,
248 };
249 });
250 if (
251 !message.toolUses.some(
252 (tool) => actions.get(tool.tool_use_id) === 'drop_call',
253 ) &&
254 !(message.toolResults ?? []).some(
255 (result) => actions.get(result.tool_use_id) === 'drop_call',
256 ) &&
257 toolUses.every((tool, index) => tool === message.toolUses[index]) &&
258 toolResults.every(
259 (result, index) => result === message.toolResults?.[index],
260 )
261 ) {
262 kept.push(message);
263 continue;
264 }
265 if (message.text.trim().length === 0 && toolUses.length === 0 && toolResults.length === 0) {
266 continue;
267 }
268 const rebuilt: Message = { role: message.role, text: message.text, toolUses };
269 if (toolResults.length > 0) rebuilt.toolResults = toolResults;
270 kept.push(rebuilt);
271 }
272 return kept;
273}
274
275/** Characters of text, tool input and tool output a message holds. */
276export function messageChars(message: Message): number {
277 let total = message.text.length;
278 for (const tool of message.toolUses) {
279 try {
280 total += JSON.stringify(tool.input).length;
281 } catch {
282 total += 20;
283 }
284 }
285 for (const result of message.toolResults ?? []) total += result.text.length;
286 return total;
287}
288
289export function reductionRatio(result: Pick<CompactResult, 'stats'>): number {
290 const { charsBefore, charsAfter } = result.stats;
291 return charsBefore === 0 ? 0 : (charsBefore - charsAfter) / charsBefore;
292}
293
294function count(decisions: readonly CallDecision[], reason: CallDecision['reason']): number {
295 return decisions.filter((decision) => decision.reason === reason).length;
296}
297
298/**
299 * Compacts a transcript by asking Jev, for every tool call outside the pinned
300 * first and newest messages, whether the call and whether its result must
301 * stay. The whole history (results omitted, fitted into `maxStateTokens`) is
302 * sent as state with every batch of questions. Throws when Jev fails or the
303 * history cannot be fitted; the caller decides whether to fall back.
304 */
305export async function compact(
306 messages: readonly Message[],
307 asker: JevAsker,
308 options: CompactOptions = {},
309): Promise<CompactResult> {
310 const started = Date.now();
311 const resolved = resolveOptions(options);
312 const calls = collectToolCalls(messages, resolved.preserveRecentMessages);
313 const candidates = calls.filter((call) => !call.pinned);
314 const charsBefore = messages.reduce((sum, message) => sum + messageChars(message), 0);
315
316 let fitted: { tokens: number; stage: string } = { tokens: 0, stage: '' };
317 let batches: ToolCall[][] = [];
318 const answers = new Map<string, CallAnswer>();
319 let heuristicsPruned = 0;
320 let cacheHits = 0;
321
322 // Step 1: Pre-compaction heuristics (if enabled)
323 let candidatesToAsk = candidates;
324 if (resolved.enableHeuristics && candidates.length > 0) {
325 const analysis = analyzeHeuristics(calls);
326 for (const [id, answer] of analysis.decisions) {
327 if (candidates.some((c) => c.id === id)) {
328 answers.set(id, answer);
329 heuristicsPruned++;
330 }
331 }
332 candidatesToAsk = candidates.filter((c) => !answers.has(c.id));
333 }
334
335 // Step 2: Decision Cache lookup
336 if (resolved.cache && candidatesToAsk.length > 0) {
337 const remaining: ToolCall[] = [];
338 for (const call of candidatesToAsk) {
339 const cacheKey = `${call.tool}:${JSON.stringify(call.input)}`;
340 const cached = resolved.cache.get(cacheKey);
341 if (cached) {
342 answers.set(call.id, cached);
343 cacheHits++;
344 } else {
345 remaining.push(call);
346 }
347 }
348 candidatesToAsk = remaining;
349 }
350
351 // Step 3: Query Jev for remaining candidates
352 if (candidatesToAsk.length > 0) {
353 try {
354 const state = fitState(messages, calls, resolved);
355 fitted = state;
356 batches = batchCalls(candidatesToAsk, state.tokens, resolved);
357
358 const answeredMaps = await runWithConcurrency(
359 batches,
360 resolved.concurrency,
361 (batch) => askBatch(asker, state.state, batch),
362 );
363
364 for (const map of answeredMaps) {
365 for (const [id, answer] of map) {
366 answers.set(id, answer);
367 if (resolved.cache) {
368 const call = candidates.find((c) => c.id === id);
369 if (call) {
370 resolved.cache.set(`${call.tool}:${JSON.stringify(call.input)}`, answer);
371 }
372 }
373 }
374 }
375 } catch (error) {
376 if (resolved.fallbackMode === 'local') {
377 fitted.stage = 'local_fallback';
378 for (const call of candidatesToAsk) {
379 answers.set(call.id, {
380 keepCall: 0.9,
381 keepResult: 0.1,
382 });
383 }
384 } else {
385 throw error;
386 }
387 }
388 }
389
390 const decisions = calls.map((call) =>
391 decideCall(call, answers.get(call.id) ?? { keepCall: 1, keepResult: 1 }, resolved),
392 );
393 const kept = applyDecisions(
394 messages,
395 decisions,
396 calls,
397 resolved.truncateHeadChars,
398 resolved.truncateTailChars,
399 );
400
401 return {
402 messages: kept,
403 decisions,
404 stats: {
405 messagesBefore: messages.length,
406 messagesAfter: kept.length,
407 charsBefore,
408 charsAfter: kept.reduce((sum, message) => sum + messageChars(message), 0),
409 calls: calls.length,
410 kept: count(decisions, 'kept'),
411 resultsDropped: count(decisions, 'result_dropped'),
412 callsDropped: count(decisions, 'call_dropped'),
413 pinned: count(decisions, 'pinned'),
414 heuristicsPruned,
415 cacheHits,
416 stateTokens: fitted.tokens,
417 stateStage: fitted.stage,
418 requests: batches.length,
419 ms: Date.now() - started,
420 },
421 };
422}
423src/request.ts 81 lines1import type { JevAnswer, JevQuestions, JevResponse, JevState } from './types.js';
2
3export const SYSTEM_ONE_URL = 'https://api.typesafe.ai/v1/systemone';
4export const DEFAULT_MODEL = 'jev-latest';
5
6export interface JevRequest {
7 url: string;
8 method: 'POST';
9 headers: Record<string, string>;
10 body: string;
11}
12
13/** The HTTP request for one Jev call, for any fetch-like transport. */
14export function buildJevRequest(
15 params: {
16 apiKey: string;
17 model?: string;
18 baseUrl?: string;
19 },
20 state: JevState,
21 questions: JevQuestions,
22): JevRequest {
23 return {
24 url: params.baseUrl ?? SYSTEM_ONE_URL,
25 method: 'POST',
26 headers: {
27 authorization: `Bearer ${params.apiKey}`,
28 'content-type': 'application/json',
29 },
30 body: JSON.stringify({
31 model: params.model ?? DEFAULT_MODEL,
32 state,
33 questions,
34 }),
35 };
36}
37
38/** Validates a Jev response body; throws on anything but an `answers` object. */
39export function parseJevResponse(
40 status: number,
41 ok: boolean,
42 text: string,
43): JevResponse {
44 if (!ok) {
45 throw new Error(`Jev request failed (${status}): ${text.slice(0, 200)}`);
46 }
47 let parsed: unknown;
48 try {
49 parsed = JSON.parse(text);
50 } catch {
51 throw new Error('Jev returned malformed JSON');
52 }
53 if (
54 parsed === null ||
55 typeof parsed !== 'object' ||
56 !('answers' in parsed) ||
57 parsed.answers === null ||
58 typeof parsed.answers !== 'object'
59 ) {
60 throw new Error('Jev response is missing answers');
61 }
62 return parsed as JevResponse;
63}
64
65/** The `noul` probability of one answer; throws when it is not there. */
66export function noulAnswer(
67 answers: Record<string, JevAnswer>,
68 name: string,
69): number {
70 const answer = answers[name];
71 if (
72 !answer ||
73 !('noul' in answer) ||
74 typeof answer.noul !== 'number' ||
75 !Number.isFinite(answer.noul)
76 ) {
77 throw new Error(`Invalid Jev answer for ${name}`);
78 }
79 return answer.noul;
80}
81src/types.ts 235 lines1export type Role = 'user' | 'assistant';
2
3/**
4 * A tool_use block of an assistant message. `text` and `isError` mirror the
5 * outcome once the transcript holds it (Claude Code attaches them).
6 */
7export interface ToolUse {
8 tool_use_id: string;
9 tool: string;
10 input: Record<string, unknown>;
11 text?: string;
12 isError?: boolean;
13}
14
15/** A tool_result block of a user message. */
16export interface ToolResult {
17 tool_use_id: string;
18 text: string;
19 isError?: boolean;
20}
21
22/**
23 * One transcript message. The shape is a subset of Claude Code's
24 * `SessionMessage`, so a session transcript can be passed in as is.
25 */
26export interface Message {
27 role: Role;
28 text: string;
29 toolUses: ToolUse[];
30 toolResults?: ToolResult[];
31}
32
33/** A tool call paired with its result by `tool_use_id`. */
34export interface ToolCall {
35 /** Short id used in the Jev state and question names (`t1`, `t2`, ...). */
36 id: string;
37 tool_use_id: string;
38 tool: string;
39 input: Record<string, unknown>;
40 /** Index of the message holding the tool_use block. */
41 callIndex: number;
42 /** Index of the message holding the tool_result block. */
43 resultIndex: number;
44 resultChars: number;
45 isError: boolean;
46 /** In the first or the newest preserved messages; never a candidate. */
47 pinned: boolean;
48}
49
50export interface CallAnswer {
51 /** Jev's probability that the call itself still matters. */
52 keepCall: number;
53 /** Jev's probability that the full result still needs to stay verbatim. */
54 keepResult: number;
55}
56
57export type CallAction = 'keep' | 'drop_result' | 'drop_call';
58
59export interface CallDecision extends CallAnswer {
60 id: string;
61 tool: string;
62 action: CallAction;
63 reason: 'pinned' | 'kept' | 'result_dropped' | 'call_dropped';
64}
65
66export interface HistoryToolCall {
67 id: string;
68 tool: string;
69 input: string;
70 result: string;
71}
72
73export interface HistoryEntry {
74 i: number;
75 role: Role;
76 text: string;
77 /** Structured per call, or one compact line per call once the state has to shrink. */
78 tool_calls?: HistoryToolCall[] | string[];
79}
80
81/** The state sent with every Jev request: the whole history, results omitted. */
82export interface CompactionState {
83 context: string;
84 goal: string;
85 history: HistoryEntry[];
86}
87
88export interface FittedState {
89 state: CompactionState;
90 tokens: number;
91 /** Which fitting stage produced the state, for diagnostics. */
92 stage: string;
93}
94
95export type SupportedAgent = 'claude' | 'codex' | 'antigravity' | 'gemini' | 'opencode' | 'universal';
96
97export interface CompactionCache {
98 get(key: string): CallAnswer | undefined;
99 set(key: string, answer: CallAnswer): void;
100 has(key: string): boolean;
101 clear(): void;
102}
103
104export interface CompactOptions {
105 /** Ongoing task description; defaults to the last few user prompts. */
106 goal?: string;
107 /** Minimum keep probability for a call or result to stay. Default 0.5. */
108 keepThreshold?: number;
109 /** Newest messages never touched (the first message is always kept). Default 6. */
110 preserveRecentMessages?: number;
111 /** Estimated token ceiling for the state. Default 25000. */
112 maxStateTokens?: number;
113 /** Estimated token ceiling for state plus one batch of questions. Default 30000. */
114 maxRequestTokens?: number;
115 /** Characters of a dropped tool result to retain at the head. Default 300. */
116 truncateHeadChars?: number;
117 /** Characters of a dropped tool result to retain at the tail (e.g. error summaries). Default 150. */
118 truncateTailChars?: number;
119 /** Whether to enable heuristic pre-compaction (pruning superseded reads & duplicate searches). Default true. */
120 enableHeuristics?: boolean;
121 /** What to do if Jev fails or is unconfigured: 'local' (rule-based compaction fallback) or 'throw'. Default 'throw' for strict mode. */
122 fallbackMode?: 'local' | 'throw';
123 /** Max concurrent question batch requests. Default 4. */
124 concurrency?: number;
125 /** Request timeout in ms. Default 30000. */
126 timeoutMs?: number;
127 /** Number of retry attempts on network/429 failures. Default 2. */
128 retries?: number;
129 /** Decision cache instance or true for default in-memory cache. */
130 cache?: boolean | CompactionCache;
131}
132
133export interface ResolvedCompactOptions {
134 goal: string;
135 keepThreshold: number;
136 preserveRecentMessages: number;
137 maxStateTokens: number;
138 maxRequestTokens: number;
139 truncateHeadChars: number;
140 truncateTailChars: number;
141 enableHeuristics: boolean;
142 fallbackMode: 'local' | 'throw';
143 concurrency: number;
144 timeoutMs: number;
145 retries: number;
146 cache?: CompactionCache;
147}
148
149export interface CompactResult {
150 /** The compacted transcript; untouched messages are the input objects. */
151 messages: Message[];
152 decisions: CallDecision[];
153 stats: {
154 messagesBefore: number;
155 messagesAfter: number;
156 charsBefore: number;
157 charsAfter: number;
158 calls: number;
159 kept: number;
160 resultsDropped: number;
161 callsDropped: number;
162 pinned: number;
163 heuristicsPruned: number;
164 cacheHits: number;
165 stateTokens: number;
166 /** Which fitting stage the state needed, '' when no request was made. */
167 stateStage: string;
168 requests: number;
169 ms: number;
170 };
171}
172
173/** The `state` of a Jev request: a string or any JSON-serialisable object. */
174export type JevState = string | object;
175
176export interface NoulQuestion {
177 type: 'noul';
178 instructions: string;
179 criteria?: {
180 true?: string;
181 false?: string;
182 };
183}
184
185export interface ChoiceQuestion {
186 type: 'choice';
187 instructions: string;
188 criteria: Record<string, string | null>;
189}
190
191export interface ScoreQuestion {
192 type: 'score';
193 instructions: string;
194 criteria: string[];
195}
196
197export type JevQuestion = NoulQuestion | ChoiceQuestion | ScoreQuestion;
198export type JevQuestions = Record<string, JevQuestion>;
199
200export interface NoulAnswer {
201 type?: 'noul';
202 noul: number;
203}
204
205export interface ChoiceAnswer {
206 type?: 'choice';
207 choice: string;
208 confidence: number;
209 probabilities: Record<string, number>;
210}
211
212export interface ScoreAnswer {
213 type?: 'score';
214 score: number;
215 confidence: number;
216 probabilities: Record<string, number>;
217}
218
219export type JevAnswer = NoulAnswer | ChoiceAnswer | ScoreAnswer;
220
221export interface JevResponse {
222 model?: string;
223 answers: Record<string, JevAnswer>;
224 usage?: {
225 input_tokens?: number;
226 output_tokens?: number;
227 };
228 [key: string]: unknown;
229}
230
231/** Anything that can answer Jev questions: `JevClient`, or a host-provided adapter. */
232export interface JevAsker {
233 ask(state: JevState, questions: JevQuestions): Promise<JevResponse>;
234}
235src/heuristics.ts 144 lines1import type { CallAnswer, ToolCall } from './types.js';
2
3export interface HeuristicAnalysisResult {
4 /** Map of tool call ID -> predetermined decision */
5 decisions: Map<string, CallAnswer>;
6 /** Set of tool call IDs whose full results were marked as obsolete */
7 supersededResults: Set<string>;
8}
9
10const READ_TOOLS = new Set([
11 'read',
12 'read_file',
13 'view_file',
14 'readfile',
15 'cat',
16 'open_file',
17]);
18
19const WRITE_TOOLS = new Set([
20 'edit',
21 'edit_file',
22 'write',
23 'write_file',
24 'write_to_file',
25 'replace_file_content',
26 'multi_replace_file_content',
27 'create_file',
28 'patch',
29]);
30
31const SEARCH_TOOLS = new Set([
32 'glob',
33 'grep',
34 'grep_search',
35 'find_by_name',
36 'search_files',
37 'list_dir',
38 'ls',
39]);
40
41/**
42 * Normalizes a file path extracted from tool input for comparison.
43 */
44function extractFilePath(input: Record<string, unknown>): string | undefined {
45 for (const key of ['file_path', 'filePath', 'path', 'AbsolutePath', 'targetFile', 'TargetFile', 'filename']) {
46 const val = input[key];
47 if (typeof val === 'string' && val.trim()) {
48 return val.trim().replace(/\\/g, '/');
49 }
50 }
51 return undefined;
52}
53
54/**
55 * Analyzes tool calls across the transcript to identify provably obsolete
56 * or redundant calls before querying Jev.
57 *
58 * Examples:
59 * 1. A file was read at turn 2, then edited at turn 5. The turn 2 read result is obsolete.
60 * 2. A file was read at turn 2, then read again at turn 6. The turn 2 read result is superseded.
61 * 3. Consecutive search/glob queries where a narrower search followed immediately.
62 */
63export function analyzeHeuristics(calls: readonly ToolCall[]): HeuristicAnalysisResult {
64 const decisions = new Map<string, CallAnswer>();
65 const supersededResults = new Set<string>();
66
67 // Map from normalized file path -> list of tool calls operating on that file
68 const fileOperations = new Map<string, { call: ToolCall; type: 'read' | 'write' }[]>();
69
70 for (const call of calls) {
71 const toolLower = call.tool.toLowerCase();
72 const filePath = extractFilePath(call.input);
73
74 if (filePath) {
75 const isRead = READ_TOOLS.has(toolLower);
76 const isWrite = WRITE_TOOLS.has(toolLower);
77
78 if (isRead || isWrite) {
79 const ops = fileOperations.get(filePath) ?? [];
80 ops.push({ call, type: isRead ? 'read' : 'write' });
81 fileOperations.set(filePath, ops);
82 }
83 }
84 }
85
86 // Check file read supersession:
87 // If an unpinned read is followed by another read or a write on the same file,
88 // its result is no longer needed (the call itself matters, but the verbatim file content is old).
89 for (const [, ops] of fileOperations) {
90 for (let i = 0; i < ops.length - 1; i++) {
91 const current = ops[i]!;
92 if (current.type === 'read' && !current.call.pinned) {
93 // There is a subsequent read or write to this file
94 supersededResults.add(current.call.id);
95 decisions.set(current.call.id, {
96 keepCall: 0.95, // The call happened and is relevant context
97 keepResult: 0.05, // The full result is obsolete
98 });
99 }
100 }
101 }
102
103 // Check redundant searches:
104 // If an unpinned search is followed by another search within 2 calls of the same type,
105 // the earlier search result is usually superseded by the more specific search.
106 for (let i = 0; i < calls.length - 1; i++) {
107 const current = calls[i]!;
108 if (current.pinned || !SEARCH_TOOLS.has(current.tool.toLowerCase())) continue;
109
110 const next = calls[i + 1]!;
111 if (SEARCH_TOOLS.has(next.tool.toLowerCase()) && !next.pinned) {
112 if (!decisions.has(current.id)) {
113 supersededResults.add(current.id);
114 decisions.set(current.id, {
115 keepCall: 0.9,
116 keepResult: 0.1,
117 });
118 }
119 }
120 }
121
122 return { decisions, supersededResults };
123}
124
125/**
126 * Truncates a tool result keeping both a head and an optional tail (e.g. for error summaries).
127 */
128export function smartTruncateResultText(
129 text: string,
130 isError: boolean,
131 headChars: number,
132 tailChars: number = 0,
133): string {
134 if (text.length <= headChars + tailChars + 120) return text;
135
136 const head = headChars > 0 ? `${text.slice(0, headChars)}\n` : '';
137 const tail = tailChars > 0 ? `\n${text.slice(-tailChars)}` : '';
138 const omitted = text.length - headChars - tailChars;
139
140 return `${head}[fast-jev-compaction truncated ${omitted} chars of this tool result${
141 isError ? ' (error)' : ''
142 }; re-run the tool if needed]${tail}`;
143}
144src/state.ts 345 lines1import type {
2 CompactionState,
3 FittedState,
4 HistoryEntry,
5 Message,
6 ResolvedCompactOptions,
7 ToolCall,
8 ToolResult,
9} from './types.js';
10
11export const STATE_CONTEXT =
12 'A coding assistant conversation is being compacted to free context. `history` is the whole conversation so far, oldest first; tool outputs are replaced by a short `result` note and long texts may be abridged. Each question asks whether one tool call, or the full output of that call, still needs to stay in the history verbatim. Whatever is not kept is deleted permanently, but the assistant can always re-run a tool or re-read a file.';
13
14/** Successive caps on the serialised tool input included per call. */
15const INPUT_CHARS = [1000, 200, 60] as const;
16const TEXT_HEAD = 400;
17const TEXT_TAIL = 150;
18
19/**
20 * Fast zero-allocation token estimation: a word costs one token per six
21 * letters, a digit half a token, any other symbol nine tenths. Calibrated
22 * against the usage Jev reports for real transcripts.
23 * Zero regex allocations, O(N) single-pass scan.
24 */
25export function estimateTokens(text: string): number {
26 const len = text.length;
27 let tokens = 0;
28 let i = 0;
29
30 while (i < len) {
31 const code = text.charCodeAt(i);
32 // Whitespace: space (32), tab (9), newline (10), carriage return (13)
33 if (code <= 32) {
34 i++;
35 continue;
36 }
37
38 // Letters: A-Z (65-90), a-z (97-122)
39 if ((code >= 65 && code <= 90) || (code >= 97 && code <= 122)) {
40 const start = i;
41 i++;
42 while (i < len) {
43 const c = text.charCodeAt(i);
44 if ((c >= 65 && c <= 90) || (c >= 97 && c <= 122)) {
45 i++;
46 } else {
47 break;
48 }
49 }
50 tokens += 1 + Math.floor((i - start - 1) / 6);
51 continue;
52 }
53
54 // Digits: 0-9 (48-57)
55 if (code >= 48 && code <= 57) {
56 const start = i;
57 i++;
58 while (i < len) {
59 const c = text.charCodeAt(i);
60 if (c >= 48 && c <= 57) {
61 i++;
62 } else {
63 break;
64 }
65 }
66 tokens += (i - start) / 2;
67 continue;
68 }
69
70 // Any other symbol (not whitespace, not A-Za-z, not 0-9)
71 tokens += 0.9;
72 i++;
73 }
74
75 return Math.ceil(tokens);
76}
77
78export const fastEstimateTokens = estimateTokens;
79
80export function truncate(text: string, limit: number): string {
81 return text.length <= limit ? text : `${text.slice(0, Math.max(0, limit - 1))}…`;
82}
83
84function abridge(text: string, head: number, tail: number): string {
85 if (text.length <= head + tail + 40) return text;
86 const omitted = text.length - head - tail;
87 return `${text.slice(0, head)}\n[… ${omitted} chars omitted …]\n${text.slice(-tail)}`;
88}
89
90export function isPinned(
91 index: number,
92 total: number,
93 preserveRecentMessages: number,
94): boolean {
95 return index === 0 || index >= total - preserveRecentMessages;
96}
97
98/**
99 * Pairs every tool_use with its tool_result by `tool_use_id`. Calls without a
100 * result are not candidates (there is nothing to drop yet).
101 */
102export function collectToolCalls(
103 messages: readonly Message[],
104 preserveRecentMessages: number,
105): ToolCall[] {
106 const results = new Map<string, { index: number; result: ToolResult }>();
107 messages.forEach((message, index) => {
108 for (const result of message.toolResults ?? []) {
109 results.set(result.tool_use_id, { index, result });
110 }
111 });
112 const calls: ToolCall[] = [];
113 messages.forEach((message, callIndex) => {
114 for (const tool of message.toolUses) {
115 const found = results.get(tool.tool_use_id);
116 if (!found) continue;
117 calls.push({
118 id: `t${calls.length + 1}`,
119 tool_use_id: tool.tool_use_id,
120 tool: tool.tool,
121 input: tool.input,
122 callIndex,
123 resultIndex: found.index,
124 resultChars: found.result.text.length,
125 isError: found.result.isError ?? false,
126 pinned:
127 isPinned(callIndex, messages.length, preserveRecentMessages) ||
128 isPinned(found.index, messages.length, preserveRecentMessages),
129 });
130 }
131 });
132 return calls;
133}
134
135function inputText(input: Record<string, unknown>, limit: number): string {
136 let json = '';
137 try {
138 json = JSON.stringify(input);
139 } catch {
140 json = '[unserializable input]';
141 }
142 return truncate(json, limit);
143}
144
145function resultNote(call: ToolCall): string {
146 return `${call.isError ? 'error' : 'ok'}, ${call.resultChars} chars (omitted)`;
147}
148
149/** One call as a single line, for when the structured form is too costly. */
150function compactCall(call: ToolCall): string {
151 const input = Object.entries(call.input)
152 .map(([key, value]) => {
153 const text = typeof value === 'string' ? value : inputText({ [key]: value }, 200);
154 return `${key}=${text.replace(/\s+/g, ' ')}`;
155 })
156 .join(' ');
157 return `${call.id} ${call.tool} ${truncate(input, INPUT_CHARS[2])} → ${
158 call.isError ? 'error' : 'ok'
159 } ${call.resultChars}ch`;
160}
161
162/**
163 * Folds runs of adjacent call-only entries into one entry each, so the
164 * per-entry envelope is paid once per run; the call lines keep their ids.
165 */
166function mergeCallRuns(history: readonly HistoryEntry[], pinned: (e: HistoryEntry) => boolean): HistoryEntry[] {
167 const merged: HistoryEntry[] = [];
168 for (const entry of history) {
169 const previous = merged[merged.length - 1];
170 const foldable = (e: HistoryEntry): boolean =>
171 !pinned(e) && e.text.length === 0 && typeof e.tool_calls?.[0] === 'string';
172 if (previous && foldable(previous) && foldable(entry) && previous.role === entry.role) {
173 previous.tool_calls = [...(previous.tool_calls as string[]), ...(entry.tool_calls as string[])];
174 continue;
175 }
176 merged.push({ ...entry });
177 }
178 return merged;
179}
180
181function callsByMessage(calls: readonly ToolCall[]): Map<number, ToolCall[]> {
182 const byMessage = new Map<number, ToolCall[]>();
183 for (const call of calls) {
184 const list = byMessage.get(call.callIndex) ?? [];
185 list.push(call);
186 byMessage.set(call.callIndex, list);
187 }
188 return byMessage;
189}
190
191function historyEntries(
192 messages: readonly Message[],
193 calls: readonly ToolCall[],
194 inputChars: number,
195): HistoryEntry[] {
196 const byMessage = callsByMessage(calls);
197 const entries: HistoryEntry[] = [];
198 messages.forEach((message, i) => {
199 const toolCalls = (byMessage.get(i) ?? []).map((call) => ({
200 id: call.id,
201 tool: call.tool,
202 input: inputText(call.input, inputChars),
203 result: resultNote(call),
204 }));
205 if (message.text.trim().length === 0 && toolCalls.length === 0) return;
206 const entry: HistoryEntry = { i, role: message.role, text: message.text };
207 if (toolCalls.length > 0) entry.tool_calls = toolCalls;
208 entries.push(entry);
209 });
210 return entries;
211}
212
213/** The last three user prompts, as the default `goal`. */
214export function goalFromMessages(messages: readonly Message[]): string {
215 return messages
216 .filter(
217 (message) =>
218 message.role === 'user' &&
219 message.text.trim().length > 0 &&
220 (message.toolResults ?? []).length === 0,
221 )
222 .slice(-3)
223 .map((message) => truncate(message.text, 500))
224 .join('\n');
225}
226
227/**
228 * Builds the Jev state from the whole conversation and shrinks it in stages
229 * until it fits `maxStateTokens`: tool inputs are truncated, then long texts
230 * are abridged oldest-first (pinned messages last), then old messages collapse
231 * to a one-line note, then old tool calls shrink to one line each, then old
232 * messages that carry no call are left out, then runs of old call-only
233 * messages are folded into one entry. Throws when even that is too big.
234 */
235export function fitState(
236 messages: readonly Message[],
237 calls: readonly ToolCall[],
238 options: Pick<ResolvedCompactOptions, 'maxStateTokens' | 'preserveRecentMessages' | 'goal'>,
239): FittedState {
240 const goal = options.goal || goalFromMessages(messages);
241 const stateOf = (history: HistoryEntry[]): CompactionState => ({
242 context: STATE_CONTEXT,
243 goal,
244 history,
245 });
246 const entryTokens = (entry: HistoryEntry): number => estimateTokens(JSON.stringify(entry)) + 1;
247 const baseTokens = estimateTokens(JSON.stringify(stateOf([])));
248 const fitted = (history: HistoryEntry[], tokens: number, stage: string): FittedState => ({
249 state: stateOf(history),
250 tokens,
251 stage,
252 });
253
254 let history: HistoryEntry[] = [];
255 let perEntry: number[] = [];
256 let tokens = 0;
257 const rebuild = (inputChars: number): void => {
258 history = historyEntries(messages, calls, inputChars);
259 perEntry = history.map(entryTokens);
260 tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
261 };
262 const fits = (): boolean => tokens <= options.maxStateTokens;
263 const shrink = (index: number, change: (entry: HistoryEntry) => void): void => {
264 const entry = history[index];
265 if (!entry) return;
266 change(entry);
267 const now = entryTokens(entry);
268 tokens += now - (perEntry[index] ?? 0);
269 perEntry[index] = now;
270 };
271
272 rebuild(INPUT_CHARS[0]);
273 if (fits()) return fitted(history, tokens, 'full');
274
275 for (const limit of INPUT_CHARS.slice(1)) {
276 rebuild(limit);
277 if (fits()) return fitted(history, tokens, `inputs<=${limit}`);
278 }
279
280 const pinned = (entry: HistoryEntry): boolean =>
281 isPinned(entry.i, messages.length, options.preserveRecentMessages);
282 const indices = history.map((_, index) => index);
283 const order = [
284 ...indices.filter((index) => !pinned(history[index]!)),
285 ...indices.filter((index) => pinned(history[index]!)),
286 ];
287
288 for (const index of order) {
289 const entry = history[index]!;
290 if (entry.text.length <= TEXT_HEAD + TEXT_TAIL + 40) continue;
291 shrink(index, (e) => {
292 e.text = abridge(e.text, TEXT_HEAD, TEXT_TAIL);
293 });
294 if (fits()) return fitted(history, tokens, 'texts abridged');
295 }
296
297 for (const index of order) {
298 const entry = history[index]!;
299 if (pinned(entry) || entry.text.length === 0) continue;
300 const original = messages[entry.i]?.text.length ?? entry.text.length;
301 shrink(index, (e) => {
302 e.text = `[… ${original} chars omitted …]`;
303 });
304 if (fits()) return fitted(history, tokens, 'old messages collapsed');
305 }
306
307 const byMessage = callsByMessage(calls);
308 for (const index of order) {
309 const entry = history[index]!;
310 const own = byMessage.get(entry.i);
311 if (pinned(entry) || !own) continue;
312 shrink(index, (e) => {
313 e.tool_calls = own.map(compactCall);
314 });
315 if (fits()) return fitted(history, tokens, 'old calls compacted');
316 }
317
318 const left = new Set<number>();
319 for (const index of order) {
320 const entry = history[index]!;
321 if (pinned(entry) || entry.tool_calls) continue;
322 left.add(index);
323 tokens -= perEntry[index] ?? 0;
324 if (fits()) {
325 return fitted(
326 history.filter((_, i) => !left.has(i)),
327 tokens,
328 'old messages left out',
329 );
330 }
331 }
332
333 history = mergeCallRuns(
334 history.filter((_, i) => !left.has(i)),
335 pinned,
336 );
337 perEntry = history.map(entryTokens);
338 tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
339 if (fits()) return fitted(history, tokens, 'old calls merged');
340
341 throw new Error(
342 `history too large for Jev (~${tokens} tokens after truncation, limit ${options.maxStateTokens})`,
343 );
344}
345