Verbatim Laya-guided compaction for Claude Code sessions (local, no API key).

Claude Code compaction that keeps your conversation verbatim and trims old tool output instead of summarizing, decided by a small model running on your own machine.
Claude Code's built-in compaction asks a model to summarize the conversation. Summaries are lossy: a file path, an exact error or a constraint can vanish even when it matters later. This plugin never rewrites anything. When the context fills up, it removes or truncates old tool calls and their outputs that are no longer needed, and leaves every user and assistant message exactly as written.
The decisions come from two places:
No API key, no network at compaction time, nothing leaves your machine. It is a port of fast-jev-compaction, which asks TypeSafe's hosted Jev model instead.
Results on three real sessions: 27–33% smaller in 19–47 s, against 79 s for a built-in compaction on the same machine. Details, and what did not work, in BENCHMARK.md.
preserveRecentMessages messages are never touched.maxScoredCalls remaining calls each get a small state for Laya: the call, its outcome, how long ago it ran, whether later calls touched the same target, the current goal and the start of its output.P(keep) ≥ keepThreshold keeps the call; otherwise P(keep) + P(truncate) ≥ keepThreshold keeps the call with a truncated output; otherwise it is removed.autoTargetReduction (50%), it also truncates outputs Laya would keep, lowest P(keep) first. Otherwise the window would fill again within a few tool calls and the next auto-compaction would fall back to a summary. /compact and the plugin's own 60% trigger keep Laya's answers as they are.minReductionRatio smaller, it replaces the history. Otherwise, or if anything fails, Claude Code's built-in summary runs as usual. When even trimming everything could not reach that minimum, Laya is not started at all.It runs on every compaction trigger: /compact, Claude Code's auto-compaction (at its threshold or when a prompt is too long), its ahead-of-time precompute, and the plugin's own trigger at compactAtPercent. The hook runs backend/laya_compact.py once per compaction with uv run --offline, so no model stays in memory between compactions; a compaction arriving while another is scoring waits its turn.
mps), an NVIDIA GPU (cuda) or CPU~/.claude/settings.json: { "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" } }
claude plugin marketplace add kartikeyaagr/fast-laya-compaction
claude plugin install fast-laya-compaction@fast-laya-compaction
laya==0.3.20 and a reviewed Hugging Face commit). The hook never downloads; until this has run, compaction falls back to the built-in summary with a toast that shows this exact command: uv run --script ~/.claude/plugins/cache/fast-laya-compaction/fast-laya-compaction/<version>/backend/laya_compact.py --warmup
It ends with laya_compact: ready (typed-decisions on mps, …).
/compact. The toast reads kept N/M messages, no summary (…), or fallback to built-in summary (<reason>).Set options with /plugin configure fast-laya-compaction@fast-laya-compaction inside Claude Code, or at install time with claude plugin install … --config keepThreshold=0.55.
| Option | Default | What it does |
|---|---|---|
keepThreshold | 0.5 | Minimum probability for a call (or its full output) to stay; higher trims more |
preserveRecentMessages | 6 | Newest messages never touched (the first is always kept) |
compactAtPercent | 60 | Context percentage at which the plugin requests compaction |
minReductionRatio | 0.25 | Minimum reduction to replace the history instead of summarizing |
autoTargetReduction | 0.5 | On auto-compaction, the reduction to reach even past Laya's answers; 0 turns it off |
truncateHeadChars | 300 | Characters of a truncated output kept before its note |
maxScoredCalls | 80 | Most calls scored per compaction, oldest first; bounds the time |
timeoutSeconds | 120 | Longest a Laya run may take before falling back |
model | typed-decisions | Laya checkpoint: typed-decisions, multilingual (faster, weaker) or english |
checkpointPath | unset | Absolute path of a fine-tuned checkpoint; overrides model |
maxCallStateChars | 1600 | Size cap on what Laya reads about one call |
goal | latest user prompt | Task description Laya weighs calls against |
uvPath | uv | uv executable, if it is not on Claude Code's PATH |
device | auto | mps, cpu or cuda |
Laya is only weakly decisive on this task: its keep probabilities mostly fall between 0.3 and 0.65, ranking edits and writes above exploratory ls, git and search calls, and it rarely chooses drop. See what it would do to your own sessions before relying on it. The dry-run changes nothing:
git clone https://github.com/kartikeyaagr/fast-laya-compaction && cd fast-laya-compaction
npm install
npm run dry-run -- ~/.claude/projects/<project>/<session>.jsonl --threshold 0.5
It prints every call with its keep/truncate/drop probabilities and action, the reduction, the timings, and whether the hook would replace the history. Add --target 0.5 to see what an auto-compaction would do.
Prompt changes move Laya's answers a lot (renaming one key in its input changed a session from 33% to 23%) without saying which way is closer to Jev, and Laya cannot read a whole conversation the way Jev does (measured). So Jev is used as a teacher:
TYPESAFE_API_KEY=… in .env (git-ignored). This sends those transcripts to TypeSafe and stores Jev's per-call answers in data/ (git-ignored). Aim for 20–40 sessions: npm run jev-labels -- ~/.claude/projects/<project>/*.jsonlscripts/variants.ts: npm run agreement -- --variant timeline. It reports decision agreement, a Jev-by-Laya confusion table, rank correlation and reduction under both. Jev's two probabilities map onto Laya's three so that matching the target reproduces Jev's decision at any threshold.npm run export-training -- --variant <best> writes rows in the format of Laya's own training notebook, with whole sessions held out.notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb, changed to read laya-train.jsonl, start from the typed-decisions checkpoint, keep max_len/head_max_len at 512/128, fit the temperature on laya-holdout.jsonl, and delete temperature_by_options from the saved rl_agent_config.json. This step has not been run here.npm run agreement -- --model /abs/path/to/checkpoint, then set the checkpointPath option.To ship your own fork or a new version:
version in .claude-plugin/plugin.json, .claude-plugin/marketplace.json and package.json.npm run typecheck && npm test && npm run test:backend && npm run validate:plugin..claude-plugin/marketplace.json), so claude plugin marketplace add <owner>/<repo> works as soon as the push lands. Only committed files are installed; data/, tasks/, docs/ and .env are git-ignored. claude plugin marketplace update fast-laya-compaction
claude plugin update fast-laya-compaction@fast-laya-compaction # then restart Claude Code
A new version installs to a new cache directory; the Laya download is shared, so setup does not need to run again unless laya or the checkpoint changes.
For local testing without publishing: CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir /path/to/fast-laya-compaction, or claude plugin marketplace add /path/to/fast-laya-compaction to install from a local directory.
The compaction logic is also a TypeScript library (not published to npm; build it with npm run build and import from dist/):
import { compactMessages, reductionRatio, type Message } from 'fast-laya-compaction';
const result = await compactMessages(transcript, { preserveRecentMessages: 4 });
console.log(result.decisions, result.stats);
if (reductionRatio(result) < 0.25) {
// not worth it: keep the original transcript, or summarize instead
}
Message is a subset of Claude Code's SessionMessage. compactMessages runs the same Laya script through uv from Node. To bring your own scorer, implement LayaScorer and call compact(messages, scorer, options); collectToolCalls, supersededBy, callState, decideCall and applyDecisions are exported too.
types/claude-code.d.ts was generated by Claude Code 2.1.274.npm install
npm run typecheck && npm test # library, hook and scripts (vitest)
npm run test:backend # Python contract tests (fake model, no torch)
npm run validate:plugin
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir .
| Path | What |
|---|---|
hooks/fast-laya.ts | The Claude Code hook: config, the Laya process, fallback, auto-compact |
src/ | The library: call pairing, the superseded rule, per-call states, decisions |
backend/laya_compact.py | One-shot Laya scorer (stdin → stdout), run with uv run --script |
scripts/ | dry-run, benchmark ceilings, and the Jev teacher tools |
BENCHMARK.md | Measurements |
MIT. See LICENSE.
hooks/fast-laya.ts 358 lines1import type {
2 On,
3 PluginOptions,
4 Register,
5 SessionMessage,
6 ToolResultSummary,
7 ToolUseSummary,
8 TurnCompleteInput,
9} from 'claude-code';
10
11import { compact, maxReduction, reductionRatio, resolveOptions } from '../src/compact.js';
12import { buildLayaRequest, DEFAULT_MODEL, layaArgv, parseLayaResponse } from '../src/laya.js';
13import type {
14 CompactOptions,
15 CompactResult,
16 LayaScorer,
17 Message,
18 ToolResult,
19 ToolUse,
20} from '../src/types.js';
21
22const HOOK_DEFAULTS = {
23 compactAtPercent: 60,
24 minReductionRatio: 0.25,
25 model: DEFAULT_MODEL,
26 uvPath: 'uv',
27 timeoutSeconds: 120,
28 autoTargetReduction: 0.5,
29};
30
31/** Triggers where the window is full (or about to be): the compaction has to make room. */
32const MAKE_ROOM: ReadonlySet<string> = new Set(['auto', 'precompute']);
33
34export type HookRunInit = {
35 stdin?: string;
36 timeoutMs?: number;
37};
38
39export type HookRunResult = {
40 exitCode: number;
41 stdout: string;
42 stderr: string;
43};
44
45/** The shape of `$.process.run`, so the hook can be driven without an engine. */
46export type HookRun = (argv: readonly string[], init?: HookRunInit) => Promise<HookRunResult>;
47
48export type HookConfig = CompactOptions & {
49 compactAtPercent: number;
50 minReductionRatio: number;
51 model: string;
52 uvPath: string;
53 device?: string;
54 timeoutSeconds: number;
55 /** `targetReduction` on auto-compaction, where leaving too little room re-triggers it at once. */
56 autoTargetReduction: number;
57};
58
59function optionNumber(options: PluginOptions, key: string, fallback: number): number {
60 const value = options[key];
61 return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
62}
63
64function optionString(options: PluginOptions, key: string): string | undefined {
65 const value = options[key];
66 return typeof value === 'string' && value.length > 0 ? value : undefined;
67}
68
69/** Reads the plugin's `userConfig` values; anything missing takes the defaults. */
70export function resolveHookConfig(options: PluginOptions): HookConfig {
71 const numbers: Partial<Omit<CompactOptions, 'goal'>> = {};
72 for (const key of [
73 'keepThreshold',
74 'preserveRecentMessages',
75 'maxCallStateChars',
76 'maxScoredCalls',
77 'truncateHeadChars',
78 ] as const) {
79 const value = options[key];
80 if (typeof value === 'number' && Number.isFinite(value)) numbers[key] = value;
81 }
82 const config: HookConfig = {
83 ...numbers,
84 compactAtPercent: optionNumber(options, 'compactAtPercent', HOOK_DEFAULTS.compactAtPercent),
85 minReductionRatio: optionNumber(
86 options,
87 'minReductionRatio',
88 HOOK_DEFAULTS.minReductionRatio,
89 ),
90 // A fine-tuned checkpoint directory, when set, replaces the published checkpoint.
91 model: optionString(options, 'checkpointPath') ?? optionString(options, 'model') ?? HOOK_DEFAULTS.model,
92 uvPath: optionString(options, 'uvPath') ?? HOOK_DEFAULTS.uvPath,
93 timeoutSeconds: Math.max(
94 1,
95 optionNumber(options, 'timeoutSeconds', HOOK_DEFAULTS.timeoutSeconds),
96 ),
97 autoTargetReduction: Math.min(
98 1,
99 Math.max(0, optionNumber(options, 'autoTargetReduction', HOOK_DEFAULTS.autoTargetReduction)),
100 ),
101 };
102 const device = optionString(options, 'device');
103 if (device) config.device = device;
104 const goal = optionString(options, 'goal');
105 if (goal) config.goal = goal;
106 return config;
107}
108
109/**
110 * A `LayaScorer` over the engine's `$.process.run`: one short-lived Laya
111 * process per compaction, the request on stdin, the scores on stdout.
112 * `$.process.run` rejects both when the command cannot start and when it
113 * outlives the timeout; the elapsed time tells the two apart.
114 */
115export function processScorer(
116 run: HookRun,
117 argv: readonly string[],
118 config: Pick<HookConfig, 'model' | 'device' | 'timeoutSeconds'>,
119): LayaScorer {
120 return {
121 async score(states, question) {
122 const timeoutMs = config.timeoutSeconds * 1000;
123 const stdin = buildLayaRequest(states, question, { model: config.model, device: config.device });
124 const started = Date.now();
125 let result: HookRunResult;
126 try {
127 result = await run(argv, { stdin, timeoutMs });
128 } catch (error) {
129 if (Date.now() - started >= timeoutMs - 1000) {
130 throw new Error(`Laya timed out after ${config.timeoutSeconds}s`);
131 }
132 throw new Error(`Laya could not start (${error instanceof Error ? error.message : String(error)})`);
133 }
134 return parseLayaResponse(result.exitCode, result.stdout, result.stderr, argv[argv.length - 1]);
135 },
136 };
137}
138
139function toolUseSummary(tool: ToolUse): ToolUseSummary {
140 const summary: ToolUseSummary = {
141 tool_use_id: tool.tool_use_id,
142 tool: tool.tool,
143 input: tool.input,
144 };
145 if (tool.text !== undefined) summary.text = tool.text;
146 if (tool.isError) summary.isError = true;
147 return summary;
148}
149
150function toolResultSummary(result: ToolResult): ToolResultSummary {
151 return {
152 tool_use_id: result.tool_use_id,
153 text: result.text,
154 isError: result.isError ?? false,
155 };
156}
157
158/**
159 * Maps the library's output back onto session messages. Whatever came back
160 * unchanged (a message, a tool use, a tool result) is the engine's own object,
161 * handle included; anything rebuilt is a fresh message without a handle, so the
162 * engine takes the edited content instead of its original.
163 */
164export function toSessionMessages(
165 input: readonly SessionMessage[],
166 output: readonly Message[],
167): SessionMessage[] {
168 const messages = new Map<Message, SessionMessage>();
169 const uses = new Map<ToolUse, ToolUseSummary>();
170 const results = new Map<ToolResult, ToolResultSummary>();
171 for (const message of input) {
172 messages.set(message, message);
173 for (const tool of message.toolUses) uses.set(tool, tool);
174 for (const result of message.toolResults ?? []) results.set(result, result);
175 }
176 return output.map((message) => {
177 const own = messages.get(message);
178 if (own) return own;
179 const rebuilt: SessionMessage = {
180 role: message.role,
181 text: message.text,
182 toolUses: message.toolUses.map((tool) => uses.get(tool) ?? toolUseSummary(tool)),
183 };
184 if (message.toolResults && message.toolResults.length > 0) {
185 rebuilt.toolResults = message.toolResults.map(
186 (result) => results.get(result) ?? toolResultSummary(result),
187 );
188 }
189 return rebuilt;
190 });
191}
192
193export type SessionCompaction = {
194 result: CompactResult;
195 messages: SessionMessage[];
196};
197
198/** Runs the library over a session transcript; throws when Laya cannot run or fails. */
199export async function compactSession(
200 messages: readonly SessionMessage[],
201 config: HookConfig,
202 run: HookRun,
203 pluginRoot: string,
204 trigger?: string,
205): Promise<SessionCompaction> {
206 // Without room to gain there is nothing for Laya to decide: say so before starting it.
207 const ceiling = maxReduction(messages, config);
208 if (ceiling < config.minReductionRatio) {
209 throw new Error(`nothing worth trimming: at most ${percent(ceiling)} could go`);
210 }
211 const options = {
212 ...config,
213 targetReduction: trigger && MAKE_ROOM.has(trigger) ? config.autoTargetReduction : 0,
214 };
215 const scorer = processScorer(run, layaArgv(config.uvPath, pluginRoot), config);
216 const result = await compact(messages, scorer, options);
217 return { result, messages: toSessionMessages(messages, result.messages) };
218}
219
220function percent(ratio: number): string {
221 return `${Math.round(ratio * 100)}%`;
222}
223
224function seconds(ms: number): string {
225 return `${(ms / 1000).toFixed(1)}s`;
226}
227
228export function summarize(result: CompactResult): string {
229 const { stats } = result;
230 const parts = [
231 stats.kept > 0 ? `${stats.kept} kept` : '',
232 stats.resultsDropped > 0 ? `${stats.resultsDropped} results truncated` : '',
233 stats.callsDropped > 0 ? `${stats.callsDropped} call_dropped` : '',
234 stats.superseded > 0 ? `${stats.superseded} superseded calls removed` : '',
235 stats.pinned > 0 ? `${stats.pinned} pinned` : '',
236 stats.unscored > 0 ? `${stats.unscored} unscored` : '',
237 stats.trimmed > 0 ? `${stats.trimmed} trimmed for room` : '',
238 ].filter(Boolean);
239 const laya = stats.model
240 ? `${stats.model} on ${stats.device}, load ${seconds(stats.loadMs)} + infer ${seconds(stats.inferMs)}`
241 : 'Laya not called';
242 return `${percent(reductionRatio(result))} reduction; ${parts.join(', ') || 'no tool calls'}; ${laya}`;
243}
244
245const UI_LOG_MAX_CHARS = 4096;
246
247export function decisionLog(result: CompactResult): string {
248 return result.decisions
249 .filter((d) => d.reason !== 'pinned' && d.reason !== 'unscored')
250 .map((d) => {
251 const p = d.probabilities;
252 const what = d.reason === 'trimmed' ? 'trimmed' : d.action;
253 return `${d.id}:${d.tool}:${what}/${p ? `keep=${p.keep.toFixed(2)}/truncate=${p.truncate.toFixed(2)}` : d.reason}`;
254 })
255 .join(' ');
256}
257
258export function decisionLogLines(
259 result: CompactResult,
260 maxChars: number = UI_LOG_MAX_CHARS,
261): string[] {
262 const entries = decisionLog(result).split(' ').filter(Boolean);
263 if (entries.length === 0) return ['decisions: (none)'];
264 const chunks: string[] = [];
265 let current = '';
266 for (const entry of entries) {
267 const next = current ? `${current} ${entry}` : entry;
268 if (current && next.length > maxChars - 24) {
269 chunks.push(current);
270 current = entry;
271 } else current = next;
272 }
273 chunks.push(current);
274 return chunks.map((chunk, index) =>
275 chunks.length === 1
276 ? `decisions: ${chunk}`
277 : `decisions (${index + 1}/${chunks.length}): ${chunk}`,
278 );
279}
280
281function notify(
282 $: {
283 ui: {
284 log: (text: string) => void;
285 toast: (text: string, options?: { timeoutMs?: number }) => void;
286 };
287 },
288 text: string,
289): void {
290 $.ui.log(text);
291 $.ui.toast(text, { timeoutMs: 15_000 });
292}
293
294export const register: Register = (on: On, options: PluginOptions) => {
295 const config = resolveHookConfig(options);
296 let compacting = false;
297 // One Laya process at a time: each holds a checkpoint (~2-3 GB). A compaction
298 // arriving while another runs (an auto-compaction behind an ahead-of-time
299 // `precompute`, or a subagent's) waits its turn instead of losing Laya.
300 let queue: Promise<void> = Promise.resolve();
301
302 on('session.compact', async ($, event, next) => {
303 const previous = queue;
304 let done!: () => void;
305 queue = new Promise<void>((resolve) => (done = resolve));
306 await previous;
307 try {
308 const { result, messages } = await compactSession(
309 event.messages,
310 config,
311 (argv, init) => $.process.run(argv, init),
312 $.plugin.root,
313 event.trigger,
314 );
315 for (const line of decisionLogLines(result)) $.ui.log(line);
316 if (reductionRatio(result) < config.minReductionRatio) {
317 notify(
318 $,
319 `fallback to built-in summary (below ${percent(config.minReductionRatio)} minimum: ${summarize(result)})`,
320 );
321 return next(event);
322 }
323 notify(
324 $,
325 `kept ${messages.length}/${event.messages.length} messages, no summary (${summarize(result)})`,
326 );
327 return { messages };
328 } catch (error) {
329 notify(
330 $,
331 `fallback to built-in summary (${error instanceof Error ? error.message : String(error)})`,
332 );
333 return next(event);
334 } finally {
335 done();
336 }
337 });
338
339 on('turn.complete', async ($, event: TurnCompleteInput, next) => {
340 if (compacting) return next(event);
341 try {
342 const { context } = await $.session.usage();
343 if ((context.percent ?? 0) < config.compactAtPercent) return next(event);
344 compacting = true;
345 await $.session.compact();
346 } catch (error) {
347 $.ui.log(
348 `auto-compact skipped (${error instanceof Error ? error.message : String(error)})`,
349 );
350 } finally {
351 compacting = false;
352 }
353 return next(event);
354 });
355};
356
357export { resolveOptions };
358src/compact.ts 356 lines1import { callState, collectToolCalls, goalFromMessages, supersededBy } from './state.js';
2import type {
3 CallDecision,
4 CallProbabilities,
5 ChoiceQuestion,
6 CompactOptions,
7 CompactResult,
8 LayaScorer,
9 LayaScores,
10 Message,
11 ResolvedCompactOptions,
12 ToolCall,
13 ToolUse,
14} from './types.js';
15
16export const DEFAULT_OPTIONS: ResolvedCompactOptions = {
17 goal: '',
18 keepThreshold: 0.5,
19 preserveRecentMessages: 6,
20 maxCallStateChars: 1600,
21 maxScoredCalls: 80,
22 truncateHeadChars: 300,
23 targetReduction: 0,
24};
25
26function finite(value: number | undefined, fallback: number): number {
27 return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
28}
29
30export function resolveOptions(options: CompactOptions = {}): ResolvedCompactOptions {
31 return {
32 goal: options.goal ?? DEFAULT_OPTIONS.goal,
33 keepThreshold: finite(options.keepThreshold, DEFAULT_OPTIONS.keepThreshold),
34 preserveRecentMessages: Math.max(
35 0,
36 Math.floor(
37 finite(options.preserveRecentMessages, DEFAULT_OPTIONS.preserveRecentMessages),
38 ),
39 ),
40 maxCallStateChars: Math.max(
41 200,
42 Math.floor(finite(options.maxCallStateChars, DEFAULT_OPTIONS.maxCallStateChars)),
43 ),
44 maxScoredCalls: Math.max(
45 0,
46 Math.floor(finite(options.maxScoredCalls, DEFAULT_OPTIONS.maxScoredCalls)),
47 ),
48 truncateHeadChars: Math.max(
49 0,
50 Math.floor(finite(options.truncateHeadChars, DEFAULT_OPTIONS.truncateHeadChars)),
51 ),
52 targetReduction: Math.min(1, Math.max(0, finite(options.targetReduction, DEFAULT_OPTIONS.targetReduction))),
53 };
54}
55
56/**
57 * The one question asked about every scored call, each with its own state.
58 * The criteria are worded positively: Laya reads negations poorly. (A bare
59 * yes/no question measured worse on real sessions: less reduction, a less
60 * sensible ranking, and the same answers when asked the opposite way.)
61 */
62export const CALL_QUESTION: ChoiceQuestion = {
63 type: 'choice',
64 instructions:
65 "A coding assistant's conversation is being compacted. The state describes one earlier tool call. What should happen to it?",
66 criteria: {
67 keep: 'the assistant still needs the full output verbatim',
68 truncate: 'only the fact that this call was made still matters',
69 drop: 'the call is obsolete: superseded, exploratory, or already acted on',
70 },
71};
72
73/** Laya's answer for one call; throws when a probability is missing or out of range. */
74export function callProbabilities(scores: LayaScores['scores'], id: string): CallProbabilities {
75 const p = scores[id];
76 const valid = (x: unknown): x is number => typeof x === 'number' && x >= 0 && x <= 1;
77 if (!p || !valid(p.keep) || !valid(p.truncate) || !valid(p.drop)) {
78 throw new Error(`Invalid Laya answer for ${id}`);
79 }
80 return { keep: p.keep, truncate: p.truncate, drop: p.drop };
81}
82
83/**
84 * One call's fate. Pinned calls stay and superseded ones go with their
85 * result. For a scored call: `P(keep) ≥ threshold` keeps it whole, else
86 * `P(keep) + P(truncate) ≥ threshold` keeps the call with a truncated output,
87 * else it goes. An unscored call stays.
88 */
89export function decideCall(
90 call: Pick<ToolCall, 'id' | 'tool' | 'pinned'>,
91 verdict: { superseded?: boolean; probabilities?: CallProbabilities },
92 options: Pick<ResolvedCompactOptions, 'keepThreshold'>,
93): CallDecision {
94 const base = { id: call.id, tool: call.tool };
95 if (call.pinned) return { ...base, action: 'keep', reason: 'pinned' };
96 if (verdict.superseded) return { ...base, action: 'drop_call', reason: 'superseded' };
97 const p = verdict.probabilities;
98 if (!p) return { ...base, action: 'keep', reason: 'unscored' };
99 if (p.keep >= options.keepThreshold) return { ...base, probabilities: p, action: 'keep', reason: 'kept' };
100 if (p.keep + p.truncate >= options.keepThreshold) {
101 return { ...base, probabilities: p, action: 'drop_result', reason: 'result_dropped' };
102 }
103 return { ...base, probabilities: p, action: 'drop_call', reason: 'call_dropped' };
104}
105
106function truncatedResultText(text: string, isError: boolean, headChars: number): string {
107 if (text.length <= headChars + 120) return text;
108 const head = headChars > 0 ? `${text.slice(0, headChars)}\n` : '';
109 return `${head}[fast-laya-compaction truncated ${text.length - headChars} chars of this tool result${
110 isError ? ' (error)' : ''
111 }; re-run the tool if needed]`;
112}
113
114/**
115 * Rebuilds the conversation from the decisions. A dropped call disappears
116 * together with its result; a dropped result keeps a bounded head and note.
117 * Messages that lose all their content are removed; untouched messages are
118 * returned as the same objects they came in as.
119 */
120export function applyDecisions(
121 messages: readonly Message[],
122 decisions: readonly CallDecision[],
123 calls: readonly ToolCall[],
124 headChars: number,
125): Message[] {
126 const byId = new Map(calls.map((call) => [call.id, call]));
127 const actions = new Map<string, CallDecision['action']>();
128 for (const decision of decisions) {
129 const call = byId.get(decision.id);
130 if (call && decision.action !== 'keep') actions.set(call.tool_use_id, decision.action);
131 }
132 const kept: Message[] = [];
133 for (const message of messages) {
134 const touched =
135 message.toolUses.some((tool) => actions.has(tool.tool_use_id)) ||
136 (message.toolResults ?? []).some((result) => actions.has(result.tool_use_id));
137 if (!touched) {
138 kept.push(message);
139 continue;
140 }
141 const toolUses = message.toolUses
142 .filter((tool) => actions.get(tool.tool_use_id) !== 'drop_call')
143 .map((tool) => {
144 if (actions.get(tool.tool_use_id) !== 'drop_result') return tool;
145 const text = truncatedResultText(
146 tool.text ?? '',
147 tool.isError ?? false,
148 headChars,
149 );
150 if ((tool.text ?? '') === text) return tool;
151 const copy: ToolUse = {
152 tool_use_id: tool.tool_use_id,
153 tool: tool.tool,
154 input: tool.input,
155 text,
156 };
157 if (tool.isError) copy.isError = true;
158 return copy;
159 });
160 const toolResults = (message.toolResults ?? [])
161 .filter((result) => actions.get(result.tool_use_id) !== 'drop_call')
162 .map((result) => {
163 if (actions.get(result.tool_use_id) !== 'drop_result') return result;
164 const text = truncatedResultText(result.text, result.isError ?? false, headChars);
165 return text === result.text
166 ? result
167 : {
168 tool_use_id: result.tool_use_id,
169 text,
170 isError: result.isError,
171 };
172 });
173 if (
174 !message.toolUses.some(
175 (tool) => actions.get(tool.tool_use_id) === 'drop_call',
176 ) &&
177 !(message.toolResults ?? []).some(
178 (result) => actions.get(result.tool_use_id) === 'drop_call',
179 ) &&
180 toolUses.every((tool, index) => tool === message.toolUses[index]) &&
181 toolResults.every(
182 (result, index) => result === message.toolResults?.[index],
183 )
184 ) {
185 kept.push(message);
186 continue;
187 }
188 if (message.text.trim().length === 0 && toolUses.length === 0 && toolResults.length === 0) {
189 continue;
190 }
191 const rebuilt: Message = { role: message.role, text: message.text, toolUses };
192 if (toolResults.length > 0) rebuilt.toolResults = toolResults;
193 kept.push(rebuilt);
194 }
195 return kept;
196}
197
198/** Characters of text, tool input and tool output a message holds. */
199export function messageChars(message: Message): number {
200 let total = message.text.length;
201 for (const tool of message.toolUses) {
202 try {
203 total += JSON.stringify(tool.input).length;
204 } catch {
205 total += 20;
206 }
207 }
208 for (const result of message.toolResults ?? []) total += result.text.length;
209 return total;
210}
211
212export function reductionRatio(result: Pick<CompactResult, 'stats'>): number {
213 const { charsBefore, charsAfter } = result.stats;
214 return charsBefore === 0 ? 0 : (charsBefore - charsAfter) / charsBefore;
215}
216
217function count(decisions: readonly CallDecision[], reason: CallDecision['reason']): number {
218 return decisions.filter((decision) => decision.reason === reason).length;
219}
220
221function charsOf(messages: readonly Message[]): number {
222 return messages.reduce((sum, message) => sum + messageChars(message), 0);
223}
224
225/**
226 * The calls a compaction looks at: all of them paired, the superseded ones
227 * (removed by rule) and the candidates Laya scores, oldest first.
228 */
229function plan(messages: readonly Message[], resolved: ResolvedCompactOptions) {
230 const calls = collectToolCalls(messages, resolved.preserveRecentMessages);
231 const superseded = new Set(
232 calls.filter((call) => !call.pinned && supersededBy(call, calls)).map((call) => call.id),
233 );
234 // Oldest first: they are the likeliest to be stale, and scoring time grows with every call.
235 const candidates = calls
236 .filter((call) => !call.pinned && !superseded.has(call.id))
237 .slice(0, resolved.maxScoredCalls);
238 return { calls, superseded, candidates };
239}
240
241/**
242 * The most a compaction could remove, before asking Laya: every superseded
243 * call gone and every candidate's output truncated. When even this is too
244 * little, there is no point starting Laya.
245 */
246export function maxReduction(messages: readonly Message[], options: CompactOptions = {}): number {
247 const resolved = resolveOptions(options);
248 const { calls, superseded, candidates } = plan(messages, resolved);
249 const scored = new Set(candidates.map((call) => call.id));
250 const decisions = calls.map((call): CallDecision => {
251 const base = { id: call.id, tool: call.tool };
252 if (superseded.has(call.id)) return { ...base, action: 'drop_call', reason: 'superseded' };
253 if (scored.has(call.id)) return { ...base, action: 'drop_result', reason: 'result_dropped' };
254 return { ...base, action: 'keep', reason: call.pinned ? 'pinned' : 'unscored' };
255 });
256 const before = charsOf(messages);
257 return before === 0 ? 0 : 1 - charsOf(applyDecisions(messages, decisions, calls, resolved.truncateHeadChars)) / before;
258}
259
260/**
261 * Truncates outputs Laya would keep, the lowest `P(keep)` first, until the
262 * reduction reaches `targetReduction`. Laya ranks calls better than it
263 * calibrates them, so when room is needed its order decides what goes.
264 */
265function trimToTarget(
266 messages: readonly Message[],
267 decisions: readonly CallDecision[],
268 calls: readonly ToolCall[],
269 resolved: ResolvedCompactOptions,
270): CallDecision[] {
271 const out = [...decisions];
272 const before = charsOf(messages);
273 const reached = () =>
274 before === 0 ||
275 1 - charsOf(applyDecisions(messages, out, calls, resolved.truncateHeadChars)) / before >= resolved.targetReduction;
276 if (resolved.targetReduction <= 0 || reached()) return out;
277 const order = out
278 .map((decision, index) => ({ decision, index }))
279 .filter(({ decision }) => decision.reason === 'kept')
280 .sort((a, b) => a.decision.probabilities!.keep - b.decision.probabilities!.keep);
281 for (const { decision, index } of order) {
282 out[index] = { ...decision, action: 'drop_result', reason: 'trimmed' };
283 if (reached()) break;
284 }
285 return out;
286}
287
288/**
289 * Compacts a transcript. Outside the pinned first and newest messages, a
290 * read-only call that a later call repeated or made stale is removed with
291 * its result; of the rest, the oldest `maxScoredCalls` are each described by
292 * their own small state (see `callState`) and Laya is asked, in one batched
293 * request, whether each should be kept, truncated or dropped. With a
294 * `targetReduction`, outputs Laya would keep are then truncated, lowest
295 * `P(keep)` first, until it is met. Throws when Laya fails; the caller decides
296 * whether to fall back.
297 */
298export async function compact(
299 messages: readonly Message[],
300 scorer: LayaScorer,
301 options: CompactOptions = {},
302): Promise<CompactResult> {
303 const started = Date.now();
304 const resolved = resolveOptions(options);
305 const { calls, superseded, candidates } = plan(messages, resolved);
306 const charsBefore = charsOf(messages);
307
308 let laya: Omit<LayaScores, 'scores'> = { model: '', device: '', loadMs: 0, inferMs: 0 };
309 const answers = new Map<string, CallProbabilities>();
310 if (candidates.length > 0) {
311 const goal = resolved.goal || goalFromMessages(messages);
312 const states = candidates.map((call) => ({
313 id: call.id,
314 state: callState(call, messages, calls, goal, resolved.maxCallStateChars),
315 }));
316 const { model, device, loadMs, inferMs, scores } = await scorer.score(states, CALL_QUESTION);
317 laya = { model, device, loadMs, inferMs };
318 for (const call of candidates) answers.set(call.id, callProbabilities(scores, call.id));
319 }
320
321 const decisions = trimToTarget(
322 messages,
323 calls.map((call) =>
324 decideCall(call, { superseded: superseded.has(call.id), probabilities: answers.get(call.id) }, resolved),
325 ),
326 calls,
327 resolved,
328 );
329 const kept = applyDecisions(
330 messages,
331 decisions,
332 calls,
333 resolved.truncateHeadChars,
334 );
335 return {
336 messages: kept,
337 decisions,
338 stats: {
339 messagesBefore: messages.length,
340 messagesAfter: kept.length,
341 charsBefore,
342 charsAfter: charsOf(kept),
343 calls: calls.length,
344 kept: count(decisions, 'kept'),
345 resultsDropped: count(decisions, 'result_dropped'),
346 callsDropped: count(decisions, 'call_dropped'),
347 superseded: count(decisions, 'superseded'),
348 pinned: count(decisions, 'pinned'),
349 unscored: count(decisions, 'unscored'),
350 trimmed: count(decisions, 'trimmed'),
351 ...laya,
352 ms: Date.now() - started,
353 },
354 };
355}
356src/laya.ts 97 lines1import type { CallState, ChoiceQuestion, LayaScores } from './types.js';
2
3/** Laya's checkpoint fine-tuned for typed decisions; the base ones are near chance zero-shot. */
4export const DEFAULT_MODEL = 'typed-decisions';
5
6/** The scoring script, relative to the plugin (and package) root. */
7export const SCRIPT_PATH = 'backend/laya_compact.py';
8
9/** What to run once before the first compaction; the hook names the installed script's absolute path. */
10export function setupHint(script: string = SCRIPT_PATH): string {
11 return `Laya is not set up, run once: uv run --script ${script} --warmup`;
12}
13
14export const SETUP_HINT = setupHint();
15
16/** The script's exit code for a checkpoint missing from the Hugging Face cache. */
17const EXIT_NOT_CACHED = 3;
18
19/** How uv reports a script environment it cannot build under `--offline`. */
20const UV_OFFLINE = /network was disabled|not found in the cache/i;
21
22export interface LayaRequestOptions {
23 /** `typed-decisions` (default), `multilingual` or `english`. */
24 model?: string;
25 /** Torch device; auto (mps on Apple Silicon) when absent. */
26 device?: string;
27 /** Tokens per question row: head (question and options) plus state. */
28 maxLen?: number;
29 headMaxLen?: number;
30 /** States per forward pass. */
31 batchSize?: number;
32}
33
34/**
35 * The command that scores one compaction. `--offline` keeps a hook from ever
36 * building the torch environment inline: without the one-time setup it fails
37 * fast and the caller falls back.
38 */
39export function layaArgv(uvPath: string, root: string): string[] {
40 return [uvPath, 'run', '--quiet', '--offline', '--script', `${root}/${SCRIPT_PATH}`];
41}
42
43/** The script's stdin for one compaction: one question over every call state. */
44export function buildLayaRequest(
45 states: readonly CallState[],
46 question: ChoiceQuestion,
47 options: LayaRequestOptions = {},
48): string {
49 return JSON.stringify({
50 model: options.model ?? DEFAULT_MODEL,
51 ...(options.device ? { device: options.device } : {}),
52 max_len: options.maxLen ?? 512,
53 head_max_len: options.headMaxLen ?? 128,
54 batch_size: options.batchSize ?? 32,
55 question,
56 states,
57 });
58}
59
60function lastLine(text: string): string {
61 const lines = text.split('\n').map((line) => line.trim()).filter(Boolean);
62 return (lines[lines.length - 1] ?? '').slice(0, 200);
63}
64
65/** Validates the script's output; throws a message fit for a toast on anything else. */
66export function parseLayaResponse(
67 exitCode: number,
68 stdout: string,
69 stderr: string,
70 script: string = SCRIPT_PATH,
71): LayaScores {
72 if (exitCode === EXIT_NOT_CACHED || (exitCode !== 0 && UV_OFFLINE.test(stderr))) {
73 throw new Error(setupHint(script));
74 }
75 if (exitCode !== 0) {
76 throw new Error(`Laya exited with ${exitCode}: ${lastLine(stderr) || 'no output'}`);
77 }
78 let parsed: unknown;
79 try {
80 // The answer is the last line; anything a library printed before it is noise.
81 parsed = JSON.parse(stdout.trim().split('\n').pop() ?? '');
82 } catch {
83 throw new Error('Laya returned malformed JSON');
84 }
85 const body = parsed as Record<string, unknown> | null;
86 if (!body || typeof body !== 'object' || !body.scores || typeof body.scores !== 'object') {
87 throw new Error('Laya response is missing scores');
88 }
89 return {
90 model: String(body.model ?? ''),
91 device: String(body.device ?? ''),
92 loadMs: Number(body.load_ms) || 0,
93 inferMs: Number(body.infer_ms) || 0,
94 scores: body.scores as LayaScores['scores'],
95 };
96}
97src/types.ts 161 lines1export type Role = 'user' | 'assistant';
2
3/**
4 * A tool_use block of an assistant message. `text` and `isError` mirror the
5 * outcome once the transcript holds it (Claude Code attaches them).
6 */
7export interface ToolUse {
8 tool_use_id: string;
9 tool: string;
10 input: Record<string, unknown>;
11 text?: string;
12 isError?: boolean;
13}
14
15/** A tool_result block of a user message. */
16export interface ToolResult {
17 tool_use_id: string;
18 text: string;
19 isError?: boolean;
20}
21
22/**
23 * One transcript message. The shape is a subset of Claude Code's
24 * `SessionMessage`, so a session transcript can be passed in as is.
25 */
26export interface Message {
27 role: Role;
28 text: string;
29 toolUses: ToolUse[];
30 toolResults?: ToolResult[];
31}
32
33/** A tool call paired with its result by `tool_use_id`. */
34export interface ToolCall {
35 /** Short id used in the Laya request and decision log (`t1`, `t2`, ...). */
36 id: string;
37 tool_use_id: string;
38 tool: string;
39 input: Record<string, unknown>;
40 /** Index of the message holding the tool_use block. */
41 callIndex: number;
42 /** Index of the message holding the tool_result block. */
43 resultIndex: number;
44 resultChars: number;
45 isError: boolean;
46 /** In the first or the newest preserved messages; never a candidate. */
47 pinned: boolean;
48}
49
50export type CallAction = 'keep' | 'drop_result' | 'drop_call';
51
52/** Laya's answer for one call: keep it whole, truncate its output, or drop it. */
53export interface CallProbabilities {
54 keep: number;
55 truncate: number;
56 drop: number;
57}
58
59export interface CallDecision {
60 id: string;
61 tool: string;
62 /** Laya's answer; only on scored calls. */
63 probabilities?: CallProbabilities;
64 action: CallAction;
65 /**
66 * `superseded`: a later call repeated it or changed its target, so it is
67 * removed without asking Laya. `kept`/`result_dropped`/`call_dropped`: Laya's
68 * answer. `trimmed`: Laya would keep it, but its output was truncated to reach
69 * `targetReduction`.
70 */
71 reason: 'pinned' | 'superseded' | 'unscored' | 'kept' | 'result_dropped' | 'call_dropped' | 'trimmed';
72}
73
74export interface CompactOptions {
75 /** Ongoing task description; defaults to the latest user prompt. */
76 goal?: string;
77 /** Minimum keep probability for a call or its full output to stay. Default 0.5. */
78 keepThreshold?: number;
79 /** Newest messages never touched (the first message is always kept). Default 6. */
80 preserveRecentMessages?: number;
81 /** Character cap on the state Laya reads for one call. Default 1600. */
82 maxCallStateChars?: number;
83 /** Most calls scored per compaction, oldest first; newer ones are kept. Bounds latency. Default 80. */
84 maxScoredCalls?: number;
85 /** Characters of a dropped tool result to retain. Default 300. */
86 truncateHeadChars?: number;
87 /**
88 * Reduction to reach even past Laya's answers: outputs Laya would keep are
89 * truncated, least likely to be needed first, until it is met. 0 (default) off.
90 */
91 targetReduction?: number;
92}
93
94export interface ResolvedCompactOptions {
95 goal: string;
96 keepThreshold: number;
97 preserveRecentMessages: number;
98 maxCallStateChars: number;
99 maxScoredCalls: number;
100 truncateHeadChars: number;
101 targetReduction: number;
102}
103
104export interface CompactResult {
105 /** The compacted transcript; untouched messages are the input objects. */
106 messages: Message[];
107 decisions: CallDecision[];
108 stats: {
109 messagesBefore: number;
110 messagesAfter: number;
111 charsBefore: number;
112 charsAfter: number;
113 calls: number;
114 kept: number;
115 resultsDropped: number;
116 /** Calls Laya dropped with their results. */
117 callsDropped: number;
118 /** Calls removed with their results because a later call superseded them. */
119 superseded: number;
120 pinned: number;
121 /** Kept without scoring: past `maxScoredCalls`. */
122 unscored: number;
123 /** Outputs Laya would keep, truncated to reach `targetReduction`. */
124 trimmed: number;
125 /** Laya checkpoint and torch device that scored the calls; '' when Laya was not called. */
126 model: string;
127 device: string;
128 /** Checkpoint load and batched inference time inside the Laya process. */
129 loadMs: number;
130 inferMs: number;
131 ms: number;
132 };
133}
134
135/** A Laya `choice` question: one label per criterion, answered with a probability each. */
136export interface ChoiceQuestion {
137 type: 'choice';
138 instructions: string;
139 criteria: Record<string, string>;
140}
141
142/** What Laya reads about one tool call; key order matters, Laya cuts from the end. */
143export interface CallState {
144 id: string;
145 state: Record<string, unknown>;
146}
147
148export interface LayaScores {
149 model: string;
150 device: string;
151 loadMs: number;
152 inferMs: number;
153 /** Per call id, the probability of every criterion label. */
154 scores: Record<string, Record<string, number>>;
155}
156
157/** Anything that answers one question over many call states: a Laya process, or a fake. */
158export interface LayaScorer {
159 score(states: readonly CallState[], question: ChoiceQuestion): Promise<LayaScores>;
160}
161src/state.ts 211 lines1import type { Message, ToolCall, ToolResult } from './types.js';
2
3/** Cap on the serialised tool input in a one-line call. */
4const INPUT_CHARS = 200;
5const GOAL_CHARS = 300;
6const INTENT_CHARS = 200;
7const OUTPUT_HEAD_CHARS = 400;
8const LATER_CALLS = 3;
9
10/** Input keys that name what a call operates on; equal targets mean the same file, command or search. */
11const TARGET_KEYS = ['file_path', 'notebook_path', 'path', 'url', 'command', 'pattern'] as const;
12
13/** Tools that change their target. A change is never superseded: the record that it was made still matters. */
14const MUTATING_TOOLS = new Set(['Edit', 'MultiEdit', 'Write', 'NotebookEdit']);
15
16/** Fields shortened, then removed, in this order when a state exceeds its cap: least decisive first. */
17const TRIM_ORDER = ['output_head', 'intent', 'goal', 'later_calls'] as const;
18
19export function truncate(text: string, limit: number): string {
20 return text.length <= limit ? text : `${text.slice(0, Math.max(0, limit - 1))}…`;
21}
22
23export function isPinned(
24 index: number,
25 total: number,
26 preserveRecentMessages: number,
27): boolean {
28 return index === 0 || index >= total - preserveRecentMessages;
29}
30
31/**
32 * Pairs every tool_use with its tool_result by `tool_use_id`. Calls without a
33 * result are not candidates (there is nothing to drop yet).
34 */
35export function collectToolCalls(
36 messages: readonly Message[],
37 preserveRecentMessages: number,
38): ToolCall[] {
39 const results = new Map<string, { index: number; result: ToolResult }>();
40 messages.forEach((message, index) => {
41 for (const result of message.toolResults ?? []) {
42 results.set(result.tool_use_id, { index, result });
43 }
44 });
45 const calls: ToolCall[] = [];
46 messages.forEach((message, callIndex) => {
47 for (const tool of message.toolUses) {
48 const found = results.get(tool.tool_use_id);
49 if (!found) continue;
50 calls.push({
51 id: `t${calls.length + 1}`,
52 tool_use_id: tool.tool_use_id,
53 tool: tool.tool,
54 input: tool.input,
55 callIndex,
56 resultIndex: found.index,
57 resultChars: found.result.text.length,
58 isError: found.result.isError ?? false,
59 pinned:
60 isPinned(callIndex, messages.length, preserveRecentMessages) ||
61 isPinned(found.index, messages.length, preserveRecentMessages),
62 });
63 }
64 });
65 return calls;
66}
67
68function inputText(value: unknown, limit: number): string {
69 let json = '';
70 try {
71 json = JSON.stringify(value) ?? String(value);
72 } catch {
73 json = '[unserializable input]';
74 }
75 return truncate(json, limit);
76}
77
78/** One call as a single line, target first, e.g. `Edit file_path=src/a.ts old_string=… replace_all=false`. */
79export function callLine(call: Pick<ToolCall, 'tool' | 'input'>): string {
80 const isTarget = (key: string) => (TARGET_KEYS as readonly string[]).includes(key);
81 const entries = Object.entries(call.input);
82 const input = [...entries.filter(([key]) => isTarget(key)), ...entries.filter(([key]) => !isTarget(key))]
83 .map(([key, value]) => {
84 const text = typeof value === 'string' ? value : inputText(value, INPUT_CHARS);
85 return `${key}=${text.replace(/\s+/g, ' ')}`;
86 })
87 .join(' ');
88 return truncate(`${call.tool} ${input}`.trim(), INPUT_CHARS);
89}
90
91/** What a call operates on (file, command, search), or undefined when its input names nothing. */
92export function callTarget(call: Pick<ToolCall, 'input'>): string | undefined {
93 const parts = TARGET_KEYS.flatMap((key) => {
94 const value = call.input[key];
95 return typeof value === 'string' && value.trim() ? [value.trim()] : [];
96 });
97 return parts.length > 0 ? parts.join(' ') : undefined;
98}
99
100/** The same tool with the same input, a free-text `description` aside. */
101function isRepeat(a: ToolCall, b: ToolCall): boolean {
102 const { description: _a, ...inputA } = a.input;
103 const { description: _b, ...inputB } = b.input;
104 return a.tool === b.tool && inputText(inputA, Infinity) === inputText(inputB, Infinity);
105}
106
107/**
108 * The later call that makes a read-only call obsolete: an exact repeat (the
109 * same read, search or command run again) or a change to the same target,
110 * after which the earlier output is stale. Such a call can be removed with
111 * its result without asking Laya.
112 */
113export function supersededBy(call: ToolCall, calls: readonly ToolCall[]): ToolCall | undefined {
114 if (MUTATING_TOOLS.has(call.tool)) return undefined;
115 const target = callTarget(call);
116 if (!target) return undefined;
117 return calls.find(
118 (later) =>
119 later.callIndex > call.callIndex &&
120 callTarget(later) === target &&
121 (MUTATING_TOOLS.has(later.tool) || isRepeat(later, call)),
122 );
123}
124
125/** The latest user prompts (tool-result messages excluded), as the default `goal`. */
126export function goalFromMessages(messages: readonly Message[], count = 1): string {
127 return messages
128 .filter(
129 (message) =>
130 message.role === 'user' &&
131 message.text.trim().length > 0 &&
132 (message.toolResults ?? []).length === 0,
133 )
134 .slice(-count)
135 .map((message) => truncate(message.text, GOAL_CHARS))
136 .join('\n');
137}
138
139/** The assistant text that led to a call: its own message's, else the nearest before it this turn. */
140function intentOf(messages: readonly Message[], callIndex: number): string {
141 for (let i = callIndex; i >= 0; i--) {
142 const message = messages[i]!;
143 if (message.role === 'user' && (message.toolResults ?? []).length === 0) break;
144 if (message.role === 'assistant' && message.text.trim()) return message.text.trim();
145 }
146 return '';
147}
148
149function resultText(messages: readonly Message[], call: ToolCall): string {
150 return (
151 messages[call.resultIndex]?.toolResults?.find((r) => r.tool_use_id === call.tool_use_id)?.text ?? ''
152 );
153}
154
155/** Later assistant text mentions the target (a file by its base name). */
156function mentionedLater(messages: readonly Message[], call: ToolCall, target: string): boolean {
157 const needle = target.includes('/') ? target.slice(target.lastIndexOf('/') + 1) : target;
158 if (needle.length < 3) return false;
159 return messages
160 .slice(call.resultIndex + 1)
161 .some((message) => message.role === 'assistant' && message.text.includes(needle));
162}
163
164function capState(state: Record<string, unknown>, maxChars: number): Record<string, unknown> {
165 for (const key of TRIM_ORDER) {
166 const over = JSON.stringify(state).length - maxChars;
167 if (over <= 0) break;
168 const value = state[key];
169 if (typeof value === 'string' && value.length > over + 20) {
170 state[key] = truncate(value, value.length - over);
171 } else delete state[key];
172 }
173 return state;
174}
175
176/**
177 * What Laya reads about one tool call. Laya keeps only the head of a long
178 * state, so the most decisive facts come first: the call, its outcome, how
179 * long ago it was, and whether later calls touched the same target.
180 * Context and a glimpse of the output follow, and are the first to go under
181 * `maxChars`.
182 */
183export function callState(
184 call: ToolCall,
185 messages: readonly Message[],
186 calls: readonly ToolCall[],
187 goal: string,
188 maxChars: number,
189): Record<string, unknown> {
190 const target = callTarget(call);
191 const later = target
192 ? calls.filter((other) => other.callIndex > call.callIndex && callTarget(other) === target)
193 : [];
194 const state: Record<string, unknown> = {
195 tool_call: callLine(call),
196 outcome: `${call.isError ? 'error' : 'ok'}, ${call.resultChars} chars`,
197 messages_since: messages.length - 1 - call.resultIndex,
198 // Laya reads key names as words: renaming this `touched_later` cut the
199 // measured reduction from 33% to 23% on one session. Keep it.
200 superseded: later.length > 0,
201 };
202 if (later.length > 0) state.later_calls = later.slice(0, LATER_CALLS).map(callLine);
203 if (target) state.mentioned_later = mentionedLater(messages, call, target);
204 if (goal) state.goal = truncate(goal, GOAL_CHARS);
205 const intent = intentOf(messages, call.callIndex);
206 if (intent) state.intent = truncate(intent, INTENT_CHARS);
207 const output = resultText(messages, call);
208 if (output) state.output_head = truncate(output, OUTPUT_HEAD_CHARS);
209 return capState(state, maxChars);
210}
211