SLOPSHOPPER

fast-laya-compaction

Verbatim Laya-guided compaction for Claude Code sessions (local, no API key).

newtoastprocess
v0.2.0MITupdated 2026-09-27kartikeyaagr/fast-laya-compaction
A shopper browsing a rack in a slop shop
README

fast-laya-compaction

Claude Code compaction that keeps your conversation verbatim and trims old tool output instead of summarizing, decided by a small model running on your own machine.

License: MIT

About

Claude Code's built-in compaction asks a model to summarize the conversation. Summaries are lossy: a file path, an exact error or a constraint can vanish even when it matters later. This plugin never rewrites anything. When the context fills up, it removes or truncates old tool calls and their outputs that are no longer needed, and leaves every user and assistant message exactly as written.

The decisions come from two places:

  • A rule removes a read-only call that a later call repeated exactly, or whose file a later edit changed.
  • Laya, a 421M-parameter classifier that runs locally, scores every other old call as keep, truncate (keep a 300-character head) or drop.

No API key, no network at compaction time, nothing leaves your machine. It is a port of fast-jev-compaction, which asks TypeSafe's hosted Jev model instead.

Results on three real sessions: 27–33% smaller in 19–47 s, against 79 s for a built-in compaction on the same machine. Details, and what did not work, in BENCHMARK.md.

How it works

  1. Tool calls are paired with their results. The first message and the newest preserveRecentMessages messages are never touched.
  2. Superseded calls are removed with their results, without a model: a read-only call (Read, Grep, Bash, WebFetch, …) repeated exactly later, or a read of a file a later Edit/Write changed. Edits and writes themselves are never removed.
  3. The oldest maxScoredCalls remaining calls each get a small state for Laya: the call, its outcome, how long ago it ran, whether later calls touched the same target, the current goal and the start of its output.
  4. Laya answers keep / truncate / drop for all of them in one batched pass. P(keep) ≥ keepThreshold keeps the call; otherwise P(keep) + P(truncate) ≥ keepThreshold keeps the call with a truncated output; otherwise it is removed.
  5. On auto-compaction the window is full, so the hook makes room: if Laya's answers free less than autoTargetReduction (50%), it also truncates outputs Laya would keep, lowest P(keep) first. Otherwise the window would fill again within a few tool calls and the next auto-compaction would fall back to a summary. /compact and the plugin's own 60% trigger keep Laya's answers as they are.
  6. If the result is at least minReductionRatio smaller, it replaces the history. Otherwise, or if anything fails, Claude Code's built-in summary runs as usual. When even trimming everything could not reach that minimum, Laya is not started at all.

It runs on every compaction trigger: /compact, Claude Code's auto-compaction (at its threshold or when a prompt is too long), its ahead-of-time precompute, and the plugin's own trigger at compactAtPercent. The hook runs backend/laya_compact.py once per compaction with uv run --offline, so no model stays in memory between compactions; a compaction arriving while another is scoring waits its turn.

Requirements

  • Claude Code 2.1.274+ (function hooks, early access)
  • uv
  • ~1.6 GB of disk for torch and the checkpoint, ~2–3 GB of free memory while compacting
  • Apple Silicon (mps), an NVIDIA GPU (cuda) or CPU

Install

  1. Enable function hooks wherever Claude Code runs, for example in ~/.claude/settings.json:
   { "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" } }
  1. Install the plugin:
   claude plugin marketplace add kartikeyaagr/fast-laya-compaction
   claude plugin install fast-laya-compaction@fast-laya-compaction
  1. Download Laya once (torch and ~843 MB of weights, pinned to laya==0.3.20 and a reviewed Hugging Face commit). The hook never downloads; until this has run, compaction falls back to the built-in summary with a toast that shows this exact command:
   uv run --script ~/.claude/plugins/cache/fast-laya-compaction/fast-laya-compaction/<version>/backend/laya_compact.py --warmup

It ends with laya_compact: ready (typed-decisions on mps, …).

  1. Check it: in a session with some tool calls, run /compact. The toast reads kept N/M messages, no summary (…), or fallback to built-in summary (<reason>).

Configuration

Set options with /plugin configure fast-laya-compaction@fast-laya-compaction inside Claude Code, or at install time with claude plugin install … --config keepThreshold=0.55.

OptionDefaultWhat it does
keepThreshold0.5Minimum probability for a call (or its full output) to stay; higher trims more
preserveRecentMessages6Newest messages never touched (the first is always kept)
compactAtPercent60Context percentage at which the plugin requests compaction
minReductionRatio0.25Minimum reduction to replace the history instead of summarizing
autoTargetReduction0.5On auto-compaction, the reduction to reach even past Laya's answers; 0 turns it off
truncateHeadChars300Characters of a truncated output kept before its note
maxScoredCalls80Most calls scored per compaction, oldest first; bounds the time
timeoutSeconds120Longest a Laya run may take before falling back
modeltyped-decisionsLaya checkpoint: typed-decisions, multilingual (faster, weaker) or english
checkpointPathunsetAbsolute path of a fine-tuned checkpoint; overrides model
maxCallStateChars1600Size cap on what Laya reads about one call
goallatest user promptTask description Laya weighs calls against
uvPathuvuv executable, if it is not on Claude Code's PATH
deviceautomps, cpu or cuda

Tune before you trust it

Laya is only weakly decisive on this task: its keep probabilities mostly fall between 0.3 and 0.65, ranking edits and writes above exploratory ls, git and search calls, and it rarely chooses drop. See what it would do to your own sessions before relying on it. The dry-run changes nothing:

git clone https://github.com/kartikeyaagr/fast-laya-compaction && cd fast-laya-compaction
npm install
npm run dry-run -- ~/.claude/projects/<project>/<session>.jsonl --threshold 0.5

It prints every call with its keep/truncate/drop probabilities and action, the reduction, the timings, and whether the hook would replace the history. Add --target 0.5 to see what an auto-compaction would do.

Getting closer to Jev

Prompt changes move Laya's answers a lot (renaming one key in its input changed a session from 33% to 23%) without saying which way is closer to Jev, and Laya cannot read a whole conversation the way Jev does (measured). So Jev is used as a teacher:

  1. Label sessions with Jev. Put TYPESAFE_API_KEY=… in .env (git-ignored). This sends those transcripts to TypeSafe and stores Jev's per-call answers in data/ (git-ignored). Aim for 20–40 sessions: npm run jev-labels -- ~/.claude/projects/<project>/*.jsonl
  2. Score Laya against Jev and compare prompt variants from scripts/variants.ts: npm run agreement -- --variant timeline. It reports decision agreement, a Jev-by-Laya confusion table, rank correlation and reduction under both. Jev's two probabilities map onto Laya's three so that matching the target reproduces Jev's decision at any threshold.
  3. Export training data with the best variant: npm run export-training -- --variant <best> writes rows in the format of Laya's own training notebook, with whole sessions held out.
  4. Fine-tune on a CUDA GPU (a free Kaggle 2×T4 session; an 8 GB Mac cannot) with Laya's notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb, changed to read laya-train.jsonl, start from the typed-decisions checkpoint, keep max_len/head_max_len at 512/128, fit the temperature on laya-holdout.jsonl, and delete temperature_by_options from the saved rl_agent_config.json. This step has not been run here.
  5. Use it: npm run agreement -- --model /abs/path/to/checkpoint, then set the checkpointPath option.

Publishing

To ship your own fork or a new version:

  1. Bump version in .claude-plugin/plugin.json, .claude-plugin/marketplace.json and package.json.
  2. npm run typecheck && npm test && npm run test:backend && npm run validate:plugin.
  3. Commit and push to GitHub. The repository is its own marketplace (.claude-plugin/marketplace.json), so claude plugin marketplace add <owner>/<repo> works as soon as the push lands. Only committed files are installed; data/, tasks/, docs/ and .env are git-ignored.
  4. Users update with:
   claude plugin marketplace update fast-laya-compaction
   claude plugin update fast-laya-compaction@fast-laya-compaction   # then restart Claude Code

A new version installs to a new cache directory; the Laya download is shared, so setup does not need to run again unless laya or the checkpoint changes.

For local testing without publishing: CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir /path/to/fast-laya-compaction, or claude plugin marketplace add /path/to/fast-laya-compaction to install from a local directory.

Library usage

The compaction logic is also a TypeScript library (not published to npm; build it with npm run build and import from dist/):

import { compactMessages, reductionRatio, type Message } from 'fast-laya-compaction';

const result = await compactMessages(transcript, { preserveRecentMessages: 4 });
console.log(result.decisions, result.stats);
if (reductionRatio(result) < 0.25) {
  // not worth it: keep the original transcript, or summarize instead
}

Message is a subset of Claude Code's SessionMessage. compactMessages runs the same Laya script through uv from Node. To bring your own scorer, implement LayaScorer and call compact(messages, scorer, options); collectToolCalls, supersededBy, callState, decideCall and applyDecisions are exported too.

Limitations

  • Only tool calls and their outputs are ever removed; messages are never shortened.
  • A probability is not proof an output is safe to trim. Truncation keeps a head and a note, and the assistant can re-run the tool.
  • Speed depends heavily on free memory (0.2–0.8 s per scored call on an 8 GB M2).
  • Function hooks are early access and may change between Claude Code releases; types/claude-code.d.ts was generated by Claude Code 2.1.274.
  • Laya is young and moves fast, so the package and weights are pinned.

Development

npm install
npm run typecheck && npm test      # library, hook and scripts (vitest)
npm run test:backend               # Python contract tests (fake model, no torch)
npm run validate:plugin
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir .
PathWhat
hooks/fast-laya.tsThe Claude Code hook: config, the Laya process, fallback, auto-compact
src/The library: call pairing, the superseded rule, per-call states, decisions
backend/laya_compact.pyOne-shot Laya scorer (stdin → stdout), run with uv run --script
scripts/dry-run, benchmark ceilings, and the Jev teacher tools
BENCHMARK.mdMeasurements

Acknowledgments

  • fast-jev-compaction (MIT): the original design and most of the compaction code.
  • Laya by Convai Innovations (Apache-2.0, code and weights).

License

MIT. See LICENSE.

Source 5 files
hooks/fast-laya.ts 358 lines
1import type {
2  On,
3  PluginOptions,
4  Register,
5  SessionMessage,
6  ToolResultSummary,
7  ToolUseSummary,
8  TurnCompleteInput,
9} from 'claude-code';
10
11import { compact, maxReduction, reductionRatio, resolveOptions } from '../src/compact.js';
12import { buildLayaRequest, DEFAULT_MODEL, layaArgv, parseLayaResponse } from '../src/laya.js';
13import type {
14  CompactOptions,
15  CompactResult,
16  LayaScorer,
17  Message,
18  ToolResult,
19  ToolUse,
20} from '../src/types.js';
21
22const HOOK_DEFAULTS = {
23  compactAtPercent: 60,
24  minReductionRatio: 0.25,
25  model: DEFAULT_MODEL,
26  uvPath: 'uv',
27  timeoutSeconds: 120,
28  autoTargetReduction: 0.5,
29};
30
31/** Triggers where the window is full (or about to be): the compaction has to make room. */
32const MAKE_ROOM: ReadonlySet<string> = new Set(['auto', 'precompute']);
33
34export type HookRunInit = {
35  stdin?: string;
36  timeoutMs?: number;
37};
38
39export type HookRunResult = {
40  exitCode: number;
41  stdout: string;
42  stderr: string;
43};
44
45/** The shape of `$.process.run`, so the hook can be driven without an engine. */
46export type HookRun = (argv: readonly string[], init?: HookRunInit) => Promise<HookRunResult>;
47
48export type HookConfig = CompactOptions & {
49  compactAtPercent: number;
50  minReductionRatio: number;
51  model: string;
52  uvPath: string;
53  device?: string;
54  timeoutSeconds: number;
55  /** `targetReduction` on auto-compaction, where leaving too little room re-triggers it at once. */
56  autoTargetReduction: number;
57};
58
59function optionNumber(options: PluginOptions, key: string, fallback: number): number {
60  const value = options[key];
61  return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
62}
63
64function optionString(options: PluginOptions, key: string): string | undefined {
65  const value = options[key];
66  return typeof value === 'string' && value.length > 0 ? value : undefined;
67}
68
69/** Reads the plugin's `userConfig` values; anything missing takes the defaults. */
70export function resolveHookConfig(options: PluginOptions): HookConfig {
71  const numbers: Partial<Omit<CompactOptions, 'goal'>> = {};
72  for (const key of [
73    'keepThreshold',
74    'preserveRecentMessages',
75    'maxCallStateChars',
76    'maxScoredCalls',
77    'truncateHeadChars',
78  ] as const) {
79    const value = options[key];
80    if (typeof value === 'number' && Number.isFinite(value)) numbers[key] = value;
81  }
82  const config: HookConfig = {
83    ...numbers,
84    compactAtPercent: optionNumber(options, 'compactAtPercent', HOOK_DEFAULTS.compactAtPercent),
85    minReductionRatio: optionNumber(
86      options,
87      'minReductionRatio',
88      HOOK_DEFAULTS.minReductionRatio,
89    ),
90    // A fine-tuned checkpoint directory, when set, replaces the published checkpoint.
91    model: optionString(options, 'checkpointPath') ?? optionString(options, 'model') ?? HOOK_DEFAULTS.model,
92    uvPath: optionString(options, 'uvPath') ?? HOOK_DEFAULTS.uvPath,
93    timeoutSeconds: Math.max(
94      1,
95      optionNumber(options, 'timeoutSeconds', HOOK_DEFAULTS.timeoutSeconds),
96    ),
97    autoTargetReduction: Math.min(
98      1,
99      Math.max(0, optionNumber(options, 'autoTargetReduction', HOOK_DEFAULTS.autoTargetReduction)),
100    ),
101  };
102  const device = optionString(options, 'device');
103  if (device) config.device = device;
104  const goal = optionString(options, 'goal');
105  if (goal) config.goal = goal;
106  return config;
107}
108
109/**
110 * A `LayaScorer` over the engine's `$.process.run`: one short-lived Laya
111 * process per compaction, the request on stdin, the scores on stdout.
112 * `$.process.run` rejects both when the command cannot start and when it
113 * outlives the timeout; the elapsed time tells the two apart.
114 */
115export function processScorer(
116  run: HookRun,
117  argv: readonly string[],
118  config: Pick<HookConfig, 'model' | 'device' | 'timeoutSeconds'>,
119): LayaScorer {
120  return {
121    async score(states, question) {
122      const timeoutMs = config.timeoutSeconds * 1000;
123      const stdin = buildLayaRequest(states, question, { model: config.model, device: config.device });
124      const started = Date.now();
125      let result: HookRunResult;
126      try {
127        result = await run(argv, { stdin, timeoutMs });
128      } catch (error) {
129        if (Date.now() - started >= timeoutMs - 1000) {
130          throw new Error(`Laya timed out after ${config.timeoutSeconds}s`);
131        }
132        throw new Error(`Laya could not start (${error instanceof Error ? error.message : String(error)})`);
133      }
134      return parseLayaResponse(result.exitCode, result.stdout, result.stderr, argv[argv.length - 1]);
135    },
136  };
137}
138
139function toolUseSummary(tool: ToolUse): ToolUseSummary {
140  const summary: ToolUseSummary = {
141    tool_use_id: tool.tool_use_id,
142    tool: tool.tool,
143    input: tool.input,
144  };
145  if (tool.text !== undefined) summary.text = tool.text;
146  if (tool.isError) summary.isError = true;
147  return summary;
148}
149
150function toolResultSummary(result: ToolResult): ToolResultSummary {
151  return {
152    tool_use_id: result.tool_use_id,
153    text: result.text,
154    isError: result.isError ?? false,
155  };
156}
157
158/**
159 * Maps the library's output back onto session messages. Whatever came back
160 * unchanged (a message, a tool use, a tool result) is the engine's own object,
161 * handle included; anything rebuilt is a fresh message without a handle, so the
162 * engine takes the edited content instead of its original.
163 */
164export function toSessionMessages(
165  input: readonly SessionMessage[],
166  output: readonly Message[],
167): SessionMessage[] {
168  const messages = new Map<Message, SessionMessage>();
169  const uses = new Map<ToolUse, ToolUseSummary>();
170  const results = new Map<ToolResult, ToolResultSummary>();
171  for (const message of input) {
172    messages.set(message, message);
173    for (const tool of message.toolUses) uses.set(tool, tool);
174    for (const result of message.toolResults ?? []) results.set(result, result);
175  }
176  return output.map((message) => {
177    const own = messages.get(message);
178    if (own) return own;
179    const rebuilt: SessionMessage = {
180      role: message.role,
181      text: message.text,
182      toolUses: message.toolUses.map((tool) => uses.get(tool) ?? toolUseSummary(tool)),
183    };
184    if (message.toolResults && message.toolResults.length > 0) {
185      rebuilt.toolResults = message.toolResults.map(
186        (result) => results.get(result) ?? toolResultSummary(result),
187      );
188    }
189    return rebuilt;
190  });
191}
192
193export type SessionCompaction = {
194  result: CompactResult;
195  messages: SessionMessage[];
196};
197
198/** Runs the library over a session transcript; throws when Laya cannot run or fails. */
199export async function compactSession(
200  messages: readonly SessionMessage[],
201  config: HookConfig,
202  run: HookRun,
203  pluginRoot: string,
204  trigger?: string,
205): Promise<SessionCompaction> {
206  // Without room to gain there is nothing for Laya to decide: say so before starting it.
207  const ceiling = maxReduction(messages, config);
208  if (ceiling < config.minReductionRatio) {
209    throw new Error(`nothing worth trimming: at most ${percent(ceiling)} could go`);
210  }
211  const options = {
212    ...config,
213    targetReduction: trigger && MAKE_ROOM.has(trigger) ? config.autoTargetReduction : 0,
214  };
215  const scorer = processScorer(run, layaArgv(config.uvPath, pluginRoot), config);
216  const result = await compact(messages, scorer, options);
217  return { result, messages: toSessionMessages(messages, result.messages) };
218}
219
220function percent(ratio: number): string {
221  return `${Math.round(ratio * 100)}%`;
222}
223
224function seconds(ms: number): string {
225  return `${(ms / 1000).toFixed(1)}s`;
226}
227
228export function summarize(result: CompactResult): string {
229  const { stats } = result;
230  const parts = [
231    stats.kept > 0 ? `${stats.kept} kept` : '',
232    stats.resultsDropped > 0 ? `${stats.resultsDropped} results truncated` : '',
233    stats.callsDropped > 0 ? `${stats.callsDropped} call_dropped` : '',
234    stats.superseded > 0 ? `${stats.superseded} superseded calls removed` : '',
235    stats.pinned > 0 ? `${stats.pinned} pinned` : '',
236    stats.unscored > 0 ? `${stats.unscored} unscored` : '',
237    stats.trimmed > 0 ? `${stats.trimmed} trimmed for room` : '',
238  ].filter(Boolean);
239  const laya = stats.model
240    ? `${stats.model} on ${stats.device}, load ${seconds(stats.loadMs)} + infer ${seconds(stats.inferMs)}`
241    : 'Laya not called';
242  return `${percent(reductionRatio(result))} reduction; ${parts.join(', ') || 'no tool calls'}; ${laya}`;
243}
244
245const UI_LOG_MAX_CHARS = 4096;
246
247export function decisionLog(result: CompactResult): string {
248  return result.decisions
249    .filter((d) => d.reason !== 'pinned' && d.reason !== 'unscored')
250    .map((d) => {
251      const p = d.probabilities;
252      const what = d.reason === 'trimmed' ? 'trimmed' : d.action;
253      return `${d.id}:${d.tool}:${what}/${p ? `keep=${p.keep.toFixed(2)}/truncate=${p.truncate.toFixed(2)}` : d.reason}`;
254    })
255    .join(' ');
256}
257
258export function decisionLogLines(
259  result: CompactResult,
260  maxChars: number = UI_LOG_MAX_CHARS,
261): string[] {
262  const entries = decisionLog(result).split(' ').filter(Boolean);
263  if (entries.length === 0) return ['decisions: (none)'];
264  const chunks: string[] = [];
265  let current = '';
266  for (const entry of entries) {
267    const next = current ? `${current} ${entry}` : entry;
268    if (current && next.length > maxChars - 24) {
269      chunks.push(current);
270      current = entry;
271    } else current = next;
272  }
273  chunks.push(current);
274  return chunks.map((chunk, index) =>
275    chunks.length === 1
276      ? `decisions: ${chunk}`
277      : `decisions (${index + 1}/${chunks.length}): ${chunk}`,
278  );
279}
280
281function notify(
282  $: {
283    ui: {
284      log: (text: string) => void;
285      toast: (text: string, options?: { timeoutMs?: number }) => void;
286    };
287  },
288  text: string,
289): void {
290  $.ui.log(text);
291  $.ui.toast(text, { timeoutMs: 15_000 });
292}
293
294export const register: Register = (on: On, options: PluginOptions) => {
295  const config = resolveHookConfig(options);
296  let compacting = false;
297  // One Laya process at a time: each holds a checkpoint (~2-3 GB). A compaction
298  // arriving while another runs (an auto-compaction behind an ahead-of-time
299  // `precompute`, or a subagent's) waits its turn instead of losing Laya.
300  let queue: Promise<void> = Promise.resolve();
301
302  on('session.compact', async ($, event, next) => {
303    const previous = queue;
304    let done!: () => void;
305    queue = new Promise<void>((resolve) => (done = resolve));
306    await previous;
307    try {
308      const { result, messages } = await compactSession(
309        event.messages,
310        config,
311        (argv, init) => $.process.run(argv, init),
312        $.plugin.root,
313        event.trigger,
314      );
315      for (const line of decisionLogLines(result)) $.ui.log(line);
316      if (reductionRatio(result) < config.minReductionRatio) {
317        notify(
318          $,
319          `fallback to built-in summary (below ${percent(config.minReductionRatio)} minimum: ${summarize(result)})`,
320        );
321        return next(event);
322      }
323      notify(
324        $,
325        `kept ${messages.length}/${event.messages.length} messages, no summary (${summarize(result)})`,
326      );
327      return { messages };
328    } catch (error) {
329      notify(
330        $,
331        `fallback to built-in summary (${error instanceof Error ? error.message : String(error)})`,
332      );
333      return next(event);
334    } finally {
335      done();
336    }
337  });
338
339  on('turn.complete', async ($, event: TurnCompleteInput, next) => {
340    if (compacting) return next(event);
341    try {
342      const { context } = await $.session.usage();
343      if ((context.percent ?? 0) < config.compactAtPercent) return next(event);
344      compacting = true;
345      await $.session.compact();
346    } catch (error) {
347      $.ui.log(
348        `auto-compact skipped (${error instanceof Error ? error.message : String(error)})`,
349      );
350    } finally {
351      compacting = false;
352    }
353    return next(event);
354  });
355};
356
357export { resolveOptions };
358
src/compact.ts 356 lines
1import { callState, collectToolCalls, goalFromMessages, supersededBy } from './state.js';
2import type {
3  CallDecision,
4  CallProbabilities,
5  ChoiceQuestion,
6  CompactOptions,
7  CompactResult,
8  LayaScorer,
9  LayaScores,
10  Message,
11  ResolvedCompactOptions,
12  ToolCall,
13  ToolUse,
14} from './types.js';
15
16export const DEFAULT_OPTIONS: ResolvedCompactOptions = {
17  goal: '',
18  keepThreshold: 0.5,
19  preserveRecentMessages: 6,
20  maxCallStateChars: 1600,
21  maxScoredCalls: 80,
22  truncateHeadChars: 300,
23  targetReduction: 0,
24};
25
26function finite(value: number | undefined, fallback: number): number {
27  return typeof value === 'number' && Number.isFinite(value) ? value : fallback;
28}
29
30export function resolveOptions(options: CompactOptions = {}): ResolvedCompactOptions {
31  return {
32    goal: options.goal ?? DEFAULT_OPTIONS.goal,
33    keepThreshold: finite(options.keepThreshold, DEFAULT_OPTIONS.keepThreshold),
34    preserveRecentMessages: Math.max(
35      0,
36      Math.floor(
37        finite(options.preserveRecentMessages, DEFAULT_OPTIONS.preserveRecentMessages),
38      ),
39    ),
40    maxCallStateChars: Math.max(
41      200,
42      Math.floor(finite(options.maxCallStateChars, DEFAULT_OPTIONS.maxCallStateChars)),
43    ),
44    maxScoredCalls: Math.max(
45      0,
46      Math.floor(finite(options.maxScoredCalls, DEFAULT_OPTIONS.maxScoredCalls)),
47    ),
48    truncateHeadChars: Math.max(
49      0,
50      Math.floor(finite(options.truncateHeadChars, DEFAULT_OPTIONS.truncateHeadChars)),
51    ),
52    targetReduction: Math.min(1, Math.max(0, finite(options.targetReduction, DEFAULT_OPTIONS.targetReduction))),
53  };
54}
55
56/**
57 * The one question asked about every scored call, each with its own state.
58 * The criteria are worded positively: Laya reads negations poorly. (A bare
59 * yes/no question measured worse on real sessions: less reduction, a less
60 * sensible ranking, and the same answers when asked the opposite way.)
61 */
62export const CALL_QUESTION: ChoiceQuestion = {
63  type: 'choice',
64  instructions:
65    "A coding assistant's conversation is being compacted. The state describes one earlier tool call. What should happen to it?",
66  criteria: {
67    keep: 'the assistant still needs the full output verbatim',
68    truncate: 'only the fact that this call was made still matters',
69    drop: 'the call is obsolete: superseded, exploratory, or already acted on',
70  },
71};
72
73/** Laya's answer for one call; throws when a probability is missing or out of range. */
74export function callProbabilities(scores: LayaScores['scores'], id: string): CallProbabilities {
75  const p = scores[id];
76  const valid = (x: unknown): x is number => typeof x === 'number' && x >= 0 && x <= 1;
77  if (!p || !valid(p.keep) || !valid(p.truncate) || !valid(p.drop)) {
78    throw new Error(`Invalid Laya answer for ${id}`);
79  }
80  return { keep: p.keep, truncate: p.truncate, drop: p.drop };
81}
82
83/**
84 * One call's fate. Pinned calls stay and superseded ones go with their
85 * result. For a scored call: `P(keep) ≥ threshold` keeps it whole, else
86 * `P(keep) + P(truncate) ≥ threshold` keeps the call with a truncated output,
87 * else it goes. An unscored call stays.
88 */
89export function decideCall(
90  call: Pick<ToolCall, 'id' | 'tool' | 'pinned'>,
91  verdict: { superseded?: boolean; probabilities?: CallProbabilities },
92  options: Pick<ResolvedCompactOptions, 'keepThreshold'>,
93): CallDecision {
94  const base = { id: call.id, tool: call.tool };
95  if (call.pinned) return { ...base, action: 'keep', reason: 'pinned' };
96  if (verdict.superseded) return { ...base, action: 'drop_call', reason: 'superseded' };
97  const p = verdict.probabilities;
98  if (!p) return { ...base, action: 'keep', reason: 'unscored' };
99  if (p.keep >= options.keepThreshold) return { ...base, probabilities: p, action: 'keep', reason: 'kept' };
100  if (p.keep + p.truncate >= options.keepThreshold) {
101    return { ...base, probabilities: p, action: 'drop_result', reason: 'result_dropped' };
102  }
103  return { ...base, probabilities: p, action: 'drop_call', reason: 'call_dropped' };
104}
105
106function truncatedResultText(text: string, isError: boolean, headChars: number): string {
107  if (text.length <= headChars + 120) return text;
108  const head = headChars > 0 ? `${text.slice(0, headChars)}\n` : '';
109  return `${head}[fast-laya-compaction truncated ${text.length - headChars} chars of this tool result${
110    isError ? ' (error)' : ''
111  }; re-run the tool if needed]`;
112}
113
114/**
115 * Rebuilds the conversation from the decisions. A dropped call disappears
116 * together with its result; a dropped result keeps a bounded head and note.
117 * Messages that lose all their content are removed; untouched messages are
118 * returned as the same objects they came in as.
119 */
120export function applyDecisions(
121  messages: readonly Message[],
122  decisions: readonly CallDecision[],
123  calls: readonly ToolCall[],
124  headChars: number,
125): Message[] {
126  const byId = new Map(calls.map((call) => [call.id, call]));
127  const actions = new Map<string, CallDecision['action']>();
128  for (const decision of decisions) {
129    const call = byId.get(decision.id);
130    if (call && decision.action !== 'keep') actions.set(call.tool_use_id, decision.action);
131  }
132  const kept: Message[] = [];
133  for (const message of messages) {
134    const touched =
135      message.toolUses.some((tool) => actions.has(tool.tool_use_id)) ||
136      (message.toolResults ?? []).some((result) => actions.has(result.tool_use_id));
137    if (!touched) {
138      kept.push(message);
139      continue;
140    }
141    const toolUses = message.toolUses
142      .filter((tool) => actions.get(tool.tool_use_id) !== 'drop_call')
143      .map((tool) => {
144        if (actions.get(tool.tool_use_id) !== 'drop_result') return tool;
145        const text = truncatedResultText(
146          tool.text ?? '',
147          tool.isError ?? false,
148          headChars,
149        );
150        if ((tool.text ?? '') === text) return tool;
151        const copy: ToolUse = {
152          tool_use_id: tool.tool_use_id,
153          tool: tool.tool,
154          input: tool.input,
155          text,
156        };
157        if (tool.isError) copy.isError = true;
158        return copy;
159      });
160    const toolResults = (message.toolResults ?? [])
161      .filter((result) => actions.get(result.tool_use_id) !== 'drop_call')
162      .map((result) => {
163        if (actions.get(result.tool_use_id) !== 'drop_result') return result;
164        const text = truncatedResultText(result.text, result.isError ?? false, headChars);
165        return text === result.text
166          ? result
167          : {
168              tool_use_id: result.tool_use_id,
169              text,
170              isError: result.isError,
171            };
172      });
173    if (
174      !message.toolUses.some(
175        (tool) => actions.get(tool.tool_use_id) === 'drop_call',
176      ) &&
177      !(message.toolResults ?? []).some(
178        (result) => actions.get(result.tool_use_id) === 'drop_call',
179      ) &&
180      toolUses.every((tool, index) => tool === message.toolUses[index]) &&
181      toolResults.every(
182        (result, index) => result === message.toolResults?.[index],
183      )
184    ) {
185      kept.push(message);
186      continue;
187    }
188    if (message.text.trim().length === 0 && toolUses.length === 0 && toolResults.length === 0) {
189      continue;
190    }
191    const rebuilt: Message = { role: message.role, text: message.text, toolUses };
192    if (toolResults.length > 0) rebuilt.toolResults = toolResults;
193    kept.push(rebuilt);
194  }
195  return kept;
196}
197
198/** Characters of text, tool input and tool output a message holds. */
199export function messageChars(message: Message): number {
200  let total = message.text.length;
201  for (const tool of message.toolUses) {
202    try {
203      total += JSON.stringify(tool.input).length;
204    } catch {
205      total += 20;
206    }
207  }
208  for (const result of message.toolResults ?? []) total += result.text.length;
209  return total;
210}
211
212export function reductionRatio(result: Pick<CompactResult, 'stats'>): number {
213  const { charsBefore, charsAfter } = result.stats;
214  return charsBefore === 0 ? 0 : (charsBefore - charsAfter) / charsBefore;
215}
216
217function count(decisions: readonly CallDecision[], reason: CallDecision['reason']): number {
218  return decisions.filter((decision) => decision.reason === reason).length;
219}
220
221function charsOf(messages: readonly Message[]): number {
222  return messages.reduce((sum, message) => sum + messageChars(message), 0);
223}
224
225/**
226 * The calls a compaction looks at: all of them paired, the superseded ones
227 * (removed by rule) and the candidates Laya scores, oldest first.
228 */
229function plan(messages: readonly Message[], resolved: ResolvedCompactOptions) {
230  const calls = collectToolCalls(messages, resolved.preserveRecentMessages);
231  const superseded = new Set(
232    calls.filter((call) => !call.pinned && supersededBy(call, calls)).map((call) => call.id),
233  );
234  // Oldest first: they are the likeliest to be stale, and scoring time grows with every call.
235  const candidates = calls
236    .filter((call) => !call.pinned && !superseded.has(call.id))
237    .slice(0, resolved.maxScoredCalls);
238  return { calls, superseded, candidates };
239}
240
241/**
242 * The most a compaction could remove, before asking Laya: every superseded
243 * call gone and every candidate's output truncated. When even this is too
244 * little, there is no point starting Laya.
245 */
246export function maxReduction(messages: readonly Message[], options: CompactOptions = {}): number {
247  const resolved = resolveOptions(options);
248  const { calls, superseded, candidates } = plan(messages, resolved);
249  const scored = new Set(candidates.map((call) => call.id));
250  const decisions = calls.map((call): CallDecision => {
251    const base = { id: call.id, tool: call.tool };
252    if (superseded.has(call.id)) return { ...base, action: 'drop_call', reason: 'superseded' };
253    if (scored.has(call.id)) return { ...base, action: 'drop_result', reason: 'result_dropped' };
254    return { ...base, action: 'keep', reason: call.pinned ? 'pinned' : 'unscored' };
255  });
256  const before = charsOf(messages);
257  return before === 0 ? 0 : 1 - charsOf(applyDecisions(messages, decisions, calls, resolved.truncateHeadChars)) / before;
258}
259
260/**
261 * Truncates outputs Laya would keep, the lowest `P(keep)` first, until the
262 * reduction reaches `targetReduction`. Laya ranks calls better than it
263 * calibrates them, so when room is needed its order decides what goes.
264 */
265function trimToTarget(
266  messages: readonly Message[],
267  decisions: readonly CallDecision[],
268  calls: readonly ToolCall[],
269  resolved: ResolvedCompactOptions,
270): CallDecision[] {
271  const out = [...decisions];
272  const before = charsOf(messages);
273  const reached = () =>
274    before === 0 ||
275    1 - charsOf(applyDecisions(messages, out, calls, resolved.truncateHeadChars)) / before >= resolved.targetReduction;
276  if (resolved.targetReduction <= 0 || reached()) return out;
277  const order = out
278    .map((decision, index) => ({ decision, index }))
279    .filter(({ decision }) => decision.reason === 'kept')
280    .sort((a, b) => a.decision.probabilities!.keep - b.decision.probabilities!.keep);
281  for (const { decision, index } of order) {
282    out[index] = { ...decision, action: 'drop_result', reason: 'trimmed' };
283    if (reached()) break;
284  }
285  return out;
286}
287
288/**
289 * Compacts a transcript. Outside the pinned first and newest messages, a
290 * read-only call that a later call repeated or made stale is removed with
291 * its result; of the rest, the oldest `maxScoredCalls` are each described by
292 * their own small state (see `callState`) and Laya is asked, in one batched
293 * request, whether each should be kept, truncated or dropped. With a
294 * `targetReduction`, outputs Laya would keep are then truncated, lowest
295 * `P(keep)` first, until it is met. Throws when Laya fails; the caller decides
296 * whether to fall back.
297 */
298export async function compact(
299  messages: readonly Message[],
300  scorer: LayaScorer,
301  options: CompactOptions = {},
302): Promise<CompactResult> {
303  const started = Date.now();
304  const resolved = resolveOptions(options);
305  const { calls, superseded, candidates } = plan(messages, resolved);
306  const charsBefore = charsOf(messages);
307
308  let laya: Omit<LayaScores, 'scores'> = { model: '', device: '', loadMs: 0, inferMs: 0 };
309  const answers = new Map<string, CallProbabilities>();
310  if (candidates.length > 0) {
311    const goal = resolved.goal || goalFromMessages(messages);
312    const states = candidates.map((call) => ({
313      id: call.id,
314      state: callState(call, messages, calls, goal, resolved.maxCallStateChars),
315    }));
316    const { model, device, loadMs, inferMs, scores } = await scorer.score(states, CALL_QUESTION);
317    laya = { model, device, loadMs, inferMs };
318    for (const call of candidates) answers.set(call.id, callProbabilities(scores, call.id));
319  }
320
321  const decisions = trimToTarget(
322    messages,
323    calls.map((call) =>
324      decideCall(call, { superseded: superseded.has(call.id), probabilities: answers.get(call.id) }, resolved),
325    ),
326    calls,
327    resolved,
328  );
329  const kept = applyDecisions(
330    messages,
331    decisions,
332    calls,
333    resolved.truncateHeadChars,
334  );
335  return {
336    messages: kept,
337    decisions,
338    stats: {
339      messagesBefore: messages.length,
340      messagesAfter: kept.length,
341      charsBefore,
342      charsAfter: charsOf(kept),
343      calls: calls.length,
344      kept: count(decisions, 'kept'),
345      resultsDropped: count(decisions, 'result_dropped'),
346      callsDropped: count(decisions, 'call_dropped'),
347      superseded: count(decisions, 'superseded'),
348      pinned: count(decisions, 'pinned'),
349      unscored: count(decisions, 'unscored'),
350      trimmed: count(decisions, 'trimmed'),
351      ...laya,
352      ms: Date.now() - started,
353    },
354  };
355}
356
src/laya.ts 97 lines
1import type { CallState, ChoiceQuestion, LayaScores } from './types.js';
2
3/** Laya's checkpoint fine-tuned for typed decisions; the base ones are near chance zero-shot. */
4export const DEFAULT_MODEL = 'typed-decisions';
5
6/** The scoring script, relative to the plugin (and package) root. */
7export const SCRIPT_PATH = 'backend/laya_compact.py';
8
9/** What to run once before the first compaction; the hook names the installed script's absolute path. */
10export function setupHint(script: string = SCRIPT_PATH): string {
11  return `Laya is not set up, run once: uv run --script ${script} --warmup`;
12}
13
14export const SETUP_HINT = setupHint();
15
16/** The script's exit code for a checkpoint missing from the Hugging Face cache. */
17const EXIT_NOT_CACHED = 3;
18
19/** How uv reports a script environment it cannot build under `--offline`. */
20const UV_OFFLINE = /network was disabled|not found in the cache/i;
21
22export interface LayaRequestOptions {
23  /** `typed-decisions` (default), `multilingual` or `english`. */
24  model?: string;
25  /** Torch device; auto (mps on Apple Silicon) when absent. */
26  device?: string;
27  /** Tokens per question row: head (question and options) plus state. */
28  maxLen?: number;
29  headMaxLen?: number;
30  /** States per forward pass. */
31  batchSize?: number;
32}
33
34/**
35 * The command that scores one compaction. `--offline` keeps a hook from ever
36 * building the torch environment inline: without the one-time setup it fails
37 * fast and the caller falls back.
38 */
39export function layaArgv(uvPath: string, root: string): string[] {
40  return [uvPath, 'run', '--quiet', '--offline', '--script', `${root}/${SCRIPT_PATH}`];
41}
42
43/** The script's stdin for one compaction: one question over every call state. */
44export function buildLayaRequest(
45  states: readonly CallState[],
46  question: ChoiceQuestion,
47  options: LayaRequestOptions = {},
48): string {
49  return JSON.stringify({
50    model: options.model ?? DEFAULT_MODEL,
51    ...(options.device ? { device: options.device } : {}),
52    max_len: options.maxLen ?? 512,
53    head_max_len: options.headMaxLen ?? 128,
54    batch_size: options.batchSize ?? 32,
55    question,
56    states,
57  });
58}
59
60function lastLine(text: string): string {
61  const lines = text.split('\n').map((line) => line.trim()).filter(Boolean);
62  return (lines[lines.length - 1] ?? '').slice(0, 200);
63}
64
65/** Validates the script's output; throws a message fit for a toast on anything else. */
66export function parseLayaResponse(
67  exitCode: number,
68  stdout: string,
69  stderr: string,
70  script: string = SCRIPT_PATH,
71): LayaScores {
72  if (exitCode === EXIT_NOT_CACHED || (exitCode !== 0 && UV_OFFLINE.test(stderr))) {
73    throw new Error(setupHint(script));
74  }
75  if (exitCode !== 0) {
76    throw new Error(`Laya exited with ${exitCode}: ${lastLine(stderr) || 'no output'}`);
77  }
78  let parsed: unknown;
79  try {
80    // The answer is the last line; anything a library printed before it is noise.
81    parsed = JSON.parse(stdout.trim().split('\n').pop() ?? '');
82  } catch {
83    throw new Error('Laya returned malformed JSON');
84  }
85  const body = parsed as Record<string, unknown> | null;
86  if (!body || typeof body !== 'object' || !body.scores || typeof body.scores !== 'object') {
87    throw new Error('Laya response is missing scores');
88  }
89  return {
90    model: String(body.model ?? ''),
91    device: String(body.device ?? ''),
92    loadMs: Number(body.load_ms) || 0,
93    inferMs: Number(body.infer_ms) || 0,
94    scores: body.scores as LayaScores['scores'],
95  };
96}
97
src/types.ts 161 lines
1export type Role = 'user' | 'assistant';
2
3/**
4 * A tool_use block of an assistant message. `text` and `isError` mirror the
5 * outcome once the transcript holds it (Claude Code attaches them).
6 */
7export interface ToolUse {
8  tool_use_id: string;
9  tool: string;
10  input: Record<string, unknown>;
11  text?: string;
12  isError?: boolean;
13}
14
15/** A tool_result block of a user message. */
16export interface ToolResult {
17  tool_use_id: string;
18  text: string;
19  isError?: boolean;
20}
21
22/**
23 * One transcript message. The shape is a subset of Claude Code's
24 * `SessionMessage`, so a session transcript can be passed in as is.
25 */
26export interface Message {
27  role: Role;
28  text: string;
29  toolUses: ToolUse[];
30  toolResults?: ToolResult[];
31}
32
33/** A tool call paired with its result by `tool_use_id`. */
34export interface ToolCall {
35  /** Short id used in the Laya request and decision log (`t1`, `t2`, ...). */
36  id: string;
37  tool_use_id: string;
38  tool: string;
39  input: Record<string, unknown>;
40  /** Index of the message holding the tool_use block. */
41  callIndex: number;
42  /** Index of the message holding the tool_result block. */
43  resultIndex: number;
44  resultChars: number;
45  isError: boolean;
46  /** In the first or the newest preserved messages; never a candidate. */
47  pinned: boolean;
48}
49
50export type CallAction = 'keep' | 'drop_result' | 'drop_call';
51
52/** Laya's answer for one call: keep it whole, truncate its output, or drop it. */
53export interface CallProbabilities {
54  keep: number;
55  truncate: number;
56  drop: number;
57}
58
59export interface CallDecision {
60  id: string;
61  tool: string;
62  /** Laya's answer; only on scored calls. */
63  probabilities?: CallProbabilities;
64  action: CallAction;
65  /**
66   * `superseded`: a later call repeated it or changed its target, so it is
67   * removed without asking Laya. `kept`/`result_dropped`/`call_dropped`: Laya's
68   * answer. `trimmed`: Laya would keep it, but its output was truncated to reach
69   * `targetReduction`.
70   */
71  reason: 'pinned' | 'superseded' | 'unscored' | 'kept' | 'result_dropped' | 'call_dropped' | 'trimmed';
72}
73
74export interface CompactOptions {
75  /** Ongoing task description; defaults to the latest user prompt. */
76  goal?: string;
77  /** Minimum keep probability for a call or its full output to stay. Default 0.5. */
78  keepThreshold?: number;
79  /** Newest messages never touched (the first message is always kept). Default 6. */
80  preserveRecentMessages?: number;
81  /** Character cap on the state Laya reads for one call. Default 1600. */
82  maxCallStateChars?: number;
83  /** Most calls scored per compaction, oldest first; newer ones are kept. Bounds latency. Default 80. */
84  maxScoredCalls?: number;
85  /** Characters of a dropped tool result to retain. Default 300. */
86  truncateHeadChars?: number;
87  /**
88   * Reduction to reach even past Laya's answers: outputs Laya would keep are
89   * truncated, least likely to be needed first, until it is met. 0 (default) off.
90   */
91  targetReduction?: number;
92}
93
94export interface ResolvedCompactOptions {
95  goal: string;
96  keepThreshold: number;
97  preserveRecentMessages: number;
98  maxCallStateChars: number;
99  maxScoredCalls: number;
100  truncateHeadChars: number;
101  targetReduction: number;
102}
103
104export interface CompactResult {
105  /** The compacted transcript; untouched messages are the input objects. */
106  messages: Message[];
107  decisions: CallDecision[];
108  stats: {
109    messagesBefore: number;
110    messagesAfter: number;
111    charsBefore: number;
112    charsAfter: number;
113    calls: number;
114    kept: number;
115    resultsDropped: number;
116    /** Calls Laya dropped with their results. */
117    callsDropped: number;
118    /** Calls removed with their results because a later call superseded them. */
119    superseded: number;
120    pinned: number;
121    /** Kept without scoring: past `maxScoredCalls`. */
122    unscored: number;
123    /** Outputs Laya would keep, truncated to reach `targetReduction`. */
124    trimmed: number;
125    /** Laya checkpoint and torch device that scored the calls; '' when Laya was not called. */
126    model: string;
127    device: string;
128    /** Checkpoint load and batched inference time inside the Laya process. */
129    loadMs: number;
130    inferMs: number;
131    ms: number;
132  };
133}
134
135/** A Laya `choice` question: one label per criterion, answered with a probability each. */
136export interface ChoiceQuestion {
137  type: 'choice';
138  instructions: string;
139  criteria: Record<string, string>;
140}
141
142/** What Laya reads about one tool call; key order matters, Laya cuts from the end. */
143export interface CallState {
144  id: string;
145  state: Record<string, unknown>;
146}
147
148export interface LayaScores {
149  model: string;
150  device: string;
151  loadMs: number;
152  inferMs: number;
153  /** Per call id, the probability of every criterion label. */
154  scores: Record<string, Record<string, number>>;
155}
156
157/** Anything that answers one question over many call states: a Laya process, or a fake. */
158export interface LayaScorer {
159  score(states: readonly CallState[], question: ChoiceQuestion): Promise<LayaScores>;
160}
161
src/state.ts 211 lines
1import type { Message, ToolCall, ToolResult } from './types.js';
2
3/** Cap on the serialised tool input in a one-line call. */
4const INPUT_CHARS = 200;
5const GOAL_CHARS = 300;
6const INTENT_CHARS = 200;
7const OUTPUT_HEAD_CHARS = 400;
8const LATER_CALLS = 3;
9
10/** Input keys that name what a call operates on; equal targets mean the same file, command or search. */
11const TARGET_KEYS = ['file_path', 'notebook_path', 'path', 'url', 'command', 'pattern'] as const;
12
13/** Tools that change their target. A change is never superseded: the record that it was made still matters. */
14const MUTATING_TOOLS = new Set(['Edit', 'MultiEdit', 'Write', 'NotebookEdit']);
15
16/** Fields shortened, then removed, in this order when a state exceeds its cap: least decisive first. */
17const TRIM_ORDER = ['output_head', 'intent', 'goal', 'later_calls'] as const;
18
19export function truncate(text: string, limit: number): string {
20  return text.length <= limit ? text : `${text.slice(0, Math.max(0, limit - 1))}…`;
21}
22
23export function isPinned(
24  index: number,
25  total: number,
26  preserveRecentMessages: number,
27): boolean {
28  return index === 0 || index >= total - preserveRecentMessages;
29}
30
31/**
32 * Pairs every tool_use with its tool_result by `tool_use_id`. Calls without a
33 * result are not candidates (there is nothing to drop yet).
34 */
35export function collectToolCalls(
36  messages: readonly Message[],
37  preserveRecentMessages: number,
38): ToolCall[] {
39  const results = new Map<string, { index: number; result: ToolResult }>();
40  messages.forEach((message, index) => {
41    for (const result of message.toolResults ?? []) {
42      results.set(result.tool_use_id, { index, result });
43    }
44  });
45  const calls: ToolCall[] = [];
46  messages.forEach((message, callIndex) => {
47    for (const tool of message.toolUses) {
48      const found = results.get(tool.tool_use_id);
49      if (!found) continue;
50      calls.push({
51        id: `t${calls.length + 1}`,
52        tool_use_id: tool.tool_use_id,
53        tool: tool.tool,
54        input: tool.input,
55        callIndex,
56        resultIndex: found.index,
57        resultChars: found.result.text.length,
58        isError: found.result.isError ?? false,
59        pinned:
60          isPinned(callIndex, messages.length, preserveRecentMessages) ||
61          isPinned(found.index, messages.length, preserveRecentMessages),
62      });
63    }
64  });
65  return calls;
66}
67
68function inputText(value: unknown, limit: number): string {
69  let json = '';
70  try {
71    json = JSON.stringify(value) ?? String(value);
72  } catch {
73    json = '[unserializable input]';
74  }
75  return truncate(json, limit);
76}
77
78/** One call as a single line, target first, e.g. `Edit file_path=src/a.ts old_string=… replace_all=false`. */
79export function callLine(call: Pick<ToolCall, 'tool' | 'input'>): string {
80  const isTarget = (key: string) => (TARGET_KEYS as readonly string[]).includes(key);
81  const entries = Object.entries(call.input);
82  const input = [...entries.filter(([key]) => isTarget(key)), ...entries.filter(([key]) => !isTarget(key))]
83    .map(([key, value]) => {
84      const text = typeof value === 'string' ? value : inputText(value, INPUT_CHARS);
85      return `${key}=${text.replace(/\s+/g, ' ')}`;
86    })
87    .join(' ');
88  return truncate(`${call.tool} ${input}`.trim(), INPUT_CHARS);
89}
90
91/** What a call operates on (file, command, search), or undefined when its input names nothing. */
92export function callTarget(call: Pick<ToolCall, 'input'>): string | undefined {
93  const parts = TARGET_KEYS.flatMap((key) => {
94    const value = call.input[key];
95    return typeof value === 'string' && value.trim() ? [value.trim()] : [];
96  });
97  return parts.length > 0 ? parts.join(' ') : undefined;
98}
99
100/** The same tool with the same input, a free-text `description` aside. */
101function isRepeat(a: ToolCall, b: ToolCall): boolean {
102  const { description: _a, ...inputA } = a.input;
103  const { description: _b, ...inputB } = b.input;
104  return a.tool === b.tool && inputText(inputA, Infinity) === inputText(inputB, Infinity);
105}
106
107/**
108 * The later call that makes a read-only call obsolete: an exact repeat (the
109 * same read, search or command run again) or a change to the same target,
110 * after which the earlier output is stale. Such a call can be removed with
111 * its result without asking Laya.
112 */
113export function supersededBy(call: ToolCall, calls: readonly ToolCall[]): ToolCall | undefined {
114  if (MUTATING_TOOLS.has(call.tool)) return undefined;
115  const target = callTarget(call);
116  if (!target) return undefined;
117  return calls.find(
118    (later) =>
119      later.callIndex > call.callIndex &&
120      callTarget(later) === target &&
121      (MUTATING_TOOLS.has(later.tool) || isRepeat(later, call)),
122  );
123}
124
125/** The latest user prompts (tool-result messages excluded), as the default `goal`. */
126export function goalFromMessages(messages: readonly Message[], count = 1): string {
127  return messages
128    .filter(
129      (message) =>
130        message.role === 'user' &&
131        message.text.trim().length > 0 &&
132        (message.toolResults ?? []).length === 0,
133    )
134    .slice(-count)
135    .map((message) => truncate(message.text, GOAL_CHARS))
136    .join('\n');
137}
138
139/** The assistant text that led to a call: its own message's, else the nearest before it this turn. */
140function intentOf(messages: readonly Message[], callIndex: number): string {
141  for (let i = callIndex; i >= 0; i--) {
142    const message = messages[i]!;
143    if (message.role === 'user' && (message.toolResults ?? []).length === 0) break;
144    if (message.role === 'assistant' && message.text.trim()) return message.text.trim();
145  }
146  return '';
147}
148
149function resultText(messages: readonly Message[], call: ToolCall): string {
150  return (
151    messages[call.resultIndex]?.toolResults?.find((r) => r.tool_use_id === call.tool_use_id)?.text ?? ''
152  );
153}
154
155/** Later assistant text mentions the target (a file by its base name). */
156function mentionedLater(messages: readonly Message[], call: ToolCall, target: string): boolean {
157  const needle = target.includes('/') ? target.slice(target.lastIndexOf('/') + 1) : target;
158  if (needle.length < 3) return false;
159  return messages
160    .slice(call.resultIndex + 1)
161    .some((message) => message.role === 'assistant' && message.text.includes(needle));
162}
163
164function capState(state: Record<string, unknown>, maxChars: number): Record<string, unknown> {
165  for (const key of TRIM_ORDER) {
166    const over = JSON.stringify(state).length - maxChars;
167    if (over <= 0) break;
168    const value = state[key];
169    if (typeof value === 'string' && value.length > over + 20) {
170      state[key] = truncate(value, value.length - over);
171    } else delete state[key];
172  }
173  return state;
174}
175
176/**
177 * What Laya reads about one tool call. Laya keeps only the head of a long
178 * state, so the most decisive facts come first: the call, its outcome, how
179 * long ago it was, and whether later calls touched the same target.
180 * Context and a glimpse of the output follow, and are the first to go under
181 * `maxChars`.
182 */
183export function callState(
184  call: ToolCall,
185  messages: readonly Message[],
186  calls: readonly ToolCall[],
187  goal: string,
188  maxChars: number,
189): Record<string, unknown> {
190  const target = callTarget(call);
191  const later = target
192    ? calls.filter((other) => other.callIndex > call.callIndex && callTarget(other) === target)
193    : [];
194  const state: Record<string, unknown> = {
195    tool_call: callLine(call),
196    outcome: `${call.isError ? 'error' : 'ok'}, ${call.resultChars} chars`,
197    messages_since: messages.length - 1 - call.resultIndex,
198    // Laya reads key names as words: renaming this `touched_later` cut the
199    // measured reduction from 33% to 23% on one session. Keep it.
200    superseded: later.length > 0,
201  };
202  if (later.length > 0) state.later_calls = later.slice(0, LATER_CALLS).map(callLine);
203  if (target) state.mentioned_later = mentionedLater(messages, call, target);
204  if (goal) state.goal = truncate(goal, GOAL_CHARS);
205  const intent = intentOf(messages, call.callIndex);
206  if (intent) state.intent = truncate(intent, INTENT_CHARS);
207  const output = resultText(messages, call);
208  if (output) state.output_head = truncate(output, OUTPUT_HEAD_CHARS);
209  return capState(state, maxChars);
210}
211