SLOPSHOPPER

sieve

Keeps unnecessary tool output out of context in Claude Code and Copilot CLI, with a shared local strands-decider for close calls.

newguardcommandtoaststatusprompt
v0.2.1no licenseupdated 2026-10-05trsdn/sieve
A shopper browsing a rack in a slop shop
README

sieve

A tool-output filter for Claude Code and GitHub Copilot CLI that lets the useful part through and keeps unnecessary output out of context. The Claude Code mod replaces context-mode; both adapters can use the shared local strands-decider for calls a fixed rule cannot make.

Copilot CLI support is new in v0.2.0 and limited to shell results. Install and scope are described below. The results, search, session and command features in the next table describe the Claude Code adapter.

What it does

PartWhat
ResultsBash, Grep, Glob, WebFetch, text MCP results and Playwright snapshot reads are rewritten in place after they ran: large ones become a structural summary, the whole output goes to a file and an SQLite FTS5 index. Nothing is refused.
Toolmcp__sieve__search (BM25 over what was cut).
DeciderOne shared strands-decider serve (MLX, 127.0.0.1:8765, a LaunchAgent, see launchd/) makes the small calls a rule cannot: lookup or overview, new task, reminder, effort. It answers within 1.5 s or counts as down for 30 s; without it only the rules decide.
SessionEdited files, commands, failures and prompts are recorded; before a compaction a resume note (max 2000 chars) is stored and added to the system prompt afterwards.
Command/sieve shows cuts, chars kept out of context, index size, decider status.

Requirements

  • Claude Code with function-hook mods (plugin-authoring), sqlite3 with FTS5, node, python3
  • strands-decider with the mlx extra at ~/.local/bin/strands-decider, run as a shared service: sh launchd/install.sh (optional; without it only the rules route)

Install

claude --plugin-dir ~/dev/sieve

Disable the old plugin: set "context-mode@context-mode": false in ~/.claude/settings.json.

Develop

claude plugin validate .
claude plugin test .
node --experimental-strip-types --test copilot/copilot.test.mjs

Pure logic lives in hooks/lib.ts with tests in hooks/lib.test.ts. Everything that takes $ stays in hooks/register.ts, as top-level functions: the validator refuses $ passed to closures.

The Copilot request preparation has bilingual development and holdout cases:

node --experimental-strip-types eval/copilot_decider_eval.mjs

This needs the existing local decider. It exits nonzero for an unsafe cut or incomplete service coverage. Its original/prepared arms both use the current precision veto and differ only in input ordering; they are not a release-versus-candidate token benchmark. The first holdout is retained as a regression set after it exposed unsafe mixed-requirement decisions; holdout2 was written before adding the precision veto.

GitHub Copilot CLI

Requires Node 22.18+ or 24+ and Copilot CLI with command-hook modifiedResult support. Tested with Copilot CLI 1.0.92-3. No npm packages are needed.

From this checkout or an extracted release:

node copilot/install.mjs /path/to/your/project
cd /path/to/your/project
copilot

The installer copies the runtime into .github/sieve/, writes only its own .github/hooks/sieve-copilot.json, and adds /.sieve/ to the project's .gitignore. Other hook files are left alone; an unrelated file using the same name is not overwritten. The runtime is copied into the project so hooks do not need read access to a separate sieve checkout. Start a new trusted Copilot session after installation; existing sessions are not modified.

The Copilot adapter:

  • Handles bash and powershell results for arbitrary commands, not just benchmark fixtures. It reuses sieve's test, build, install and git log filters; unrecognized output, code and patches remain unchanged.
  • Captures the current request via userPromptSubmitted. Clear detail/lookup requests keep the original output. Clear summary requests allow a recognized filter.
  • Uses the existing shared decider for unknown intent when a recognized filter could help. The question and the 0.5 lookup threshold are exactly the ones used by the Claude adapter. A probability of 0.5 or higher keeps the original.
  • Before the Copilot decider call, a strictly recognized leading block of execution instructions is moved behind the task and its answer requirements. No text is deleted or truncated by this preparation. Mixed clauses and positional references stay in their original order; the raw-request detail veto still runs first, and the overlong-request safeguard is unchanged. The Claude adapter is unaffected.
  • Precision requirements such as named test/metric identifiers, durations, measurements, rankings or neighbour comparisons keep the original before any decider call, including when mixed into an overview request. This conservative veto can miss optimization opportunities; it prevents a broad health-check interpretation from dropping explicitly needed data.
  • Never starts the decider. It contacts only 127.0.0.1:8765, aborts after 1.5 seconds, and applies a 30-second session cooldown after an unavailable, malformed or timed-out reply. The failure is logged; the original output stays intact.
  • With SIEVE_DECIDER=0, only explicit rule-approved summaries are filtered. Missing request state or an unknown request truncated beyond 500 characters is kept rather than guessed.
  • Reads Copilot's full persisted file when a large shell result has already been spilled. Full originals remain inside the project's .sieve/sessions/; totals, failure diagnostics and shell completion metadata are preserved.
  • Writes size/verdict metrics to .sieve/usage.jsonl, with no output content in that log. Up to 500 characters of the current prompt are stored locally in the session state for the local decider. .sieve/ must remain gitignored; originals are retained until you remove them.

The shared service is optional. On macOS, its existing setup is sh launchd/install.sh; install/start it separately, not from a hook. Disable it for Copilot with SIEVE_DECIDER=0 copilot.

This is not yet parity with the Claude adapter: no MCP-result filtering, FTS search tool, /sieve command, learned-keep/repeat recovery, compaction notes, task-switch hints or effort changes are provided for Copilot. Permissions are never granted or bypassed. To disable the adapter, remove .github/hooks/sieve-copilot.json; captured files are left in place.

Copilot measurements

The local prototype was evaluated in 60 fresh sessions (six workloads, five paired repetitions, GPT-6 Luna/high). All 60 answers were correct, but filtering a small exact lookup cost 40.1% more cumulative input tokens because it removed the sought detail. The request-aware fix was then evaluated in 30 new sessions (three workloads, five paired repetitions); all 30 answers were correct:

TaskMedian paired cumulative-input changeTool calls, baseline / sieve
Small exact lookupapproximately 0%1 / 1
Small test summary-8.7%1 / 1
Large test summary-35.2%2 / 1

These are synthetic, rules-only prototype measurements, not a benchmark of the newly added Copilot decider path or a real repository. Cumulative input includes cached input on every model request; it is not the same as billing. Aggregate metrics are in eval/copilot_benchmark.json. Do not infer general speed/cost savings from this small task mix.

The release adapter's installer, safe fallback and real shared-decider path are checked separately:

The v0.2.1 Copilot candidate was also compared with v0.2.0 in 50 fresh sessions: five synthetic workloads, five paired repetitions per version, with the shared decider enabled. All 50 answers were correct; one run carried a policy warning and remains in the totals. Previously unknown overview requests were filtered in 5/5 candidate runs versus 0/5 released-version runs, with 10.9% less median paired cumulative input. All 15 precision-output pairs were byte-identical. Across this particular task mix, cumulative input fell 2.2% and credits 6.2%; credits are cache-sensitive, and this is not a general speed or cost claim. Ordering alone had unsafe cuts in the first local holdout, so the candidate includes the precision veto as well. A final separate local evaluation retained one timeout and was marked incomplete for availability; unavailable decisions keep the original.

SIEVE_LIVE_DECIDER=1 node --experimental-strip-types --test copilot/copilot.test.mjs
node eval/copilot_smoke.mjs

The smoke test uses fresh projects under ~/dev/sieve-bench-runs/, exercises real Copilot command hooks, and removes those projects' hook configurations afterward. It does not change global Copilot settings.

How it works

A short fixed guide in the system prompt (about 110 tokens, cached after the first request) asks the model to plan commands so only the answer comes back, with one concrete pattern: write test, build and log output to a file and print the summary and failures. This is the part of context-mode's start-up text that changes behaviour, without its 5,000 characters and its note on every tool call. SIEVE_GUIDE=0 turns it off.

Results of Bash, Grep, Glob, WebFetch, any MCP tool that returns text, and Playwright snapshot files read back with Read are rewritten in place after they ran. Nothing is refused and the model has nothing to learn; the only tool is mcp__sieve__search.

  • Up to 4000 chars (Glob: 150 paths) a result is left alone. Other Reads are only cut above 80000 chars: code must stay whole.
  • In between, two checks, both must agree before anything is cut:
  • A rule (isRepetitive, no model): many lines of few shapes (digits and words blanked), such as a listing, a log, progress output or a table dump. Code, config and prose never pass.
  • The decider reads the request next to the output and answers one question: is it a lookup or count (how many, which, list all, find, exact) or an overview (what is this, does it look ok)? Cut only on overview. Anything at 0.5 or above for lookup keeps the output whole; no answer keeps it whole.
  • Above 30000 chars Bash and Grep results are cut without asking (the harness caps them anyway). MCP results and snapshot reads are never cut blind, up to 1 MB: a blind cut gave wrong answers in the browser test.
  • Known commands get their own filter (the RTK idea, inside the mod): test runners keep failures with their block, skips, warnings and totals and drop passing lines; git log becomes one line per commit (a patch is never filtered); installs and builds keep warnings, errors and the closing summary. Unnamed test scripts are recognised by their output. Applied from 2000 chars, without the decider.
  • A command that fails comes back from Claude Code as text already cut in the middle (about 10,000 chars), where failures usually are; sieve cannot recover that part, it can only filter what is left. The guide exists to avoid this case.
  • Learned keep: a call that is repeated right after a cut counts as a wrong cut for its type (git log, python -m pytest, a tool name); after two, that type is never cut again, across sessions ($.store). An Edit, Write or NotebookEdit in between makes the repeat a new measurement, not a wrong cut (pytest, fix, pytest is the normal loop).
  • JSON (a Bash or MCP result that parses) becomes a schema: fields with type, presence and number ranges, value counts for fields with few values, the first rows and the rows with rare values (status: failed among thousands of ok). For an object, its keys and its largest array of objects as such a table. It counts as data, so the decider weighs it like a repetitive result.
  • A cut result becomes a summary of its structure: a directory tree with counts by extension and folder, a per-file match table for Grep, or a table of line shapes with the rare and failure lines kept verbatim. The footer names the file with the whole output, so the model can grep/wc it; the output is also indexed (mcp__sieve__search, BM25).
  • A repeated call within two calls of a cut gets the whole output: the repeat is the signal that the cut was wrong.
  • The full output is written to ~/.claude/projects/<project>/<session>/tool-results/sieve-*.txt, the harness's own folder, because the model can read it there with narrow permissions (tested with Bash(grep:*) only: 3.5 MB of Grep output became 3 KB and the model counted 39,946 matches in the file).
  • The index is one SQLite file per project (~/.claude/sieve/<folder>-<hash of the path>.db; the hash avoids the slug collision of /a/b.c and /a/b-c); rows and files older than 14 days are deleted at session start.
  • Session capture and resume note: edited files, commands, failures, prompts and what was indexed; before a compaction a note (max 2000 chars) is stored and added to the system prompt after it.
  • Measurement: ~/.claude/sieve/usage.jsonl (tool, size, verdict; no content), eval/usage_report.py; /sieve shows cuts, chars kept out, repeated calls restored, decider use; the status line shows chars kept out. SIEVE_DECIDER=0 turns the decider off (the rule and size limits still apply, but then nothing in the middle band is cut).
  • There is no execute tool any more: it was never called in any test.

The decider's jobs

The decider makes small, fast, local decisions that a rule cannot make and that would be too slow or too expensive for a model call. Each job was tested on its own prompt set before it was built (eval/), with a hold-out part that was not used to pick the threshold, and each errs on the safe side.

JobQuestionActionTest (dev / hold-out)
Lookup or overviewIs the request a lookup or count, or an overview?cut a repetitive mid-sized result only for an overviewkept 125/125 needed outputs; harder set 40/40 kept, 35/40 cut
New taskDoes the request continue the recent work or start an unrelated task?a toast suggesting /clear or /compact (never automatic)at 0.8: 68/72 and 35/36 new tasks found, no follow-up ever called new
ReminderWill the request run tests, builds, installs or read logs?one reminder line in that turn (new text only, so the cache is untouched)at 0.4: 10/10 and 5/5 found, no false alarm
EffortDoes the session's first request need real reasoning?low effort for the whole session when it is clearly simpleat 0.7: 18/18 simple found, no hard request called simple; 0.8 is used

Effort is set once per session and kept (a resumed session reads it back), because the API invalidates the prompt cache whenever effort or thinking settings change; switching per step would cost more than it saves. The new-task check reads earlier prompts from the session record, so it also works when a session is resumed in a new process. SIEVE_EFFORT=0 turns the effort job off.

Measured effects:

  • Effort, single quick questions on the real repository (12 sessions per setup): output tokens median 165 with the effort job, 194 without, 234 for no plugin; all 36 answers right; cost about 3% lower (short sessions are dominated by input).
  • Effort, the five-step repository session (3 sessions each, side by side): all 15 steps right in both; cost $0.154 against $0.196. With 5 sessions on a harder bug (below) the difference was gone, so this was most likely variance. What holds: lowering the effort did no harm in 8 sessions that included a real bug fix.
  • Reminder, the verbose test step (5 runs): no gain over the guide alone (38k against 37k window-turns); the guide already does that work there. Kept because it costs one line.
  • New task, live over a resumed five-prompt session: only the unrelated prompt (a haiku in the middle of a bug fix) was flagged; the follow-ups before and after were not.

What was measured, and what was not

EvaluationResult
eval/rule_eval.mjs: repetitive-or-not rule vs 36 real outputs35/36 right; the miss is a stack trace repeated three times
eval/need_eval2.py: four question wordings, three prompt setsThe first wording (every line / a sample) kept the output whole every time but cut only 11-24 of 25-30 skim cases at a safe threshold. The wording "kind of question: lookup or count / overview" kept 85/85 whole at 0.5 and cut 83-85 of 85
eval/need_eval3.py: a harder set with no stock phrases (edits, checks on one line, worries)lookup/overview at 0.5: kept whole 40/40, cut 35/40. The first wording at its safe threshold: kept 40/40, cut 20/40. Higher thresholds for the new wording lose needed output (0.7: 30/40 kept)
eval/decider_eval.py, eval/task_eval.pyearlier designs; the record of why "repetitive" became a rule and the prompt-class factor was dropped

Over four prompt sets the chosen setting kept all 125 requests that need the whole output; the 0.5 threshold was picked on those same sets, and the first three sets share phrasing with the criteria, so read the cut rate (83-100%) as optimistic and the harder set (88%) as the better estimate. All prompts and outputs are small, hand-written and from one person.

Reproducing the benchmarks

Everything the benchmarks read is rebuilt from nothing by one script, at pinned versions:

sh eval/setup_bench.sh                  # fixtures, web pages, both repository templates, Playwright config
sh eval/setup_bench.sh --context-mode   # also a working copy of context-mode for the comparisons

It writes to .scratch/ (git-ignored; SCRATCH=<dir> to put it elsewhere) and stops if a template does not fail exactly as described below.

InputHow it is made
fixture/eval/make_fixture.py, fixed seed: logs, JSON, source files, a deep file tree, a test script, 150 git commits. Expected answers in fixture_expected.json
web/eval/make_web.py, fixed seed: an order ledger (450 rows) and a changelog (140 entries). Expected answers in web_expected.json
realrepo-template/more-itertools at 1ea82a7, with ilen() changed to count pairs (uncommitted). 26 tests fail
realrepo-hard/the same commit, with windowed() padding one value too many, committed as "Tidy up windowed padding" so git diff does not show it. 10 tests fail
context-mode/context-mode at 5a92b7c (1.0.169), runtime dependencies installed with a trimmed package.json (its own start-up install fails, see below)
playwright-mcp.jsonnpx @playwright/mcp@latest --headless --isolated

Versions the README's numbers come from: Claude Code 2.1.289, model claude-opus-5-5 (now pinned in every script, BENCH_MODEL to change it; earlier runs did not record the model, they used the account default, which was this one when checked), strands-decider build 6d5dec6 with checkpoint StrandsAgents/strands-decider-2B-hobson-v19 on MLX, Python 3.14, macOS on Apple silicon. Each run gets its own copy of a template under ~/dev/sieve-bench-runs (BENCH_RUNS).

What can still make a rerun differ: the model's own variance (about ±15% between runs of the same session, so compare setups side by side and use 5 or more sessions per setup), prompt-cache state (costs swing with it; tokens and calls are steadier), a newer Claude Code (it changes how it cuts and stores large outputs), and @playwright/mcp@latest.

Real use (eval/real_report.py)

Benchmarks are written by the people who build the tool; real sessions are the test that counts. Load sieve in every session and keep a control group:

"env": { "CLAUDE_CODE_PLUGIN_DIRS": "~/dev/sieve", "SIEVE_HOLDOUT": "0.2" }

in ~/.claude/settings.json. With SIEVE_HOLDOUT=0.2 one session in five is a holdout (set to 0.5 from 2026-10-05: equal groups reach a usable comparison faster): no cuts, no guide, no effort change, no reminder, no toast; it only logs what it would have done. Every line of usage.jsonl carries the session id and a project key; bench runs (cwd under sieve-bench-runs or .scratch, or SIEVE_BENCH=1) are marked and left out. Logged per session: each prompt's p(simple) and p(reminder) (one decider call for both), cuts and would-be cuts, follow-ups (a later call that reads a cut output's file), searches, restored repeats.

python3 eval/real_report.py [--days N] [--sessions] joins that log with Claude Code's own transcripts (~/.claude/projects/*/<session>.jsonl, token usage per model call) and compares active sessions with the holdout: window-turns, per call, output tokens, calls; how often the model went back to a cut output; and how many low-effort sessions later had a request with p(simple) < 0.3 (the effort job's known risk).

Benchmark (eval/bench.py, eval/report.py)

13 tasks on a generated project (eval/make_fixture.py) with large logs, data, file trees, test output and git history, headless claude -p, medians. Incomplete: 151 of 260 planned runs (eval/bench_v2.jsonl; the script resumes where it stopped).

Setup (about 43 runs each, same tasks)CorrectTokensContext at the end
no plugin100%42,75618,844
sieve, rules only100%42,91318,934
sieve, rules + decider100%37,84418,937
  • Where the output is large and the request does not need all of it, sieve saves tokens (mid-find-skim 37.8k vs 42.7k; run-log 37.1k vs 57.9k in the earlier run). Where the model already asks for a small output, nothing changes, which is most tasks.
  • The decider's share: about 5k tokens (-12%) over rules alone on the tasks run so far; one more check with 5 repetitions per cell is still due.
  • context-mode, 5-task check after repairing its install (1 run per cell, so no more than a sanity check): all correct, about 6.4k tokens more per session than no plugin (tool descriptions), same turns.
  • Single runs vary a lot (tests took 3 to 7 turns for the same setup because the model explores differently), so differences under about 10% are noise. Cost is noisy too (prompt-cache hits); tokens and turns are steadier.
  • Not measured: interactive sessions (approval dialogs), sessions long enough to compact, other models, the approval path of execute, wrong cuts in real use (/sieve counts them).

Real repository session (eval/repobench.py, eval/repo_report.py)

The closest to normal work: a copy of more-itertools (933 tests, 2,503 commits) with one injected bug, and a five-step session: run the tests verbosely and count skipped and failed, find and fix the bug (judged by running the suite afterwards, tests untouched), find the most-changed file in the last 200 commits, count the public functions in a 175 KB module, recall the bug. Each run gets its own copy. "Window-turns" sums the context size over every model call: what the window costs over the session.

RunSetupRightWindow-turnsUncached inputCost
first, 4 setups at onceno plugin15/15254,88311,340$0.200
context-mode15/15275,05921,995$0.288
sieve without the guide15/15258,50111,872$0.211
side by side, guide added (3 sessions)no plugin15/15259,3185,319$0.163
sieve with the guide15/15207,1734,125$0.130
side by side, ablationsieve15/15229,3474,255$0.135
sieve, decider off15/15256,4745,096$0.161
  • Repeated with Claude Code 2.1.289 (3 sessions each, side by side, eval/repo_results_v3.jsonl): no plugin 202,387 window-turns, sieve 204,484 (+1%), both 15/15 right, same calls. The guide's earlier gain is gone: the no-plugin run itself now keeps the verbose test run small (step 1: 2 calls, 36k window-turns, as with the guide). Session cost is bimodal in bo
Source 2 files
hooks/register.ts 538 lines
1import type { Register } from 'claude-code'
2import {
3  LIMITS,
4  LOOKUP_AT,
5  NEED_QUESTION,
6  buildSnapshot,
7  chunkText,
8  compactText,
9  filterCommand,
10  ftsQuery,
11  isRepetitive,
12  judgeSize,
13  needState,
14  projectKey,
15  signature,
16  sqlQuote,
17  summarize,
18  summarizeJson,
19} from './lib'
20
21const PORT = 8765
22const DECIDER = `http://127.0.0.1:${PORT}/v1/systemone`
23const SEARCH = 'mcp__sieve__search'
24const KEEP_DAYS = 14
25// New topic: suggest a reset only when the decider is this sure (no follow-up was ever called new at 0.8,
26// eval/roles_eval.py). Reminder: a missed reminder costs more than a needless line, so the bar is low.
27const SWITCH_AT = 0.8
28// Lower the effort only for a session whose first request is clearly simple (no hard request was called
29// simple at 0.7, eval/effort_eval.py). Set once: changing effort invalidates the prompt cache.
30const SIMPLE_AT = 0.8
31const REMIND_AT = 0.4
32const REMINDER = 'sieve: this request will likely run tests, builds or logs. Write their output to a file and print only the summary and failures (cmd > /tmp/out.log 2>&1; echo exit=$?; tail -n 15 /tmp/out.log; grep -E "FAIL|ERROR" /tmp/out.log).'
33// The one thing context-mode's start-up text gets right, in a few words: plan commands so only
34// the answer comes back. Fixed text in the system prompt, so it is cached after the first request.
35const GUIDE = 'Tool output stays in every later request, so ask commands for the answer, not the data. For test runs, builds and long logs, write the output to a file and print only what you need, for example: cmd > /tmp/out.log 2>&1; echo exit=$?; tail -n 15 /tmp/out.log; grep -E "FAIL|ERROR|skipped" /tmp/out.log. This matters most when a command fails: Claude Code then cuts the middle of its output, where the failures usually are. Long results may come back summarised by sieve with the path of the full output: query that file instead of running the command again.'
36// Command filters apply from this size on; a type of call cut wrongly this often is never cut again.
37const FILTER_MIN = 2000
38const WRONG_LIMIT = 2
39// A decider that does not answer in time counts as down: it must never hold up a prompt or a tool call.
40const DECIDER_MS = 1500
41const MUTATING = ['Edit', 'Write', 'NotebookEdit']
42
43const text = (t: string) => ({ result: [{ type: 'text', text: t }] })
44
45const S = {
46  checkedAt: 0,
47  useDecider: true,
48  guide: true,
49  session: '',
50  db: '',
51  deciderReady: false,
52  dir: '',
53  out: '',
54  cuts: [] as { what: string; sig: string; left: number }[],
55  recut: 0,
56  restored: 0,
57  prompt: '',
58  prompts: 0,
59  effort: undefined as undefined | 'low',
60  useEffort: true,
61  switches: 0,
62  reminders: 0,
63  seen: 0,
64  kept: 0,
65  compacted: 0,
66  asked: 0,
67  askedYes: 0,
68  indexed: 0,
69  // Real-use measurement: a holdout session only logs what it would have done; bench runs are marked.
70  holdout: false,
71  cutPaths: [] as string[],
72  bench: false,
73  project: '',
74}
75
76async function sql($: any, script: string) {
77  const ran = await $.process.run(['sqlite3', S.db], { stdin: script, timeoutMs: 60000 })
78  if (ran.exitCode !== 0) throw new Error(`sqlite3: ${ran.stderr.trim()}`)
79  return ran.stdout as string
80}
81
82async function record($: any, kind: string, data: string) {
83  if (!S.session || !data) return
84  const ts = await $.clock.now()
85  await sql($, `insert into events values (${sqlQuote(S.session)}, ${ts}, ${sqlQuote(kind)}, ${sqlQuote(data.slice(0, 300))});`)
86}
87
88async function snapshot($: any): Promise<string> {
89  const raw = await sql($, `.mode json\nselect kind, data from events where session = ${sqlQuote(S.session)} order by ts;`)
90  const events = raw.trim() ? JSON.parse(raw) : []
91  const compactions = await sql($, `select count from resume where session = ${sqlQuote(S.session)};`)
92  return events.length ? buildSnapshot(events, Number(compactions.trim() || 0) + 1) : ''
93}
94
95async function index($: any, source: string, content: string): Promise<number> {
96  const chunks = chunkText(content, source)
97  const rows = chunks
98    .map(c => `insert into chunks values (${sqlQuote(source)}, ${sqlQuote(c.title)}, ${sqlQuote(c.body)});`)
99    .join('\n')
100  const ts = await $.clock.now()
101  await sql($, `begin;\ndelete from chunks where source = ${sqlQuote(source)};\n${rows}\ninsert or replace into sources values (${sqlQuote(source)}, ${ts});\ncommit;`)
102  S.indexed += chunks.length
103  return chunks.length
104}
105
106async function search($: any, queries: string[], limit = 3): Promise<string> {
107  const out: string[] = []
108  for (const q of queries) {
109    if (!ftsQuery(q)) continue
110    let rows: { source: string; title: string; hit: string }[] = []
111    // All terms first; any term only when nothing holds them all together.
112    for (const join of ['AND', 'OR'] as const) {
113      const raw = await sql(
114        $,
115        `.mode json\nselect source, title, snippet(chunks, 2, '[', ']', '…', 40) as hit from chunks where chunks match ${sqlQuote(ftsQuery(q, join))} order by bm25(chunks) limit ${limit};`,
116      )
117      rows = raw.trim() ? JSON.parse(raw) : []
118      if (rows.length) break
119    }
120    out.push(
121      `## ${q}\n` +
122        (rows.map(r => `- ${r.source} › ${r.title}\n  ${r.hit.replaceAll('\n', ' ')}`).join('\n') || '(no matches)'),
123    )
124  }
125  return out.join('\n\n')
126}
127
128// $.http.fetch has no timeout of its own: undefined when the answer takes longer than DECIDER_MS.
129async function fetchWithin($: any, url: string, init?: Record<string, unknown>) {
130  return Promise.race([$.http.fetch(url, init), $.clock.sleep(DECIDER_MS).then(() => undefined)])
131}
132
133async function markDown($: any) {
134  S.deciderReady = false
135  S.checkedAt = await $.clock.now()
136}
137
138async function deciderUp($: any): Promise<boolean> {
139  if (!S.useDecider) return false
140  const now = await $.clock.now()
141  if (S.deciderReady || now - S.checkedAt < 30000) return S.deciderReady
142  S.checkedAt = now
143  try {
144    S.deciderReady = (await fetchWithin($, `http://127.0.0.1:${PORT}/health`))?.ok === true
145  } catch {
146    S.deciderReady = false
147  }
148  return S.deciderReady
149}
150
151async function ask($: any, state: string, questions: Record<string, unknown>) {
152  if (!(await deciderUp($))) return undefined
153  try {
154    const res = await fetchWithin($, DECIDER, {
155      method: 'POST',
156      headers: { 'content-type': 'application/json' },
157      body: JSON.stringify({ state, questions }),
158    })
159    if (!res) {
160      // hung or slow: off for 30 s, then checked again
161      await markDown($)
162      return undefined
163    }
164    return res.ok ? (JSON.parse(res.text).answers as Record<string, any>) : undefined
165  } catch {
166    await markDown($)
167    return undefined
168  }
169}
170
171// Which size limits apply to a call: any MCP result, a Playwright snapshot file read back, or the built-in tool.
172function limitKey(e: any): string {
173  if (e.tool.startsWith('mcp__')) return 'mcp'
174  if (e.tool === 'Read' && /\/\.playwright-mcp\//.test(String(e.file_path))) return 'ReadSnapshot'
175  return e.tool
176}
177
178const blocksOf = (r: any): any[] | undefined => (Array.isArray(r) ? r : Array.isArray(r?.content) ? r.content : undefined)
179
180// What a result carries as text, and the same result with that text replaced.
181function textOf(e: any, r: any): string | undefined {
182  if (e.tool.startsWith('mcp__')) {
183    const blocks = blocksOf(r)
184    return blocks && blocks.length && blocks.every(b => b?.type === 'text' && typeof b.text === 'string') ? blocks.map(b => b.text).join('\n\n') : undefined
185  }
186  switch (e.tool) {
187    // A command that failed comes back as one string, already cut in the middle by the harness.
188    case 'Bash': return typeof r === 'string' ? r : r.stdout
189    case 'Grep': return r.content
190    case 'WebFetch': return r.result
191    case 'Read': return r.file?.content
192    case 'Glob': return Array.isArray(r.filenames) ? r.filenames.join('\n') : undefined
193    default: return undefined
194  }
195}
196
197function withText(e: any, r: any, t: string): any {
198  if (e.tool.startsWith('mcp__')) return Array.isArray(r) ? [{ type: 'text', text: t }] : { ...r, content: [{ type: 'text', text: t }] }
199  switch (e.tool) {
200    case 'Bash': return typeof r === 'string' ? { stdout: t, stderr: '', interrupted: false } : { ...r, stdout: t, persistedOutputPath: undefined, persistedOutputSize: undefined }
201    case 'Grep': return { ...r, content: t }
202    case 'WebFetch': return { ...r, result: t }
203    case 'Read': return { ...r, file: { ...r.file, content: t } }
204    default: return { ...r, filenames: t.split('\n'), truncated: true }
205  }
206}
207
208// A short, stable label for what was called: how a repeat is recognised.
209function describe(e: any): string {
210  const { tool, tool_use_id, consent, ...input } = e
211  return String(e.command ?? e.pattern ?? e.url ?? e.file_path ?? `${tool} ${JSON.stringify(input).slice(0, 160)}`)
212}
213
214// What a learned rule is keyed on: the command's tool and subcommand, or the tool.
215function sigOf(e: any): string {
216  return e.tool === 'Bash' ? `Bash:${signature(String(e.command))}` : e.tool
217}
218
219async function wrongCuts($: any, sig: string): Promise<number> {
220  const all = ((await $.store.get('wrongCuts')) ?? {}) as Record<string, number>
221  return all[sig] ?? 0
222}
223
224async function learnWrongCut($: any, sig: string) {
225  const all = ((await $.store.get('wrongCuts')) ?? {}) as Record<string, number>
226  all[sig] = (all[sig] ?? 0) + 1
227  await $.store.set('wrongCuts', all)
228}
229
230// The decider's one job on a result: does the request need every line of it? The rule has
231// already said the output is repetitive; code and prose never get here.
232async function needsEveryLine($: any, e: any, full: string): Promise<boolean> {
233  if (!S.prompt) return true
234  const answers = await ask($, needState(S.prompt, e.tool, describe(e), full), NEED_QUESTION)
235  const p = answers?.kind?.probabilities?.['lookup or count']
236  S.asked += 1
237  // No answer means keep it whole: a wrong cut costs a round trip, a missed cut only some bytes.
238  if (typeof p !== 'number' || p >= LOOKUP_AT) return true
239  S.askedYes += 1
240  return false
241}
242
243// What the session has been doing, for the decider to compare a new request against.
244async function recentWork($: any): Promise<string> {
245  const raw = await sql($, `.mode json\nselect kind, data from events where session = ${sqlQuote(S.session)} and kind in ('prompt','file','command') order by ts desc limit 12;`)
246  const rows: { kind: string; data: string }[] = raw.trim() ? JSON.parse(raw) : []
247  const of = (k: string, n: number) => rows.filter(r => r.kind === k).slice(0, n).map(r => r.data.slice(0, 160))
248  return `Recent requests: ${of('prompt', 3).map(p => `'${p}'`).join(', ') || 'none'}. Recently edited: ${of('file', 4).join(', ') || 'nothing'}. Recently ran: ${of('command', 3).join('; ') || 'nothing'}.`
249}
250
251async function isNewTask($: any, text: string): Promise<boolean> {
252  const answers = await ask($, `${await recentWork($)}\nNew request: ${text.slice(0, 600)}`, {
253    q: {
254      type: 'choice',
255      instructions: 'Does the new request continue the recent work, or start a different, unrelated task?',
256      criteria: {
257        continue: 'a follow-up, fix, extension, question or action about the same work',
258        'new task': 'a different topic, project or kind of work that does not need the recent context',
259      },
260    },
261  })
262  const p = answers?.q?.probabilities?.['new task']
263  return typeof p === 'number' && p >= SWITCH_AT
264}
265
266// Both questions read the same state (the request alone), so they share one decider call.
267async function promptSignals($: any, text: string): Promise<{ simple?: number; remind?: number }> {
268  const answers = await ask($, text.slice(0, 600), {
269    simple: {
270      type: 'choice',
271      instructions: 'How much reasoning does this request need?',
272      criteria: {
273        simple: 'a lookup, a count, a one-line change, a rename, a quick factual question or running one command',
274        complex: 'debugging, designing, a change across several files, an unclear cause, or anything that needs careful thought',
275      },
276    },
277    remind: {
278      type: 'choice',
279      instructions: 'Will answering this request involve running tests, builds, installs or reading long logs?',
280      criteria: {
281        yes: 'it runs a test suite, a build or compile, an install, a linter, or reads logs or CI output',
282        no: 'it reads or edits code, explains, writes text, or answers a question',
283      },
284    },
285  })
286  const num = (v: unknown) => (typeof v === 'number' ? v : undefined)
287  return { simple: num(answers?.simple?.probabilities?.simple), remind: num(answers?.remind?.probabilities?.yes) }
288}
289
290async function logUsage($: any, line: Record<string, unknown>) {
291  try {
292    const tag = { s: S.session, p: S.project, ...(S.holdout ? { holdout: true } : {}), ...(S.bench ? { bench: true } : {}) }
293    await $.process.run(['sh', '-c', 'cat >> "$0"', `${S.dir}/usage.jsonl`], { stdin: `${JSON.stringify({ t: await $.clock.now(), ...tag, ...line })}\n` })
294  } catch {
295    // measurement must never get in the way
296  }
297}
298
299// A cut followed within two calls by the same call again was a wrong cut: the repeat gets the
300// whole output, and the type of call is remembered across sessions. An edit in between makes the
301// repeat a new measurement (pytest, Edit, pytest), not a wrong cut.
302function watchRepeat(e: any): string | undefined {
303  if (MUTATING.includes(e.tool)) {
304    S.cuts = []
305    return undefined
306  }
307  const what = describe(e)
308  let hit: string | undefined
309  for (const c of S.cuts) {
310    if (c.left <= 0) continue
311    if (c.what === what) hit = c.sig
312    c.left -= 1
313  }
314  S.cuts = S.cuts.filter(c => c.left > 0 && c.what !== what)
315  if (hit) S.recut += 1
316  return hit
317}
318
319// Cuts a large result in place: the model gets a summary of its structure, the whole output
320// stays in a file and in the index. Undefined means leave the result as it is.
321async function compact($: any, e: any, ran: any): Promise<any | undefined> {
322  const r = ran.result
323  const key = limitKey(e)
324  if (ran.deny !== undefined || !r || !LIMITS[key]) return undefined
325  let full = textOf(e, r)
326  if (typeof full !== 'string') {
327    await logUsage($, { tool: e.tool, unreadable: typeof r, keys: r && typeof r === 'object' ? Object.keys(r).slice(0, 12) : [], isError: ran.isError === true, textLen: String(ran.text ?? '').length })
328    return undefined
329  }
330  // Bash keeps a long output in a file and hands back a preview: use the file, not the preview.
331  let path: string | undefined = typeof r === 'object' ? r.persistedOutputPath : undefined
332  if (path) {
333    try {
334      full = await $.fs.read(path)
335    } catch {
336      // over 4 MiB or gone: the preview is what there is
337    }
338  }
339  const sig = sigOf(e)
340  if ((await wrongCuts($, sig)) >= WRONG_LIMIT) {
341    await logUsage($, { tool: key === 'mcp' ? 'mcp' : e.tool, size: full.length, verdict: 'learned-keep' })
342    return undefined
343  }
344  // A known command gets its own filter: failures, totals and warnings stay, the rest is noise.
345  const filtered = e.tool === 'Bash' && full.length > FILTER_MIN ? filterCommand(String(e.command), full) : undefined
346  const size = e.tool === 'Glob' ? r.filenames.length : full.length
347  const verdict = filtered ? 'filter' : judgeSize(key, size, ran.isError === true, 1)
348  // JSON is data, never code: it gets a schema summary and counts as repetitive for the decider.
349  const json = filtered || size <= (LIMITS[key]?.soft ?? Infinity) ? undefined : summarizeJson(full)
350  const repetitive = e.tool === 'Glob' || json !== undefined || isRepetitive(full)
351  let cut = verdict === 'compact' || verdict === 'filter'
352  if (verdict === 'ask' && repetitive) cut = !(await needsEveryLine($, e, full))
353  await logUsage($, { tool: key === 'mcp' ? 'mcp' : e.tool, sig, size, verdict, repetitive, json: json !== undefined, cut })
354  if (!cut || S.holdout) return undefined
355
356  const source = `${e.tool}:${(await $.clock.now()).toString(36)}${++S.seen}`.replace(/[^\w.:-]/g, '_')
357  if (!path) {
358    path = `${S.out}/sieve-${source.replace(/:/g, '-')}.txt`
359    await $.fs.write(path, full)
360  }
361  S.cutPaths.push(path.split('/').pop()!)
362  const foot = `[sieve: ${full.length} chars summarised. Full output: ${path}. For exact counts or lookups run grep/wc/awk on that file; ${SEARCH} finds passages. Repeat the same call to get everything.]`
363  const short = filtered ? `${filtered}\n${foot}` : json ? `${json}\n${foot}` : repetitive ? `${summarize(e.tool, full)}\n${foot}` : `${compactText(full, { head: 1800, tail: 1200, signal: 15 })}\n${foot}`
364  await index($, source, full)
365  await record($, 'cut', `${source} ${describe(e)}`)
366  S.cuts.push({ what: describe(e), sig, left: 2 })
367  S.kept += full.length - short.length
368  S.compacted += 1
369  $.ui.status(`sieve: ${Math.round(S.kept / 1000)}k chars kept out`)
370  return { result: withText(e, r, short) }
371}
372
373export const register: Register = on => {
374  on('session.start', async ($, e, next) => {
375    const home = (await $.env.get('HOME')) ?? ''
376    S.useDecider = (await $.env.get('SIEVE_DECIDER')) !== '0'
377    S.guide = (await $.env.get('SIEVE_GUIDE')) !== '0'
378    S.useEffort = (await $.env.get('SIEVE_EFFORT')) !== '0'
379    // SIEVE_HOLDOUT=0.2: one session in five changes nothing and only logs, the control group for real use.
380    S.holdout = Math.random() < Number((await $.env.get('SIEVE_HOLDOUT')) ?? 0)
381    S.bench = (await $.env.get('SIEVE_BENCH')) === '1' || /\/(sieve-bench-runs|\.scratch)\//.test(`${e.cwd}/`)
382    S.project = projectKey(e.cwd)
383    // The harness's own slug for its projects folder (S.out must match it); the index gets a collision-free key.
384    const slug = e.cwd.replace(/[/.]/g, '-')
385    S.dir = `${home}/.claude/sieve`
386    S.db = `${S.dir}/${projectKey(e.cwd)}.db`
387    S.session = await $.session.id()
388    // The harness's own folder for long outputs: the model may read it without asking.
389    S.out = `${home}/.claude/projects/${slug}/${S.session}/tool-results`
390    await $.process.run(['mkdir', '-p', S.dir, S.out])
391    const cutoff = (await $.clock.now()) - KEEP_DAYS * 86400000
392    await sql(
393      $,
394      `create virtual table if not exists chunks using fts5(source, title, body, tokenize='porter unicode61');
395create table if not exists sources (source text primary key, ts integer);
396create table if not exists events (session text, ts integer, kind text, data text);
397create table if not exists resume (session text primary key, snapshot text, count integer);
398delete from chunks where source in (select source from sources where ts < ${cutoff});
399delete from sources where ts < ${cutoff};
400delete from events where ts < ${cutoff};`,
401    )
402    await $.process.run(['find', `${home}/.claude/projects`, '-name', 'sieve-*.txt', '-mtime', `+${KEEP_DAYS}`, '-delete'])
403
404    await $.tool.register({
405      name: 'search',
406      description: 'Search output that was cut from earlier results (BM25). Pass several related queries at once.',
407      inputSchema: {
408        type: 'object',
409        properties: { queries: { type: 'array', items: { type: 'string' } }, limit: { type: 'number' } },
410        required: ['queries'],
411      },
412    })
413    await $.command.register({ name: 'sieve', description: 'sieve: what was kept out of the context' })
414    void deciderUp($)
415    return next(e)
416  })
417
418  on('command.run', { command: 'sieve' }, async $ => {
419    const raw = await sql($, `.mode json\nselect count(*) as chunks, count(distinct source) as sources from chunks;`)
420    const { chunks, sources } = JSON.parse(raw)[0]
421    return {
422      text: `sieve: ${S.compacted} results cut this session, ~${Math.round(S.kept / 1000)}k chars kept out of context; ${S.restored} repeated calls got the whole output; ${S.reminders} reminders, ${S.switches} new-task hints; ${chunks} chunks from ${sources} sources indexed (this project); decider ${S.deciderReady ? `ready, asked ${S.asked}x, allowed a cut ${S.askedYes}x` : 'off'}.`,
423    }
424  })
425
426  on('tool.call', { tool: 'mcp__sieve__search' }, async ($, e: any) => {
427    await logUsage($, { event: 'search' })
428    return text(await search($, e.queries, e.limit ?? 3))
429  })
430
431  // Every result passes here: recorded for the resume note, cut when it is large.
432  on('tool.call', async ($, e: any, next) => {
433    if (e.tool.startsWith('mcp__sieve__')) return next(e)
434    // The model going back to a cut output: the summary was not enough on its own.
435    const called = JSON.stringify(e)
436    if (S.cutPaths.some(f => called.includes(f))) await logUsage($, { event: 'followup', tool: e.tool })
437    const repeat = watchRepeat(e)
438    const ran = await next(e)
439    if (!LIMITS[limitKey(e)]) await logUsage($, { tool: e.tool, size: String(ran.text ?? '').length })
440    try {
441      if (['Edit', 'Write', 'NotebookEdit'].includes(e.tool) && ran.deny === undefined) await record($, 'file', e.file_path ?? e.notebook_path)
442      else if (e.tool === 'Bash' && ran.deny === undefined) await record($, ran.isError ? 'error' : 'command', e.command)
443      if (repeat) {
444        S.restored += 1
445        await learnWrongCut($, repeat)
446        await logUsage($, { tool: e.tool, sig: repeat, restored: true })
447        return ran
448      }
449      return (await compact($, e, ran)) ?? ran
450    } catch (err) {
451      // never break a tool call; but say why nothing was cut
452      await logUsage($, { tool: e.tool, error: String(err).slice(0, 300) })
453      return ran
454    }
455  })
456
457  on('session.compact', async ($, e, next) => {
458    try {
459      const note = await snapshot($)
460      if (note) {
461        const n = await sql($, `select count from resume where session = ${sqlQuote(S.session)};`)
462        await sql($, `insert or replace into resume values (${sqlQuote(S.session)}, ${sqlQuote(note)}, ${Number(n.trim() || 0) + 1});`)
463      }
464    } catch {
465      // a failed snapshot only costs the resume note
466    }
467    return next(e)
468  })
469
470  on('prompt.compose', async ($, e, next) => {
471    const composed = await next(e)
472    try {
473      const raw = await sql($, `.mode json\nselect snapshot from resume where session = ${sqlQuote(S.session)};`)
474      const note = raw.trim() ? JSON.parse(raw)[0]?.snapshot : ''
475      const guide = S.guide && !S.holdout ? [{ id: 'sieve:guide', text: GUIDE, scope: 'session' as const }] : []
476      const resume = note ? [{ id: 'sieve:resume', text: `Before the last compaction, this session had:\n${note}`, scope: 'session' as const }] : []
477      return { sections: [...composed.sections, ...guide, ...resume] }
478    } catch {
479      // no note, no section
480    }
481    return composed
482  })
483
484  // The main loop's requests carry the session's effort; subagents keep their own.
485  on('turn.step', async function* ($, e, next) {
486    if (S.effort && !e.agentId) return yield* next({ ...e, effort: S.effort })
487    return yield* next(e)
488  })
489
490  // The request is what the decider weighs results against; it also decides whether this turn
491  // starts a new topic (suggest a reset) and whether it needs the one-line reminder.
492  on('prompt.submit', async ($, e, next) => {
493    const text = e.text
494    S.prompt = text.slice(0, 500)
495    let context = e.context ?? []
496    try {
497      // earlier prompts of this session, from the record: a resumed session starts a new process
498      const prior = Number((await sql($, `select count(*) from events where session = ${sqlQuote(S.session)} and kind = 'prompt';`)).trim() || 0)
499      const sig = await promptSignals($, text)
500      await logUsage($, { event: 'prompt', n: prior, simple: sig.simple, remind: sig.remind })
501      // Effort is decided once, at the first request, and kept for the session (a resumed one included).
502      if (prior === 0) {
503        if (S.useEffort && (sig.simple ?? 0) >= SIMPLE_AT) {
504          await logUsage($, { event: 'effort-low' })
505          if (!S.holdout) {
506            S.effort = 'low'
507            await record($, 'effort', 'low')
508          }
509        }
510      } else {
511        const kept = (await sql($, `select data from events where session = ${sqlQuote(S.session)} and kind = 'effort' order by ts desc limit 1;`)).trim()
512        S.effort = kept === 'low' ? 'low' : undefined
513      }
514      if (prior >= 2 && (await isNewTask($, text))) {
515        S.switches += 1
516        await logUsage($, { event: 'new-task' })
517        if (!S.holdout)
518        // Effort stays low for the session (changing it breaks the cache); /clear is the clean place to reset it.
519        $.ui.toast(
520          S.effort === 'low'
521            ? 'sieve: this looks like a new task, and this session runs at low effort. /clear keeps the earlier work out of every request and restores normal effort.'
522            : 'sieve: this looks like a new task. /clear (or /compact) would keep the earlier work out of every request.',
523        )
524      }
525      if (S.guide && (sig.remind ?? 0) >= REMIND_AT) {
526        S.reminders += 1
527        await logUsage($, { event: 'reminder' })
528        if (!S.holdout) context = [...context, REMINDER]
529      }
530    } catch {
531      // the decider is optional
532    }
533    S.prompts += 1
534    await record($, 'prompt', text.slice(0, 200)).catch(() => {})
535    return next(context === e.context ? e : { ...e, context })
536  })
537}
538
hooks/lib.ts 460 lines
1export const RESULT_LIMIT = 6000
2export const AUTO_INDEX_LIMIT = 12000
3
4// Chosen on four prompt sets: at or above this, a lookup needs the full output.
5export const LOOKUP_AT = 0.5
6export const NEED_QUESTION = {
7  kind: {
8    type: 'choice',
9    instructions: 'Which kind of question is the request?',
10    criteria: {
11      'lookup or count': 'how many, which one, list all, find, exists, exact',
12      overview: 'what is this, summarize, describe, does it look ok, any sign of trouble',
13    },
14  },
15}
16
17export const needState = (prompt: string, tool: string, description: string, full: string): string =>
18  `Request: ${prompt}\n${tool}: ${description}\n---\n${full.slice(0, 1500)}\n…\n${full.slice(-500)}`
19
20export const sqlQuote = (s: string): string => `'${s.replaceAll("'", "''")}'`
21
22export const ftsQuery = (q: string, join: 'AND' | 'OR' = 'OR'): string =>
23  (q.match(/[\p{L}\p{N}_]{2,}/gu) ?? []).map(t => `"${t}"`).join(` ${join} `)
24
25export type Chunk = { title: string; body: string }
26
27export const chunkText = (text: string, fallbackTitle: string, max = 1800): Chunk[] => {
28  const chunks: Chunk[] = []
29  let title = fallbackTitle
30  let buf: string[] = []
31  let size = 0
32  const flush = () => {
33    const body = buf.join('\n').trim()
34    if (body) chunks.push({ title, body })
35    buf = []
36    size = 0
37  }
38  for (const line of text.split('\n')) {
39    const heading = /^#{1,4}\s+(.*)/.exec(line)
40    if (heading) {
41      flush()
42      title = heading[1]!.slice(0, 120)
43    }
44    buf.push(line)
45    size += line.length + 1
46    if (size >= max) flush()
47  }
48  flush()
49  return chunks
50}
51
52export const headTail = (text: string, limit = RESULT_LIMIT): string =>
53  text.length <= limit
54    ? text
55    : `${text.slice(0, limit / 2)}\n… [${text.length - limit} chars omitted] …\n${text.slice(-limit / 2)}`
56
57export const interpreter = (language: string, code: string): string[] | undefined => {
58  switch (language) {
59    case 'shell':
60    case 'bash':
61    case 'sh':
62      return ['/bin/sh', '-c', code]
63    case 'python':
64      return ['python3', '-c', code]
65    case 'javascript':
66    case 'js':
67      return ['node', '-e', code]
68    default:
69      return undefined
70  }
71}
72
73// Commands whose output is large by nature, so no model call is needed to say so.
74const BULKY =
75  /^\s*(git\s+(log|diff|show)(?!.*(-n\s*\d|--stat|--oneline.*-\d|-\d+\b))|find\s|ls\s+-\w*R|tree\b|docker\s+logs|kubectl\s+(logs|get\s+.*-o\s*(yaml|json))|npm\s+(ls|list)\b|pip\s+freeze|cat\s+\S+\.(log|json|csv|lock))/
76const HAS_LIMIT = /\|\s*(head|tail|wc|grep|rg|jq|awk|sed|sort\s.*\|\s*head)\b|>\s*\S+|--max-count|-n\s*\d+/
77
78export const isBulkyCommand = (command: string): boolean =>
79  BULKY.test(command) && !HAS_LIMIT.test(command)
80
81export const isRawFetch = (command: string): boolean =>
82  /^\s*(curl|wget)\b/.test(command) && !/(\s-o\s|\s-O\b|--output|>\s*\S+|\|\s*(head|jq|grep|wc))/.test(command)
83
84export type SessionEvent = { kind: string; data: string }
85
86// A resume note small enough to ride in the system prompt after a compaction.
87export const buildSnapshot = (events: SessionEvent[], compactCount: number, max = 2000): string => {
88  const of = (kind: string) => events.filter(e => e.kind === kind).map(e => e.data)
89  const last = (xs: string[], n: number) => [...new Set(xs)].slice(-n)
90  const lines = [
91    `<resume compactions="${compactCount}">`,
92    ...last(of('prompt'), 4).map(p => `  <ask>${p}</ask>`),
93    ...last(of('file'), 15).map(f => `  <edited>${f}</edited>`),
94    ...last(of('error'), 4).map(c => `  <failed>${c}</failed>`),
95    ...last(of('command'), 5).map(c => `  <ran>${c}</ran>`),
96    ...last(of('cut'), 8).map(c => `  <indexed>${c}</indexed>`),
97    '</resume>',
98  ]
99  let out = lines.join('\n')
100  while (out.length > max && lines.length > 2) {
101    lines.splice(1, 1)
102    out = lines.join('\n')
103  }
104  return out
105}
106
107// Lines in a cut-out middle that must survive: failures are what the model reads logs for.
108const SIGNAL = /\b(error|fail(ed|ure|ing)?|fatal|panic|exception|traceback|warn(ing)?|denied|refused|timeout|not found|cannot|unable)\b/i
109
110export type CompactOptions = { head: number; tail: number; signal: number }
111
112export const compactText = (text: string, { head, tail, signal }: CompactOptions): string => {
113  if (text.length <= head + tail) return text
114  const lines = text.split('\n')
115  let start = 0
116  let used = 0
117  while (start < lines.length && used + lines[start]!.length < head) used += lines[start++]!.length + 1
118  let end = lines.length
119  used = 0
120  while (end > start && used + lines[end - 1]!.length < tail) {
121    end -= 1
122    used += lines[end]!.length + 1
123  }
124  const middle = lines.slice(start, end)
125  const hits = middle.filter(l => SIGNAL.test(l)).slice(0, signal).map(l => l.slice(0, 200))
126  const note = `… [${middle.length} lines / ${middle.join('\n').length} chars omitted${hits.length ? `; ${hits.length} signal line(s) kept below` : ''}] …`
127  return [...lines.slice(0, start), note, ...hits, ...(hits.length ? ['…'] : []), ...lines.slice(end)].join('\n')
128}
129
130// Characters a result may carry before it is cut (`soft`), and above which it is cut without
131// asking the decider (`hard`). Glob counts paths, not characters.
132export const LIMITS: Record<string, { soft: number; hard: number }> = {
133  Bash: { soft: 4000, hard: 30000 },
134  Grep: { soft: 4000, hard: 30000 },
135  WebFetch: { soft: 3000, hard: 12000 },
136  Glob: { soft: 150, hard: 300 },
137  Read: { soft: 80000, hard: 80000 },
138  // Any MCP result made of text, and a Playwright snapshot file read back: data, not code.
139  // No blind cut here: the decider always weighs the request first (a blind cut at 30000 gave wrong answers in the browser test).
140  mcp: { soft: 4000, hard: 1000000 },
141  ReadSnapshot: { soft: 4000, hard: 1000000 },
142}
143
144export type Verdict = 'pass' | 'ask' | 'compact'
145
146// Pure size rule; the decider only ever sees the `ask` band.
147export const judgeSize = (tool: string, size: number, isError: boolean, factor = 1): Verdict => {
148  const limit = LIMITS[tool]
149  if (!limit) return 'pass'
150  const soft = limit.soft * factor
151  const hard = limit.hard * factor
152  if (size <= soft) return 'pass'
153  if (size > hard) return 'compact'
154  return isError ? 'pass' : 'ask'
155}
156
157// ---- structure-aware summaries -------------------------------------------------------------
158
159// A line with its variable parts blanked: "GET /items/17 200 12ms" and "GET /items/9 200 7ms" share one.
160export const template = (line: string): string =>
161  line
162    .replace(/[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}/gi, '<id>')
163    .replace(/\b[0-9a-f]{12,}\b/gi, '<hex>')
164    .replace(/\d+/g, '#')
165    .slice(0, 120)
166
167// Coarser still: every word is "w", so "numpy 2.1.0" and "torch 2.14.0" are one shape. Lists of
168// names (a directory, installed packages) repeat at this level though no two lines are alike.
169const coarse = (line: string): string =>
170  line.replace(/[A-Za-z_][\w.-]*/g, 'w').replace(/\d+/g, '#').replace(/\s+/g, ' ').replace(/(w[ ./-]?)+/g, 'W').slice(0, 60)
171
172export type Shape = { lines: number; distinct: number; top3: number; coarseTop3: number }
173
174export const lineShape = (text: string): Shape => {
175  const counts = new Map<string, number>()
176  const rough = new Map<string, number>()
177  const lines = text.split('\n').filter(l => l.trim())
178  for (const l of lines) {
179    counts.set(template(l), (counts.get(template(l)) ?? 0) + 1)
180    rough.set(coarse(l), (rough.get(coarse(l)) ?? 0) + 1)
181  }
182  const share = (m: Map<string, number>) => ([...m.values()].sort((a, b) => b - a).slice(0, 3).reduce((a, b) => a + b, 0) / Math.max(lines.length, 1))
183  return { lines: lines.length, distinct: counts.size, top3: share(counts), coarseTop3: share(rough) }
184}
185
186// Many lines built from a few shapes: a listing, a log, progress output. Code and prose are not.
187export const isRepetitive = (text: string): boolean => {
188  const s = lineShape(text)
189  return s.lines >= 30 && (s.distinct / s.lines <= 0.3 || s.top3 >= 0.6 || s.coarseTop3 >= 0.9)
190}
191
192const SIGNAL_LINE = /\b(fail(ed|ure|ing)?|fatal|panic|exception|traceback|denied|refused|timed? ?out)\b|Error\b|error:/i
193
194const clip = (l: string, n = 200) => (l.length > n ? `${l.slice(0, n)}…` : l)
195
196const isPath = (l: string) => /^[\w@.~/-][^\s:]*\/[^\s:]*$/.test(l.trim())
197const isGrepLine = (l: string) => /^[^\s:]+:\d+[:-]/.test(l)
198
199export const kindOf = (tool: string, text: string): 'listing' | 'grep' | 'lines' => {
200  const lines = text.split('\n').filter(l => l.trim())
201  const share = (f: (l: string) => boolean) => lines.filter(f).length / Math.max(lines.length, 1)
202  if (tool === 'Glob' || share(isPath) >= 0.8) return 'listing'
203  if (tool === 'Grep' || share(isGrepLine) >= 0.8) return 'grep'
204  return 'lines'
205}
206
207const top = <T>(m: Map<T, number>, n: number): [T, number][] => [...m.entries()].sort((a, b) => b[1] - a[1]).slice(0, n)
208
209const treeSummary = (text: string): string => {
210  const paths = text.split('\n').filter(l => l.trim())
211  const ext = new Map<string, number>()
212  for (const p of paths) {
213    const m = /\.([A-Za-z0-9]+)$/.exec(p)
214    ext.set(m ? `.${m[1]}` : '(none)', (ext.get(m ? `.${m[1]}` : '(none)') ?? 0) + 1)
215  }
216  // The deepest directory level that still gives at most 40 groups.
217  let groups = new Map<string, number>()
218  for (let depth = 1; depth <= 12; depth++) {
219    const g = new Map<string, number>()
220    for (const p of paths) {
221      const dir = p.split('/').slice(0, depth).join('/')
222      g.set(dir, (g.get(dir) ?? 0) + 1)
223    }
224    if (g.size > 40 && depth > 1) break
225    groups = g
226  }
227  return [
228    `${paths.length} paths. By extension: ${top(ext, 10).map(([e, n]) => `${e} ${n}`).join(', ')}`,
229    'By directory:',
230    ...[...groups.entries()].map(([d, n]) => `  ${d}  ${n}`),
231    `first: ${paths[0]}`,
232    `last: ${paths[paths.length - 1]}`,
233  ].join('\n')
234}
235
236const grepSummary = (text: string): string => {
237  const files = new Map<string, number>()
238  const first = new Map<string, string>()
239  for (const l of text.split('\n')) {
240    const m = /^([^\s:]+):\d+[:-]/.exec(l)
241    if (!m) continue
242    files.set(m[1]!, (files.get(m[1]!) ?? 0) + 1)
243    if (!first.has(m[1]!)) first.set(m[1]!, clip(l, 160))
244  }
245  const total = [...files.values()].reduce((a, b) => a + b, 0)
246  return [
247    `${total} matches in ${files.size} files. Per file (most first), with its first hit:`,
248    ...top(files, 30).map(([f, n]) => `  ${n}×  ${first.get(f)}`),
249  ].join('\n')
250}
251
252// Frequent shapes as a table; shapes seen once or twice, and failure lines, verbatim.
253const linesSummary = (text: string, rareCap = 25): string => {
254  const lines = text.split('\n')
255  const counts = new Map<string, number>()
256  const example = new Map<string, string>()
257  for (const l of lines) {
258    if (!l.trim()) continue
259    const t = template(l)
260    counts.set(t, (counts.get(t) ?? 0) + 1)
261    if (!example.has(t)) example.set(t, clip(l, 140))
262  }
263  const keep = new Set<number>()
264  lines.forEach((l, i) => {
265    if (!l.trim()) return
266    if (i < 4 || i >= lines.length - 4) keep.add(i)
267    else if (counts.get(template(l))! <= 2 || SIGNAL_LINE.test(l)) keep.add(i)
268  })
269  const kept = [...keep].sort((a, b) => a - b)
270  const picked = kept.length > rareCap + 8 ? [...kept.slice(0, 4), ...kept.slice(4, -4).slice(0, rareCap), ...kept.slice(-4)] : kept
271  return [
272    `${lines.filter(l => l.trim()).length} lines, ${counts.size} distinct shapes. Most frequent:`,
273    ...top(counts, 10).map(([t, n]) => `  ${n}×  ${example.get(t)}`),
274    `Rare and failure lines, in order (${picked.length} of ${kept.length}):`,
275    ...picked.map(i => `  ${clip(lines[i]!)}`),
276  ].join('\n')
277}
278
279export const summarize = (tool: string, text: string): string => {
280  const kind = kindOf(tool, text)
281  return kind === 'listing' ? treeSummary(text) : kind === 'grep' ? grepSummary(text) : linesSummary(text)
282}
283
284// ---- command-specific filters (the RTK idea, inside the mod) --------------------------------
285
286// "cd x && git log --stat -n 5" -> "git log": what a learned rule or a filter is keyed on.
287export const signature = (command: string): string => {
288  const last = command.split(/&&|;|\|\|/).map(s => s.trim()).filter(Boolean).filter(s => !/^cd\s/.test(s))[0] ?? ''
289  const words = last.replace(/^(sudo|time|env(\s+\w+=\S+)+)\s+/, '').split(/\s+/)
290  const tool = (words[0] ?? '').split('/').pop() ?? ''
291  const sub = words.slice(1).find(w => !w.startsWith('-')) ?? ''
292  const withSub = ['git', 'npm', 'pnpm', 'yarn', 'uv', 'pip', 'pip3', 'cargo', 'go', 'docker', 'kubectl', 'poetry', 'bun']
293  if (/^python3?$/.test(tool) && words[1] === '-m') return `python -m ${words[2] ?? ''}`
294  return withSub.includes(tool) && sub ? `${tool} ${sub}` : tool
295}
296
297const TEST_CMD = /^(pytest|py\.test|python -m (pytest|unittest)|jest|vitest|mocha|cargo test|go test|npm test|pnpm test|yarn test|bun test|npm run|pnpm run|yarn run|mvn|gradle|rspec|phpunit|tox|nox|make)$/
298const PASS_LINE = /^\s*(PASS\b|✓|✔|ok\b|\.+$|test \S+ \.\.\. ok$|\S+\s+\.\.\.\s+ok$|.*\bPASSED\b|=== RUN\b|--- PASS\b)/
299const FAIL_LINE = /\b(FAIL(ED)?|ERROR|Error|Exception|Traceback|panic|AssertionError|assert|✗|✕|×)\b|^\s*E\s{2,}/
300const SKIP_LINE = /\b(skip(ped)?|xfail|xpass|expected failure|warn(ing)?|deprecat\w*)\b/i
301const SUMMARY_LINE = /\b(\d+ (passed|failed|errors?|skipped|tests?|suites?)|Ran \d+ tests?|^OK\b|^FAILED\b|Tests?:|Test Suites:|test result:|Summary)\b/i
302
303// Passing tests and progress dots go; failures keep their block, totals stay.
304export const filterTests = (text: string): string => {
305  const lines = text.split('\n')
306  const keep: string[] = []
307  let dropped = 0
308  let block = 0
309  for (let i = 0; i < lines.length; i++) {
310    const l = lines[i]!
311    const passing = PASS_LINE.test(l) && !FAIL_LINE.test(l)
312    if (FAIL_LINE.test(l) && !PASS_LINE.test(l)) block = 25
313    // a passing line never belongs to a failure block, wherever it stands
314    if (passing && !SUMMARY_LINE.test(l) && i < lines.length - 6) {
315      dropped++
316      continue
317    }
318    if (block > 0 || SUMMARY_LINE.test(l) || SKIP_LINE.test(l) || i >= lines.length - 6) {
319      keep.push(l)
320      block = l.trim() === '' && block < 20 ? 0 : block - 1
321    } else if (PASS_LINE.test(l) || l.trim() === '') dropped++
322    else if (i < 4) keep.push(l)
323    else dropped++
324  }
325  return `${keep.join('\n')}\n[sieve: ${dropped} passing or progress lines left out]`
326}
327
328// One line per commit: hash, date, author, subject (and the stat line, if any).
329export const filterGitLog = (text: string): string | undefined => {
330  if (/^diff --git /m.test(text)) return undefined // a patch is code: never on size alone
331  const out: string[] = []
332  let cur: { h: string; a: string; d: string; s: string; st: string } | undefined
333  const flush = () => cur && out.push(`${cur.h.slice(0, 9)} ${cur.d} ${cur.a}: ${cur.s}${cur.st ? `  (${cur.st})` : ''}`)
334  for (const l of text.split('\n')) {
335    const c = /^commit ([0-9a-f]{7,40})/.exec(l)
336    if (c) { flush(); cur = { h: c[1]!, a: '', d: '', s: '', st: '' }; continue }
337    if (!cur) continue
338    const a = /^Author:\s+(.*?)\s*<.*>$/.exec(l) ?? /^Author:\s+(.*)$/.exec(l)
339    if (a) cur.a = a[1]!.trim()
340    else if (/^Date:\s+/.test(l)) cur.d = l.replace(/^Date:\s+/, '').trim().split(' ').slice(1, 5).join(' ')
341    else if (!cur.s && /^\s{4}\S/.test(l)) cur.s = l.trim()
342    else if (/\d+ files? changed/.test(l)) cur.st = l.trim()
343  }
344  flush()
345  return out.length ? `${out.length} commits\n${out.join('\n')}` : undefined
346}
347
348// Installs and builds: warnings, errors and the closing summary; downloads and progress go.
349export const filterNoise = (text: string): string => {
350  const lines = text.split('\n')
351  const keep = new Set<number>()
352  lines.forEach((l, i) => {
353    if (/\b(warn(ing)?|error|ERR!|fail(ed)?|denied|not found|conflict|deprecated|vulnerab|added \d+|removed \d+|changed \d+|up to date|Successfully|Installed \d+|Resolved \d+|Finished|Compiled|built in|Done in)\b/i.test(l)) {
354      for (let k = Math.max(0, i - 1); k <= Math.min(lines.length - 1, i + 2); k++) keep.add(k)
355    }
356    if (i < 3 || i >= lines.length - 8) keep.add(i)
357  })
358  const kept = [...keep].sort((a, b) => a - b).map(i => lines[i]!)
359  return `${kept.join('\n')}\n[sieve: ${lines.length - kept.length} progress lines left out]`
360}
361
362// The filter for a command, if one is known and it actually shrinks the output.
363export const filterCommand = (command: string, text: string): string | undefined => {
364  const sig = signature(command)
365  let out: string | undefined
366  if (sig === 'git log') out = filterGitLog(text)
367  else if (TEST_CMD.test(sig) && (sig !== 'make' && sig !== 'npm run' && sig !== 'pnpm run' && sig !== 'yarn run' || /\b(test|spec|check)\b/.test(command))) out = filterTests(text)
368  else if (/^(npm|pnpm|yarn|bun) (install|i|add|ci|update)$|^(pip|pip3|uv|poetry) (install|add|sync|lock|update)$|^python -m pip$|^(make|tsc|webpack|vite|cargo build|go build|gradle|mvn|docker build)$/.test(sig) || /^(npm|pnpm|yarn) run$/.test(sig) && /\bbuild\b/.test(command)) out = filterNoise(text)
369  // An unknown command whose output reads like a test run (many passing lines) gets the test filter.
370  if (!out) {
371    const lines = text.split('\n').filter(l => l.trim())
372    if (lines.length >= 30 && lines.filter(l => PASS_LINE.test(l)).length / lines.length >= 0.5) out = filterTests(text)
373  }
374  return out && out.length < text.length * 0.7 ? out : undefined
375}
376
377// ---- project key ----------------------------------------------------------------------------
378
379// "/a/b.c" and "/a/b-c" give the same slug; the folder name plus a hash of the whole path do not.
380export const projectKey = (cwd: string): string => {
381  let h = 0x811c9dc5
382  for (let i = 0; i < cwd.length; i++) h = Math.imul(h ^ cwd.charCodeAt(i), 0x01000193) >>> 0
383  const base = (cwd.split('/').filter(Boolean).pop() ?? 'root').replace(/[^\w.-]/g, '_').slice(0, 40)
384  return `${base}-${h.toString(16).padStart(8, '0')}`
385}
386
387// ---- JSON: schema, counts and outliers instead of lines --------------------------------------
388
389const typeOf = (v: unknown): string => (v === null ? 'null' : Array.isArray(v) ? 'array' : typeof v)
390
391const show = (v: unknown, n = 160): string => clip(JSON.stringify(v), n)
392
393// An array of objects as a table: fields with type and presence, value counts for fields with few
394// values, ranges for numbers, a few rows from the start and the rows that stand out.
395const describeRows = (rows: Record<string, unknown>[], label: string): string => {
396  const fields = new Map<string, { types: Set<string>; n: number; values: Map<string, number>; min: number; max: number }>()
397  for (const r of rows) {
398    for (const [k, v] of Object.entries(r)) {
399      let f = fields.get(k)
400      if (!f) fields.set(k, (f = { types: new Set(), n: 0, values: new Map(), min: Infinity, max: -Infinity }))
401      f.types.add(typeOf(v))
402      f.n += 1
403      if (typeof v === 'number') {
404        f.min = Math.min(f.min, v)
405        f.max = Math.max(f.max, v)
406      }
407      if (typeof v === 'string' || typeof v === 'boolean' || v === null) {
408        const key = String(v).slice(0, 60)
409        if (f.values.size <= 30 || f.values.has(key)) f.values.set(key, (f.values.get(key) ?? 0) + 1)
410      }
411    }
412  }
413  const lines = [`${label}: ${rows.length} objects, ${fields.size} fields`]
414  // A field's rare values mark the rows worth showing (status "failed" among thousands of "ok").
415  const rare: [string, string, number][] = []
416  for (const [k, f] of [...fields.entries()].slice(0, 40)) {
417    const presence = f.n < rows.length ? `, in ${f.n}` : ''
418    const range = f.min <= f.max ? `, ${f.min}..${f.max}` : ''
419    lines.push(`  ${k}: ${[...f.types].join('|')}${presence}${range}`)
420    if (f.values.size >= 2 && f.values.size <= 12) {
421      const counts = top(f.values, 12)
422      lines.push(`    ${counts.map(([v, n]) => `${v} ${n}`).join(', ')}`)
423      for (const [v, n] of counts) if (n <= Math.max(1, rows.length * 0.05)) rare.push([k, v, n])
424    }
425  }
426  lines.push('first rows:', ...rows.slice(0, 3).map(r => `  ${show(r)}`))
427  // Rarest values first, every one of them shown at least once: 3 "failed" must not lose to 40 "pending".
428  const odd = new Set<Record<string, unknown>>()
429  for (const [k, v] of rare.sort((a, b) => a[2] - b[2])) {
430    const hits = rows.filter((r, i) => i >= 3 && String(r[k]).slice(0, 60) === v)
431    for (const r of hits.slice(0, hits.length <= 5 ? 5 : 2)) if (odd.size < 12) odd.add(r)
432  }
433  if (odd.size) lines.push('rows with rare values:', ...[...odd].map(r => `  ${show(r)}`))
434  return lines.join('\n')
435}
436
437const isRowArray = (v: unknown): v is Record<string, unknown>[] =>
438  Array.isArray(v) && v.length > 0 && v.every(x => typeOf(x) === 'object')
439
440// A summary of a JSON document, or undefined when the text is not JSON (or is small and flat).
441export const summarizeJson = (text: string): string | undefined => {
442  const t = text.trim()
443  if (!/^[[{]/.test(t)) return undefined
444  let doc: unknown
445  try {
446    doc = JSON.parse(t)
447  } catch {
448    return undefined
449  }
450  if (isRowArray(doc)) return describeRows(doc, 'JSON array')
451  if (Array.isArray(doc)) return `JSON array: ${doc.length} items of ${[...new Set(doc.map(typeOf))].join('|')}\nfirst: ${show(doc.slice(0, 5), 400)}\nlast: ${show(doc.slice(-3), 300)}`
452  if (typeOf(doc) !== 'object') return undefined
453  // An object: its keys, and the largest array of objects in it as a table (a list response's "items").
454  const obj = doc as Record<string, unknown>
455  const keys = Object.entries(obj).map(([k, v]) => `  ${k}: ${Array.isArray(v) ? `array(${v.length})` : typeOf(v) === 'object' ? `object(${Object.keys(v as object).length} keys)` : show(v, 80)}`)
456  const arrays = Object.entries(obj).filter(([, v]) => isRowArray(v)).sort((a, b) => (b[1] as unknown[]).length - (a[1] as unknown[]).length)
457  const head = `JSON object, ${keys.length} keys:\n${keys.slice(0, 40).join('\n')}`
458  return arrays.length ? `${head}\n${describeRows(arrays[0]![1] as Record<string, unknown>[], `.${arrays[0]![0]}`)}` : head
459}
460