A calibrated context sieve for Claude Code. Every large tool result is judged by a System One model before it enters context; irrelevant blocks are replaced…

A calibrated context sieve for Claude Code.
Every large Read, Bash, or Grep result is judged before it enters Claude's context. Blocks the judge is confident you don't need are replaced with a three-line stub: what was hidden, a one-paragraph summary from a cheap model, and a key that restores the full text on demand. Nothing is lost; it just stops costing tokens until you ask for it.
Terms used below. The judge is the model that answers one yes/no question per block ("is this block needed for the current task?") with a probability. By default that is Jev, TypeSafe AI's System One model: a model that returns calibrated probabilities for typed questions instead of generating text, so a hundred questions come back in one call in a few hundred milliseconds. Jev is in early access. The adapter is TypeSafe's system-one-adapter package, which answers the same questions by prompting Claude Haiku 4.5; it is not calibrated, but it lets the whole pipeline run today.
Read big.py ──► Claude Code ──► tool.call hook ──► winnow
│
split into ~25-line blocks ◄─────────────────────────┘
one call to the judge: "is block N needed for the current task?" ×N, in parallel
keep confident-yes and uncertain blocks verbatim
hide confident-no blocks: cache full text ─► summarize ─► stub
│
Claude sees ◄── { result } ◄───────────────────────────────┘
A stub looks like this:
[winnow] Lines 41-188 (148 lines) hidden: judged unlikely to matter for the current task (relevance <= 0.22).
[winnow] Summary: Argparse setup for the --export and --format flags, plus the license header.
[winnow] Full text cached as key a1b2c3d4e5f6. Call winnow_recall(key="a1b2c3d4e5f6", start=41, end=188) if you need it.
Without a summarizer the middle line is a deterministic digest instead, so Claude still knows what kind of thing it lost: for a search result, which files the hidden matches came from and how many each; for anything else, the line count, whether it was mostly comments, imports or repetition, and the first line.
Two safety rules are built in. If the judge thinks the output shows an error, nothing is hidden. If a block's probability is merely uncertain (between WINNOW_DROP and WINNOW_KEEP), it is kept. Both thresholds are tunable; the rules themselves are not optional. The default WINNOW_DROP of 0.1 is the bin that came back clean on hand-labeled replay (see Measured); raise it only with your own evidence.
The hook is a function-hook module (see How it hooks in) that talks to a small resident server (winnow serve) on loopback, started at session start, so a judged call costs about 16 ms plus the judge call rather than a Python startup. winnow never judges its own files or its own commands, so recalls and labeling sheets always come back whole.
A second hook runs at prompt time. It ranks the memory files Claude Code keeps for the project (~/.claude/projects/<project>/memory/*.md, everything except the MEMORY.md index, which Claude already loads) plus any directories in WINNOW_CONTEXT_DIRS against your prompt, and injects the relevant ones so Claude reads what it needs without a round of Read calls.
What winnow changes is only what Claude sees. Files on disk, the commands that ran, and Claude Code's own transcript are untouched.
Requirements: Python 3.10+, uv, Claude Code 2.1.260 or newer with function hooks enabled (early access; one line in settings, below).
First turn on function hooks in ~/.claude/settings.json (winnow does nothing without this, and winnow doctor checks it):
{ "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" } }
Then install:
git clone https://github.com/GhalebDweikat/winnow.git
claude plugin marketplace add ./winnow
claude plugin install winnow@winnow
Then add a key (next section), open a new Claude Code session, and read any file longer than about 1,500 characters. If a [winnow] line appears in the result, it's working. If not, see Troubleshooting.
Before you have any key, you can still see what it does:
uv run --project winnow/sidecar winnow demo --fake
That runs a synthetic 130-line file through the real pipeline with a keyword judge and prints what Claude would have seen. Once a key is set, drop --fake and the same command makes the first real judge call.
The repo is its own plugin marketplace, so it also installs straight from GitHub without cloning:
claude plugin marketplace add GhalebDweikat/winnow
claude plugin install winnow@winnow
GitHub shorthand clones over SSH by default; set CLAUDE_CODE_PLUGIN_PREFER_HTTPS=1 if you don't have an SSH key on this machine.
Installing applies everywhere that shares your ~/.claude config: the CLI, the desktop app, and IDE extensions. New sessions pick the plugin up; running sessions don't. The first session after install syncs the sidecar's environment and starts the resident server, which takes a few seconds once; after that, sessions share the running server and start instantly.
Installed plugins are copied to ~/.claude/plugins/cache/, not linked, so after pulling changes run claude plugin update winnow@winnow. For a hot-reload loop while developing, load the checkout for one session instead:
claude --plugin-dir ./winnow
To scope the plugin to one project rather than your whole account, add --scope project to the marketplace add command.
winnow needs one key for the judge and, optionally, one for summaries. Nothing runs until at least the judge key is in place; until then every hook passes results through untouched and, once per session, tells you so.
1. Get a Jev key. Jev is in early access. Join the waitlist at typesafe.ai, and once you're admitted create a key at console.typesafe.ai/settings/keys. No key yet? Skip to step 3.
2. Put the keys where hooks can see them. A hook runs with the environment of whatever launched Claude Code. A key exported in one terminal is invisible to the desktop app and to IDE sessions. Either of these works everywhere:
~/.winnow/env (on Windows, %USERPROFILE%\.winnow\env), one KEY=VALUE per line. winnow reads it on every hook call. Keep it private; it is outside the repo. TYPESAFE_API_KEY=ts-...
ANTHROPIC_API_KEY=sk-ant-...
env block of ~/.claude/settings.json, which Claude Code applies to every session and every subprocess it starts: { "env": { "TYPESAFE_API_KEY": "ts-...", "ANTHROPIC_API_KEY": "sk-ant-..." } }
A variable already in the environment wins over the file, so a plain shell export still works for CLI use.
3. No Jev key yet? Use the adapter. Add WINNOW_JUDGE=adapter to the same file. The adapter sends the identical request to Claude Haiku 4.5 through your Anthropic credentials (ANTHROPIC_API_KEY, or a profile from the ant CLI's ant auth login). Its probabilities are not calibrated, but the whole pipeline works, and switching to Jev later is one line. This path bills your Anthropic account; see cost.
4. Verify.
uv run --project winnow/sidecar winnow doctor
It prints which keys were found, where they came from, and whether each backend initializes. Then winnow demo (without --fake) makes one real judge call and shows the result.
Summaries use the Anthropic credentials. Set WINNOW_SUMMARY=0 to turn them off; stubs then say "Summary unavailable" and everything else still works.
winnow commandsFrom the clone, every command is uv run --project winnow/sidecar winnow <command>. To have winnow on your PATH anywhere:
uv tool install ./winnow/sidecar
winnow doctor
Commands: doctor, demo [--fake], stats, recall <key> [--start N --end M], replay {extract,judge,score,run,sample,label,import-labels,agreement}, serve [--ensure|--status|--stop], bench [--http], clean, mcp, hook <event>.
winnow's job is reading everything Claude reads, so be clear about where it goes.
| Data | Sent to | When |
|---|---|---|
| The tool output being judged, in blocks, plus a short task description from the transcript (last user request, last assistant sentence) and the tool's arguments | TypeSafe (judge typesafe) or Anthropic (judge adapter) | every judged result over WINNOW_MIN_CHARS |
| The hidden blocks only | Anthropic | when summaries are on and something was hidden |
| Your prompt and the first 600 characters of each candidate memory file | the judge | every prompt, when candidate files exist |
Nothing is sent when the judge is off, and nothing is sent for outputs under the size threshold. The full text of every hidden output is kept locally in ~/.winnow/cache/ for recall; there is no eviction yet, so clear it when you like.
Approximate cost per judged result, for a 10,000-token output:
| Judge | Judge call | Summaries (up to 4, Haiku 4.5) | Total |
|---|---|---|---|
| Jev at $0.042 per million input tokens | $0.0004 | about $0.005 | under a cent |
| Adapter on Haiku 4.5 at $1 per million input tokens | about $0.01 | about $0.005 | a few cents |
winnow stats reports the judge's actual token usage and cost after the fact.
All settings are environment variables (or lines in ~/.winnow/env). Defaults are deliberately conservative.
| Variable | Default | Meaning |
|---|---|---|
WINNOW_MODE | active | active rewrites tool results; shadow judges and logs only |
WINNOW_JUDGE | typesafe | typesafe, adapter, or off |
WINNOW_MODEL | jev-latest | Jev model id |
WINNOW_JUDGE_TIMEOUT | 15 | Seconds per judge call, including one retry |
WINNOW_ADAPTER_PROVIDER | anthropic | Provider behind the adapter (anthropic or openai) |
WINNOW_ADAPTER_MODEL | claude-haiku-4-5 | Model behind the adapter |
WINNOW_TOOLS | Read,Bash,Grep | Tools whose output is judged. This can only narrow the set; the module wraps the tools listed in TOOLS in hooks/winnow.ts, so to add a tool edit that list too |
WINNOW_QUESTIONS | structured | Question set the judge is asked with: structured, default, or strict |
WINNOW_EXCLUDE_PATHS | WINNOW_HOME | Reads under these directories are never judged (path-separator delimited) |
WINNOW_EXCLUDE_COMMANDS | \bwinnow\b | Bash commands matching this regex are never judged |
WINNOW_PORT | 47311 | Sidecar port; PORT in hooks/winnow.ts must match |
WINNOW_MIN_CHARS | 1500 | Outputs shorter than this are never touched |
WINNOW_DROP | 0.1 | Hide a block only when P(needed) is below this |
WINNOW_KEEP | 0.5 | Error-gate threshold; also the line between "confident keep" and "uncertain keep" |
WINNOW_MIN_PRUNE_RATIO | 0.2 | Skip the rewrite unless at least this fraction of the text would be hidden |
WINNOW_BLOCK_LINES | 25 | Target lines per block |
WINNOW_MAX_BLOCKS | 200 | Cap on questions per call; block size grows to fit |
WINNOW_MAX_STATE_CHARS | 120000 | Blocks beyond this budget are kept unjudged |
WINNOW_SUMMARY | 1 | Summarize hidden groups |
WINNOW_SUMMARY_MODEL | claude-haiku-4-5 | Summarizer model |
WINNOW_SUMMARY_MAX_GROUPS | 4 | Summarize at most this many hidden groups per result; the rest get "Summary unavailable" |
WINNOW_SUMMARY_MAX_CHARS | 20000 | Characters of a hidden group sent to the summarizer |
WINNOW_CONTEXT_DIRS | Extra directories of .md files for the prompt-time selector. Path-separator delimited: : on macOS and Linux, ; on Windows | |
WINNOW_CONTEXT_GATE | 0.5 | Minimum P(relevant) to inject a file |
WINNOW_CONTEXT_TOP_K | 3 | Max files injected per prompt |
WINNOW_CONTEXT_MAX_CHARS | 8000 | Total characters injected per prompt |
WINNOW_CONTEXT_MAX_CANDIDATES | 60 | Max files considered per prompt |
WINNOW_HOME | ~/.winnow | Cache, decision log, env file |
WINNOW_MODE=shadow
In shadow mode every hook does its full job, judging, caching and logging, but never changes what Claude sees and never calls the summarizer. Use it for the first week with any judge. winnow stats then reports how many results it would have rewritten and how many characters it would have saved, and each decision line in ~/.winnow/decisions.jsonl carries the per-block probabilities and a cache key, so winnow recall <key> shows you exactly what would have been hidden. Switch to WINNOW_MODE=active when the decisions look right.
Your Claude Code transcripts already hold hundreds of large tool results, each followed by what Claude did next. winnow replay turns that into a labeled benchmark and scores a judge against it, offline.
winnow replay run --judge lexical # every transcript under ~/.claude/projects
winnow replay run --judge lexical --limit 200 path/to/session.jsonl
The label is weak but free: a block counts as needed if, later in the same turn, Claude reused one of its lines in an edit, write or command, or mentioned a distinctive identifier that appears in few other blocks. Blocks that Claude read, understood, and never quoted get labeled not needed, so treat the reported regret as an upper bound.
The report gives, for each WINNOW_DROP threshold, how much would be hidden and what share of needed blocks that would cost (regret), plus a calibration table and expected calibration error. The lexical judge is a keyless word-overlap baseline; any real judge has to beat it. When you have a key:
winnow replay judge --judge adapter # re-judge the same cases
winnow replay judge --judge typesafe
winnow replay score --judged ~/.winnow/replay/judged-typesafe.jsonl
Cases, judged files and scores live in ~/.winnow/replay/. Nothing leaves the machine unless you pick a judge that calls an API.
Weak labels are good enough to rank judges and not good enough to trust a regret number. To get real labels, draw a blind sample and label it:
winnow replay sample --judged ~/.winnow/replay/judged-typesafe.jsonl # 100 blocks, stratified by judge probability
winnow replay label --sample ~/.winnow/replay/sample.jsonl --labeler you # interactive: y needed, x not needed, u unsure
winnow replay agreement --sample ~/.winnow/replay/sample.jsonl # weak vs you, labeler vs labeler
winnow replay score --judged ~/.winnow/replay/judged-typesafe.jsonl --labels ~/.winnow/replay/labels.jsonl --labeler you
The sample also comes as a Markdown sheet (sample.md) if you would rather read it in an editor and import answers from a text file with winnow replay import-labels. A second labeler on a subset (--limit 20 --seed 2) gives an agreement number.
The words the judge is asked with matter. WINNOW_QUESTIONS selects a set, and winnow replay judge --questions <name> scores one against the same cases as any other:
| set | what it is | result on hand labels |
|---|---|---|
structured (default) | criteria as what / not_for / examples objects | clean below 0.1, hides ~5% of text there |
default | one-line criteria | clean below 0.1, hides ~2% |
strict | "directly about the task" | overconfident: 23% of its bottom bin was needed |
First results on 300 real cases, 97 blind hand labels, three question sets and the sidecar's latency are in docs/DESIGN.md, with the raw score files under docs/results/ and a draft write-up in docs/WRITEUP.md. Reports print calibration (ECE) and ordering (ROC AUC) side by side, because a judge that answers the base rate for every block scores a fine ECE and can hide nothing; an experiment with jevlike, an open Jev-shaped model, is what made that necessary (see docs/DESIGN.md).
winnow is a Claude Code function-hook plugin: a TypeScript module, hooks/winnow.ts, that the engine loads in-process. Its tool.call handler wraps every Read, Bash and Grep call, hands the result to the sidecar (which reads the task from the session's transcript), and returns the sidecar's rewrite as the tool's result; its prompt.submit handler appends the selected context files to the prompt. Function hooks are early access, behind a flag; winnow is tested on Claude Code 2.1.277 and later. On 2.1.277 a subagent's tool calls don't reach the module, so they pass through untouched; on 2.1.286 they are judged like any other. Without the flag the module never loads and winnow does nothing; winnow doctor says so.
When a result is rewritten you see a toast: winnow: hid 3 of 8 blocks of Read (5.1k to 1.8k chars; winnow_recall ab12). Small results never leave the process; the judging, thresholds and cache are all in the Python sidecar, so nothing measured below changes with the hook mechanism.
The module's tests run under Claude Code's own kit, with no key and no sidecar (the kit has no network, so they cover task reconstruction and the pass-through path):
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude plugin test .
For editor types, run /plugin-types ./.claude/types inside a Claude Code session with the flag on; the tsconfig.json at the repo root already includes that folder, hooks/ and tests/. The surface may change between releases; the plugin pins nothing and is tested against the current one in CI.
The module posts every large result to winnow serve on 127.0.0.1:47311. The SessionStart hook runs winnow serve --ensure, which starts a detached server if none is answering and replaces one left over from an older plugin version. The server keeps the SDK loaded and the judge's connection warm, re-reads ~/.winnow/env whenever it changes, and exits after 45 idle minutes. When a session needs it again, the module starts it (winnow serve --revive) and resends.
winnow serve --status # is it up, how many requests, is the judge built
winnow serve --stop # off until the next session start; sessions already open won't restart it
winnow serve --ensure # what SessionStart runs; prints nothing
winnow bench --http # 381 ms for a Python start per call vs 16 ms through the resident sidecar, on the machine this was built on
If the server is down and cannot be started, results pass through unjudged, winnow says so once, and winnow doctor says why. After a failed start the module waits a minute before trying again, so a dead sidecar costs nothing per call. Set WINNOW_PORT and PORT in hooks/winnow.ts together if the port is taken.
winnow bench # hook overhead with the judge off
winnow clean # drop cache entries older than 30 days, then trim to 200 MB
winnow stats
Reports outputs judged and rewritten, characters and estimated tokens saved, judge latency and cost, and the regret rate: the share of hidden outputs that Claude later asked to recall.
The better regret number comes from you. winnow review walks the stubs from your recent sessions, newest first, and shows the whole picture a person needs: what was asked (for a subagent, its delegation prompt), which subagent ran it, what Claude kept and what it lost, what Claude did in its next few actions, and an automatic check for whether any of those actions reused a line from the hidden text. Then one question per stub: was hiding that fine? Ten of these while the session is still fresh in your head are worth more than a hundred labels on old transcripts, and winnow stats reports the result as human regret.
winnow review --limit 10 # y fine, x should have been kept, u unsure
Recall from the shell:
winnow recall a1b2c3d4e5f6 --start 41 --end 188
Working and silently disabled look the same from inside a session, so check in this order.
claude plugin list should show winnow@winnow as enabled. Enable with claude plugin enable winnow@winnow and start a new session.winnow serve --status. If not, winnow serve --ensure starts it; the SessionStart hook does the same, and an open session starts it again when it next needs it. When it can't be started, results pass through, winnow shows one message saying so, and claude --debug logs why.winnow doctor. The common failure is a missing key, or a key set in a terminal that the desktop app never sees. When the judge can't start, winnow also posts one message per session saying so.tail -1 ~/.winnow/decisions.jsonl after reading a large file. A line with "rewritten": true and a key means a stub went to Claude. "reason": "nothing_to_prune" means the judge thought every block mattered; "below_min_prune_ratio" means it would have hidden less than WINNOW_MIN_PRUNE_RATIO of the text, so the rewrite was skipped (the usual outcome on ordinary source files at a conservative WINNOW_DROP). No line at all means the hook didn't run: winnow doctor checks the function-hooks flag; then ~/.winnow/errors.log; then claude --debug, which logs hooks module winnow@winnow loaded when the module is in.judge_error. Read ~/.winnow/errors.log; it has the traceback. Timeouts show up as TypeSafeAPITimeoutError; raise WINNOW_JUDGE_TIMEOUT or lower WINNOW_MAX_STATE_CHARS.winnow doctor shows whether they were found.winnow_recall, or read the range it names with offset/limit. To keep a directory out of winnow's reach entirely, add it to WINNOW_EXCLUDE_PATHS.claude plugin disable winnow@winnow
winnow serve --stop
New sessions won't load winnow. Sessions already open keep the module until they end, but after --stop they leave the sidecar off and pass results through untouched. claude plugin enable winnow@winnow turns it back on from the next session.
claude plugin uninstall winnow@winnow
claude plugin marketplace remove winnow
Then delete ~/.winnow (cache, decision log, and your env file) if you don't want it kept.
The SessionStart hook that starts the sidecar runs under Git Bash when it is installed, otherwise PowerShell; the command works in both. Paths from Claude Code arrive with backslashes, which winnow handles. The env file lives at %USERPROFILE%\.winnow\env. WINNOW_CONTEXT_DIRS uses ; between directories. uv installs with winget install astral-sh.uv or from astral.sh.
winnow/
├── .claude-plugin/plugin.json plugin manifest
├── .claude-plugin/marketplace.json hooks/winnow.ts 248 lines1/**
2 * winnow's hooks: a Claude Code function-hook module ("Claude Mods", early access).
3 *
4 * Claude Code loads this file when it runs with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1;
5 * without the flag winnow does nothing, which `winnow doctor` reports. The module
6 * wraps every Read, Bash and Grep call in-process: the result goes to the resident
7 * sidecar (`winnow serve`) and the sidecar's rewrite comes back as the tool's
8 * result. At prompt time it asks the sidecar which context files the prompt needs
9 * and appends them to the prompt's context.
10 *
11 * The sidecar reads the task a call serves from the tail of the session's
12 * transcript, a subagent's own by its agent id, so a large result never waits on
13 * the engine copying the whole conversation into the module. When the sidecar has
14 * stopped answering (it exits after 45 idle minutes), the module starts it again
15 * and resends, instead of passing everything through until the next session start.
16 */
17import type { EngineInterface, Register } from 'claude-code'
18
19/** The sidecar's port. Keep it equal to WINNOW_PORT. */
20export const PORT = 47311
21export const TOOLS = ['Read', 'Bash', 'Grep'] as const
22/** Results shorter than this are never rewritten (the sidecar's WINNOW_MIN_CHARS default); skip the round trip. */
23const MIN_CHARS = 1500
24/** After the sidecar could not be reached or started, results pass through without a POST for this long. */
25export const RETRY_MS = 60_000
26/** The same, when it was stopped on purpose with `winnow serve --stop`. */
27export const STOPPED_RETRY_MS = 10 * 60_000
28/** `winnow serve --revive` exit codes; keep them equal to serve.py's REVIVE_* values. */
29const REVIVE_STARTED = 0
30const REVIVE_STOPPED = 3
31const REVIVE_RUNNING = 4
32
33/** What the sidecar puts in its X-Winnow response header when it rewrote a result. */
34export type Meta = { hidden?: number; blocks?: number; before?: number; after?: number; key?: string }
35
36type HookOutput = {
37 hookSpecificOutput?: { updatedToolOutput?: unknown; additionalContext?: string }
38 systemMessage?: string // the sidecar's once-per-session notice when its judge cannot start
39}
40
41type ToolCallEvent = { tool: string; tool_use_id?: string; agentId?: string } & Record<string, unknown>
42
43export function describeMeta(tool: string, meta: Meta): string {
44 const k = (n: number | undefined) =>
45 n === undefined ? '?' : n >= 10_000 ? `${Math.round(n / 1000)}k` : n >= 1000 ? `${(n / 1000).toFixed(1)}k` : String(n)
46 const key = meta.key === undefined ? '' : `; winnow_recall ${meta.key}`
47 return `winnow: hid ${meta.hidden ?? '?'} of ${meta.blocks ?? '?'} blocks of ${tool} (${k(meta.before)} to ${k(meta.after)} chars${key})`
48}
49
50type Answer = { output?: HookOutput; meta?: Meta }
51
52/** What one POST came to: nobody answered, or the sidecar's answer (an empty one means pass through). */
53export type Sent = { reached: false } | ({ reached: true } & Answer)
54
55/**
56 * What `winnow serve --revive` came to: a sidecar started, one that was answering
57 * all along (so the POST failed for some other reason), one stopped on purpose
58 * with `winnow serve --stop`, or nothing that could run.
59 */
60export type Revival = 'started' | 'running' | 'stopped' | 'failed'
61
62/**
63 * What the module knows about reaching the sidecar, apart from the engine: whether
64 * to try at all right now, the restart already under way, and whether a failed
65 * restart has been reported. A stopped sidecar must cost nothing per call: a
66 * refused connection is not free (on Windows the stack retries the SYN first).
67 */
68export class Link {
69 private quietUntil = 0
70 private starting: Promise<Revival> | undefined
71 private warned = false
72 private readonly now: () => number
73
74 constructor(now: () => number = () => Date.now()) {
75 this.now = now
76 }
77
78 /** True while results pass through without trying the sidecar. */
79 quiet(): boolean {
80 return this.now() < this.quietUntil
81 }
82
83 /** The restart a call that failed alongside this one already started, if any. */
84 pending(): Promise<Revival> | undefined {
85 return this.starting
86 }
87
88 /** Makes `start` the restart every call that fails while it runs waits on. */
89 track(start: Promise<Revival>): Promise<Revival> {
90 const shared: Promise<Revival> = start
91 .catch((): Revival => 'failed')
92 .finally(() => {
93 if (this.starting === shared) this.starting = undefined
94 })
95 this.starting = shared
96 return shared
97 }
98
99 /** After a restart that did not bring the sidecar back: pass results through untried for a while. */
100 settle(revival: Revival): void {
101 if (revival === 'running') return // it answers; that POST failed for some other reason
102 this.quietUntil = this.now() + (revival === 'stopped' ? STOPPED_RETRY_MS : RETRY_MS)
103 }
104
105 /** True the first time a restart fails, so the person hears about it once, not on every call. */
106 firstFailure(revival: Revival): boolean {
107 if (revival !== 'failed' || this.warned) return false
108 this.warned = true
109 return true
110 }
111}
112
113async function send($: EngineInterface, path: string, payload: unknown): Promise<Sent> {
114 const url = `http://127.0.0.1:${PORT}${path}`
115 let res
116 try {
117 res = await $.http.fetch(url, {
118 method: 'POST',
119 headers: { 'content-type': 'application/json' },
120 body: JSON.stringify(payload),
121 })
122 } catch (err) {
123 $.ui.log(`winnow: sidecar not answering at ${url}: ${String(err)}`, { to: 'debug' })
124 return { reached: false }
125 }
126 if (!res.ok) {
127 $.ui.log(`winnow: sidecar answered ${res.status} at ${url}`, { to: 'debug' })
128 return { reached: true }
129 }
130 let meta: Meta | undefined
131 const raw = res.headers['x-winnow']
132 if (raw !== undefined) {
133 try {
134 meta = JSON.parse(raw) as Meta
135 } catch {
136 meta = undefined
137 }
138 }
139 if (res.text.trim() === '') return { reached: true, meta } // empty 2xx: pass-through
140 try {
141 return { reached: true, output: JSON.parse(res.text) as HookOutput, meta }
142 } catch {
143 $.ui.log('winnow: sidecar sent something that is not JSON', { to: 'debug' })
144 return { reached: true }
145 }
146}
147
148/**
149 * Runs `winnow serve --revive` with the sidecar's own interpreter: the venv the
150 * SessionStart hook's `uv run` made beside the plugin, or `uv run` itself when
151 * there is none yet. No shell, so nothing depends on what a shell adds to PATH.
152 */
153async function startSidecar($: EngineInterface): Promise<Revival> {
154 const project = `${$.plugin.root}/sidecar`
155 const revive = ['-m', 'winnow', 'serve', '--revive']
156 let argv = ['uv', 'run', '-q', '--project', project, 'python', ...revive]
157 for (const python of [`${project}/.venv/Scripts/python.exe`, `${project}/.venv/bin/python`]) {
158 if (await $.fs.stat(python).then(() => true, () => false)) {
159 argv = [python, ...revive]
160 break
161 }
162 }
163 try {
164 const { exitCode } = await $.process.run(argv, { timeoutMs: 30_000 })
165 if (exitCode === REVIVE_STARTED) return 'started'
166 if (exitCode === REVIVE_RUNNING) return 'running'
167 if (exitCode === REVIVE_STOPPED) return 'stopped'
168 $.ui.log(`winnow: serve --revive exited ${exitCode}`, { to: 'debug' })
169 return 'failed'
170 } catch (err) {
171 $.ui.log(`winnow: could not run ${argv[0]}: ${String(err)}`, { to: 'debug' })
172 return 'failed'
173 }
174}
175
176async function whereami($: EngineInterface): Promise<{ session_id: string; cwd: string }> {
177 const [session_id, cwd] = await Promise.all([$.session.id().catch(() => ''), $.session.cwd().catch(() => '')])
178 return { session_id, cwd }
179}
180
181/**
182 * One request to the sidecar. A POST that reaches nobody starts the sidecar again,
183 * once for every call that failed alongside it, and is sent once more; when that
184 * does not bring it back, results pass through untried for a while.
185 */
186async function ask($: EngineInterface, link: Link, path: string, payload: unknown): Promise<Answer | undefined> {
187 if (link.quiet()) return undefined
188 const first = await send($, path, payload)
189 if (first.reached) return first
190 const revival = await (link.pending() ?? link.track(startSidecar($)))
191 if (link.firstFailure(revival)) {
192 $.ui.toast('winnow: the sidecar is not running and could not be started, so results pass through untouched. Run `winnow doctor` to see why.')
193 }
194 if (revival === 'started') {
195 const again = await send($, path, payload)
196 if (again.reached) return again
197 }
198 link.settle(revival)
199 return undefined
200}
201
202export const register: Register = (on) => {
203 const link = new Link()
204
205 for (const tool of TOOLS) {
206 on('tool.call', { tool }, async ($, e, next) => {
207 const answer = await next(e)
208 if (!('result' in answer) || answer.result === undefined) return answer
209 if (JSON.stringify(answer.result).length < MIN_CHARS) return answer
210 if (link.quiet()) return answer
211
212 const { tool: toolName, tool_use_id, agentId, ...input } = e as ToolCallEvent
213 // No task here: the sidecar reads it from the transcript's tail. `$.session.messages()` would
214 // copy the whole conversation on every call (seconds, in a long session), and with no
215 // argument it is the main conversation even inside a subagent.
216 const res = await ask($, link, '/hook/post-tool-use', {
217 hook_event_name: 'PostToolUse',
218 source: 'function-hook',
219 ...(await whereami($)),
220 agent_id: agentId,
221 tool_name: toolName,
222 tool_input: input,
223 tool_response: answer.result,
224 tool_use_id,
225 })
226 if (typeof res?.output?.systemMessage === 'string') $.ui.toast(res.output.systemMessage)
227 const updated = res?.output?.hookSpecificOutput?.updatedToolOutput
228 if (updated === undefined) return answer
229 if (res?.meta !== undefined) $.ui.toast(describeMeta(toolName, res.meta))
230 return { ...answer, result: updated }
231 })
232 }
233
234 on('prompt.submit', async ($, e, next) => {
235 if (e.text.trim().length < 12 || link.quiet()) return next(e)
236 const res = await ask($, link, '/hook/user-prompt-submit', {
237 hook_event_name: 'UserPromptSubmit',
238 source: 'function-hook',
239 ...(await whereami($)),
240 prompt: e.text,
241 })
242 if (typeof res?.output?.systemMessage === 'string') $.ui.toast(res.output.systemMessage)
243 const extra = res?.output?.hookSpecificOutput?.additionalContext
244 if (typeof extra !== 'string' || extra === '') return next(e)
245 return next({ ...e, context: [...(e.context ?? []), extra] })
246 })
247}
248