Harness engineering for Claude Code: attribute every prompt/agent change, keep every candidate, diagnose from traces, hold out a split - and search over Claude…

Filesystem-backed end-to-end optimization of executable LLM harnesses, following Meta-Harness: End-to-End Optimization of Model Harnesses (paper.pdf).
A harness is the code around a fixed model: what it stores, retrieves, and shows the model at each step. This repository searches over that code. A coding-agent proposer reads the full experience filesystem — every prior candidate's source, scores, execution traces, and the reasoning that produced it — and writes new candidates. Candidates are evaluated on a search split; the Pareto frontier over (accuracy, context tokens) is scored once on a held-out split the proposer never sees.
The repository is also a Claude Code plugin. Not a wrapper around the CLI — model-invoked skills that change how Claude works on prompts, agent scaffolds and retrieval, plus a hook layer that observes failures and enforces what has been learned from them:
claude --plugin-dir .
optimizing-harnesses carries the discipline (one change per candidate, every candidate kept, no score without a trace, held-out split read once). building-eval-sets, reading-execution-traces and running-harness-search handle the pieces, optimizing-claude-code points the search at Claude Code's own scaffold, and learning-from-failures drives the loop below. The engine is the escalation path when hand-tuning stalls; the skills need nothing installed.
See docs/plugin.md for the skill list, the baseline testing behind it, and why Haiku 4.5 is the default harness model.
python -m meta_harness learn # failure -> artifact -> replay-scored -> staged
python -m meta_harness learn --tasks tasks.json # also require no regression on a task set
python -m meta_harness learn --status
python -m meta_harness learn --accept <id>
Three layers. Observe: a tool.call hook records tool errors and repeated calls to a per-session log. Learn: the CLI above ranks those failures together with ones mined from past transcripts, proposes one artifact at the strongest layer that can carry it, and replays it. Enforce: a tool.check hook applies accepted rules and a prompt.submit hook attaches accepted injections to the user's turn as context. The plugin listens only on tool.call, tool.check and prompt.submit; nothing rewrites the system prompt.
An artifact is staged only if it fixes the failure it was born from; one whose origin replay still fails is archived with the verdict that killed it. With --tasks, it must also score no worse than the baseline on that task set, with context cost as the tiebreak; without --tasks, that check does not run and the recorded verdict says so. Nothing installs itself.
Artifacts sit at four layers, strongest first: rule (a tool.check, costing no standing tokens and not ignorable) > injection (context attached to a matching turn, paid when it fires) > skill > doctrine. Only rule and injection are enforced by the hooks today. That ordering is this project's own design position — it appears nowhere in paper.pdf. tools/prose_vs_rule.py is the experiment that would test it against a measured baseline, and it has not been run here.
What the paper does establish (Table 3, online text classification, median/best score): a proposer given scores only reaches 34.6/41.3, scores plus an LLM summary reaches 34.9/38.7, and full raw execution traces reach 50.0/56.7. Raw trace access is the paper's key ingredient — a summary "may even hurt by compressing away diagnostically useful details." Appendix A.2 adds a second, independently evidenced principle: on TerminalBench-2 the proposer regressed six consecutive iterations while editing prompts and control flow, diagnosed the shared prompt edit as the confound, and then won with a purely additive change. Additive beats invasive.
One measured negative result of our own: the paper's winning TerminalBench-2 discovery was an environment snapshot injected before the first model call. That does not transfer to a developer's own repository, where the environment is already known — 50 orientation calls across 344 sessions on this machine. We do not build it.
A receipt is one line the hooks write each time a learned rule, an injection, a session rule or the drift note acts on a call or a prompt: which artifact, what it did (acted, or held), and the call number in the session. It carries no message text and no tool input. Receipts go to <harness_home>/receipts-<session>.jsonl, next to the observations, which now record the same call number so the two can be matched.
To find out whether an artifact does anything, the harness needs something to compare against. So a learned rule, an injection or the drift note is held out on a random 10% of the occasions it would have acted: the hook records a held receipt and does nothing. The cost is real: a held-out learned rule does not deny that call, so about one time in ten the failure it exists to prevent is allowed through. Session rules and rejection memory are never held out: they are something you said, and they always apply.
Set the rate in <harness_home>/receipts.json:
{"holdout_rate": 0.1}
The value is clamped to 0 through 0.5; a missing or unreadable file means 0.1. Set it to 0 to switch the holdout off entirely (receipts are still written, but nothing is held, so learned artifacts stay "not enough data").
python -m meta_harness receipts report # add --json for the raw rows
python -m meta_harness receipts export --out receipts.json
The report compares how often the failure signature recurred in later calls of the same session with the artifact acting versus held out. Verdicts:
learn --status); it never retires one itself.Evidence accumulates slowly. At a 10% holdout an artifact needs many occasions before either arm has 5 receipts, which takes weeks of ordinary use. Until then, expect "not enough data".
python -m meta_harness waste (add --json, --this-project, --since, --limit) reads your transcripts and reports stretches of tool calls that ended with you correcting course, the calls they burned, and identical failures retried. It also writes <harness_home>/waste.json. Its correction count is a keyword heuristic: in a hand-labelled audit, 11 of the 21 flagged stretches audited were genuine corrections, and recall is unmeasured. The report carries that caveat itself; treat the numbers as an estimate, not an audit.
Beyond enforcing artifacts, the hooks add (details in hooks/README.md):
meta-harness waste run; nothing mines your history in the background.drift.json enables it, and "enabled" then means a plain call-count threshold: the config's judge field is validated but never run./meta-harness:harness waste | pending | why - the report, what is staged or pending, and why the last denial happened.Mined failures are weighted by recency and by Claude Code version (version_weight). On the current data the version weight is inert: every session shares Claude Code minor 2.1.
The hooks need Claude Code's function-hook surface, which is early access:
export CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 # PowerShell: $env:CLAUDE_CODE_ENABLE_FUNCTION_HOOKS="1"
claude --plugin-dir .
See hooks/README.md for the hook layer and the design for the rest.
python -m venv venv
venv/Scripts/activate # Windows; source venv/bin/activate elsewhere
pip install -e .
pip install -e ".[dev]" # adds pytest
No runtime dependencies. Python 3.10+.
All model traffic goes through Unikey, an OpenAI-compatible gateway (https://www.getunikey.ai/v1, Authorization: Bearer $UNIKEY_API_KEY).
export UNIKEY_API_KEY=sk-... # PowerShell: $env:UNIKEY_API_KEY="sk-..."
meta-harness models --provider unikey # GET /v1/models
The proposer is a coding agent, billed separately. Unikey also exposes an Anthropic-compatible /v1/messages, so Claude Code can be routed through it:
export ANTHROPIC_BASE_URL=https://www.getunikey.ai
export ANTHROPIC_AUTH_TOKEN=$UNIKEY_API_KEY
--provider openai|anthropic|compatible still work, reading OPENAI_API_KEY / ANTHROPIC_API_KEY and their *_BASE_URL overrides.
meta-harness run \
--tasks data/classification.jsonl \
--task-type classification \
--provider unikey --model gpt-5.2 \
--iterations 20 --candidates 2 --repeats 3 --max-workers 4 \
--proposer-model claude-sonnet-4-6 \
--root .meta-harness
Dataset formats: .jsonl, .json, .csv.
--task-type | required fields | optional | metric |
|---|---|---|---|
classification | input, label | labels | exact match |
math | problem | answer | \boxed{} + numeric equivalence |
terminal | instruction | test_command, image, workdir, timeout | test command exit code |
Without --test-tasks, --tasks is split 70/30 (--search-fraction, --split-seed). Without --baseline, the seeds in baselines/ for that task type are used. Without --proposer-command, the proposer is the claude CLI; use --proposer-command to drive any other agent, or python tools/llm_proposer.py for a plain-API proposer that needs no agent installed. See docs/using-with-coding-agents.md.
Two sample datasets ship in data/.
.meta-harness/
run.json config + the search-split tasks
frontier.json Pareto frontier over (score, context_cost)
candidates/<id>/harness.py candidate source
candidates/<id>/scores.json score, context_cost, repeats, score_std, valid, error
candidates/<id>/traces.jsonl every model call, task start/end, harness event
candidates/<id>/traces-N.jsonl additional repeats
candidates/<id>/proposer_reasoning.md
proposals/iteration-NNNN/ proposer stdout/stderr
views/iteration-NNNN/ what the proposer was allowed to read (ablation modes)
.meta-harness-test/
test_results.json held-out scores, written once, outside the proposer's reach
.meta-harness-cache/ model-call cache, keyed by (model, prompt, kwargs)
Inspect a finished run: meta-harness inspect .meta-harness
meta-harness run ... --proposer-view scores # source + scores only
meta-harness run ... --proposer-view summary # + LLM summaries, no raw traces
meta-harness run ... --proposer-view full # everything (default)
Terminal tasks run the model's commands inside Docker:
meta-harness run --task-type terminal --tasks data/terminal.jsonl \
--provider unikey --model claude-sonnet-4-6
Each row needs an image. Rows without one are refused unless you pass --allow-local-shell, which executes model-authored commands on your machine — only do that in a throwaway environment.
class Harness:
def run(self, task, model, trace):
prompt = f"Classify: {task['input']}"
trace.event("prompt", {"prompt": prompt})
return model(prompt)
The instance persists across the tasks of one evaluation (that is your memory) and is recreated per evaluation. model(prompt) -> str; model.call(prompt) -> (str, usage). Every call is priced in input tokens and written to the trace. meta_harness.retrieval provides TfidfIndex, BM25Index, and reciprocal_rank_fusion.
Seven working examples live in baselines/; the exact contract handed to the proposer is meta_harness/skill/SKILL.md.
meta_harness/core.py — the outer loop, experience filesystem, Pareto frontier, metering.meta_harness/sandbox.py — interface validation in a subprocess under a timeout, so an LLM-written while True cannot hang the search.meta_harness/agent_proposer.py — the coding-agent proposer and the view ablation.meta_harness/cache.py — disk cache so --repeats and re-runs do not re-bill.tools/llm_proposer.py — proposer that needs only an API key, no coding-agent CLI.meta_harness/replay.py — failure signatures, and the replay a proposed artifact is scored on.meta_harness/learn.py — ranking, proposal, and the retention decision.meta_harness/harness_store.py — staged and installed artifacts, and the accept/reject gate.hooks/ — the function-hook layer: observe, enforce, inject, rejection memory, session rules, first-run line, drift note (disabled). Fails open by construction.meta_harness/waste.py — the waste report behind meta-harness waste.tools/prose_vs_rule.py — the prose-versus-mechanism experiment (built, not yet run).docs/plugin.md — the Claude Code plugin: skills, agent, install, Haiku defaults.docs/using-with-coding-agents.md — Claude Code as proposer, optimizing agent harnesses, and shipping a discovered harness back into your own agent.docs/superpowers/specs/ and docs/superpowers/plans/ — the spec and task-by-task plan this implementation follows.python -m pytest -q
No network calls. The hook layer is covered by Node tests under hooks/ that the Python suite shells out to; they are skipped, with the reason stated, when node is absent.
hooks/harness.ts 844 lines1/**
2 * Meta-Harness function hooks: observe failures, enforce learned artifacts.
3 *
4 * Every handler fails open: most run inside `guardBefore`, `guardAfter` or `guardAfterMap`, all of which swallow a
5 * throw and fall through to `next` rather than break the turn — and none ever calls `next` a
6 * second time once it has been called. `registerBootstrap` does its fail-open handling by hand,
7 * so its post-`next` marker write is swallowed locally instead of reaching any wrapper at all. A learning system that can break a session, or submit the
8 * human's prompt twice, is worse than no learning system.
9 *
10 * Per-turn features (rejection memory, session rules, injections, the first-run message) live on
11 * `prompt.submit`, never `prompt.section`. `prompt.section` fires once per NAMED SECTION of the
12 * system prompt, its sections are cached for the whole session, and returning `{ text }` REPLACES
13 * that section — there is no `event.prompt`/`event.text` carrying the user's words on it, and a
14 * handler with no section-name filter fires for and can overwrite every section that exists,
15 * including ones it has never heard of. `prompt.submit` carries the user's actual turn as
16 * `e.text`, and a hook adds anything the model should see via `next({ ...e, context: [...] })`
17 * without touching the system prompt at all.
18 */
19
20import type { On, PluginOptions, Register } from 'claude-code';
21import {
22 addSessionRule,
23 callKey,
24 evaluateRule,
25 loadInstalled,
26 loadRules,
27 parseStopInstruction,
28 rememberRejection,
29 stableStringify,
30 toolArgs,
31 type Rule,
32 type SessionState,
33 wasRejected,
34} from './rules.js';
35import { type DriftConfig, driftNote, loadDriftConfig, shouldWarn, userSpoke } from './drift.js';
36import { bumpCall, holdoutRate, shouldHold, writeReceipt } from './receipts.js';
37
38export type Fallible<E, R> = (io: any, event: E, next: (e: E) => Promise<R>) => Promise<R>;
39
40/** Log a skip the same way everywhere, without ever risking a second throw of its own. */
41function logSkip(io: any, name: string, error: unknown): void {
42 try {
43 io.ui.log(`meta-harness ${name} skipped (${error instanceof Error ? error.message : String(error)})`);
44 } catch {
45 // logging must never be the thing that breaks the turn either
46 }
47}
48
49/**
50 * Why the guards below take `(name, io, event, next, handler)` and are CALLED from inside a
51 * function literal instead of returning a hook: Claude Code's hooks loader rejects any `on(...)`
52 * whose hook argument is not a function literal (or the name of one) written in the call — a
53 * wrapper call such as `on('tool.call', afterCall(...))` fails the whole module, so no hook runs.
54 * Every registration is therefore `on('<event>', async ($, e, next) => guardX(..., $, e, next, ...))`;
55 * tests/test_hook_literals.py enforces that statically.
56 */
57
58/**
59 * Runs a handler that may call `next` itself, so a throw becomes a pass-through instead of a
60 * broken turn. `next` is the rest of the chain (for prompt.submit, the human's prompt reaching
61 * core; for tool.check, the tool running). Recovering a throw by calling it again is only safe if
62 * the handler never reached it: once it has, a second call submits the prompt twice or runs the
63 * tool twice. So once `next` has been called, its own outcome stands on every path.
64 */
65export async function guardBefore<E, R>(
66 name: string,
67 io: any,
68 event: E,
69 next: (e: E) => Promise<R>,
70 handler: (next: (e: E) => Promise<R>) => Promise<R>,
71): Promise<R> {
72 let called = false;
73 let settled: { ok: true; value: R } | { ok: false; error: unknown } | null = null;
74 const once = async (e: E): Promise<R> => {
75 called = true;
76 try {
77 const value = await next(e);
78 settled = { ok: true, value };
79 return value;
80 } catch (error) {
81 settled = { ok: false, error };
82 throw error;
83 }
84 };
85 try {
86 return await handler(once);
87 } catch (error) {
88 logSkip(io, name, error);
89 if (!called) return next(event);
90 const outcome = settled as { ok: true; value: R } | { ok: false; error: unknown } | null;
91 if (outcome?.ok) return outcome.value;
92 throw outcome ? outcome.error : error;
93 }
94}
95
96/**
97 * Runs a handler whose work happens AFTER the tool has already run. `guardBefore` recovers a
98 * throw that happened before `next` by calling `next(event)`: for a post-`next` handler that
99 * would run the tool a SECOND time, duplicating the side effect of a Bash or Write. Here `next` is
100 * called exactly once, up front (a rejection propagates once, untouched), and the fallible work is
101 * what gets swallowed — so an observer that throws costs an observation, never a repeated command.
102 */
103export async function guardAfter<E, R>(
104 name: string,
105 io: any,
106 event: E,
107 next: (e: E) => Promise<R>,
108 handler: (outcome: R) => Promise<void>,
109): Promise<R> {
110 const outcome = await next(event);
111 try {
112 await handler(outcome);
113 } catch (error) {
114 logSkip(io, name, error);
115 }
116 return outcome;
117}
118
119/**
120 * Like `guardAfter`, for a handler that may return a REPLACEMENT outcome (a copy of `next`'s with
121 * something appended). `next` is called exactly once, up front; if the handler throws, core's
122 * outcome is returned as it came, never a second call to `next`.
123 */
124export async function guardAfterMap<E, R>(
125 name: string,
126 io: any,
127 event: E,
128 next: (e: E) => Promise<R>,
129 handler: (outcome: R) => Promise<R>,
130): Promise<R> {
131 const outcome = await next(event);
132 try {
133 return await handler(outcome);
134 } catch (error) {
135 logSkip(io, name, error);
136 return outcome;
137 }
138}
139
140/** `guardBefore` as a hook-returning wrapper, for tests. Never pass its result to `on` directly. */
141export function safely<E, R>(name: string, handler: Fallible<E, R>): Fallible<E, R> {
142 return (io, event, next) => guardBefore(name, io, event, next, (once) => handler(io, event, once));
143}
144
145/** `guardAfter` as a hook-returning wrapper, for tests. Never pass its result to `on` directly. */
146export function afterCall<E, R>(
147 name: string,
148 handler: (io: any, event: E, outcome: R) => Promise<void>,
149): Fallible<E, R> {
150 return (io, event, next) => guardAfter(name, io, event, next, (outcome) => handler(io, event, outcome));
151}
152
153/** `guardAfterMap` as a hook-returning wrapper, for tests. Never pass its result to `on` directly. */
154export function afterCallMap<E, R>(
155 name: string,
156 handler: (io: any, event: E, outcome: R) => Promise<R>,
157): Fallible<E, R> {
158 return (io, event, next) => guardAfterMap(name, io, event, next, (outcome) => handler(io, event, outcome));
159}
160
161const REPEAT_WINDOW = 6;
162const REPEAT_THRESHOLD = 4;
163
164type Recent = { tool: string; key: string };
165
166/** The harness's home directory: overridable for tests, otherwise under the user's profile. */
167async function harnessHome(io: any): Promise<string> {
168 // `$.env.get` answers with a Promise in Claude Code; awaiting also accepts a plain value.
169 const override = await io.env?.get?.('META_HARNESS_HOME');
170 if (override) return override;
171 const profile = (await io.env?.get?.('USERPROFILE')) || (await io.env?.get?.('HOME'));
172 return `${profile}/.claude/harness`;
173}
174
175/**
176 * Mirrors meta_harness.replay.ERROR_PATTERNS: same regex text and slug, in the same order, so a
177 * `cause` recorded here dedupes against what Python's `_cause()` computes from the same text.
178 *
179 * Every backslash in these patterns is doubled (`\\d`, `\\[`, `\\(`, ...). These are plain JS
180 * string literals fed to `new RegExp(pattern, 'i')` below, and a JS string literal silently
181 * drops a backslash in front of any character it does not recognize as an escape -- `\d`, `\S`,
182 * `\[`, `\]`, `\(` and `\)` are NOT recognized string escapes (only `\n`, `\r`, `\t`, `\'`, `\\`,
183 * etc. are), so a single backslash here compiles to a regex with a missing metacharacter escape
184 * (`\d+` -> a literal-`d` character class, `\(...\)` -> an unescaped group, and so on). A prior
185 * version of this file used single backslashes throughout, which type-checked and passed a
186 * test that only compared source text, while classifying 235 of 1,167 real failure texts
187 * differently from the Python side. Regex literals (`/.../i`) would not have this problem, but
188 * would need a parallel array-of-RegExp construction; doubling the backslash here keeps this
189 * array's shape (string, string) identical to ERROR_PATTERNS's (str, str), which is what the
190 * parity test below relies on to generate its own fixtures programmatically in the future.
191 *
192 * `tests/test_hook_assets.py::test_ts_and_python_cause_agree_on_fixtures` actually EXECUTES both
193 * classifiers over `tests/cause_fixtures.py::CAUSE_FIXTURES` (via node) and asserts identical
194 * slugs -- not a source-text comparison, which cannot detect this class of bug.
195 *
196 * UnicodeEncodeError must stay ordered before the decode/charmap/cp1252 pattern: an encode
197 * failure's message can also contain the word "decode" in surrounding text, so checking decode
198 * first would misclassify it. That ordering was fixed on the Python side in Task 2.
199 */
200const CAUSE_PATTERNS: ReadonlyArray<readonly [string, string]> = [
201 ['UnicodeEncodeError', 'unicode-encode'],
202 ['UnicodeDecodeError|charmap|cp1252', 'unicode-decode'],
203 ['No such file or directory|cannot find the (file|path)', 'missing-path'],
204 ['Permission denied|EACCES', 'permission'],
205 ['command not found|is not recognized as', 'missing-command'],
206 ['timed out|TimeoutExpired', 'timeout'],
207 ['has not been read yet|must read.*before', 'unread-edit'],
208 ['String to replace not found|old_string', 'edit-mismatch'],
209 ['SyntaxError|unterminated', 'syntax'],
210 ['contains multiple operations|Compound command changes working directory', 'compound-shell'],
211 ['requires approval|denied by the Claude Code auto mode classifier', 'needs-approval'],
212 ['doesn\'t want to proceed|tool use was rejected', 'user-rejected'],
213 ['Blocked:|blocked by a deny rule', 'blocked-policy'],
214 ['unexpected EOF while looking|simple_expansion|expansion obfuscation', 'shell-quoting'],
215 ['not in Claude\'s tab group|determine which page this action targets', 'tab-target'],
216 ['modified since read', 'stale-read'],
217 ['EISDIR|illegal operation on a directory', 'is-directory'],
218 ['Traceback \\(most recent call last\\)', 'python-traceback'],
219 ['Permission to (use|read).*has been denied', 'permission-denied-tool'],
220 ['File does not exist', 'missing-path'],
221 ['is temporarily unavailable', 'model-unavailable'],
222 ['InputValidationError|Workflow script file not found|No task found with ID|Invalid workflow script|scriptPath must be a script path|Unknown skill:|Task ID is required', 'workflow-error'],
223 ['Failed to execute JavaScript|JavaScript execution error', 'js-error'],
224 ['Error capturing screenshot|actions\\[\\d+\\][^\\n]*failed|Failed to find element|Failed to execute action|Error capturing zoomed screenshot|is not a supported form input|Can\'t interact with browser-internal', 'browser-action-failed'],
225 ['No such tool available', 'unknown-tool'],
226 ['hook did not respond before|tool did not respond in time', 'hook-timeout'],
227 ['Found \\d+ matches of the string', 'edit-mismatch'],
228 ['Python was not found|pdftoppm is not installed', 'missing-command'],
229 ['node:internal/modules/(package_json_reader|run_main)|Cannot find module|ERR_MODULE_NOT_FOUND', 'module-not-found'],
230 ['"error":\\{"name":"(HttpException|McpError)"|already exists in local config', 'api-error'],
231 ['fatal: (pathspec|detected dubious ownership|ambiguous argument|.*is outside repository)|ignored by one of your \\.gitignore|docker: Error response from daemon', 'git-error'],
232 ['On branch \\S+\\r?\\nYour branch is (up to date|ahead of)|warning: in the working copy of', 'git-noise'],
233 ['npm error code|npm warn exec', 'npm-error'],
234 ['exceeds maximum allowed tokens', 'output-too-large'],
235 ['ConnectionRefusedError|connection refused|ECONNREFUSED', 'connection-refused'],
236 ['was blocked\\. For security|is blocked\\. This path is protected|denied by your permission', 'blocked-policy'],
237 ['=+ FAILURES =+|ERROR at setup of|\\bAssertionError\\b|FAILED \\S+::', 'test-failure'],
238 ['tab group no longer exists|Missing required parameter tabId', 'tab-target'],
239 ['needs design-system authorization', 'needs-approval'],
240 ['ENAMETOOLONG', 'path-too-long'],
241 ['error TS\\d+|imported but unused', 'ts-error'],
242 ['not logged into any GitHub hosts', 'gh-auth-error'],
243];
244
245/** Classifies failure text into the same slug meta_harness.replay._cause() would produce. */
246export function cause(text: string): string {
247 const value = text ?? '';
248 for (const [pattern, slug] of CAUSE_PATTERNS) {
249 if (new RegExp(pattern, 'i').test(value)) return slug;
250 }
251 return 'other';
252}
253
254/** This session's id, or 'unknown' when the engine cannot supply one; never throws. */
255async function currentSessionId(io: any): Promise<string> {
256 try {
257 const id = await io.session?.id?.();
258 return id ? String(id) : 'unknown';
259 } catch {
260 return 'unknown';
261 }
262}
263
264/**
265 * Appends one JSON line to this session's own observation file.
266 *
267 * One file per session, not one shared file: `$.fs` exposes no append or lock primitive, so a
268 * shared file would need a non-atomic exists/read/write cycle that two concurrent sessions can
269 * race and clobber each other on. Per-session files remove the race instead of trying to guard
270 * it; a reader globs `observed-*.jsonl` under the harness home. `$.fs.write` creates missing
271 * parent directories itself, so a fresh harness home on a new machine needs no separate mkdir.
272 */
273async function observe(io: any, sessionId: string, record: Record<string, unknown>): Promise<void> {
274 const path = `${(await harnessHome(io))}/observed-${sessionId}.jsonl`;
275 const line = `${JSON.stringify({ ts: new Date().toISOString(), ...record })}\n`;
276 const existing = (await io.fs.exists(path)) ? await io.fs.read(path) : '';
277 await io.fs.write(path, existing + line);
278}
279
280function resultText(result: unknown): string {
281 return typeof result === 'string' ? result : JSON.stringify(result ?? '');
282}
283
284/**
285 * The text of a real `tool.call` outcome (ToolCallResult) worth checking for a rejection
286 * announcement — and ONLY when the outcome actually says the call was refused or errored.
287 *
288 * The `deny` branch's `deny` string always counts: a hook or core only sets it on an actual
289 * refusal. The answered branch's `text`/`result` count ONLY when `isError` is true — `text` and
290 * `result` are present on every SUCCESSFUL call too (a normal Read's file contents are `result`,
291 * and its `text` is the same content joined for the model), so checking them unconditionally
292 * means a Read of any file that happens to CONTAIN the phrase "tool use was rejected" — this very
293 * file, for one — gets recorded as a rejection and denied for the rest of the session. Gating on
294 * `isError`/`deny` is what keeps a successful result's mere text out of consideration entirely.
295 */
296function rejectionAnnouncement(outcome: any): string {
297 if (typeof outcome?.deny === 'string' && outcome.deny.length > 0) return outcome.deny;
298 if (outcome?.isError !== true) return '';
299 const parts = [outcome?.text, resultText(outcome?.result)];
300 return parts.filter((part): part is string => typeof part === 'string' && part.length > 0).join(' ');
301}
302
303/**
304 * Unambiguous tool-error phrases: strings a tool emits when it refuses, which do not plausibly
305 * appear as the FIRST thing in a successful result. Anything weaker (a bare `error:` anywhere in
306 * the text) matched a successful `Grep` for the word "error:" and recorded it as a failure;
307 * those counts feed selection ranking, so the noise became the thing the learn loop chased.
308 */
309const ERROR_PHRASES =
310 /has not been read yet|String to replace not found|is not recognized as an internal or external command|No such file or directory/i;
311
312/** True when the tool call actually failed. */
313function isError(result: unknown): boolean {
314 if (result && typeof result === 'object') {
315 const flag = (result as any).is_error ?? (result as any).isError;
316 // The engine's own verdict is authoritative; never second-guess it by grepping the text.
317 if (typeof flag === 'boolean') return flag;
318 }
319 const text = resultText(result);
320 // Otherwise only an error ANNOUNCED at the start of the result counts, plus a short list of
321 // phrases a tool only ever emits when it refused.
322 return /^\s*"?(error|[A-Za-z.]*Error:|Traceback \(most recent call last\))/i.test(text)
323 || ERROR_PHRASES.test(text);
324}
325
326/**
327 * Launchers stripped from the front of a Bash command before a session rule's pattern is
328 * compared: `python -m`, `python3 -m`, `py -m`, `npx`, `uv run`, `poetry run`, `pipx run`, plus a
329 * leading `env` and any `VAR=value` assignments. Takes lower-cased whitespace tokens.
330 */
331export function stripLaunchers(tokens: string[]): string[] {
332 let rest = tokens;
333 for (;;) {
334 if (rest[0] === 'env' || /^[a-z_][a-z0-9_]*=/.test(rest[0] ?? '')) {
335 rest = rest.slice(1);
336 } else if (['python', 'python3', 'py'].includes(rest[0] ?? '') && rest[1] === '-m') {
337 rest = rest.slice(2);
338 } else if (rest[0] === 'npx') {
339 rest = rest.slice(1);
340 } else if (['uv', 'poetry', 'pipx'].includes(rest[0] ?? '') && rest[1] === 'run') {
341 rest = rest.slice(2);
342 } else {
343 return rest;
344 }
345 }
346}
347
348/** The deny reason for a session rule, saying what the matching actually does. */
349export function sessionRuleReason(tool: string, rule: { pattern: string; requires?: string; flag?: string }): string {
350 const head = 'Session rule from this conversation: this blocks any';
351 if (tool !== 'Bash') {
352 return `${head} ${tool} call whose input contains "${rule.pattern}" [session-rule]`;
353 }
354 const what = `${head} Bash command whose first words are "${rule.pattern}" `
355 + '(after any launcher such as python -m, npx, uv run, poetry run, pipx run, env or VAR=x)';
356 if (rule.requires) return `${what} unless "${rule.requires}" is one of its words [session-rule]`;
357 if (rule.flag) return `${what} and "${rule.flag}" is one of its words [session-rule]`;
358 return `${what}, whatever its arguments [session-rule]`;
359}
360
361/** Observes tool.call outcomes: records errors and repeated identical calls to observed-<session>.jsonl. */
362export function registerObserver(add: On): void {
363 const recent: Recent[] = [];
364
365 // guardAfter, not guardBefore: this handler's work runs after the tool has already executed, so a
366 // recovery that re-entered next() would run the tool twice.
367 add('tool.call', async (io: any, event: any, next: any) => guardAfter('tool.call', io, event, next, async (outcome: any) => {
368 // The one place the session's call counter advances: once per tool.call, error or not.
369 const sessionId = await currentSessionId(io);
370 const call = bumpCall(sessionId);
371 const tool = String(event?.tool ?? 'unknown');
372 const args = toolArgs(event);
373 const key = `${tool}:${stableStringify(args).slice(0, 200)}`;
374
375 recent.push({ tool, key });
376 if (recent.length > REPEAT_WINDOW) recent.shift();
377 const repeats = recent.filter((entry) => entry.key === key).length;
378
379 if (repeats >= REPEAT_THRESHOLD) {
380 await observe(io, sessionId, { kind: 'repeat', call, tool, cause: 'repeat', input: args });
381 }
382 const text = resultText((outcome as any)?.result);
383 if (isError((outcome as any)?.result)) {
384 await observe(io, sessionId, {
385 kind: 'tool_error',
386 call,
387 tool,
388 cause: cause(text),
389 input: args,
390 text: text.slice(0, 400),
391 });
392 }
393 }));
394}
395
396/** Unambiguous phrases the engine emits when the human declines a tool call outright, as opposed
397 * to the tool itself failing. Recording on these two phrases only (never a bare "no" in the
398 * conversation) keeps rejection memory from firing on an ordinary declined suggestion in prose. */
399const REJECTION_PATTERN = /doesn't want to proceed|tool use was rejected/i;
400
401/**
402 * Enforces installed `rule` artifacts at tool.check: the layer that costs no standing tokens and
403 * cannot be talked around, because it runs before the tool call, not as prose in a prompt.
404 *
405 * Also carries this session's rejection memory (task 13) and session-scoped "stop doing X" rules
406 * (task 14), both stored on the same `SessionState` so `tool.check` can consult them ahead of the
407 * installed-rule loop. Rejection detection lives in its own `tool.call` handler below; there is
408 * no read-tracking handler here, since read-before-edit is the engine's job, not this plugin's
409 * (see hooks/rules.ts). Claude Code refuses a second registration of one event without a matcher, so these
410 * hooks are added to `register`'s per-event list and chained inside its single `on` per event.
411 * `prompt.submit` (not `prompt.section`) is where the human's own words
412 * are read, since only `prompt.submit`'s event carries `text`.
413 */
414export function registerRules(add: On): void {
415 const state: SessionState = {
416 callCounts: new Map<string, number>(),
417 rejected: new Map<string, string>(),
418 sessionRules: [],
419 };
420 let rules: Rule[] | null = null;
421
422 // guardAfter, not guardBefore: rejection detection reads the outcome, which only exists AFTER the
423 // tool has already run (or been denied), so a `guardBefore` recovery re-entering next() would run
424 // the tool a second time. `next` resolves before this handler's own body runs at all.
425 //
426 // There is no read-tracking hook here any more: a live test proved Claude Code's own engine
427 // already denies an Edit/Write of a file not read this session ("File has not been read yet"),
428 // and it runs before this plugin ever sees the call. The plugin-side copy of that check (a
429 // `read-before-edit` rule plus this handler's read-tracking) is retired -- see hooks/rules.ts.
430 add('tool.call', async (io: any, event: any, next: any) => guardAfter('tool.call:rejection-memory', io, event, next, async (outcome: any) => {
431 const text = rejectionAnnouncement(outcome);
432 if (text && REJECTION_PATTERN.test(text)) {
433 rememberRejection(state, String(event?.tool ?? ''), toolArgs(event));
434 }
435 }));
436
437 // prompt.submit, not prompt.section: only prompt.submit's event carries the human's actual
438 // words (`e.text`). This handler mutates session state only (never the model-visible prompt),
439 // so it always passes `event` through to `next` unchanged.
440 add('prompt.submit', async (io: any, event: any, next: any) => guardBefore('prompt.submit:nlrules', io, event, next, async (next) => {
441 const text = String(event?.text ?? '');
442 // Only a human's own prompt may create a denying rule: a peer, plugin, scheduled trigger or
443 // notification saying "don't run pytest" is not the user asking. Same origin rule the drift
444 // note uses to decide who spoke (`userSpoke`, hooks/drift.ts).
445 const instruction = userSpoke(event) ? parseStopInstruction(text) : null;
446 if (instruction) {
447 addSessionRule(state, instruction);
448 const path = `${(await harnessHome(io))}/pending-session-rules.json`;
449 await io.fs.write(path, JSON.stringify(state.sessionRules, null, 2));
450 }
451 if (text) {
452 const lower = text.toLowerCase();
453 for (const [key, snippet] of [...(state.rejected ?? new Map<string, string>()).entries()]) {
454 if (snippet && lower.includes(snippet.toLowerCase())) {
455 state.rejected!.delete(key);
456 }
457 }
458 }
459 return next(event);
460 }));
461
462 add('tool.check', async (io: any, event: any, next: any) => guardBefore('tool.check', io, event, next, async (next) => {
463 const tool = String(event?.tool ?? '');
464 const home = await harnessHome(io);
465 const sessionId = await currentSessionId(io);
466 if (wasRejected(state, tool, event?.input)) {
467 // The user already said no: never held out, but still receipted for reporting.
468 await writeReceipt(io, home, sessionId, {
469 event: 'tool.check', source: 'rejection-memory', artifact: null, signature: null, tool, decision: 'acted',
470 });
471 return {
472 decision: 'deny',
473 reason: 'This exact call was already rejected earlier this session [rejection-memory]',
474 };
475 }
476 for (const rule of state.sessionRules ?? []) {
477 if (rule.tool !== tool) continue;
478 // Bash rules match the command string, tokenised on whitespace — never a substring of the
479 // whole JSON input, or a rule for "git push" would also fire on `git status` (which merely
480 // contains the token "git") or on an unrelated call whose description happens to mention
481 // the pattern. Non-Bash rules (Edit-like verbs) have no `command` field to tokenise, so
482 // they fall back to the same JSON-substring check as before.
483 const command = tool === 'Bash' ? String((event?.input as any)?.command ?? '') : '';
484 const commandTokens = command.split(/\s+/).filter(Boolean).map((t) => t.toLowerCase());
485 const patternTokens = rule.pattern.toLowerCase().split(/\s+/).filter(Boolean);
486
487 // The pattern is matched against the command's first words AFTER any launcher, so a rule
488 // about `pytest` also covers `python -m pytest`, `uv run pytest`, `FOO=1 pytest`, ...
489 const programTokens = stripLaunchers(commandTokens);
490 let matches: boolean;
491 if (tool === 'Bash') {
492 matches = patternTokens.length > 0
493 && patternTokens.every((t, i) => programTokens[i] === t);
494 } else {
495 const haystack = JSON.stringify(event?.input ?? {}).toLowerCase();
496 matches = haystack.includes(rule.pattern.toLowerCase());
497 }
498 if (!matches) continue;
499
500 // A "without <flag>" qualifier IS representable: the rule allows the call when that flag
501 // is present, so "stop running pytest without -q" denies `pytest tests/` but ALLOWS
502 // `pytest -q` — the exact command the human asked to keep, not the command they asked to
503 // stop. Matched as a whole token of the command string, not a substring of the JSON input
504 // (a substring match would let "--quick" satisfy a "-q" requirement it does not).
505 if (rule.requires) {
506 const required = String(rule.requires).toLowerCase();
507 const present = tool === 'Bash'
508 ? commandTokens.includes(required)
509 : JSON.stringify(event?.input ?? {}).toLowerCase().includes(required);
510 if (present) continue;
511 }
512
513 // The opposite polarity: a flag named directly in the object ("git push --force") means
514 // the rule denies ONLY calls that also carry that flag as a whole token — a plain
515 // "git push origin main" must stay allowed.
516 if (rule.flag) {
517 const flag = String(rule.flag).toLowerCase();
518 const present = tool === 'Bash'
519 ? commandTokens.includes(flag)
520 : JSON.stringify(event?.input ?? {}).toLowerCase().includes(flag);
521 if (!present) continue;
522 }
523
524 // The user's direct order: never held out, but still receipted for reporting.
525 await writeReceipt(io, home, sessionId, {
526 event: 'tool.check', source: 'session-rule', artifact: null, signature: null, tool, decision: 'acted',
527 });
528 return { decision: 'deny', reason: sessionRuleReason(tool, rule) };
529 }
530 if (rules === null) rules = await loadRules(io, home);
531 for (const rule of rules) {
532 const verdict = evaluateRule(rule, event, state);
533 if (!verdict.deny) continue;
534 // A learned rule may be held out: a held rule behaves exactly as if it were not installed.
535 const held = shouldHold(await holdoutRate(io, home));
536 await writeReceipt(io, home, sessionId, {
537 event: 'tool.check', source: 'learned-rule', artifact: rule.artifactId,
538 signature: rule.signature, tool, decision: held ? 'held' : 'acted',
539 });
540 if (held) continue;
541 return { decision: 'deny', reason: verdict.reason };
542 }
543 // Counted here, before the call proceeds, so a `repeat-call` rule sees how many times this
544 // exact call has already been allowed. Counting at tool.check rather than after the result
545 // keeps the observer the only post-`next` handler on this path.
546 const key = callKey(event);
547 state.callCounts.set(key, (state.callCounts.get(key) ?? 0) + 1);
548 return next(event);
549 }));
550}
551
552type Injection = { artifactId: string; triggers: string[]; text: string; signature: string };
553
554/**
555 * Per-artifact injected-text cap, matching the `reason` cap the rules layer already applies
556 * (rules.ts:~50). An `injection` is paid in standing tokens on every turn it fires, unlike a
557 * `rule`, which is exactly why it sits below `rule` in the layer ordering — an uncapped
558 * injection would out-cost the stronger, free layer above it.
559 */
560const INJECTION_TEXT_CAP = 400;
561
562/**
563 * Cap on the joined text of ALL matched injections for one turn, so N installed injections
564 * cannot add up past a bound even though each is individually capped. Set to three artifacts'
565 * worth of INJECTION_TEXT_CAP: enough for a few unrelated matches to coexist, not enough for an
566 * unbounded number of installs to dominate the prompt.
567 */
568const INJECTION_TOTAL_CAP = INJECTION_TEXT_CAP * 3;
569
570const TRUNCATION_MARKER = '… [truncated]';
571
572/** Truncates visibly rather than silently, so a capped injection cannot be mistaken for a short one. */
573function truncate(text: string, limit: number): string {
574 return text.length <= limit ? text : `${text.slice(0, limit)}${TRUNCATION_MARKER}`;
575}
576
577/** Installed `injection` artifacts only: a `rule`, `skill` or `doctrine` row is never surfaced here. */
578async function loadInjections(io: any, home: string): Promise<Injection[]> {
579 const artifacts = await loadInstalled(io, home, 'injection');
580 return artifacts.map((artifact) => ({
581 artifactId: String(artifact.id),
582 triggers: (artifact.origin as any)?.triggers ?? [],
583 text: truncate(String(artifact.payload ?? ''), INJECTION_TEXT_CAP),
584 signature: String(artifact.signature ?? ''),
585 }));
586}
587
588/**
589 * Injects installed `injection` artifacts as model-only context on `prompt.submit`: the layer
590 * below `rule`, paid in standing tokens only on the turns where it actually fires.
591 *
592 * This used to live on `prompt.section` and return `{ text }` to replace a system-prompt
593 * section — with no section-name filter, that both matched triggers against the WRONG text (the
594 * system prompt's own section content, not the human's turn) and, the moment one injection
595 * artifact was ever installed, deleted or overwrote every other section of the system prompt on
596 * every turn it didn't match. `prompt.submit` fixes both: triggers match `e.text` (the human's
597 * actual words) and a match is ATTACHED via `next({ ...e, context: [...] })`, never a
598 * replacement. The hook itself knows what it attached, so use is recorded by construction via
599 * `observe` instead of asking the model to self-report a retrieval it might forget.
600 */
601export function registerInjection(add: On): void {
602 let injections: Injection[] | null = null;
603
604 add('prompt.submit', async (io: any, event: any, next: any) => guardBefore('prompt.submit:injection', io, event, next, async (next) => {
605 const home = await harnessHome(io);
606 if (injections === null) injections = await loadInjections(io, home);
607 if (injections.length === 0) return next(event);
608
609 const haystack = String(event?.text ?? '').toLowerCase();
610 const matched = injections.filter((injection) =>
611 injection.triggers.some((trigger) => haystack.includes(String(trigger).toLowerCase())));
612 if (matched.length === 0) return next(event);
613
614 const sessionId = await currentSessionId(io);
615 // Each matched injection draws independently; a held one behaves as if it were not installed.
616 const acted: Injection[] = [];
617 for (const injection of matched) {
618 const held = shouldHold(await holdoutRate(io, home));
619 await writeReceipt(io, home, sessionId, {
620 event: 'prompt.submit', source: 'learned-injection', artifact: injection.artifactId,
621 signature: injection.signature, tool: '', decision: held ? 'held' : 'acted',
622 });
623 if (!held) acted.push(injection);
624 }
625 if (acted.length === 0) return next(event);
626 for (const injection of acted) {
627 await observe(io, sessionId, { kind: 'injected', artifactId: injection.artifactId });
628 }
629 const joined = truncate(acted.map((injection) => injection.text).join('\n\n'), INJECTION_TOTAL_CAP);
630 return next({ ...event, context: [...(event?.context ?? []), joined] });
631 }));
632}
633
634/**
635 * Says once per stretch, on the result of the tool call that crosses `min_calls`, that a stretch
636 * of tool calls has run long: the note rides in the result's `context` ("what the model reads
637 * after the tool's result and the user never sees"), so the model reads it mid-stretch, BEFORE
638 * the user would notice, not after the user has already spoken.
639 *
640 * Disabled unless `<harnessHome>/drift.json` explicitly enables it (see `hooks/drift.ts`); the
641 * ship gate refused every judge, so by default this counts calls and says nothing.
642 *
643 * `tool.call` uses `guardAfterMap`: `next` is called exactly once, up front, and the note is
644 * appended to a COPY of the answered outcome; a throw anywhere after `next` returns core's
645 * outcome untouched, never re-running the tool. A `{ deny }` outcome is never given context (its
646 * type forbids it): past the threshold the note waits for the next answered call. The latch is
647 * set once the note is decided: attached, or held out by the receipts holdout (a held note
648 * returns core's outcome unchanged and is still receipted).
649 *
650 * `prompt.submit` only resets the stretch when the user speaks (see `userSpoke`); it attaches
651 * nothing and passes the event through unchanged, under `guardBefore`, which calls `next` once.
652 * Nothing here listens on `prompt.section`, whose return replaces a system-prompt section.
653 */
654export function registerDrift(add: On): void {
655 let count = 0;
656 let warned = false;
657 let config: DriftConfig | null | undefined;
658
659 add('tool.call', async (io: any, event: any, next: any) => guardAfterMap('tool.call:drift', io, event, next, async (outcome: any) => {
660 count += 1;
661 if (warned) return outcome;
662 if (outcome === null || typeof outcome !== 'object' || 'deny' in outcome) return outcome;
663 const home = await harnessHome(io);
664 if (config === undefined) config = await loadDriftConfig(io, home);
665 if (!shouldWarn(count, config)) return outcome;
666 // The latch is set whether the note acts or is held: one decision per stretch either way.
667 warned = true;
668 const held = shouldHold(await holdoutRate(io, home));
669 await writeReceipt(io, home, await currentSessionId(io), {
670 event: 'tool.call', source: 'drift-note', artifact: null, signature: null,
671 tool: String(event?.tool ?? ''), decision: held ? 'held' : 'acted',
672 });
673 if (held) return outcome;
674 const context = Array.isArray(outcome.context) ? outcome.context : [];
675 return { ...outcome, context: [...context, driftNote(count)] };
676 }));
677
678 add('prompt.submit', async (io: any, event: any, next: any) => guardBefore('prompt.submit:drift', io, event, next, async (next) => {
679 if (userSpoke(event)) {
680 count = 0;
681 warned = false;
682 config = undefined; // re-read per stretch, so an edited drift.json takes effect
683 }
684 return next(event);
685 }));
686}
687
688/**
689 * Says once, on the first `prompt.submit` after install, what wrong-direction work has cost so
690 * far, then never again: `<harnessHome>/bootstrap.json` is the marker. Presence of that file
691 * (not any in-memory flag) is the only thing that gates the message, so it stays correct across
692 * every session after the first, not just within the process that happened to write it.
693 *
694 * The numbers come from the report `meta-harness waste` (with or without `--json`) writes to
695 * `<harnessHome>/waste.json` (`{ sessions, corrections: { calls_burned, ... }, ... }`, per
696 * `meta_harness/waste.py`'s `waste_report()`). This hook never runs that command and nothing
697 * runs it in the background: until someone has run `meta-harness waste` (or `/meta-harness:harness waste`)
698 * once, there is no report and no line. A report with `sessions: 0` is skipped, unmarked.
699 *
700 * Handles its own failures rather than relying on `safely`: this handler calls `next` in the
701 * MIDDLE, then does one more fallible thing afterward (writing the marker). Under the old
702 * `safely`, a marker write that threw, or `next(...)` itself rejecting, reached a catch that
703 * called `next` a SECOND time — submitting the human's prompt twice, the second time with no
704 * line attached (every turn, on a read-only harness home). `safely` no longer does that either,
705 * but here the rule is explicit: `next` is called exactly once on every path, a failure before
706 * it falls back to the unmodified event, and a failed marker write is swallowed locally.
707 *
708 * The marker is written ONLY after the line has actually been attached via `next(...)` and that
709 * call has resolved: a missing `waste.json` means "no report yet, try again next turn", not
710 * "never again", and a transient read/parse failure before `next` is swallowed locally (falling
711 * back to the unmodified event) rather than suppressing the line for good — either way nothing
712 * is marked done, so the question keeps being asked (at most once per turn, which costs nothing
713 * on a turn with no report to show) until it can actually be answered once.
714 */
715export function registerBootstrap(add: On): void {
716 add('prompt.submit', async (io: any, event: any, next: any) => {
717 let outboundEvent = event;
718 let markDone = false;
719
720 try {
721 const home = (await harnessHome(io));
722 const bootstrapPath = `${home}/bootstrap.json`;
723 if (!(await io.fs.exists(bootstrapPath))) {
724 const wastePath = `${home}/waste.json`;
725 if (await io.fs.exists(wastePath)) {
726 const waste: any = JSON.parse(await io.fs.read(wastePath));
727 const sessions = waste?.sessions;
728 const callsBurned = waste?.corrections?.calls_burned;
729 // A report over zero sessions (an empty or mis-pointed transcript store) says nothing
730 // worth saying once and for good: skip it and leave the marker unwritten.
731 if (typeof sessions === 'number' && sessions > 0 && callsBurned != null) {
732 // Hedged on purpose: the count comes from a keyword heuristic whose hand-labelled
733 // precision is about half, so the line is an estimate, never a statement of fact.
734 const line = `Across ${sessions} past sessions, an estimated ~${callsBurned} tool calls may have gone `
735 + 'to work you later corrected (a rough heuristic, roughly half of its flags are genuine). '
736 + '`/meta-harness:harness waste` for the breakdown and how reliable it is.';
737 outboundEvent = { ...event, context: [...(event?.context ?? []), line] };
738 markDone = true;
739 }
740 }
741 }
742 } catch (error) {
743 logSkip(io, 'prompt.submit:bootstrap', error);
744 outboundEvent = event; // fail open: send the turn through exactly as it arrived
745 markDone = false;
746 }
747
748 // Called exactly once, unconditionally, on every path above — the one and only next() call
749 // this handler ever makes.
750 const result = await next(outboundEvent);
751
752 if (markDone) {
753 try {
754 await io.fs.write(`${(await harnessHome(io))}/bootstrap.json`, JSON.stringify({ shown: new Date().toISOString() }));
755 } catch (error) {
756 // Swallowed locally, never recovered by calling next() again: a failed marker write
757 // just means the question is asked again next turn, not that the turn breaks or the
758 // prompt is submitted twice.
759 logSkip(io, 'prompt.submit:bootstrap', error);
760 }
761 }
762 return result;
763 });
764}
765
766type AnyHook = (io: any, event: any, next: (e: any) => Promise<any>) => Promise<any>;
767
768/**
769 * Runs `hooks` as one chain, first outermost, the way separate registrations would nest: each
770 * hook's `next` is the rest of the chain, and the last one's `next` is core.
771 */
772export function chain(hooks: AnyHook[], io: any, event: any, next: (e: any) => Promise<any>): Promise<any> {
773 const run = (index: number, e: any): Promise<any> =>
774 index === hooks.length ? next(e) : hooks[index](io, e, (inner: any) => run(index + 1, inner));
775 return run(0, event);
776}
777
778/**
779 * Registers each event ONCE. Claude Code's loader refuses a module that registers one event twice
780 * without a matcher ("on(\"tool.call\") is registered twice without a matcher"), so the features
781 * above add their hooks to a per-event list here and each event gets a single literal that runs
782 * that list in the order the features were added (observer, rules, injection, bootstrap, drift).
783 */
784export const register: Register = (on: On, options: PluginOptions) => {
785 void options;
786 const hooks: Record<string, AnyHook[]> = {};
787 const add = ((event: string, hook: AnyHook) => {
788 (hooks[event] ??= []).push(hook);
789 }) as unknown as On;
790 registerObserver(add);
791 registerRules(add);
792 registerInjection(add);
793 registerBootstrap(add);
794 registerDrift(add);
795 on('tool.call', async ($: any, event: any, next: any) => chain(hooks['tool.call'] ?? [], {
796 fs: {
797 read: (path: string) => $.fs.read(path),
798 write: (path: string, text: string) => $.fs.write(path, text),
799 exists: (path: string) => $.fs.exists(path),
800 },
801 // $.env.get takes a literal name (the loader lists what a module reads), so one call per name.
802 env: {
803 get: (name: string) => (name === 'META_HARNESS_HOME' ? $.env.get('META_HARNESS_HOME')
804 : name === 'USERPROFILE' ? $.env.get('USERPROFILE')
805 : name === 'HOME' ? $.env.get('HOME')
806 : Promise.resolve(undefined)),
807 },
808 session: { id: () => $.session.id() },
809 ui: { log: (text: string) => $.ui.log(text) },
810 }, event, next));
811 on('tool.check', async ($: any, event: any, next: any) => chain(hooks['tool.check'] ?? [], {
812 fs: {
813 read: (path: string) => $.fs.read(path),
814 write: (path: string, text: string) => $.fs.write(path, text),
815 exists: (path: string) => $.fs.exists(path),
816 },
817 // $.env.get takes a literal name (the loader lists what a module reads), so one call per name.
818 env: {
819 get: (name: string) => (name === 'META_HARNESS_HOME' ? $.env.get('META_HARNESS_HOME')
820 : name === 'USERPROFILE' ? $.env.get('USERPROFILE')
821 : name === 'HOME' ? $.env.get('HOME')
822 : Promise.resolve(undefined)),
823 },
824 session: { id: () => $.session.id() },
825 ui: { log: (text: string) => $.ui.log(text) },
826 }, event, next));
827 on('prompt.submit', async ($: any, event: any, next: any) => chain(hooks['prompt.submit'] ?? [], {
828 fs: {
829 read: (path: string) => $.fs.read(path),
830 write: (path: string, text: string) => $.fs.write(path, text),
831 exists: (path: string) => $.fs.exists(path),
832 },
833 // $.env.get takes a literal name (the loader lists what a module reads), so one call per name.
834 env: {
835 get: (name: string) => (name === 'META_HARNESS_HOME' ? $.env.get('META_HARNESS_HOME')
836 : name === 'USERPROFILE' ? $.env.get('USERPROFILE')
837 : name === 'HOME' ? $.env.get('HOME')
838 : Promise.resolve(undefined)),
839 },
840 session: { id: () => $.session.id() },
841 ui: { log: (text: string) => $.ui.log(text) },
842 }, event, next));
843};
844hooks/rules.ts 352 lines1/**
2 * Installed `rule` artifacts, applied at tool.check.
3 *
4 * A rule is the strongest layer available: it costs no standing tokens and cannot be talked
5 * around. There is no built-in read-before-edit rule here: a live test proved Claude Code's own
6 * engine already refuses an Edit/Write of a file that has not been read this session ("File has
7 * not been read yet"), and it runs before this plugin ever sees the call. A plugin-side copy of
8 * that check only added failure modes on top of an enforcement that already existed -- every
9 * Edit denied, every new-file Write denied, and a Write-then-Edit denied even though the Write
10 * itself supplied the read the plugin's tracker never recorded. This module now only carries
11 * rules the engine does not already enforce (installed artifacts, `repeat-call`).
12 */
13
14export type Rule = {
15 artifactId: string;
16 kind: string;
17 tools: string[];
18 reason: string;
19 threshold?: number;
20 /** The failure signature from the artifact's `installed.json` row ('' when absent), for receipts. */
21 signature: string;
22};
23
24export type SessionRule = {
25 tool: string;
26 /** The command PREFIX the instruction names, e.g. "git push" from "stop using git push
27 * --force", or "pytest" from "don't run pytest". Matched against the START of the tokenised
28 * command string (whole tokens), never a substring of the whole JSON input — otherwise a
29 * rule for "git push" would also fire on an unrelated call that merely mentions "git push"
30 * inside a description or a file path. */
31 pattern: string;
32 /** The one flag/token that, if PRESENT in the call, means the rule allows it. Captures a
33 * "without <flag>" qualifier ("stop running pytest without -q" denies pytest unless the call
34 * also contains "-q"): the qualifier is representable, so it narrows the rule instead of
35 * being dropped into a blanket denial. Absent when the human's qualifier (if any) could not
36 * be represented this way, in which case the rule denies its pattern unconditionally. */
37 requires?: string;
38 /** The one flag/token that, if PRESENT in the call, means the rule DENIES it — the opposite
39 * polarity of `requires`. Captures a flag named directly in the instruction's object ("stop
40 * using git push --force" denies calls that start with "git push" AND carry "--force" as a
41 * whole token, but allows a plain "git push origin main"). Mutually exclusive with
42 * `requires` in practice, since an instruction supplies at most one qualifier. */
43 flag?: string;
44};
45
46export type SessionState = {
47 /** How many times each exact tool+input call has already been allowed this session. */
48 callCounts: Map<string, number>;
49 /** Calls the human has explicitly rejected this session, keyed by callKey, valued by a
50 * searchable snippet of the input so a later mention of the same command can clear it. */
51 rejected?: Map<string, string>;
52 /** "Stop doing X" rules added from the human's own words, in effect for this session only. */
53 sessionRules?: SessionRule[];
54};
55
56/** Keys a `tool.call` event carries beside the tool's own arguments (ToolCallReserved, AgentLoop). */
57const RESERVED_TOOL_CALL_KEYS = new Set(['tool', 'tool_use_id', 'agentId', 'consent']);
58
59/**
60 * The tool's own arguments on a `tool.call` event. On `tool.call` Claude Code puts them at the TOP
61 * LEVEL of the event (`e.command` for Bash, `e.file_path` for Read/Edit) beside the reserved keys;
62 * only `tool.check` nests them under `input`. An object `input` is still honoured, for fakes and in
63 * case a future engine nests them. Every `tool.call` reader goes through this, never `event.input`.
64 */
65export function toolArgs(event: unknown): Record<string, unknown> {
66 if (!event || typeof event !== 'object') return {};
67 const record = event as Record<string, unknown>;
68 if (record.input && typeof record.input === 'object' && !Array.isArray(record.input)) {
69 return record.input as Record<string, unknown>;
70 }
71 const args: Record<string, unknown> = {};
72 for (const [key, value] of Object.entries(record)) {
73 if (!RESERVED_TOOL_CALL_KEYS.has(key)) args[key] = value;
74 }
75 return args;
76}
77
78/** JSON with object keys sorted, so the same arguments give the same key whatever their order. */
79export function stableStringify(value: unknown): string {
80 return JSON.stringify(value, (_key, inner) => {
81 if (inner && typeof inner === 'object' && !Array.isArray(inner)) {
82 return Object.fromEntries(Object.keys(inner).sort().map((k) => [k, (inner as any)[k]]));
83 }
84 return inner;
85 });
86}
87
88/** The identity of one tool call, for counting exact repeats. Input order is whatever the
89 * engine sends; the same call in the same session serializes the same way, which is all the
90 * repeat counter needs. */
91export function callKey(event: { tool?: string; input?: Record<string, unknown> }): string {
92 let input = '';
93 try {
94 input = stableStringify(event.input ?? {});
95 } catch {
96 input = String(event.input ?? '');
97 }
98 return `${String(event.tool ?? '')}:${input.slice(0, 400)}`;
99}
100
101/**
102 * Default repeat count at which a `repeat-call` rule denies. Mirrors
103 * meta_harness.replay.THRASH_THRESHOLD, so the rule that enforces a thrash fix and the replay
104 * expectation that verifies it are talking about the same number.
105 */
106export const DEFAULT_REPEAT_THRESHOLD = 4;
107
108/** No built-in rules ship: read-before-edit is the engine's job now (see the module doc above),
109 * and nothing else has earned a built-in yet. Kept as an array (rather than removed outright) so
110 * `loadRules` has one shape to spread regardless of how many built-ins eventually exist. */
111export const BUILT_IN_RULES: Rule[] = [];
112
113export type InstalledArtifact = {
114 id: string;
115 type: string;
116 origin?: Record<string, unknown>;
117 payload?: unknown;
118 /** Copied from the artifact's `installed.json` row; '' when that row has none. */
119 signature?: string;
120};
121
122/**
123 * Reads `installed.json`, keeps only the rows of the given `type`, and loads each one's
124 * `artifact.json`. The single reader every layer (`rule`, `injection`, and whichever of
125 * `skill`/`doctrine` follows) shares, so the fail-open semantics live in exactly one place: a
126 * missing or corrupt registry, or a missing/malformed artifact.json, yields fewer rows rather
127 * than throwing, and every caller gets that behaviour identically instead of re-deriving it.
128 */
129export async function loadInstalled(io: any, home: string, type: string): Promise<InstalledArtifact[]> {
130 const registry = `${home}/installed.json`;
131 if (!(await io.fs.exists(registry))) return [];
132 let entries: Array<{ id: string; type: string; signature?: unknown }> = [];
133 try {
134 entries = JSON.parse(await io.fs.read(registry));
135 } catch {
136 return [];
137 }
138 const out: InstalledArtifact[] = [];
139 for (const entry of entries) {
140 if (entry.type !== type) continue;
141 const path = `${home}/artifacts/${entry.id}/artifact.json`;
142 if (!(await io.fs.exists(path))) continue;
143 try {
144 const artifact = JSON.parse(await io.fs.read(path));
145 out.push({ ...artifact, signature: typeof entry.signature === 'string' ? entry.signature : '' });
146 } catch {
147 continue;
148 }
149 }
150 return out;
151}
152
153/** Installed rule artifacts, plus the built-ins. Missing or malformed files are ignored. */
154export async function loadRules(io: any, home: string): Promise<Rule[]> {
155 const rules = [...BUILT_IN_RULES];
156 const artifacts = await loadInstalled(io, home, 'rule');
157 for (const artifact of artifacts) {
158 rules.push({
159 artifactId: String(artifact.id),
160 // `origin.rule_kind` is what meta_harness.learn.propose_artifact emits for the matcher to
161 // dispatch on; `origin.kind` is the EPISODE kind ('tool_error'/'thrash'), kept for
162 // provenance and never a matcher name. Falling back to it keeps a hand-written artifact
163 // working, and an unknown kind simply never denies.
164 kind: (artifact.origin as any)?.rule_kind ?? (artifact.origin as any)?.kind ?? 'custom',
165 tools: (artifact.origin as any)?.tools ?? [],
166 reason: String(artifact.payload ?? '').slice(0, 400),
167 threshold: Number((artifact.origin as any)?.threshold) || undefined,
168 signature: String(artifact.signature ?? ''),
169 });
170 }
171 return rules;
172}
173
174/** Evaluates one rule against one tool.check event. `deny: true` means the call must not proceed. */
175export function evaluateRule(
176 rule: Rule,
177 event: { tool?: string; input?: Record<string, unknown> },
178 state: SessionState,
179): { deny: boolean; reason?: string } {
180 const tool = String(event.tool ?? '');
181 if (!rule.tools.includes(tool)) return { deny: false };
182
183 if (rule.kind === 'repeat-call') {
184 // The only condition a learned rule can decide from the call and the session state alone:
185 // this exact call has already been made enough times to be the thrash the artifact was born
186 // from. A first attempt is never blocked, so the rule cannot break work that is going fine.
187 const threshold = rule.threshold && rule.threshold > 1 ? rule.threshold : DEFAULT_REPEAT_THRESHOLD;
188 const seen = state.callCounts.get(callKey(event)) ?? 0;
189 if (seen >= threshold - 1) {
190 return {
191 deny: true,
192 reason: `${rule.reason} [${rule.artifactId}] (identical ${tool} call already made ${seen} times this session)`,
193 };
194 }
195 return { deny: false };
196 }
197
198 return { deny: false };
199}
200
201/** A short, human-searchable string standing in for a tool call's input: the command or path
202 * a person would actually type or say when talking about this call, falling back to its JSON. */
203function commandSnippet(input: unknown): string {
204 if (input && typeof input === 'object') {
205 const record = input as Record<string, unknown>;
206 if (typeof record.command === 'string') return record.command;
207 if (typeof record.file_path === 'string') return record.file_path;
208 }
209 try {
210 return JSON.stringify(input ?? '');
211 } catch {
212 return String(input ?? '');
213 }
214}
215
216/** Records that this exact call was rejected by the human this session. */
217export function rememberRejection(state: SessionState, tool: string, input: unknown): void {
218 if (!state.rejected) state.rejected = new Map<string, string>();
219 const key = callKey({ tool, input: input as Record<string, unknown> });
220 state.rejected.set(key, commandSnippet(input));
221}
222
223/** True when this exact call was rejected earlier this session and has not since been cleared. */
224export function wasRejected(state: SessionState, tool: string, input: unknown): boolean {
225 if (!state.rejected) return false;
226 const key = callKey({ tool, input: input as Record<string, unknown> });
227 return state.rejected.has(key);
228}
229
230/**
231 * A conservative leading-verb parse: text must OPEN with "stop"/"don't"/"never" AND be
232 * immediately followed by one of a fixed, small set of ACTION verbs — a gerund (running/using/
233 * calling/doing/touching/editing/deleting/pushing) OR the plain imperative (run/use/call/do/
234 * touch/edit/delete/push) — before anything counts as an instruction's object.
235 *
236 * Matching anywhere in the text, or treating the action verb as optional, both let ordinary
237 * prose through as if it were an instruction: "don't worry about it", "don't know why this
238 * fails", "never mind" and "Don't forget to update the README" all open with a trigger word but
239 * name no actionable target, and "stop" alone or "stop, that's wrong" have no object at all.
240 * Requiring one of these specific verbs, immediately after the trigger word, is what tells an
241 * instruction ("stop running X") apart from those. Accepting the bare imperative alongside the
242 * gerund is what lets "don't use git push --force", "never call the deploy script" and "don't
243 * run pytest" parse at all — an earlier version only accepted the -ing form and silently missed
244 * every plain-imperative instruction.
245 */
246const STOP_VERB =
247 /^\s*(?:stop|don'?t|never)\s+(running|using|calling|doing|touching|editing|deleting|pushing|run|use|call|do|touch|edit|delete|push)\s+(.+)/i;
248
249/** Verbs whose object is a file or a piece of code rather than a shell command. */
250const EDIT_LIKE_VERBS = new Set(['editing', 'touching', 'deleting', 'edit', 'touch', 'delete']);
251
252/** A determiner is never itself the object ("the deploy script" -> "deploy script"). */
253const LEADING_DETERMINER = /^(?:the|a|an)\s+(?=\S)/i;
254
255/** Function words — pronouns, determiners, catch-all nouns — that name no actual command, tool
256 * or file: "Stop doing that" and "stop using the" must deny nothing, not deny every call whose
257 * input happens to contain the English word "that" or "the". */
258const FUNCTION_WORDS = new Set([
259 'that', 'this', 'it', 'the', 'a', 'an', 'those', 'these', 'them', 'they', 'there', 'here',
260 'everything', 'something', 'anything', 'nothing', 'things', 'stuff', 'that\'s', 'it\'s',
261]);
262
263/** A plausible command name, flag or path: word/path characters only, and not a bare function
264 * word — "pytest", "git", "-q" and "deploy" all pass; "that" and "the" do not. */
265function isPlausibleObject(token: string): boolean {
266 if (!token) return false;
267 if (FUNCTION_WORDS.has(token.toLowerCase())) return false;
268 return /^[A-Za-z0-9\-][\w.\-/]*$/.test(token);
269}
270
271/** Words introducing a qualifier this mechanism cannot represent as a `requires` flag ("pytest
272 * UNLESS it's urgent"): the pattern is cut before the qualifier rather than including words that
273 * would make the pattern match nothing real, and `tool.check`'s denial reason says so. */
274const UNREPRESENTABLE_QUALIFIER = /\s+(?:unless|except|only if|if)\b/i;
275
276/** A "without <flag>" qualifier IS representable: it becomes `requires`, so `tool.check` can
277 * deny the pattern only when that flag is absent, instead of denying it in every form — denying
278 * `pytest -q` outright, the exact command the human asked to KEEP, was the wrong call. */
279const WITHOUT_QUALIFIER = /\s+without\s+(\S+)/i;
280
281/** The rule's object: the command PREFIX up to its first flag (a plausible command/path only —
282 * see `isPlausibleObject`), plus whichever one qualifier the human's phrasing carried:
283 *
284 * - a "without X" qualifier becomes `requires` (X must be PRESENT to ALLOW the call) — unchanged
285 * from fix round 2.
286 * - a flag named directly in the object ("git push --force") becomes `flag` (X must be PRESENT
287 * to DENY the call) — new in fix round 3, so "don't use git push --force" denies only calls
288 * that both start with "git push" AND carry "--force", not every "git" command.
289 *
290 * When the object has no flag at all ("never run git push"), the whole object (up to any
291 * qualifier) is the prefix and the rule denies that subcommand unconditionally.
292 *
293 * Returns null when no plausible object can be found at all, or when the object is nothing but
294 * a bare flag with no leading subcommand token. */
295function parseObject(rest: string): { pattern: string; requires?: string; flag?: string } | null {
296 const cleaned = rest.trim().replace(LEADING_DETERMINER, '');
297
298 const withoutMatch = WITHOUT_QUALIFIER.exec(cleaned);
299 const requires = withoutMatch ? withoutMatch[1].replace(/["'.,?!]+$/g, '') : undefined;
300 const beforeQualifier = withoutMatch
301 ? cleaned.slice(0, withoutMatch.index)
302 : (cleaned.split(UNREPRESENTABLE_QUALIFIER)[0] ?? cleaned);
303
304 const rawTokens = beforeQualifier.trim().split(/\s+/).filter(Boolean);
305 if (rawTokens.length === 0) return null;
306 const tokens = rawTokens.map((token, i) =>
307 i === rawTokens.length - 1 ? token.replace(/["'.,?!]+$/g, '') : token);
308
309 if (!isPlausibleObject(tokens[0])) return null;
310
311 const prefixTokens: string[] = [];
312 let flag: string | undefined;
313 for (const token of tokens) {
314 if (token.length > 1 && token.startsWith('-')) {
315 flag = token;
316 break;
317 }
318 prefixTokens.push(token);
319 }
320 if (prefixTokens.length === 0) return null;
321
322 const pattern = prefixTokens.join(' ');
323 const result: { pattern: string; requires?: string; flag?: string } = { pattern };
324 if (requires) result.requires = requires;
325 else if (flag) result.flag = flag;
326 return result;
327}
328
329/** Parses a "stop doing X" / "don't run X again" instruction out of free text, or returns null
330 * for anything that is not unambiguously such an instruction (a question, a description, an
331 * acknowledgement with no actionable target, an object that is a pronoun/determiner/catch-all
332 * rather than a command, etc). The returned `pattern` is deliberately just the object's leading
333 * token (e.g. "pytest", not "pytest without -q"): a "without X" qualifier becomes `requires`
334 * instead (see `parseObject`), and any other unrepresentable qualifier is dropped rather than
335 * baked into a pattern that would deny nothing real — `tool.check`'s denial reason says so
336 * explicitly either way, so the human sees exactly what is actually blocked. */
337export function parseStopInstruction(text: string): SessionRule | null {
338 const match = STOP_VERB.exec(text ?? '');
339 if (!match) return null;
340 const verb = match[1].toLowerCase();
341 const object = parseObject(match[2] ?? '');
342 if (!object) return null;
343 const tool = EDIT_LIKE_VERBS.has(verb) ? 'Edit' : 'Bash';
344 return { tool, ...object };
345}
346
347/** Adds a session-scoped rule parsed from the human's own words. Never touches the installed store. */
348export function addSessionRule(state: SessionState, rule: SessionRule): void {
349 if (!state.sessionRules) state.sessionRules = [];
350 state.sessionRules.push(rule);
351}
352hooks/drift.ts 73 lines1/**
2 * Live drift surfacing: counts tool calls in the current stretch (the calls since the user last
3 * spoke) and, on the call that reaches `min_calls`, says so ONCE as model-only `context` on that
4 * tool's result, so the model reads it mid-stretch rather than after the user has spoken.
5 *
6 * Configured by `<harnessHome>/drift.json`, shaped
7 * `{"enabled": boolean, "min_calls": number, "judge": "knn"|"overlap"|"model"}`. A missing,
8 * empty, unreadable, unparseable or ill-shaped file means DISABLED. That is how the ship gate in
9 * `tools/tune_drift.py` is honoured: it refused every judge, so nothing writes an enabling file
10 * and the feature ships built, tested, and off.
11 *
12 * The live hook only has the call count to go on; the `judge` field names which offline judge
13 * justified turning it on and is validated, not executed here.
14 */
15
16export type DriftConfig = {
17 enabled: true;
18 min_calls: number;
19 judge?: 'knn' | 'overlap' | 'model';
20};
21
22const JUDGES = new Set(['knn', 'overlap', 'model']);
23
24/**
25 * The drift config, or null (disabled) for anything short of an explicit, well-formed
26 * `{"enabled": true, "min_calls": <positive number>}`. Never throws.
27 */
28export async function loadDriftConfig(io: any, home: string): Promise<DriftConfig | null> {
29 try {
30 const path = `${home}/drift.json`;
31 if (!(await io.fs.exists(path))) return null;
32 const raw = await io.fs.read(path);
33 if (typeof raw !== 'string' || raw.trim() === '') return null;
34 const parsed: any = JSON.parse(raw);
35 if (parsed === null || typeof parsed !== 'object' || Array.isArray(parsed)) return null;
36 if (parsed.enabled !== true) return null;
37 const minCalls = parsed.min_calls;
38 if (typeof minCalls !== 'number' || !Number.isFinite(minCalls) || minCalls < 1) return null;
39 if (parsed.judge !== undefined && !JUDGES.has(parsed.judge)) return null;
40 return { enabled: true, min_calls: minCalls, judge: parsed.judge };
41 } catch {
42 return null;
43 }
44}
45
46/** Whether a stretch of `count` calls is long enough to mention under `config`. */
47export function shouldWarn(count: number, config: DriftConfig | null): boolean {
48 return config !== null && config.enabled === true && count >= config.min_calls;
49}
50
51/** The one line the model reads when a stretch has run long. */
52export function driftNote(count: number): string {
53 return `meta-harness drift: ${count} tool calls have run since the user last spoke. `
54 + 'Before continuing, check the work still matches what was asked.';
55}
56
57/**
58 * Origins that mean the user themself spoke, which ends the stretch. A notification, peer
59 * message, schedule or plugin submission does not: the stretch it lands in keeps counting, and
60 * its once-per-stretch latch stays set. An absent or unrecognised
61 * origin is treated as the user speaking, the choice that produces fewer notes, not more.
62 */
63const NON_USER_ORIGINS = new Set([
64 'task-notification', 'scheduled-trigger', 'peer', 'peer-send-message', 'projects-relay',
65 'channel', 'coordinator', 'observer', 'observer-activity', 'auto-continuation', 'slack-ping',
66 'plugin',
67]);
68
69export function userSpoke(event: any): boolean {
70 const kind = event?.origin?.kind;
71 return !(typeof kind === 'string' && NON_USER_ORIGINS.has(kind));
72}
73hooks/receipts.ts 54 lines1/**
2 * Decision receipts (spec: docs/superpowers/specs/2026-09-29-decision-receipts-design.md).
3 * A receipt is written when the plugin acts - or, for a held-out case, would have acted - and
4 * the outcome is attributed offline by meta_harness/receipts.py. Helpers take `io`, never `$`.
5 */
6export type ReceiptFields = {
7 event: string; source: string; artifact: string | null;
8 signature: string | null; tool: string; decision: 'acted' | 'held';
9};
10
11const DEFAULT_RATE = 0.1;
12const MAX_RATE = 0.5;
13const calls = new Map<string, number>();
14let drawFn: () => number = Math.random;
15
16export function bumpCall(sessionId: string): number {
17 const next = (calls.get(sessionId) ?? 0) + 1;
18 calls.set(sessionId, next);
19 return next;
20}
21
22export function currentCall(sessionId: string): number {
23 return calls.get(sessionId) ?? 0;
24}
25
26export function setDraw(fn: () => number): void { drawFn = fn; }
27export function draw(): number { return drawFn(); }
28export function shouldHold(rate: number): boolean { return rate > 0 && draw() < rate; }
29
30export async function holdoutRate(io: any, home: string): Promise<number> {
31 try {
32 const path = `${home}/receipts.json`;
33 if (!(await io.fs.exists(path))) return DEFAULT_RATE;
34 const value = Number(JSON.parse(await io.fs.read(path))?.holdout_rate);
35 if (!Number.isFinite(value)) return DEFAULT_RATE;
36 return Math.min(MAX_RATE, Math.max(0, value));
37 } catch {
38 return DEFAULT_RATE;
39 }
40}
41
42export async function writeReceipt(io: any, home: string, sessionId: string,
43 record: ReceiptFields): Promise<void> {
44 try {
45 const path = `${home}/receipts-${sessionId}.jsonl`;
46 const line = JSON.stringify({ ts: new Date().toISOString(), session: sessionId,
47 call: currentCall(sessionId), ...record }) + '\n';
48 const prior = (await io.fs.exists(path)) ? await io.fs.read(path) : '';
49 await io.fs.write(path, prior + line);
50 } catch (error) {
51 try { io.log?.(`meta-harness: receipt not written: ${String(error)}`); } catch { /* fail open */ }
52 }
53}
54