SLOPSHOPPER

Meta-Harness

Harness engineering for Claude Code: attribute every prompt/agent change, keep every candidate, diagnose from traces, hold out a split - and search over Claude…

newguardprompt
v0.8.0MITupdated 2026-09-29Vatsa10/Meta-Harness
A shopper browsing a rack in a slop shop
README

Meta-Harness

Filesystem-backed end-to-end optimization of executable LLM harnesses, following Meta-Harness: End-to-End Optimization of Model Harnesses (paper.pdf).

A harness is the code around a fixed model: what it stores, retrieves, and shows the model at each step. This repository searches over that code. A coding-agent proposer reads the full experience filesystem — every prior candidate's source, scores, execution traces, and the reasoning that produced it — and writes new candidates. Candidates are evaluated on a search split; the Pareto frontier over (accuracy, context tokens) is scored once on a held-out split the proposer never sees.

Use it from Claude Code

The repository is also a Claude Code plugin. Not a wrapper around the CLI — model-invoked skills that change how Claude works on prompts, agent scaffolds and retrieval, plus a hook layer that observes failures and enforces what has been learned from them:

claude --plugin-dir .

optimizing-harnesses carries the discipline (one change per candidate, every candidate kept, no score without a trace, held-out split read once). building-eval-sets, reading-execution-traces and running-harness-search handle the pieces, optimizing-claude-code points the search at Claude Code's own scaffold, and learning-from-failures drives the loop below. The engine is the escalation path when hand-tuning stalls; the skills need nothing installed.

See docs/plugin.md for the skill list, the baseline testing behind it, and why Haiku 4.5 is the default harness model.

Learning from failures

python -m meta_harness learn            # failure -> artifact -> replay-scored -> staged
python -m meta_harness learn --tasks tasks.json   # also require no regression on a task set
python -m meta_harness learn --status
python -m meta_harness learn --accept <id>

Three layers. Observe: a tool.call hook records tool errors and repeated calls to a per-session log. Learn: the CLI above ranks those failures together with ones mined from past transcripts, proposes one artifact at the strongest layer that can carry it, and replays it. Enforce: a tool.check hook applies accepted rules and a prompt.submit hook attaches accepted injections to the user's turn as context. The plugin listens only on tool.call, tool.check and prompt.submit; nothing rewrites the system prompt.

An artifact is staged only if it fixes the failure it was born from; one whose origin replay still fails is archived with the verdict that killed it. With --tasks, it must also score no worse than the baseline on that task set, with context cost as the tiebreak; without --tasks, that check does not run and the recorded verdict says so. Nothing installs itself.

Artifacts sit at four layers, strongest first: rule (a tool.check, costing no standing tokens and not ignorable) > injection (context attached to a matching turn, paid when it fires) > skill > doctrine. Only rule and injection are enforced by the hooks today. That ordering is this project's own design position — it appears nowhere in paper.pdf. tools/prose_vs_rule.py is the experiment that would test it against a measured baseline, and it has not been run here.

What the paper does establish (Table 3, online text classification, median/best score): a proposer given scores only reaches 34.6/41.3, scores plus an LLM summary reaches 34.9/38.7, and full raw execution traces reach 50.0/56.7. Raw trace access is the paper's key ingredient — a summary "may even hurt by compressing away diagnostically useful details." Appendix A.2 adds a second, independently evidenced principle: on TerminalBench-2 the proposer regressed six consecutive iterations while editing prompts and control flow, diagnosed the shared prompt edit as the confound, and then won with a purely additive change. Additive beats invasive.

One measured negative result of our own: the paper's winning TerminalBench-2 discovery was an environment snapshot injected before the first model call. That does not transfer to a developer's own repository, where the environment is already known — 50 orientation calls across 344 sessions on this machine. We do not build it.

Decision receipts, and the holdout

A receipt is one line the hooks write each time a learned rule, an injection, a session rule or the drift note acts on a call or a prompt: which artifact, what it did (acted, or held), and the call number in the session. It carries no message text and no tool input. Receipts go to <harness_home>/receipts-<session>.jsonl, next to the observations, which now record the same call number so the two can be matched.

To find out whether an artifact does anything, the harness needs something to compare against. So a learned rule, an injection or the drift note is held out on a random 10% of the occasions it would have acted: the hook records a held receipt and does nothing. The cost is real: a held-out learned rule does not deny that call, so about one time in ten the failure it exists to prevent is allowed through. Session rules and rejection memory are never held out: they are something you said, and they always apply.

Set the rate in <harness_home>/receipts.json:

{"holdout_rate": 0.1}

The value is clamped to 0 through 0.5; a missing or unreadable file means 0.1. Set it to 0 to switch the holdout off entirely (receipts are still written, but nothing is held, so learned artifacts stay "not enough data").

python -m meta_harness receipts report            # add --json for the raw rows
python -m meta_harness receipts export --out receipts.json

The report compares how often the failure signature recurred in later calls of the same session with the artifact acting versus held out. Verdicts:

  • helps - the failure recurred measurably less when the artifact acted.
  • no measurable effect - no difference either way. The harness proposes retiring the artifact (see learn --status); it never retires one itself.
  • not enough data - fewer than 5 receipts in an arm, or no recurrence at all to compare. No verdict is given. The drift note cannot reach a verdict yet either: its receipts carry no failure signature to match against.
  • no control arm - a session rule or rejection memory, which are never held out, so there is nothing to compare against.

Evidence accumulates slowly. At a 10% holdout an artifact needs many occasions before either arm has 5 receipts, which takes weeks of ordinary use. Until then, expect "not enough data".

Waste, and what the hooks do inside a session

python -m meta_harness waste (add --json, --this-project, --since, --limit) reads your transcripts and reports stretches of tool calls that ended with you correcting course, the calls they burned, and identical failures retried. It also writes <harness_home>/waste.json. Its correction count is a keyword heuristic: in a hand-labelled audit, 11 of the 21 flagged stretches audited were genuine corrections, and recall is unmeasured. The report carries that caveat itself; treat the numbers as an estimate, not an audit.

Beyond enforcing artifacts, the hooks add (details in hooks/README.md):

  • Rejection memory - a call you rejected is denied if retried identically in the same session.
  • "Stop doing X" session rules - a human's own prompt like "stop running pytest" denies matching calls immediately, for the rest of that session only. Prompts from plugins, peers, schedules or notifications never create one. Nothing makes a rule permanent yet. Session rules and rejection memory are never held out.
  • A first-run line - once, a hedged estimate of calls that went to later-corrected work. It needs a prior meta-harness waste run; nothing mines your history in the background.
  • A drift note - built, tested, and shipped disabled. It would say once when a long stretch of calls has run since you last spoke. The ship gate refused every judge (word-overlap precision about 0.01; the kNN judge never fired; the model judge was never evaluated), so it is off unless drift.json enables it, and "enabled" then means a plain call-count threshold: the config's judge field is validated but never run.
  • /meta-harness:harness waste | pending | why - the report, what is staged or pending, and why the last denial happened.

Mined failures are weighted by recency and by Claude Code version (version_weight). On the current data the version weight is inert: every session shares Claude Code minor 2.1.

The hooks need Claude Code's function-hook surface, which is early access:

export CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1   # PowerShell: $env:CLAUDE_CODE_ENABLE_FUNCTION_HOOKS="1"
claude --plugin-dir .

See hooks/README.md for the hook layer and the design for the rest.

Install

python -m venv venv
venv/Scripts/activate          # Windows;  source venv/bin/activate elsewhere
pip install -e .
pip install -e ".[dev]"        # adds pytest

No runtime dependencies. Python 3.10+.

Configure the model gateway (Unikey)

All model traffic goes through Unikey, an OpenAI-compatible gateway (https://www.getunikey.ai/v1, Authorization: Bearer $UNIKEY_API_KEY).

export UNIKEY_API_KEY=sk-...            # PowerShell: $env:UNIKEY_API_KEY="sk-..."
meta-harness models --provider unikey   # GET /v1/models

The proposer is a coding agent, billed separately. Unikey also exposes an Anthropic-compatible /v1/messages, so Claude Code can be routed through it:

export ANTHROPIC_BASE_URL=https://www.getunikey.ai
export ANTHROPIC_AUTH_TOKEN=$UNIKEY_API_KEY

--provider openai|anthropic|compatible still work, reading OPENAI_API_KEY / ANTHROPIC_API_KEY and their *_BASE_URL overrides.

Run a search

meta-harness run \
  --tasks data/classification.jsonl \
  --task-type classification \
  --provider unikey --model gpt-5.2 \
  --iterations 20 --candidates 2 --repeats 3 --max-workers 4 \
  --proposer-model claude-sonnet-4-6 \
  --root .meta-harness

Dataset formats: .jsonl, .json, .csv.

--task-typerequired fieldsoptionalmetric
classificationinput, labellabelsexact match
mathproblemanswer\boxed{} + numeric equivalence
terminalinstructiontest_command, image, workdir, timeouttest command exit code

Without --test-tasks, --tasks is split 70/30 (--search-fraction, --split-seed). Without --baseline, the seeds in baselines/ for that task type are used. Without --proposer-command, the proposer is the claude CLI; use --proposer-command to drive any other agent, or python tools/llm_proposer.py for a plain-API proposer that needs no agent installed. See docs/using-with-coding-agents.md.

Two sample datasets ship in data/.

What a run produces

.meta-harness/
  run.json                       config + the search-split tasks
  frontier.json                  Pareto frontier over (score, context_cost)
  candidates/<id>/harness.py     candidate source
  candidates/<id>/scores.json    score, context_cost, repeats, score_std, valid, error
  candidates/<id>/traces.jsonl   every model call, task start/end, harness event
  candidates/<id>/traces-N.jsonl additional repeats
  candidates/<id>/proposer_reasoning.md
  proposals/iteration-NNNN/      proposer stdout/stderr
  views/iteration-NNNN/          what the proposer was allowed to read (ablation modes)
.meta-harness-test/
  test_results.json              held-out scores, written once, outside the proposer's reach
.meta-harness-cache/             model-call cache, keyed by (model, prompt, kwargs)

Inspect a finished run: meta-harness inspect .meta-harness

Proposer-view ablation (paper Table 3)

meta-harness run ... --proposer-view scores    # source + scores only
meta-harness run ... --proposer-view summary   # + LLM summaries, no raw traces
meta-harness run ... --proposer-view full      # everything (default)

Agentic coding domain

Terminal tasks run the model's commands inside Docker:

meta-harness run --task-type terminal --tasks data/terminal.jsonl \
  --provider unikey --model claude-sonnet-4-6

Each row needs an image. Rows without one are refused unless you pass --allow-local-shell, which executes model-authored commands on your machine — only do that in a throwaway environment.

Writing your own harness

class Harness:
    def run(self, task, model, trace):
        prompt = f"Classify: {task['input']}"
        trace.event("prompt", {"prompt": prompt})
        return model(prompt)

The instance persists across the tasks of one evaluation (that is your memory) and is recreated per evaluation. model(prompt) -> str; model.call(prompt) -> (str, usage). Every call is priced in input tokens and written to the trace. meta_harness.retrieval provides TfidfIndex, BM25Index, and reciprocal_rank_fusion.

Seven working examples live in baselines/; the exact contract handed to the proposer is meta_harness/skill/SKILL.md.

Design notes

  • meta_harness/core.py — the outer loop, experience filesystem, Pareto frontier, metering.
  • meta_harness/sandbox.py — interface validation in a subprocess under a timeout, so an LLM-written while True cannot hang the search.
  • meta_harness/agent_proposer.py — the coding-agent proposer and the view ablation.
  • meta_harness/cache.py — disk cache so --repeats and re-runs do not re-bill.
  • tools/llm_proposer.py — proposer that needs only an API key, no coding-agent CLI.
  • meta_harness/replay.py — failure signatures, and the replay a proposed artifact is scored on.
  • meta_harness/learn.py — ranking, proposal, and the retention decision.
  • meta_harness/harness_store.py — staged and installed artifacts, and the accept/reject gate.
  • hooks/ — the function-hook layer: observe, enforce, inject, rejection memory, session rules, first-run line, drift note (disabled). Fails open by construction.
  • meta_harness/waste.py — the waste report behind meta-harness waste.
  • tools/prose_vs_rule.py — the prose-versus-mechanism experiment (built, not yet run).
  • docs/plugin.md — the Claude Code plugin: skills, agent, install, Haiku defaults.
  • docs/using-with-coding-agents.md — Claude Code as proposer, optimizing agent harnesses, and shipping a discovered harness back into your own agent.
  • docs/superpowers/specs/ and docs/superpowers/plans/ — the spec and task-by-task plan this implementation follows.

Tests

python -m pytest -q

No network calls. The hook layer is covered by Node tests under hooks/ that the Python suite shells out to; they are skipped, with the reason stated, when node is absent.

Source 4 files
hooks/harness.ts 844 lines
1/**
2 * Meta-Harness function hooks: observe failures, enforce learned artifacts.
3 *
4 * Every handler fails open: most run inside `guardBefore`, `guardAfter` or `guardAfterMap`, all of which swallow a
5 * throw and fall through to `next` rather than break the turn — and none ever calls `next` a
6 * second time once it has been called. `registerBootstrap` does its fail-open handling by hand,
7 * so its post-`next` marker write is swallowed locally instead of reaching any wrapper at all. A learning system that can break a session, or submit the
8 * human's prompt twice, is worse than no learning system.
9 *
10 * Per-turn features (rejection memory, session rules, injections, the first-run message) live on
11 * `prompt.submit`, never `prompt.section`. `prompt.section` fires once per NAMED SECTION of the
12 * system prompt, its sections are cached for the whole session, and returning `{ text }` REPLACES
13 * that section — there is no `event.prompt`/`event.text` carrying the user's words on it, and a
14 * handler with no section-name filter fires for and can overwrite every section that exists,
15 * including ones it has never heard of. `prompt.submit` carries the user's actual turn as
16 * `e.text`, and a hook adds anything the model should see via `next({ ...e, context: [...] })`
17 * without touching the system prompt at all.
18 */
19
20import type { On, PluginOptions, Register } from 'claude-code';
21import {
22  addSessionRule,
23  callKey,
24  evaluateRule,
25  loadInstalled,
26  loadRules,
27  parseStopInstruction,
28  rememberRejection,
29  stableStringify,
30  toolArgs,
31  type Rule,
32  type SessionState,
33  wasRejected,
34} from './rules.js';
35import { type DriftConfig, driftNote, loadDriftConfig, shouldWarn, userSpoke } from './drift.js';
36import { bumpCall, holdoutRate, shouldHold, writeReceipt } from './receipts.js';
37
38export type Fallible<E, R> = (io: any, event: E, next: (e: E) => Promise<R>) => Promise<R>;
39
40/** Log a skip the same way everywhere, without ever risking a second throw of its own. */
41function logSkip(io: any, name: string, error: unknown): void {
42  try {
43    io.ui.log(`meta-harness ${name} skipped (${error instanceof Error ? error.message : String(error)})`);
44  } catch {
45    // logging must never be the thing that breaks the turn either
46  }
47}
48
49/**
50 * Why the guards below take `(name, io, event, next, handler)` and are CALLED from inside a
51 * function literal instead of returning a hook: Claude Code's hooks loader rejects any `on(...)`
52 * whose hook argument is not a function literal (or the name of one) written in the call — a
53 * wrapper call such as `on('tool.call', afterCall(...))` fails the whole module, so no hook runs.
54 * Every registration is therefore `on('<event>', async ($, e, next) => guardX(..., $, e, next, ...))`;
55 * tests/test_hook_literals.py enforces that statically.
56 */
57
58/**
59 * Runs a handler that may call `next` itself, so a throw becomes a pass-through instead of a
60 * broken turn. `next` is the rest of the chain (for prompt.submit, the human's prompt reaching
61 * core; for tool.check, the tool running). Recovering a throw by calling it again is only safe if
62 * the handler never reached it: once it has, a second call submits the prompt twice or runs the
63 * tool twice. So once `next` has been called, its own outcome stands on every path.
64 */
65export async function guardBefore<E, R>(
66  name: string,
67  io: any,
68  event: E,
69  next: (e: E) => Promise<R>,
70  handler: (next: (e: E) => Promise<R>) => Promise<R>,
71): Promise<R> {
72  let called = false;
73  let settled: { ok: true; value: R } | { ok: false; error: unknown } | null = null;
74  const once = async (e: E): Promise<R> => {
75    called = true;
76    try {
77      const value = await next(e);
78      settled = { ok: true, value };
79      return value;
80    } catch (error) {
81      settled = { ok: false, error };
82      throw error;
83    }
84  };
85  try {
86    return await handler(once);
87  } catch (error) {
88    logSkip(io, name, error);
89    if (!called) return next(event);
90    const outcome = settled as { ok: true; value: R } | { ok: false; error: unknown } | null;
91    if (outcome?.ok) return outcome.value;
92    throw outcome ? outcome.error : error;
93  }
94}
95
96/**
97 * Runs a handler whose work happens AFTER the tool has already run. `guardBefore` recovers a
98 * throw that happened before `next` by calling `next(event)`: for a post-`next` handler that
99 * would run the tool a SECOND time, duplicating the side effect of a Bash or Write. Here `next` is
100 * called exactly once, up front (a rejection propagates once, untouched), and the fallible work is
101 * what gets swallowed — so an observer that throws costs an observation, never a repeated command.
102 */
103export async function guardAfter<E, R>(
104  name: string,
105  io: any,
106  event: E,
107  next: (e: E) => Promise<R>,
108  handler: (outcome: R) => Promise<void>,
109): Promise<R> {
110  const outcome = await next(event);
111  try {
112    await handler(outcome);
113  } catch (error) {
114    logSkip(io, name, error);
115  }
116  return outcome;
117}
118
119/**
120 * Like `guardAfter`, for a handler that may return a REPLACEMENT outcome (a copy of `next`'s with
121 * something appended). `next` is called exactly once, up front; if the handler throws, core's
122 * outcome is returned as it came, never a second call to `next`.
123 */
124export async function guardAfterMap<E, R>(
125  name: string,
126  io: any,
127  event: E,
128  next: (e: E) => Promise<R>,
129  handler: (outcome: R) => Promise<R>,
130): Promise<R> {
131  const outcome = await next(event);
132  try {
133    return await handler(outcome);
134  } catch (error) {
135    logSkip(io, name, error);
136    return outcome;
137  }
138}
139
140/** `guardBefore` as a hook-returning wrapper, for tests. Never pass its result to `on` directly. */
141export function safely<E, R>(name: string, handler: Fallible<E, R>): Fallible<E, R> {
142  return (io, event, next) => guardBefore(name, io, event, next, (once) => handler(io, event, once));
143}
144
145/** `guardAfter` as a hook-returning wrapper, for tests. Never pass its result to `on` directly. */
146export function afterCall<E, R>(
147  name: string,
148  handler: (io: any, event: E, outcome: R) => Promise<void>,
149): Fallible<E, R> {
150  return (io, event, next) => guardAfter(name, io, event, next, (outcome) => handler(io, event, outcome));
151}
152
153/** `guardAfterMap` as a hook-returning wrapper, for tests. Never pass its result to `on` directly. */
154export function afterCallMap<E, R>(
155  name: string,
156  handler: (io: any, event: E, outcome: R) => Promise<R>,
157): Fallible<E, R> {
158  return (io, event, next) => guardAfterMap(name, io, event, next, (outcome) => handler(io, event, outcome));
159}
160
161const REPEAT_WINDOW = 6;
162const REPEAT_THRESHOLD = 4;
163
164type Recent = { tool: string; key: string };
165
166/** The harness's home directory: overridable for tests, otherwise under the user's profile. */
167async function harnessHome(io: any): Promise<string> {
168  // `$.env.get` answers with a Promise in Claude Code; awaiting also accepts a plain value.
169  const override = await io.env?.get?.('META_HARNESS_HOME');
170  if (override) return override;
171  const profile = (await io.env?.get?.('USERPROFILE')) || (await io.env?.get?.('HOME'));
172  return `${profile}/.claude/harness`;
173}
174
175/**
176 * Mirrors meta_harness.replay.ERROR_PATTERNS: same regex text and slug, in the same order, so a
177 * `cause` recorded here dedupes against what Python's `_cause()` computes from the same text.
178 *
179 * Every backslash in these patterns is doubled (`\\d`, `\\[`, `\\(`, ...). These are plain JS
180 * string literals fed to `new RegExp(pattern, 'i')` below, and a JS string literal silently
181 * drops a backslash in front of any character it does not recognize as an escape -- `\d`, `\S`,
182 * `\[`, `\]`, `\(` and `\)` are NOT recognized string escapes (only `\n`, `\r`, `\t`, `\'`, `\\`,
183 * etc. are), so a single backslash here compiles to a regex with a missing metacharacter escape
184 * (`\d+` -> a literal-`d` character class, `\(...\)` -> an unescaped group, and so on). A prior
185 * version of this file used single backslashes throughout, which type-checked and passed a
186 * test that only compared source text, while classifying 235 of 1,167 real failure texts
187 * differently from the Python side. Regex literals (`/.../i`) would not have this problem, but
188 * would need a parallel array-of-RegExp construction; doubling the backslash here keeps this
189 * array's shape (string, string) identical to ERROR_PATTERNS's (str, str), which is what the
190 * parity test below relies on to generate its own fixtures programmatically in the future.
191 *
192 * `tests/test_hook_assets.py::test_ts_and_python_cause_agree_on_fixtures` actually EXECUTES both
193 * classifiers over `tests/cause_fixtures.py::CAUSE_FIXTURES` (via node) and asserts identical
194 * slugs -- not a source-text comparison, which cannot detect this class of bug.
195 *
196 * UnicodeEncodeError must stay ordered before the decode/charmap/cp1252 pattern: an encode
197 * failure's message can also contain the word "decode" in surrounding text, so checking decode
198 * first would misclassify it. That ordering was fixed on the Python side in Task 2.
199 */
200const CAUSE_PATTERNS: ReadonlyArray<readonly [string, string]> = [
201  ['UnicodeEncodeError', 'unicode-encode'],
202  ['UnicodeDecodeError|charmap|cp1252', 'unicode-decode'],
203  ['No such file or directory|cannot find the (file|path)', 'missing-path'],
204  ['Permission denied|EACCES', 'permission'],
205  ['command not found|is not recognized as', 'missing-command'],
206  ['timed out|TimeoutExpired', 'timeout'],
207  ['has not been read yet|must read.*before', 'unread-edit'],
208  ['String to replace not found|old_string', 'edit-mismatch'],
209  ['SyntaxError|unterminated', 'syntax'],
210  ['contains multiple operations|Compound command changes working directory', 'compound-shell'],
211  ['requires approval|denied by the Claude Code auto mode classifier', 'needs-approval'],
212  ['doesn\'t want to proceed|tool use was rejected', 'user-rejected'],
213  ['Blocked:|blocked by a deny rule', 'blocked-policy'],
214  ['unexpected EOF while looking|simple_expansion|expansion obfuscation', 'shell-quoting'],
215  ['not in Claude\'s tab group|determine which page this action targets', 'tab-target'],
216  ['modified since read', 'stale-read'],
217  ['EISDIR|illegal operation on a directory', 'is-directory'],
218  ['Traceback \\(most recent call last\\)', 'python-traceback'],
219  ['Permission to (use|read).*has been denied', 'permission-denied-tool'],
220  ['File does not exist', 'missing-path'],
221  ['is temporarily unavailable', 'model-unavailable'],
222  ['InputValidationError|Workflow script file not found|No task found with ID|Invalid workflow script|scriptPath must be a script path|Unknown skill:|Task ID is required', 'workflow-error'],
223  ['Failed to execute JavaScript|JavaScript execution error', 'js-error'],
224  ['Error capturing screenshot|actions\\[\\d+\\][^\\n]*failed|Failed to find element|Failed to execute action|Error capturing zoomed screenshot|is not a supported form input|Can\'t interact with browser-internal', 'browser-action-failed'],
225  ['No such tool available', 'unknown-tool'],
226  ['hook did not respond before|tool did not respond in time', 'hook-timeout'],
227  ['Found \\d+ matches of the string', 'edit-mismatch'],
228  ['Python was not found|pdftoppm is not installed', 'missing-command'],
229  ['node:internal/modules/(package_json_reader|run_main)|Cannot find module|ERR_MODULE_NOT_FOUND', 'module-not-found'],
230  ['"error":\\{"name":"(HttpException|McpError)"|already exists in local config', 'api-error'],
231  ['fatal: (pathspec|detected dubious ownership|ambiguous argument|.*is outside repository)|ignored by one of your \\.gitignore|docker: Error response from daemon', 'git-error'],
232  ['On branch \\S+\\r?\\nYour branch is (up to date|ahead of)|warning: in the working copy of', 'git-noise'],
233  ['npm error code|npm warn exec', 'npm-error'],
234  ['exceeds maximum allowed tokens', 'output-too-large'],
235  ['ConnectionRefusedError|connection refused|ECONNREFUSED', 'connection-refused'],
236  ['was blocked\\. For security|is blocked\\. This path is protected|denied by your permission', 'blocked-policy'],
237  ['=+ FAILURES =+|ERROR at setup of|\\bAssertionError\\b|FAILED \\S+::', 'test-failure'],
238  ['tab group no longer exists|Missing required parameter tabId', 'tab-target'],
239  ['needs design-system authorization', 'needs-approval'],
240  ['ENAMETOOLONG', 'path-too-long'],
241  ['error TS\\d+|imported but unused', 'ts-error'],
242  ['not logged into any GitHub hosts', 'gh-auth-error'],
243];
244
245/** Classifies failure text into the same slug meta_harness.replay._cause() would produce. */
246export function cause(text: string): string {
247  const value = text ?? '';
248  for (const [pattern, slug] of CAUSE_PATTERNS) {
249    if (new RegExp(pattern, 'i').test(value)) return slug;
250  }
251  return 'other';
252}
253
254/** This session's id, or 'unknown' when the engine cannot supply one; never throws. */
255async function currentSessionId(io: any): Promise<string> {
256  try {
257    const id = await io.session?.id?.();
258    return id ? String(id) : 'unknown';
259  } catch {
260    return 'unknown';
261  }
262}
263
264/**
265 * Appends one JSON line to this session's own observation file.
266 *
267 * One file per session, not one shared file: `$.fs` exposes no append or lock primitive, so a
268 * shared file would need a non-atomic exists/read/write cycle that two concurrent sessions can
269 * race and clobber each other on. Per-session files remove the race instead of trying to guard
270 * it; a reader globs `observed-*.jsonl` under the harness home. `$.fs.write` creates missing
271 * parent directories itself, so a fresh harness home on a new machine needs no separate mkdir.
272 */
273async function observe(io: any, sessionId: string, record: Record<string, unknown>): Promise<void> {
274  const path = `${(await harnessHome(io))}/observed-${sessionId}.jsonl`;
275  const line = `${JSON.stringify({ ts: new Date().toISOString(), ...record })}\n`;
276  const existing = (await io.fs.exists(path)) ? await io.fs.read(path) : '';
277  await io.fs.write(path, existing + line);
278}
279
280function resultText(result: unknown): string {
281  return typeof result === 'string' ? result : JSON.stringify(result ?? '');
282}
283
284/**
285 * The text of a real `tool.call` outcome (ToolCallResult) worth checking for a rejection
286 * announcement — and ONLY when the outcome actually says the call was refused or errored.
287 *
288 * The `deny` branch's `deny` string always counts: a hook or core only sets it on an actual
289 * refusal. The answered branch's `text`/`result` count ONLY when `isError` is true — `text` and
290 * `result` are present on every SUCCESSFUL call too (a normal Read's file contents are `result`,
291 * and its `text` is the same content joined for the model), so checking them unconditionally
292 * means a Read of any file that happens to CONTAIN the phrase "tool use was rejected" — this very
293 * file, for one — gets recorded as a rejection and denied for the rest of the session. Gating on
294 * `isError`/`deny` is what keeps a successful result's mere text out of consideration entirely.
295 */
296function rejectionAnnouncement(outcome: any): string {
297  if (typeof outcome?.deny === 'string' && outcome.deny.length > 0) return outcome.deny;
298  if (outcome?.isError !== true) return '';
299  const parts = [outcome?.text, resultText(outcome?.result)];
300  return parts.filter((part): part is string => typeof part === 'string' && part.length > 0).join(' ');
301}
302
303/**
304 * Unambiguous tool-error phrases: strings a tool emits when it refuses, which do not plausibly
305 * appear as the FIRST thing in a successful result. Anything weaker (a bare `error:` anywhere in
306 * the text) matched a successful `Grep` for the word "error:" and recorded it as a failure;
307 * those counts feed selection ranking, so the noise became the thing the learn loop chased.
308 */
309const ERROR_PHRASES =
310  /has not been read yet|String to replace not found|is not recognized as an internal or external command|No such file or directory/i;
311
312/** True when the tool call actually failed. */
313function isError(result: unknown): boolean {
314  if (result && typeof result === 'object') {
315    const flag = (result as any).is_error ?? (result as any).isError;
316    // The engine's own verdict is authoritative; never second-guess it by grepping the text.
317    if (typeof flag === 'boolean') return flag;
318  }
319  const text = resultText(result);
320  // Otherwise only an error ANNOUNCED at the start of the result counts, plus a short list of
321  // phrases a tool only ever emits when it refused.
322  return /^\s*"?(error|[A-Za-z.]*Error:|Traceback \(most recent call last\))/i.test(text)
323    || ERROR_PHRASES.test(text);
324}
325
326/**
327 * Launchers stripped from the front of a Bash command before a session rule's pattern is
328 * compared: `python -m`, `python3 -m`, `py -m`, `npx`, `uv run`, `poetry run`, `pipx run`, plus a
329 * leading `env` and any `VAR=value` assignments. Takes lower-cased whitespace tokens.
330 */
331export function stripLaunchers(tokens: string[]): string[] {
332  let rest = tokens;
333  for (;;) {
334    if (rest[0] === 'env' || /^[a-z_][a-z0-9_]*=/.test(rest[0] ?? '')) {
335      rest = rest.slice(1);
336    } else if (['python', 'python3', 'py'].includes(rest[0] ?? '') && rest[1] === '-m') {
337      rest = rest.slice(2);
338    } else if (rest[0] === 'npx') {
339      rest = rest.slice(1);
340    } else if (['uv', 'poetry', 'pipx'].includes(rest[0] ?? '') && rest[1] === 'run') {
341      rest = rest.slice(2);
342    } else {
343      return rest;
344    }
345  }
346}
347
348/** The deny reason for a session rule, saying what the matching actually does. */
349export function sessionRuleReason(tool: string, rule: { pattern: string; requires?: string; flag?: string }): string {
350  const head = 'Session rule from this conversation: this blocks any';
351  if (tool !== 'Bash') {
352    return `${head} ${tool} call whose input contains "${rule.pattern}" [session-rule]`;
353  }
354  const what = `${head} Bash command whose first words are "${rule.pattern}" `
355    + '(after any launcher such as python -m, npx, uv run, poetry run, pipx run, env or VAR=x)';
356  if (rule.requires) return `${what} unless "${rule.requires}" is one of its words [session-rule]`;
357  if (rule.flag) return `${what} and "${rule.flag}" is one of its words [session-rule]`;
358  return `${what}, whatever its arguments [session-rule]`;
359}
360
361/** Observes tool.call outcomes: records errors and repeated identical calls to observed-<session>.jsonl. */
362export function registerObserver(add: On): void {
363  const recent: Recent[] = [];
364
365  // guardAfter, not guardBefore: this handler's work runs after the tool has already executed, so a
366  // recovery that re-entered next() would run the tool twice.
367  add('tool.call', async (io: any, event: any, next: any) => guardAfter('tool.call', io, event, next, async (outcome: any) => {
368    // The one place the session's call counter advances: once per tool.call, error or not.
369    const sessionId = await currentSessionId(io);
370    const call = bumpCall(sessionId);
371    const tool = String(event?.tool ?? 'unknown');
372    const args = toolArgs(event);
373    const key = `${tool}:${stableStringify(args).slice(0, 200)}`;
374
375    recent.push({ tool, key });
376    if (recent.length > REPEAT_WINDOW) recent.shift();
377    const repeats = recent.filter((entry) => entry.key === key).length;
378
379    if (repeats >= REPEAT_THRESHOLD) {
380      await observe(io, sessionId, { kind: 'repeat', call, tool, cause: 'repeat', input: args });
381    }
382    const text = resultText((outcome as any)?.result);
383    if (isError((outcome as any)?.result)) {
384      await observe(io, sessionId, {
385        kind: 'tool_error',
386        call,
387        tool,
388        cause: cause(text),
389        input: args,
390        text: text.slice(0, 400),
391      });
392    }
393  }));
394}
395
396/** Unambiguous phrases the engine emits when the human declines a tool call outright, as opposed
397 * to the tool itself failing. Recording on these two phrases only (never a bare "no" in the
398 * conversation) keeps rejection memory from firing on an ordinary declined suggestion in prose. */
399const REJECTION_PATTERN = /doesn't want to proceed|tool use was rejected/i;
400
401/**
402 * Enforces installed `rule` artifacts at tool.check: the layer that costs no standing tokens and
403 * cannot be talked around, because it runs before the tool call, not as prose in a prompt.
404 *
405 * Also carries this session's rejection memory (task 13) and session-scoped "stop doing X" rules
406 * (task 14), both stored on the same `SessionState` so `tool.check` can consult them ahead of the
407 * installed-rule loop. Rejection detection lives in its own `tool.call` handler below; there is
408 * no read-tracking handler here, since read-before-edit is the engine's job, not this plugin's
409 * (see hooks/rules.ts). Claude Code refuses a second registration of one event without a matcher, so these
410 * hooks are added to `register`'s per-event list and chained inside its single `on` per event.
411 * `prompt.submit` (not `prompt.section`) is where the human's own words
412 * are read, since only `prompt.submit`'s event carries `text`.
413 */
414export function registerRules(add: On): void {
415  const state: SessionState = {
416    callCounts: new Map<string, number>(),
417    rejected: new Map<string, string>(),
418    sessionRules: [],
419  };
420  let rules: Rule[] | null = null;
421
422  // guardAfter, not guardBefore: rejection detection reads the outcome, which only exists AFTER the
423  // tool has already run (or been denied), so a `guardBefore` recovery re-entering next() would run
424  // the tool a second time. `next` resolves before this handler's own body runs at all.
425  //
426  // There is no read-tracking hook here any more: a live test proved Claude Code's own engine
427  // already denies an Edit/Write of a file not read this session ("File has not been read yet"),
428  // and it runs before this plugin ever sees the call. The plugin-side copy of that check (a
429  // `read-before-edit` rule plus this handler's read-tracking) is retired -- see hooks/rules.ts.
430  add('tool.call', async (io: any, event: any, next: any) => guardAfter('tool.call:rejection-memory', io, event, next, async (outcome: any) => {
431    const text = rejectionAnnouncement(outcome);
432    if (text && REJECTION_PATTERN.test(text)) {
433      rememberRejection(state, String(event?.tool ?? ''), toolArgs(event));
434    }
435  }));
436
437  // prompt.submit, not prompt.section: only prompt.submit's event carries the human's actual
438  // words (`e.text`). This handler mutates session state only (never the model-visible prompt),
439  // so it always passes `event` through to `next` unchanged.
440  add('prompt.submit', async (io: any, event: any, next: any) => guardBefore('prompt.submit:nlrules', io, event, next, async (next) => {
441    const text = String(event?.text ?? '');
442    // Only a human's own prompt may create a denying rule: a peer, plugin, scheduled trigger or
443    // notification saying "don't run pytest" is not the user asking. Same origin rule the drift
444    // note uses to decide who spoke (`userSpoke`, hooks/drift.ts).
445    const instruction = userSpoke(event) ? parseStopInstruction(text) : null;
446    if (instruction) {
447      addSessionRule(state, instruction);
448      const path = `${(await harnessHome(io))}/pending-session-rules.json`;
449      await io.fs.write(path, JSON.stringify(state.sessionRules, null, 2));
450    }
451    if (text) {
452      const lower = text.toLowerCase();
453      for (const [key, snippet] of [...(state.rejected ?? new Map<string, string>()).entries()]) {
454        if (snippet && lower.includes(snippet.toLowerCase())) {
455          state.rejected!.delete(key);
456        }
457      }
458    }
459    return next(event);
460  }));
461
462  add('tool.check', async (io: any, event: any, next: any) => guardBefore('tool.check', io, event, next, async (next) => {
463    const tool = String(event?.tool ?? '');
464    const home = await harnessHome(io);
465    const sessionId = await currentSessionId(io);
466    if (wasRejected(state, tool, event?.input)) {
467      // The user already said no: never held out, but still receipted for reporting.
468      await writeReceipt(io, home, sessionId, {
469        event: 'tool.check', source: 'rejection-memory', artifact: null, signature: null, tool, decision: 'acted',
470      });
471      return {
472        decision: 'deny',
473        reason: 'This exact call was already rejected earlier this session [rejection-memory]',
474      };
475    }
476    for (const rule of state.sessionRules ?? []) {
477      if (rule.tool !== tool) continue;
478      // Bash rules match the command string, tokenised on whitespace — never a substring of the
479      // whole JSON input, or a rule for "git push" would also fire on `git status` (which merely
480      // contains the token "git") or on an unrelated call whose description happens to mention
481      // the pattern. Non-Bash rules (Edit-like verbs) have no `command` field to tokenise, so
482      // they fall back to the same JSON-substring check as before.
483      const command = tool === 'Bash' ? String((event?.input as any)?.command ?? '') : '';
484      const commandTokens = command.split(/\s+/).filter(Boolean).map((t) => t.toLowerCase());
485      const patternTokens = rule.pattern.toLowerCase().split(/\s+/).filter(Boolean);
486
487      // The pattern is matched against the command's first words AFTER any launcher, so a rule
488      // about `pytest` also covers `python -m pytest`, `uv run pytest`, `FOO=1 pytest`, ...
489      const programTokens = stripLaunchers(commandTokens);
490      let matches: boolean;
491      if (tool === 'Bash') {
492        matches = patternTokens.length > 0
493          && patternTokens.every((t, i) => programTokens[i] === t);
494      } else {
495        const haystack = JSON.stringify(event?.input ?? {}).toLowerCase();
496        matches = haystack.includes(rule.pattern.toLowerCase());
497      }
498      if (!matches) continue;
499
500      // A "without <flag>" qualifier IS representable: the rule allows the call when that flag
501      // is present, so "stop running pytest without -q" denies `pytest tests/` but ALLOWS
502      // `pytest -q` — the exact command the human asked to keep, not the command they asked to
503      // stop. Matched as a whole token of the command string, not a substring of the JSON input
504      // (a substring match would let "--quick" satisfy a "-q" requirement it does not).
505      if (rule.requires) {
506        const required = String(rule.requires).toLowerCase();
507        const present = tool === 'Bash'
508          ? commandTokens.includes(required)
509          : JSON.stringify(event?.input ?? {}).toLowerCase().includes(required);
510        if (present) continue;
511      }
512
513      // The opposite polarity: a flag named directly in the object ("git push --force") means
514      // the rule denies ONLY calls that also carry that flag as a whole token — a plain
515      // "git push origin main" must stay allowed.
516      if (rule.flag) {
517        const flag = String(rule.flag).toLowerCase();
518        const present = tool === 'Bash'
519          ? commandTokens.includes(flag)
520          : JSON.stringify(event?.input ?? {}).toLowerCase().includes(flag);
521        if (!present) continue;
522      }
523
524      // The user's direct order: never held out, but still receipted for reporting.
525      await writeReceipt(io, home, sessionId, {
526        event: 'tool.check', source: 'session-rule', artifact: null, signature: null, tool, decision: 'acted',
527      });
528      return { decision: 'deny', reason: sessionRuleReason(tool, rule) };
529    }
530    if (rules === null) rules = await loadRules(io, home);
531    for (const rule of rules) {
532      const verdict = evaluateRule(rule, event, state);
533      if (!verdict.deny) continue;
534      // A learned rule may be held out: a held rule behaves exactly as if it were not installed.
535      const held = shouldHold(await holdoutRate(io, home));
536      await writeReceipt(io, home, sessionId, {
537        event: 'tool.check', source: 'learned-rule', artifact: rule.artifactId,
538        signature: rule.signature, tool, decision: held ? 'held' : 'acted',
539      });
540      if (held) continue;
541      return { decision: 'deny', reason: verdict.reason };
542    }
543    // Counted here, before the call proceeds, so a `repeat-call` rule sees how many times this
544    // exact call has already been allowed. Counting at tool.check rather than after the result
545    // keeps the observer the only post-`next` handler on this path.
546    const key = callKey(event);
547    state.callCounts.set(key, (state.callCounts.get(key) ?? 0) + 1);
548    return next(event);
549  }));
550}
551
552type Injection = { artifactId: string; triggers: string[]; text: string; signature: string };
553
554/**
555 * Per-artifact injected-text cap, matching the `reason` cap the rules layer already applies
556 * (rules.ts:~50). An `injection` is paid in standing tokens on every turn it fires, unlike a
557 * `rule`, which is exactly why it sits below `rule` in the layer ordering — an uncapped
558 * injection would out-cost the stronger, free layer above it.
559 */
560const INJECTION_TEXT_CAP = 400;
561
562/**
563 * Cap on the joined text of ALL matched injections for one turn, so N installed injections
564 * cannot add up past a bound even though each is individually capped. Set to three artifacts'
565 * worth of INJECTION_TEXT_CAP: enough for a few unrelated matches to coexist, not enough for an
566 * unbounded number of installs to dominate the prompt.
567 */
568const INJECTION_TOTAL_CAP = INJECTION_TEXT_CAP * 3;
569
570const TRUNCATION_MARKER = '… [truncated]';
571
572/** Truncates visibly rather than silently, so a capped injection cannot be mistaken for a short one. */
573function truncate(text: string, limit: number): string {
574  return text.length <= limit ? text : `${text.slice(0, limit)}${TRUNCATION_MARKER}`;
575}
576
577/** Installed `injection` artifacts only: a `rule`, `skill` or `doctrine` row is never surfaced here. */
578async function loadInjections(io: any, home: string): Promise<Injection[]> {
579  const artifacts = await loadInstalled(io, home, 'injection');
580  return artifacts.map((artifact) => ({
581    artifactId: String(artifact.id),
582    triggers: (artifact.origin as any)?.triggers ?? [],
583    text: truncate(String(artifact.payload ?? ''), INJECTION_TEXT_CAP),
584    signature: String(artifact.signature ?? ''),
585  }));
586}
587
588/**
589 * Injects installed `injection` artifacts as model-only context on `prompt.submit`: the layer
590 * below `rule`, paid in standing tokens only on the turns where it actually fires.
591 *
592 * This used to live on `prompt.section` and return `{ text }` to replace a system-prompt
593 * section — with no section-name filter, that both matched triggers against the WRONG text (the
594 * system prompt's own section content, not the human's turn) and, the moment one injection
595 * artifact was ever installed, deleted or overwrote every other section of the system prompt on
596 * every turn it didn't match. `prompt.submit` fixes both: triggers match `e.text` (the human's
597 * actual words) and a match is ATTACHED via `next({ ...e, context: [...] })`, never a
598 * replacement. The hook itself knows what it attached, so use is recorded by construction via
599 * `observe` instead of asking the model to self-report a retrieval it might forget.
600 */
601export function registerInjection(add: On): void {
602  let injections: Injection[] | null = null;
603
604  add('prompt.submit', async (io: any, event: any, next: any) => guardBefore('prompt.submit:injection', io, event, next, async (next) => {
605    const home = await harnessHome(io);
606    if (injections === null) injections = await loadInjections(io, home);
607    if (injections.length === 0) return next(event);
608
609    const haystack = String(event?.text ?? '').toLowerCase();
610    const matched = injections.filter((injection) =>
611      injection.triggers.some((trigger) => haystack.includes(String(trigger).toLowerCase())));
612    if (matched.length === 0) return next(event);
613
614    const sessionId = await currentSessionId(io);
615    // Each matched injection draws independently; a held one behaves as if it were not installed.
616    const acted: Injection[] = [];
617    for (const injection of matched) {
618      const held = shouldHold(await holdoutRate(io, home));
619      await writeReceipt(io, home, sessionId, {
620        event: 'prompt.submit', source: 'learned-injection', artifact: injection.artifactId,
621        signature: injection.signature, tool: '', decision: held ? 'held' : 'acted',
622      });
623      if (!held) acted.push(injection);
624    }
625    if (acted.length === 0) return next(event);
626    for (const injection of acted) {
627      await observe(io, sessionId, { kind: 'injected', artifactId: injection.artifactId });
628    }
629    const joined = truncate(acted.map((injection) => injection.text).join('\n\n'), INJECTION_TOTAL_CAP);
630    return next({ ...event, context: [...(event?.context ?? []), joined] });
631  }));
632}
633
634/**
635 * Says once per stretch, on the result of the tool call that crosses `min_calls`, that a stretch
636 * of tool calls has run long: the note rides in the result's `context` ("what the model reads
637 * after the tool's result and the user never sees"), so the model reads it mid-stretch, BEFORE
638 * the user would notice, not after the user has already spoken.
639 *
640 * Disabled unless `<harnessHome>/drift.json` explicitly enables it (see `hooks/drift.ts`); the
641 * ship gate refused every judge, so by default this counts calls and says nothing.
642 *
643 * `tool.call` uses `guardAfterMap`: `next` is called exactly once, up front, and the note is
644 * appended to a COPY of the answered outcome; a throw anywhere after `next` returns core's
645 * outcome untouched, never re-running the tool. A `{ deny }` outcome is never given context (its
646 * type forbids it): past the threshold the note waits for the next answered call. The latch is
647 * set once the note is decided: attached, or held out by the receipts holdout (a held note
648 * returns core's outcome unchanged and is still receipted).
649 *
650 * `prompt.submit` only resets the stretch when the user speaks (see `userSpoke`); it attaches
651 * nothing and passes the event through unchanged, under `guardBefore`, which calls `next` once.
652 * Nothing here listens on `prompt.section`, whose return replaces a system-prompt section.
653 */
654export function registerDrift(add: On): void {
655  let count = 0;
656  let warned = false;
657  let config: DriftConfig | null | undefined;
658
659  add('tool.call', async (io: any, event: any, next: any) => guardAfterMap('tool.call:drift', io, event, next, async (outcome: any) => {
660    count += 1;
661    if (warned) return outcome;
662    if (outcome === null || typeof outcome !== 'object' || 'deny' in outcome) return outcome;
663    const home = await harnessHome(io);
664    if (config === undefined) config = await loadDriftConfig(io, home);
665    if (!shouldWarn(count, config)) return outcome;
666    // The latch is set whether the note acts or is held: one decision per stretch either way.
667    warned = true;
668    const held = shouldHold(await holdoutRate(io, home));
669    await writeReceipt(io, home, await currentSessionId(io), {
670      event: 'tool.call', source: 'drift-note', artifact: null, signature: null,
671      tool: String(event?.tool ?? ''), decision: held ? 'held' : 'acted',
672    });
673    if (held) return outcome;
674    const context = Array.isArray(outcome.context) ? outcome.context : [];
675    return { ...outcome, context: [...context, driftNote(count)] };
676  }));
677
678  add('prompt.submit', async (io: any, event: any, next: any) => guardBefore('prompt.submit:drift', io, event, next, async (next) => {
679    if (userSpoke(event)) {
680      count = 0;
681      warned = false;
682      config = undefined; // re-read per stretch, so an edited drift.json takes effect
683    }
684    return next(event);
685  }));
686}
687
688/**
689 * Says once, on the first `prompt.submit` after install, what wrong-direction work has cost so
690 * far, then never again: `<harnessHome>/bootstrap.json` is the marker. Presence of that file
691 * (not any in-memory flag) is the only thing that gates the message, so it stays correct across
692 * every session after the first, not just within the process that happened to write it.
693 *
694 * The numbers come from the report `meta-harness waste` (with or without `--json`) writes to
695 * `<harnessHome>/waste.json` (`{ sessions, corrections: { calls_burned, ... }, ... }`, per
696 * `meta_harness/waste.py`'s `waste_report()`). This hook never runs that command and nothing
697 * runs it in the background: until someone has run `meta-harness waste` (or `/meta-harness:harness waste`)
698 * once, there is no report and no line. A report with `sessions: 0` is skipped, unmarked.
699 *
700 * Handles its own failures rather than relying on `safely`: this handler calls `next` in the
701 * MIDDLE, then does one more fallible thing afterward (writing the marker). Under the old
702 * `safely`, a marker write that threw, or `next(...)` itself rejecting, reached a catch that
703 * called `next` a SECOND time — submitting the human's prompt twice, the second time with no
704 * line attached (every turn, on a read-only harness home). `safely` no longer does that either,
705 * but here the rule is explicit: `next` is called exactly once on every path, a failure before
706 * it falls back to the unmodified event, and a failed marker write is swallowed locally.
707 *
708 * The marker is written ONLY after the line has actually been attached via `next(...)` and that
709 * call has resolved: a missing `waste.json` means "no report yet, try again next turn", not
710 * "never again", and a transient read/parse failure before `next` is swallowed locally (falling
711 * back to the unmodified event) rather than suppressing the line for good — either way nothing
712 * is marked done, so the question keeps being asked (at most once per turn, which costs nothing
713 * on a turn with no report to show) until it can actually be answered once.
714 */
715export function registerBootstrap(add: On): void {
716  add('prompt.submit', async (io: any, event: any, next: any) => {
717    let outboundEvent = event;
718    let markDone = false;
719
720    try {
721      const home = (await harnessHome(io));
722      const bootstrapPath = `${home}/bootstrap.json`;
723      if (!(await io.fs.exists(bootstrapPath))) {
724        const wastePath = `${home}/waste.json`;
725        if (await io.fs.exists(wastePath)) {
726          const waste: any = JSON.parse(await io.fs.read(wastePath));
727          const sessions = waste?.sessions;
728          const callsBurned = waste?.corrections?.calls_burned;
729          // A report over zero sessions (an empty or mis-pointed transcript store) says nothing
730          // worth saying once and for good: skip it and leave the marker unwritten.
731          if (typeof sessions === 'number' && sessions > 0 && callsBurned != null) {
732            // Hedged on purpose: the count comes from a keyword heuristic whose hand-labelled
733            // precision is about half, so the line is an estimate, never a statement of fact.
734            const line = `Across ${sessions} past sessions, an estimated ~${callsBurned} tool calls may have gone `
735              + 'to work you later corrected (a rough heuristic, roughly half of its flags are genuine). '
736              + '`/meta-harness:harness waste` for the breakdown and how reliable it is.';
737            outboundEvent = { ...event, context: [...(event?.context ?? []), line] };
738            markDone = true;
739          }
740        }
741      }
742    } catch (error) {
743      logSkip(io, 'prompt.submit:bootstrap', error);
744      outboundEvent = event; // fail open: send the turn through exactly as it arrived
745      markDone = false;
746    }
747
748    // Called exactly once, unconditionally, on every path above — the one and only next() call
749    // this handler ever makes.
750    const result = await next(outboundEvent);
751
752    if (markDone) {
753      try {
754        await io.fs.write(`${(await harnessHome(io))}/bootstrap.json`, JSON.stringify({ shown: new Date().toISOString() }));
755      } catch (error) {
756        // Swallowed locally, never recovered by calling next() again: a failed marker write
757        // just means the question is asked again next turn, not that the turn breaks or the
758        // prompt is submitted twice.
759        logSkip(io, 'prompt.submit:bootstrap', error);
760      }
761    }
762    return result;
763  });
764}
765
766type AnyHook = (io: any, event: any, next: (e: any) => Promise<any>) => Promise<any>;
767
768/**
769 * Runs `hooks` as one chain, first outermost, the way separate registrations would nest: each
770 * hook's `next` is the rest of the chain, and the last one's `next` is core.
771 */
772export function chain(hooks: AnyHook[], io: any, event: any, next: (e: any) => Promise<any>): Promise<any> {
773  const run = (index: number, e: any): Promise<any> =>
774    index === hooks.length ? next(e) : hooks[index](io, e, (inner: any) => run(index + 1, inner));
775  return run(0, event);
776}
777
778/**
779 * Registers each event ONCE. Claude Code's loader refuses a module that registers one event twice
780 * without a matcher ("on(\"tool.call\") is registered twice without a matcher"), so the features
781 * above add their hooks to a per-event list here and each event gets a single literal that runs
782 * that list in the order the features were added (observer, rules, injection, bootstrap, drift).
783 */
784export const register: Register = (on: On, options: PluginOptions) => {
785  void options;
786  const hooks: Record<string, AnyHook[]> = {};
787  const add = ((event: string, hook: AnyHook) => {
788    (hooks[event] ??= []).push(hook);
789  }) as unknown as On;
790  registerObserver(add);
791  registerRules(add);
792  registerInjection(add);
793  registerBootstrap(add);
794  registerDrift(add);
795  on('tool.call', async ($: any, event: any, next: any) => chain(hooks['tool.call'] ?? [], {
796    fs: {
797      read: (path: string) => $.fs.read(path),
798      write: (path: string, text: string) => $.fs.write(path, text),
799      exists: (path: string) => $.fs.exists(path),
800    },
801    // $.env.get takes a literal name (the loader lists what a module reads), so one call per name.
802    env: {
803      get: (name: string) => (name === 'META_HARNESS_HOME' ? $.env.get('META_HARNESS_HOME')
804        : name === 'USERPROFILE' ? $.env.get('USERPROFILE')
805        : name === 'HOME' ? $.env.get('HOME')
806        : Promise.resolve(undefined)),
807    },
808    session: { id: () => $.session.id() },
809    ui: { log: (text: string) => $.ui.log(text) },
810  }, event, next));
811  on('tool.check', async ($: any, event: any, next: any) => chain(hooks['tool.check'] ?? [], {
812    fs: {
813      read: (path: string) => $.fs.read(path),
814      write: (path: string, text: string) => $.fs.write(path, text),
815      exists: (path: string) => $.fs.exists(path),
816    },
817    // $.env.get takes a literal name (the loader lists what a module reads), so one call per name.
818    env: {
819      get: (name: string) => (name === 'META_HARNESS_HOME' ? $.env.get('META_HARNESS_HOME')
820        : name === 'USERPROFILE' ? $.env.get('USERPROFILE')
821        : name === 'HOME' ? $.env.get('HOME')
822        : Promise.resolve(undefined)),
823    },
824    session: { id: () => $.session.id() },
825    ui: { log: (text: string) => $.ui.log(text) },
826  }, event, next));
827  on('prompt.submit', async ($: any, event: any, next: any) => chain(hooks['prompt.submit'] ?? [], {
828    fs: {
829      read: (path: string) => $.fs.read(path),
830      write: (path: string, text: string) => $.fs.write(path, text),
831      exists: (path: string) => $.fs.exists(path),
832    },
833    // $.env.get takes a literal name (the loader lists what a module reads), so one call per name.
834    env: {
835      get: (name: string) => (name === 'META_HARNESS_HOME' ? $.env.get('META_HARNESS_HOME')
836        : name === 'USERPROFILE' ? $.env.get('USERPROFILE')
837        : name === 'HOME' ? $.env.get('HOME')
838        : Promise.resolve(undefined)),
839    },
840    session: { id: () => $.session.id() },
841    ui: { log: (text: string) => $.ui.log(text) },
842  }, event, next));
843};
844
hooks/rules.ts 352 lines
1/**
2 * Installed `rule` artifacts, applied at tool.check.
3 *
4 * A rule is the strongest layer available: it costs no standing tokens and cannot be talked
5 * around. There is no built-in read-before-edit rule here: a live test proved Claude Code's own
6 * engine already refuses an Edit/Write of a file that has not been read this session ("File has
7 * not been read yet"), and it runs before this plugin ever sees the call. A plugin-side copy of
8 * that check only added failure modes on top of an enforcement that already existed -- every
9 * Edit denied, every new-file Write denied, and a Write-then-Edit denied even though the Write
10 * itself supplied the read the plugin's tracker never recorded. This module now only carries
11 * rules the engine does not already enforce (installed artifacts, `repeat-call`).
12 */
13
14export type Rule = {
15  artifactId: string;
16  kind: string;
17  tools: string[];
18  reason: string;
19  threshold?: number;
20  /** The failure signature from the artifact's `installed.json` row ('' when absent), for receipts. */
21  signature: string;
22};
23
24export type SessionRule = {
25  tool: string;
26  /** The command PREFIX the instruction names, e.g. "git push" from "stop using git push
27   * --force", or "pytest" from "don't run pytest". Matched against the START of the tokenised
28   * command string (whole tokens), never a substring of the whole JSON input — otherwise a
29   * rule for "git push" would also fire on an unrelated call that merely mentions "git push"
30   * inside a description or a file path. */
31  pattern: string;
32  /** The one flag/token that, if PRESENT in the call, means the rule allows it. Captures a
33   * "without <flag>" qualifier ("stop running pytest without -q" denies pytest unless the call
34   * also contains "-q"): the qualifier is representable, so it narrows the rule instead of
35   * being dropped into a blanket denial. Absent when the human's qualifier (if any) could not
36   * be represented this way, in which case the rule denies its pattern unconditionally. */
37  requires?: string;
38  /** The one flag/token that, if PRESENT in the call, means the rule DENIES it — the opposite
39   * polarity of `requires`. Captures a flag named directly in the instruction's object ("stop
40   * using git push --force" denies calls that start with "git push" AND carry "--force" as a
41   * whole token, but allows a plain "git push origin main"). Mutually exclusive with
42   * `requires` in practice, since an instruction supplies at most one qualifier. */
43  flag?: string;
44};
45
46export type SessionState = {
47  /** How many times each exact tool+input call has already been allowed this session. */
48  callCounts: Map<string, number>;
49  /** Calls the human has explicitly rejected this session, keyed by callKey, valued by a
50   * searchable snippet of the input so a later mention of the same command can clear it. */
51  rejected?: Map<string, string>;
52  /** "Stop doing X" rules added from the human's own words, in effect for this session only. */
53  sessionRules?: SessionRule[];
54};
55
56/** Keys a `tool.call` event carries beside the tool's own arguments (ToolCallReserved, AgentLoop). */
57const RESERVED_TOOL_CALL_KEYS = new Set(['tool', 'tool_use_id', 'agentId', 'consent']);
58
59/**
60 * The tool's own arguments on a `tool.call` event. On `tool.call` Claude Code puts them at the TOP
61 * LEVEL of the event (`e.command` for Bash, `e.file_path` for Read/Edit) beside the reserved keys;
62 * only `tool.check` nests them under `input`. An object `input` is still honoured, for fakes and in
63 * case a future engine nests them. Every `tool.call` reader goes through this, never `event.input`.
64 */
65export function toolArgs(event: unknown): Record<string, unknown> {
66  if (!event || typeof event !== 'object') return {};
67  const record = event as Record<string, unknown>;
68  if (record.input && typeof record.input === 'object' && !Array.isArray(record.input)) {
69    return record.input as Record<string, unknown>;
70  }
71  const args: Record<string, unknown> = {};
72  for (const [key, value] of Object.entries(record)) {
73    if (!RESERVED_TOOL_CALL_KEYS.has(key)) args[key] = value;
74  }
75  return args;
76}
77
78/** JSON with object keys sorted, so the same arguments give the same key whatever their order. */
79export function stableStringify(value: unknown): string {
80  return JSON.stringify(value, (_key, inner) => {
81    if (inner && typeof inner === 'object' && !Array.isArray(inner)) {
82      return Object.fromEntries(Object.keys(inner).sort().map((k) => [k, (inner as any)[k]]));
83    }
84    return inner;
85  });
86}
87
88/** The identity of one tool call, for counting exact repeats. Input order is whatever the
89 * engine sends; the same call in the same session serializes the same way, which is all the
90 * repeat counter needs. */
91export function callKey(event: { tool?: string; input?: Record<string, unknown> }): string {
92  let input = '';
93  try {
94    input = stableStringify(event.input ?? {});
95  } catch {
96    input = String(event.input ?? '');
97  }
98  return `${String(event.tool ?? '')}:${input.slice(0, 400)}`;
99}
100
101/**
102 * Default repeat count at which a `repeat-call` rule denies. Mirrors
103 * meta_harness.replay.THRASH_THRESHOLD, so the rule that enforces a thrash fix and the replay
104 * expectation that verifies it are talking about the same number.
105 */
106export const DEFAULT_REPEAT_THRESHOLD = 4;
107
108/** No built-in rules ship: read-before-edit is the engine's job now (see the module doc above),
109 * and nothing else has earned a built-in yet. Kept as an array (rather than removed outright) so
110 * `loadRules` has one shape to spread regardless of how many built-ins eventually exist. */
111export const BUILT_IN_RULES: Rule[] = [];
112
113export type InstalledArtifact = {
114  id: string;
115  type: string;
116  origin?: Record<string, unknown>;
117  payload?: unknown;
118  /** Copied from the artifact's `installed.json` row; '' when that row has none. */
119  signature?: string;
120};
121
122/**
123 * Reads `installed.json`, keeps only the rows of the given `type`, and loads each one's
124 * `artifact.json`. The single reader every layer (`rule`, `injection`, and whichever of
125 * `skill`/`doctrine` follows) shares, so the fail-open semantics live in exactly one place: a
126 * missing or corrupt registry, or a missing/malformed artifact.json, yields fewer rows rather
127 * than throwing, and every caller gets that behaviour identically instead of re-deriving it.
128 */
129export async function loadInstalled(io: any, home: string, type: string): Promise<InstalledArtifact[]> {
130  const registry = `${home}/installed.json`;
131  if (!(await io.fs.exists(registry))) return [];
132  let entries: Array<{ id: string; type: string; signature?: unknown }> = [];
133  try {
134    entries = JSON.parse(await io.fs.read(registry));
135  } catch {
136    return [];
137  }
138  const out: InstalledArtifact[] = [];
139  for (const entry of entries) {
140    if (entry.type !== type) continue;
141    const path = `${home}/artifacts/${entry.id}/artifact.json`;
142    if (!(await io.fs.exists(path))) continue;
143    try {
144      const artifact = JSON.parse(await io.fs.read(path));
145      out.push({ ...artifact, signature: typeof entry.signature === 'string' ? entry.signature : '' });
146    } catch {
147      continue;
148    }
149  }
150  return out;
151}
152
153/** Installed rule artifacts, plus the built-ins. Missing or malformed files are ignored. */
154export async function loadRules(io: any, home: string): Promise<Rule[]> {
155  const rules = [...BUILT_IN_RULES];
156  const artifacts = await loadInstalled(io, home, 'rule');
157  for (const artifact of artifacts) {
158    rules.push({
159      artifactId: String(artifact.id),
160      // `origin.rule_kind` is what meta_harness.learn.propose_artifact emits for the matcher to
161      // dispatch on; `origin.kind` is the EPISODE kind ('tool_error'/'thrash'), kept for
162      // provenance and never a matcher name. Falling back to it keeps a hand-written artifact
163      // working, and an unknown kind simply never denies.
164      kind: (artifact.origin as any)?.rule_kind ?? (artifact.origin as any)?.kind ?? 'custom',
165      tools: (artifact.origin as any)?.tools ?? [],
166      reason: String(artifact.payload ?? '').slice(0, 400),
167      threshold: Number((artifact.origin as any)?.threshold) || undefined,
168      signature: String(artifact.signature ?? ''),
169    });
170  }
171  return rules;
172}
173
174/** Evaluates one rule against one tool.check event. `deny: true` means the call must not proceed. */
175export function evaluateRule(
176  rule: Rule,
177  event: { tool?: string; input?: Record<string, unknown> },
178  state: SessionState,
179): { deny: boolean; reason?: string } {
180  const tool = String(event.tool ?? '');
181  if (!rule.tools.includes(tool)) return { deny: false };
182
183  if (rule.kind === 'repeat-call') {
184    // The only condition a learned rule can decide from the call and the session state alone:
185    // this exact call has already been made enough times to be the thrash the artifact was born
186    // from. A first attempt is never blocked, so the rule cannot break work that is going fine.
187    const threshold = rule.threshold && rule.threshold > 1 ? rule.threshold : DEFAULT_REPEAT_THRESHOLD;
188    const seen = state.callCounts.get(callKey(event)) ?? 0;
189    if (seen >= threshold - 1) {
190      return {
191        deny: true,
192        reason: `${rule.reason} [${rule.artifactId}] (identical ${tool} call already made ${seen} times this session)`,
193      };
194    }
195    return { deny: false };
196  }
197
198  return { deny: false };
199}
200
201/** A short, human-searchable string standing in for a tool call's input: the command or path
202 * a person would actually type or say when talking about this call, falling back to its JSON. */
203function commandSnippet(input: unknown): string {
204  if (input && typeof input === 'object') {
205    const record = input as Record<string, unknown>;
206    if (typeof record.command === 'string') return record.command;
207    if (typeof record.file_path === 'string') return record.file_path;
208  }
209  try {
210    return JSON.stringify(input ?? '');
211  } catch {
212    return String(input ?? '');
213  }
214}
215
216/** Records that this exact call was rejected by the human this session. */
217export function rememberRejection(state: SessionState, tool: string, input: unknown): void {
218  if (!state.rejected) state.rejected = new Map<string, string>();
219  const key = callKey({ tool, input: input as Record<string, unknown> });
220  state.rejected.set(key, commandSnippet(input));
221}
222
223/** True when this exact call was rejected earlier this session and has not since been cleared. */
224export function wasRejected(state: SessionState, tool: string, input: unknown): boolean {
225  if (!state.rejected) return false;
226  const key = callKey({ tool, input: input as Record<string, unknown> });
227  return state.rejected.has(key);
228}
229
230/**
231 * A conservative leading-verb parse: text must OPEN with "stop"/"don't"/"never" AND be
232 * immediately followed by one of a fixed, small set of ACTION verbs — a gerund (running/using/
233 * calling/doing/touching/editing/deleting/pushing) OR the plain imperative (run/use/call/do/
234 * touch/edit/delete/push) — before anything counts as an instruction's object.
235 *
236 * Matching anywhere in the text, or treating the action verb as optional, both let ordinary
237 * prose through as if it were an instruction: "don't worry about it", "don't know why this
238 * fails", "never mind" and "Don't forget to update the README" all open with a trigger word but
239 * name no actionable target, and "stop" alone or "stop, that's wrong" have no object at all.
240 * Requiring one of these specific verbs, immediately after the trigger word, is what tells an
241 * instruction ("stop running X") apart from those. Accepting the bare imperative alongside the
242 * gerund is what lets "don't use git push --force", "never call the deploy script" and "don't
243 * run pytest" parse at all — an earlier version only accepted the -ing form and silently missed
244 * every plain-imperative instruction.
245 */
246const STOP_VERB =
247  /^\s*(?:stop|don'?t|never)\s+(running|using|calling|doing|touching|editing|deleting|pushing|run|use|call|do|touch|edit|delete|push)\s+(.+)/i;
248
249/** Verbs whose object is a file or a piece of code rather than a shell command. */
250const EDIT_LIKE_VERBS = new Set(['editing', 'touching', 'deleting', 'edit', 'touch', 'delete']);
251
252/** A determiner is never itself the object ("the deploy script" -> "deploy script"). */
253const LEADING_DETERMINER = /^(?:the|a|an)\s+(?=\S)/i;
254
255/** Function words — pronouns, determiners, catch-all nouns — that name no actual command, tool
256 * or file: "Stop doing that" and "stop using the" must deny nothing, not deny every call whose
257 * input happens to contain the English word "that" or "the". */
258const FUNCTION_WORDS = new Set([
259  'that', 'this', 'it', 'the', 'a', 'an', 'those', 'these', 'them', 'they', 'there', 'here',
260  'everything', 'something', 'anything', 'nothing', 'things', 'stuff', 'that\'s', 'it\'s',
261]);
262
263/** A plausible command name, flag or path: word/path characters only, and not a bare function
264 * word — "pytest", "git", "-q" and "deploy" all pass; "that" and "the" do not. */
265function isPlausibleObject(token: string): boolean {
266  if (!token) return false;
267  if (FUNCTION_WORDS.has(token.toLowerCase())) return false;
268  return /^[A-Za-z0-9\-][\w.\-/]*$/.test(token);
269}
270
271/** Words introducing a qualifier this mechanism cannot represent as a `requires` flag ("pytest
272 * UNLESS it's urgent"): the pattern is cut before the qualifier rather than including words that
273 * would make the pattern match nothing real, and `tool.check`'s denial reason says so. */
274const UNREPRESENTABLE_QUALIFIER = /\s+(?:unless|except|only if|if)\b/i;
275
276/** A "without <flag>" qualifier IS representable: it becomes `requires`, so `tool.check` can
277 * deny the pattern only when that flag is absent, instead of denying it in every form — denying
278 * `pytest -q` outright, the exact command the human asked to KEEP, was the wrong call. */
279const WITHOUT_QUALIFIER = /\s+without\s+(\S+)/i;
280
281/** The rule's object: the command PREFIX up to its first flag (a plausible command/path only —
282 * see `isPlausibleObject`), plus whichever one qualifier the human's phrasing carried:
283 *
284 * - a "without X" qualifier becomes `requires` (X must be PRESENT to ALLOW the call) — unchanged
285 *   from fix round 2.
286 * - a flag named directly in the object ("git push --force") becomes `flag` (X must be PRESENT
287 *   to DENY the call) — new in fix round 3, so "don't use git push --force" denies only calls
288 *   that both start with "git push" AND carry "--force", not every "git" command.
289 *
290 * When the object has no flag at all ("never run git push"), the whole object (up to any
291 * qualifier) is the prefix and the rule denies that subcommand unconditionally.
292 *
293 * Returns null when no plausible object can be found at all, or when the object is nothing but
294 * a bare flag with no leading subcommand token. */
295function parseObject(rest: string): { pattern: string; requires?: string; flag?: string } | null {
296  const cleaned = rest.trim().replace(LEADING_DETERMINER, '');
297
298  const withoutMatch = WITHOUT_QUALIFIER.exec(cleaned);
299  const requires = withoutMatch ? withoutMatch[1].replace(/["'.,?!]+$/g, '') : undefined;
300  const beforeQualifier = withoutMatch
301    ? cleaned.slice(0, withoutMatch.index)
302    : (cleaned.split(UNREPRESENTABLE_QUALIFIER)[0] ?? cleaned);
303
304  const rawTokens = beforeQualifier.trim().split(/\s+/).filter(Boolean);
305  if (rawTokens.length === 0) return null;
306  const tokens = rawTokens.map((token, i) =>
307    i === rawTokens.length - 1 ? token.replace(/["'.,?!]+$/g, '') : token);
308
309  if (!isPlausibleObject(tokens[0])) return null;
310
311  const prefixTokens: string[] = [];
312  let flag: string | undefined;
313  for (const token of tokens) {
314    if (token.length > 1 && token.startsWith('-')) {
315      flag = token;
316      break;
317    }
318    prefixTokens.push(token);
319  }
320  if (prefixTokens.length === 0) return null;
321
322  const pattern = prefixTokens.join(' ');
323  const result: { pattern: string; requires?: string; flag?: string } = { pattern };
324  if (requires) result.requires = requires;
325  else if (flag) result.flag = flag;
326  return result;
327}
328
329/** Parses a "stop doing X" / "don't run X again" instruction out of free text, or returns null
330 * for anything that is not unambiguously such an instruction (a question, a description, an
331 * acknowledgement with no actionable target, an object that is a pronoun/determiner/catch-all
332 * rather than a command, etc). The returned `pattern` is deliberately just the object's leading
333 * token (e.g. "pytest", not "pytest without -q"): a "without X" qualifier becomes `requires`
334 * instead (see `parseObject`), and any other unrepresentable qualifier is dropped rather than
335 * baked into a pattern that would deny nothing real — `tool.check`'s denial reason says so
336 * explicitly either way, so the human sees exactly what is actually blocked. */
337export function parseStopInstruction(text: string): SessionRule | null {
338  const match = STOP_VERB.exec(text ?? '');
339  if (!match) return null;
340  const verb = match[1].toLowerCase();
341  const object = parseObject(match[2] ?? '');
342  if (!object) return null;
343  const tool = EDIT_LIKE_VERBS.has(verb) ? 'Edit' : 'Bash';
344  return { tool, ...object };
345}
346
347/** Adds a session-scoped rule parsed from the human's own words. Never touches the installed store. */
348export function addSessionRule(state: SessionState, rule: SessionRule): void {
349  if (!state.sessionRules) state.sessionRules = [];
350  state.sessionRules.push(rule);
351}
352
hooks/drift.ts 73 lines
1/**
2 * Live drift surfacing: counts tool calls in the current stretch (the calls since the user last
3 * spoke) and, on the call that reaches `min_calls`, says so ONCE as model-only `context` on that
4 * tool's result, so the model reads it mid-stretch rather than after the user has spoken.
5 *
6 * Configured by `<harnessHome>/drift.json`, shaped
7 * `{"enabled": boolean, "min_calls": number, "judge": "knn"|"overlap"|"model"}`. A missing,
8 * empty, unreadable, unparseable or ill-shaped file means DISABLED. That is how the ship gate in
9 * `tools/tune_drift.py` is honoured: it refused every judge, so nothing writes an enabling file
10 * and the feature ships built, tested, and off.
11 *
12 * The live hook only has the call count to go on; the `judge` field names which offline judge
13 * justified turning it on and is validated, not executed here.
14 */
15
16export type DriftConfig = {
17  enabled: true;
18  min_calls: number;
19  judge?: 'knn' | 'overlap' | 'model';
20};
21
22const JUDGES = new Set(['knn', 'overlap', 'model']);
23
24/**
25 * The drift config, or null (disabled) for anything short of an explicit, well-formed
26 * `{"enabled": true, "min_calls": <positive number>}`. Never throws.
27 */
28export async function loadDriftConfig(io: any, home: string): Promise<DriftConfig | null> {
29  try {
30    const path = `${home}/drift.json`;
31    if (!(await io.fs.exists(path))) return null;
32    const raw = await io.fs.read(path);
33    if (typeof raw !== 'string' || raw.trim() === '') return null;
34    const parsed: any = JSON.parse(raw);
35    if (parsed === null || typeof parsed !== 'object' || Array.isArray(parsed)) return null;
36    if (parsed.enabled !== true) return null;
37    const minCalls = parsed.min_calls;
38    if (typeof minCalls !== 'number' || !Number.isFinite(minCalls) || minCalls < 1) return null;
39    if (parsed.judge !== undefined && !JUDGES.has(parsed.judge)) return null;
40    return { enabled: true, min_calls: minCalls, judge: parsed.judge };
41  } catch {
42    return null;
43  }
44}
45
46/** Whether a stretch of `count` calls is long enough to mention under `config`. */
47export function shouldWarn(count: number, config: DriftConfig | null): boolean {
48  return config !== null && config.enabled === true && count >= config.min_calls;
49}
50
51/** The one line the model reads when a stretch has run long. */
52export function driftNote(count: number): string {
53  return `meta-harness drift: ${count} tool calls have run since the user last spoke. `
54    + 'Before continuing, check the work still matches what was asked.';
55}
56
57/**
58 * Origins that mean the user themself spoke, which ends the stretch. A notification, peer
59 * message, schedule or plugin submission does not: the stretch it lands in keeps counting, and
60 * its once-per-stretch latch stays set. An absent or unrecognised
61 * origin is treated as the user speaking, the choice that produces fewer notes, not more.
62 */
63const NON_USER_ORIGINS = new Set([
64  'task-notification', 'scheduled-trigger', 'peer', 'peer-send-message', 'projects-relay',
65  'channel', 'coordinator', 'observer', 'observer-activity', 'auto-continuation', 'slack-ping',
66  'plugin',
67]);
68
69export function userSpoke(event: any): boolean {
70  const kind = event?.origin?.kind;
71  return !(typeof kind === 'string' && NON_USER_ORIGINS.has(kind));
72}
73
hooks/receipts.ts 54 lines
1/**
2 * Decision receipts (spec: docs/superpowers/specs/2026-09-29-decision-receipts-design.md).
3 * A receipt is written when the plugin acts - or, for a held-out case, would have acted - and
4 * the outcome is attributed offline by meta_harness/receipts.py. Helpers take `io`, never `$`.
5 */
6export type ReceiptFields = {
7  event: string; source: string; artifact: string | null;
8  signature: string | null; tool: string; decision: 'acted' | 'held';
9};
10
11const DEFAULT_RATE = 0.1;
12const MAX_RATE = 0.5;
13const calls = new Map<string, number>();
14let drawFn: () => number = Math.random;
15
16export function bumpCall(sessionId: string): number {
17  const next = (calls.get(sessionId) ?? 0) + 1;
18  calls.set(sessionId, next);
19  return next;
20}
21
22export function currentCall(sessionId: string): number {
23  return calls.get(sessionId) ?? 0;
24}
25
26export function setDraw(fn: () => number): void { drawFn = fn; }
27export function draw(): number { return drawFn(); }
28export function shouldHold(rate: number): boolean { return rate > 0 && draw() < rate; }
29
30export async function holdoutRate(io: any, home: string): Promise<number> {
31  try {
32    const path = `${home}/receipts.json`;
33    if (!(await io.fs.exists(path))) return DEFAULT_RATE;
34    const value = Number(JSON.parse(await io.fs.read(path))?.holdout_rate);
35    if (!Number.isFinite(value)) return DEFAULT_RATE;
36    return Math.min(MAX_RATE, Math.max(0, value));
37  } catch {
38    return DEFAULT_RATE;
39  }
40}
41
42export async function writeReceipt(io: any, home: string, sessionId: string,
43                                   record: ReceiptFields): Promise<void> {
44  try {
45    const path = `${home}/receipts-${sessionId}.jsonl`;
46    const line = JSON.stringify({ ts: new Date().toISOString(), session: sessionId,
47                                  call: currentCall(sessionId), ...record }) + '\n';
48    const prior = (await io.fs.exists(path)) ? await io.fs.read(path) : '';
49    await io.fs.write(path, prior + line);
50  } catch (error) {
51    try { io.log?.(`meta-harness: receipt not written: ${String(error)}`); } catch { /* fail open */ }
52  }
53}
54