SLOPSHOPPER

YWR Labs Harness

Claude Code platform guards, the adversarial code-review standard, and a docs-as-code scaffold — portable across repos and stacks

newagents
v0.64.0NOASSERTIONupdated 2026-10-09orcait-co/ywr-harness/plugins/ywr-harness
A shopper browsing a rack in a slop shop
README

YWR Labs Harness (plugin)

Portable Claude Code platform guards plus the adversarial code-review standard. Everything here is coupled to Claude Code's own hook payloads and runtime behavior — not to any repo's tech stack, directory layout, or conventions. That is why it ships as a plugin: the knowledge is identical in every repo, so it should be maintained in one place and consumed, not re-derived.

Defects in anything here are fixed in this repo, never patched in a consuming repo — see docs/adr/0010-harness-defects-fixed-in-canon.md. The sanctioned escape hatch for urgency is claude plugin disable, which is reversible and visible; a local fork is neither.

Install

Managed settings register the ywrlabs marketplace on every org machine (ADR 0010). Installing stays each member's choice:

/plugin install ywr-harness@ywrlabs

Everything here is namespaced

Plugin components resolve as ywr-harness:<name>. A bare name does not resolve:

/ywr-harness:verify   /ywr-harness:slice-close   /ywr-harness:harness-init   /ywr-harness:feedback   /ywr-harness:artifact-publish
Workflow({name: 'ywr-harness:adversarial-review', args: {...}})

This is enforced — manifest-gate.ps1 fails on a shipped instruction that names a component bare, in either the Workflow({name: ...}) or the backticked-slash form. The gate exists because a live session found exactly that defect in shipped text while every selftest passed: the suite calls the scripts directly and never goes through the host's component registry.

Platform guards (hooks)

HookEventContract
agent-model-warn.mjsPreToolUse (matcher Agent)Warn-only (ADR 0086). Speaks exactly when an Agent-tool spawn passes an explicit model naming the opus or fable family (opus, fable, or a full id such as claude-opus-5-5): the member sees a Korean banner, the model gets the org-guide rule (ADR 0109) — implementation and research workers run sonnet in every session, so the Agent-tool route is ywr-harness:worker with no per-call model; mechanical work runs haiku; opus · low is for review fan-out on an Opus session; any other opus/fable worker needs a demonstrable reason in the spawn's description; a per-call model overrides a pinned agent's frontmatter; ultracode does not lift the pins (ADR 0108). Never returns a permission decision, never blocks, and is silent on an omitted model — it names per-call overrides only; an unpinned type such as general-purpose spawned without a model inherits the session model, which the org guide's explicit-model rule governs, not this hook. No subagent_type is exempt: no plugin agent pins opus. The hook cannot see the session model.
config-change-audit.mjsConfigChangeVisibility only. Surfaces mid-session permission/hook self-modification. Never blocks.
directory-added-guard.mjsDirectoryAddedVisibility only by construction — the event carries no decision control and fires after the permission refresh. Its banner never reaches you: an /add-dir message is next-turn context for CLAUDE, an SDK register_repo_root one goes to the debug log. So it is English and tells the model that the added tree is editable and not to be assumed covered by this project's gates, which surfaces it contributes (.claude/skills, .claude/commands, .claude/agents, the two settings keys), and whether its instruction files (CLAUDE.md, .claude/CLAUDE.md, .claude/rules, CLAUDE.local.md) are in context — and asks it to tell you in one sentence (ADR 0097). It speaks on every mid-session add whose payload parses; an unusable directory gets a SCHEMA DRIFT banner instead.
session-start-githooks-nudge.mjsSessionStartSuggest-only (ADR 0029). Speaks exactly when the work tree carries .githooks/ and this clone's core.hooksPath is unset — the state where no git hook runs and nothing else says so until slice close or CI. Names the one-line fix; never sets it. A wired clone, a repo without .githooks/, and a deliberate foreign hooksPath are all silent.
session-start-model-route-guard.mjsSessionStart (matcher `startup\resume\fork`)Warn-only (ADR 0107, ADR 0130). Speaks exactly when a setting can move a model alias to an older model, or a CLAUDE_CODE_EFFORT_LEVEL overrides every worker effort pin: an ANTHROPIC_DEFAULT_*_MODEL, a CLAUDE_CODE_SUBAGENT_MODEL (any value), a full id in ANTHROPIC_MODEL / ANTHROPIC_DEFAULT_MODEL / a settings model, a non-empty modelOverrides, or a CLAUDE_CODE_USE_* cloud provider; or a non-empty CLAUDE_CODE_EFFORT_LEVEL, which beats an agent's frontmatter effort, a per-call Agent effort (both per the docs) and a Workflow agent() effort (measured on 2.1.295) — read from the environment and the user, project and local settings files. One Korean line naming each setting and where it is set; never model context, never a change. Silent otherwise. node session-start-model-route-guard.mjs --preflight is the environment-only half the eval runner uses to refuse a paid run (one line per finding, exit 1; preflight: clear, exit 0); it never reads the effort variable, which the runner's wrapper sets itself.
session-start-node-check.ps1SessionStart (matcher `startup\resume\fork`)Suggest-only. Speaks exactly when no node is found on PATH — the state in which the eight Node hooks cannot start and the host shows a hook error at session start, on a config change or added directory, on every Agent call and on every subagent stop, which reads like a defect in your repo. It names that cause, the fix (install Node.js LTS and restart Claude Code; winget / brew / the distro package), and the installed-but-invisible case (nvm, or Claude Code launched from a GUI/Dock that did not inherit the shell's PATH). Probes PATH entries by file, as the host launches an exec-form node: on Windows only node.exe counts (a node.cmd/.bat/.ps1 shim cannot be launched and is named in the notice); elsewhere an executable node file (a non-executable one is named). Found: byte-silent. Stays PowerShell — it detects node's absence — and calls no cmdlet. Never blocks.
session-start-scaffold-refresh-nudge.mjsSessionStartSuggest-only (ADR 0033). Speaks exactly when the work tree carries a ywr-harness scaffold whose TOOLCHAIN placements differ from the installed plugin's templates — the stale-vendor state ADR 0014 recorded as undetected. Byte comparison, EOL-insensitive (the seed .gitattributes makes CRLF/LF checkout variance legitimate); the placement map is read from init.ps1's own literals by a small lexer (anything it cannot read as a bare string constant or a literal table is EXTRACTION DRIFT), so no second copy exists to drift. Comparison is exact apart from line endings — letter case, a BOM, U+00AD and NUL all count. Names the count, the files (capped list, cap stated), and the remedy; writes nothing. Seeds are never compared; a marker-less post-commit is skipped exactly as the scaffold refuses it. The advice is DIRECTION-AWARE (ADR 0042) via the repo's .harness-version stamp (written by harness-init on every successful run): repo ahead of this install → update the plugin, harness-init forbidden; repo behind → refresh, direction stated as measured; same version → hand-edit named; no readable stamp → the direction-blind caveat on the human banner AND the model context (a canon working tree mid-slice, or a multi-writer repo refreshed by a newer plugin, is newer than the installed copy and a re-run would revert it). When the running copy is itself a superseded cache install (a session that outlived a plugin update — ADR 0039), the advice flips: same file list, but the basis is named STALE, /reload-plugins (or a restart) is instructed, and harness-init is forbidden from that session — the loaded skill would place the old templates. The registry probe is best-effort; any failure returns the normal nudge.
session-start-version-announce.mjsSessionStartAnnounce-once-per-version (ADR 0030). At the first session that loads a new plugin version it says so once — old → new, up to three bullets from CHANGELOG.md (the member release-notes canon, Korean), and the onboarding artifact's release-notes tab — then records the version in ~/.claude/ywr-harness/announced-version, the plugin's only user-scope write. A machine's very first run gets a one-time link-only welcome instead (ADR 0031) — "업데이트됨" is claimed only when a previous version was recorded. Steady state and downgrades are silent; manifest-gate.ps1 refuses a release whose top CHANGELOG entry does not match plugin.json.
subagent-telemetry.mjsSubagentStopAppends a per-agent JSONL ledger to <project>/.claude/telemetry/ — documented SubagentStop fields only (no model: the event carries none). Fail-open.

Hooks module (Claude Mods)

hooks/delegation-ledger.mjs is a function-hooks module, named in hooks.json under modules (ADR 0117, ADR 0118). The host runs it in its own environment — not on node — from Claude Code 2.1.287. It hooks agent.spawn, turn.step and turn.complete and is observe-only: it passes every event on unchanged, never rewrites a model or effort, never denies and never draws. Each finished model loop (a main-loop turn, an Agent-tool subagent, a Workflow agent() worker) becomes one row in its session's file, one of a ring of twenty under <project>/.claude/telemetry/delegations/ (slot-00.json..slot-19.json): the model, effort and usage of every request, the spawn's type and per-call model when the Agent tool started it, and the loop's duration. A new session takes a free slot or the oldest one; each file keeps its session's last 500 loops within 360 KiB, so the slot files never pass ~7 MiB per repo. A workflow worker shows as loop: "unspawned"; its steps show the model it inherited. No prompt, task description or answer is written. It writes only in a repo whose root holds .harness.json. Hosts below 2.1.287, disableAllHooks, --safe-mode and --bare drop it, and the nine hooks above keep running.

Adversarial review (workflow)

workflows/adversarial-review.js — lens finders → semantic dedupe → severity-gated skeptic verification. Invoke it as a workflow, or by name from the skill listing (a workflow's meta.whenToUse is what surfaces there; there is no separate skill file).

Workflow({ name: 'ywr-harness:adversarial-review', args: { scope: {
  files: [<the emitter's file list>], context: '<what the slice does>',
  invariants: [<from REVIEW.md>], gates_passed: '<the gate commands and results, verbatim>' } } })

scope is an object. A string scope still runs, but it loses sharding, the out-of-scope bucket, and skeptic #2's check against gates_passed, without saying so.

Lens defaults are deliberately repo-agnostic. Two knobs keep them that way:

argEffect
lensExtraA house-specific angle appended to every lens prompt. Use this for repo vocabulary — how tenancy isolation is implemented, naming rules, a particular determinism boundary.
lensesFull override, [{key, prompt}]. An array with no valid entry throws: zero lenses is not a review.

Prefer lensExtra. Redefining the lens set to add one angle means later improvements to the canonical defaults never reach that repo — the same outcome as a local fork.

root is optional. When omitted the "repo root" line is dropped from the prompts entirely and agents use the session's working directory; baking one repo's absolute path in as a default would be knowledge that is false everywhere else.

The scope block cites the consuming repo's own root REVIEW.md, seeded by /ywr-harness:harness-init (ADR 0054). The plugin's REVIEW.md is the canon repo's own review invariants, not a shared canon.

Local execution layer (git hooks)

/ywr-harness:harness-init places two hooks into a consuming repo and wires core.hooksPath conditionally — set when unset, left alone (byte-identical) when it already names .githooks in any spelling — the literal, ./, an absolute path, the MSYS /c/… form, a ~/ value, or a link's real target — refused when it points anywhere else, and a worktree-scoped value that outranks the local one is reported as such (ADR 0015).

HookScopeContract
.githooks/pre-commitstaged filesRuns the emitter's file-scoped gates. Whole-program gates (tests, typecheck) are deferred to CI and the deferral is reported. A failing gate blocks the commit.
.githooks/pre-pushadded lines of the pushed rangeRegex secret scan. Pre-existing secrets outside the range are not re-flagged; a false positive is exempted per line with harness:allow-secret.
.githooks/post-committhe commit just madeSlice retro gate (ADR 0017) — seven deterministic docs-drift checks, zero tokens, silent when clean. Advisory: it never blocks.

post-commit is placed under a third mode, GUARDED: written when absent, refreshed when the existing file carries the ywr-harness:post-commit marker, refused otherwise. It is the one hook filename repos commonly already use — this repo's own post-commit republishes the docs artifact — so TOOLCHAIN would destroy working automation while SEED would silently deny the retro to every repo that has one.

The retro's checks (DEP · MIGRATION · SPEC · BUILD · FEAT · UNMAPPED · DEADMAP) read their scope from the retro block in .harness.json. An empty list disables the checks it drives, and --coverage says so — silence must mean clean, never "not configured".

python scripts/harness/harness_retro.py main~3..HEAD   # whole slice — absorbs mid-slice splits
python scripts/harness/harness_retro.py --coverage     # unowned files + dead implements_in
SLICE_RETRO=0 git commit ...                           # skip once

core.hooksPath lives in .git/config, which is per-clone and never committed — so the scaffold wires exactly the machine it ran on. That gap is closed by reporting, not by wiring harder: harness_gates.py prints a hooks: line on every /ywr-harness:slice-close and CI run, so an unwired clone can never look identical to a wired one.

The emitter also checks any declared claude.ai Artifact (ADR 0032): the artifacts section of .harness.json names each url and title, and the emitter verifies the README carries the link and the title starts with the repo name (convention <repo> · <purpose>). The emitter only reports (artifact: ok | VIOLATION | none declared); the vendored CI is the enforcement point — it fails on artifact: VIOLATION. The gate cannot see claude.ai, so an undeclared Artifact is unchecked, and that state is reported rather than silent.

Both hooks degrade the way everything else here does. A missing python, awk, or emitter is a reported skip, never a silent pass — an unparsed emitter would otherwise read as "no gates matched", which is the hardest kind of gate failure to notice.

Status line (user scope)

pwsh -NoProfile -File "${CLAUDE_PLUGIN_ROOT}/statusline/install.ps1"     # add -DryRun to preview

Renders location · model · effort · ctx N%/SIZE · 5h N% · 7d N% · ywr-harness vX.Y.Z, dropping the model's (1M context) suffix — a session constant costing width on every render — in favour of how much of that window is actually gone.

The trailing segment is the plugin version installed on this machine, read from Claude Code's install registry (~/.claude/plugins/installed_plugins.json). It is deliberately the on-disk value, never the repo's or the marketplace's: with marketplace auto-update on (ADR 0026), disk moves ahead of a live session, and that visible mismatch is the "restart to apply" signal. A machine without the plugin installed renders no segment. The org-guide version segment this replaces is retired (ADR 0027) — guide drift detection is spec 0003 §7's job, where it always really lived.

A plugin cannot contribute the main status line: a plugin's settings.json supports only agent and subagentStatusLine. So this ships as canon plus an installer that writes the user's own ~/.claude/ (ADR 0016). Run it once per machine; it is not applied automatically, and nothing detects a machine that never ran it.

The installer separates its two effects because they carry different risk. The script is TOOLCHAIN, overwritten on every run. The statusLine setting follows ADR 0015's rule — written when absent, left alone when already ours, refused when it points anywhere else. settings.json is parsed and re-emitted so unrelated keys survive; a file that does not parse is a hard refusal that prints the snippet to add by hand.

A key the payload does not carry removes its segment — never 0%. Unmeasured and zero are different states, and rate-limit keys are legitimately absent on a session's first render, before the first API response. Context and quota use different threshold curves: 50% context is ordinary working state, 50% of a rate limit is already worth watching.

Upstream feedback (skill)

/ywr-harness:feedback <description> — send a defect or request to the canon (ADR 0064). The canon repo is private, so the report is filed as an issue on the PUBLIC dist repo orcait-co/ywr-harness, and the canon's session-start inbox lists every open dist issue (ADR 0085 — no label is set, because GitHub drops one set by a filer without push access). The skill drafts the body (running + registered plugin versions, claude --version, OS, owner/repo + .harness-version, the refresh nudge's verdict and init.ps1 -DryRun output quoted verbatim, git log --oneline -3 of every file a re-run would change, a dedupe fingerprint), shows it, and files it only after ONE confirmation — the reviewed file is what gets filed. It never includes file contents or diffs (the tracker is public; the canon asks in-thread), never runs harness-init, and without gh keeps the body and prints the by-hand URL (NOT FILED, exit 2).

Artifact publish (skill)

/ywr-harness:artifact-publish — republish a declared claude.ai Artifact safely (ADR 0068). A repo that commits a GENERATED Artifact source declares source + check on its artifacts.items[] entry; the emitter enforces that the source exists and rides the drift check through CI, and this skill closes the loop: list the declared items, run each check, ONE confirmation, then the Artifact-tool publish with the enforced read-before-republish sequence and a byte-level lockstep proof against the committed file. Never headless (headless sessions have no Artifact tool — measured), never from a hook, never past a VIOLATION or a failed check; an ownership refusal is reported, not worked around.

Apply-now update (skill)

/ywr-harness:update — for the member who does not want to wait for the background auto-update (ADR 0026 converges every machine by its next session start). It reads the plugin's marketplace version and scope from claude plugin list --json, runs the two CLI commands with --json (claude plugin marketplace update, then claude plugin update -s <scope>), reports the on-disk old → new from the result line, falls back to the text forms on a host without --json, never accepts a marketplace-declared command (-y, --accept-command), and hands off the two steps no skill can perform: /reload-plugins (a REPL built-in — the remedy for the update CLI's own "restart required to apply"), and /ywr-harness:harness-init only when the ADR 0033 nudge reports drifted toolchain files. It never runs the scaffold refresh itself — the byte comparison alone cannot prove direction (the .harness-version stamp orients the nudge's advice since ADR 0042, but a stampless repo stays ambiguous), and an automatic re-run could revert a working tree deliberately newer than the installed plugin (ADR 0034).

CI-invoked assets

scripts/workflow-gates.mjs (parse + behavioral gate over a workflow corpus; --dir selects the corpus, default .claude/workflows) and scripts/resolve-base.sh (CI diff-range base resolution — since ADR 0043 vendored by the scaffold as scripts/harness/resolve-base.sh and invoked by the vendored workflow to resolve the PR base once, MODE=blocking). lib/selftest-lib.ps1 is the shared assertion core the selftests dot-source — dot-source it, do not run it.

Prerequisites

  • node — runs eight of the nine hooks (hooks/*.mjs, with the shared hooks/hook-lib.mjs; Node costs ~70 ms per hook call against ~780 ms for pwsh, ADR 0116) and the gates scripts/workflow-gates.mjs and the workflow corpus gate. Absent node: those hooks cannot start on a member machine — the host reports a hook error and never a blocked call, since no hook can block — and session-start-node-check (a PowerShell hook, so it can run without node) says so once at each session start, with the fix; the suites report a skip locally and a hard failure on CI, where they must actually run.
  • pwsh (PowerShell 7+) on PATH — session-start-node-check.ps1, the skills' scripts (harness-init, feedback) and every selftest. Without it the Node hooks still run, but the node notice itself shows a hook error naming pwsh at each session start.
  • python 3.9+ (stdlib only) — every gate script under scripts/ requires it: harness_gates.py (invoked by the pre-commit hook), harness_retro.py (by the post-commit hook), and verify_map.py (by ywr-harness:verify).

Hooks run wherever the session runs; the gates and selftests are also exercised on Linux pwsh, so nothing here may assume Windows.

Gates

selftest.ps1 is the single entry point — the manifest/wiring gate, the workflow corpus gate, then every shipped PowerShell selftest:

pwsh -NoProfile -File ./selftest.ps1

The two gates run first; the suites then run side by side (-Jobs N, default the CPU count capped at 8; -Jobs 1 runs them one at a time). Each finished suite prints one done line, then the results come back in discovery order: a failed suite in full, a passing suite as one line plus any SKIP lines. -Full prints every suite's whole output. To iterate on one component, run its own *.selftest.ps1 directly.

manifest-gate.ps1 is deterministic and CLI-free. It fails on the defects with distribution blast radius: version missing (omitted version falls back to the git commit SHA, so every commit ships as a new version to every consumer), a hook path that does not resolve, a regression from exec form to shell form where a path placeholder is used, and a release-notes canon that lies (top CHANGELOG.md entry ≠ plugin.json version, or the artifact link diverging between the CHANGELOG and the announce hook — ADR 0030). manifest-gate.selftest.ps1 proves it can fail — one or more mutations per check class (the suite prints the count) plus an unmutated control, because a suite that only ever passes and a gate that fails on everything score identically without the control. The control has earned its place: it caught a broken identity gate while the negative suite was reporting every mutation caught.

Eval suite (host-coupled behavioral probes)

evals/ is a claude plugin eval suite (Claude Code ≥ 2.1.269; ADR 0076, spec 0014). Where the selftests above are hermetic — they call the scripts and never the host — each eval case runs a fresh, isolated claude -p child with ONLY this plugin loaded and grades what came out: a SessionStart hook's additionalContext reaching the model, ywr-harness:verify triggering on natural phrasing, the five disable-model-invocation skills staying uninvoked, and ywr-harness:reviewer resolving by its namespaced name with a SubagentStop ledger line behind it. Every case pins model: to one worker model alias (opus, ADR 0103/0107), and manifest-gate.ps1 refuses a case that does not (or

Source 1 files
hooks/delegation-ledger.mjs 346 lines
1// Hooks module (Claude Mods, Claude Code 2.1.287+; ADR 0117, ADR 0118) — the OBSERVE-ONLY per-request
2// delegation ledger. `hooks.json` names this file under `modules`; the settings hooks beside it are
3// unaffected.
4//
5// What it records, per finished model loop (the main loop's turn, or one run of a subagent's loop):
6// every request the loop made (`turn.step`: the model the engine resolved, the effort it sends, the
7// usage the API reported), the loop's end (`turn.complete`: reason, duration, summed usage) and, for
8// a spawned agent, the spawn itself (`agent.spawn`: type, the per-call model as given, the parent
9// model, the resolved model). An agent-team teammate raises `agent.spawn` too from Claude Code 2.1.289
10// (`e.isTeammate`); its loop keeps the `agent-tool` kind and its spawn row says `teammate: true`, which
11// is what the reader splits on (schema 2 unchanged — the field is additive). A Workflow `agent()` worker
12// raises `agent.spawn` from Claude Code 2.1.292 with `e.workflow` ({runId, agentIndex}); its spawn row
13// carries `workflow: {run_id, agent_index}` and its loop is kind `workflow` (ADR 0126), and a row with
14// no `model_param` proves the call gave no model, so the worker inherited the session's. On 2.1.287 to
15// 2.1.291 such a worker raises no `agent.spawn` (fact 89): its loop is written with `spawn: null` and
16// `loop: 'unspawned'`, whose steps still name the model and effort it ran on. On 2.1.292+ `unspawned`
17// means an evicted spawn row.
18//
19// Observe-only, by contract: every hook passes the event on unchanged (`next(e)`, `yield* next(e)`),
20// never rewrites `model`/`effort`, never denies, never draws, and swallows its own failures — the
21// recording code runs after `next` resolved and inside try/catch, because a hook that fails
22// mid-stream is "left where it stood". It never carries a gate: hosts < 2.1.287, `disableAllHooks`,
23// `--safe-mode` and a managed `allowManagedModsOnly` all drop it (ADR 0117).
24//
25// What it never writes: a prompt, a task description, an answer or any other message text (the
26// secret-adjacent rule `subagent-telemetry` keeps). Where it writes (ADR 0118): a RING of twenty
27// session files, `<project root>/.claude/telemetry/delegations/slot-00.json`..`slot-19.json` —
28// gitignored by the scaffold and the canon — and only where the root holds `.harness.json` (a repo
29// that adopted the harness), so a session in a home directory or a stranger's checkout leaves nothing
30// behind. `$.fs` has no delete, so the bound is by construction: a session claims a missing slot or
31// the oldest by `mtimeMs`, and its file keeps its last 500 loops within 360 KiB. `$.fs.write` is not
32// atomic and replaces the whole file, so a session's writes are serialized, and a slot another
33// session took (its `mtimeMs`/`size` moved since this session's last write) is given up for a new one.
34//
35// The pure half (`createLedger`, `createSessionLog`, `pickSlot`, `stepRow`, `spawnRow`, `fingerprint`)
36// is exported so `delegation-ledger.selftest.ps1` drives it under plain node with a fake engine; the
37// engine imports only `register`.
38
39export const SCHEMA = 2
40export const LEDGER_DIR = '.claude/telemetry/delegations'
41export const RING = 20
42// One session file: its last MAX_LOOPS loops within MAX_FILE_BYTES (UTF-8, counted as an upper
43// bound). The ring's worst case is RING x MAX_FILE_BYTES, ~7 MiB per adopted repo.
44export const MAX_LOOPS = 500
45export const MAX_FILE_BYTES = 360 * 1024
46// Room left in MAX_FILE_BYTES for the file's own fields around `loops` (a session id is ~40 B).
47const HEADER_RESERVE = 1024
48// A loop keeps its first MAX_STEPS step rows; the rest are counted. Bounds one row (~67 B a step)
49// and the memory of an open loop.
50export const MAX_STEPS = 200
51// Open loops are held in module memory until their loop ends. A loop that never ends (killed agent,
52// crash) would otherwise leak; past this many, the oldest is dropped unwritten.
53export const MAX_OPEN = 256
54// Spawn rows outlive their loop (a resumed agent runs again under the same id), so they are capped
55// separately and higher (a row is ~300 B). An evicted row would make that agent's later run read as
56// `unspawned` — the workflow signal ADR 0117 Decision 6 counts on hosts before 2.1.292, and a plain
57// eviction after — so each record carries `spawn_rows_evicted`, the session's running count, and a
58// reader discounts `unspawned` when it is > 0.
59export const MAX_SPAWNS = 4096
60// Session loop lists held at once (a `/clear` starts a new session in the same module).
61export const MAX_SESSIONS = 4
62
63const MAIN = 'main'
64
65export function slotName(i) {
66  return `slot-${String(i).padStart(2, '0')}.json`
67}
68
69// The slot a new session claims from a `$.fs.list` of the ledger directory: the lowest-numbered slot
70// with no entry, else the regular slot file with the oldest `mtimeMs` (ties: the lowest number), else
71// null. A slot name that is a directory, a link or anything but a file is never claimed.
72export function pickSlot(entries, ring = RING) {
73  const byName = new Map()
74  for (const en of Array.isArray(entries) ? entries : []) if (en && typeof en.name === 'string') byName.set(en.name, en)
75  let oldest = null
76  for (let i = 0; i < ring; i++) {
77    const en = byName.get(slotName(i))
78    if (!en) return i
79    if (en.kind !== 'file' || en.isLink || typeof en.mtimeMs !== 'number') continue
80    if (oldest === null || en.mtimeMs < oldest.mtimeMs) oldest = { i, mtimeMs: en.mtimeMs }
81  }
82  return oldest ? oldest.i : null
83}
84
85// An upper bound on the UTF-8 length of a string: a code unit above U+007F counts 3 bytes (a
86// surrogate pair counts 6 for its 4).
87export function utf8Bound(s) {
88  let n = s.length
89  for (let i = 0; i < s.length; i++) if (s.charCodeAt(i) > 0x7f) n += 2
90  return n
91}
92
93function usageRow(u) {
94  if (!u || typeof u !== 'object') return null
95  return {
96    model: u.model ?? null,
97    input_tokens: u.input_tokens ?? 0,
98    output_tokens: u.output_tokens ?? 0,
99    cache_read_input_tokens: u.cache_read_input_tokens ?? 0,
100    cache_creation_input_tokens: u.cache_creation_input_tokens ?? 0,
101  }
102}
103
104export function spawnRow(e, r) {
105  const wf = e.workflow && typeof e.workflow === 'object' ? e.workflow : null
106  return {
107    tool_use_id: e.tool_use_id ?? null,
108    subagent_type: e.subagentType ?? null,
109    provider: e.provider ? `${e.provider.plugin}/${e.provider.tier}` : null,
110    model_param: e.model ?? null,
111    parent_model: e.parentModel ?? null,
112    model: r && typeof r === 'object' && !('deny' in r && r.deny) ? (r.model ?? null) : null,
113    denied: !!(r && typeof r === 'object' && r.deny),
114    fork: !!e.fork,
115    background: !!e.background,
116    teammate: !!e.isTeammate,
117    workflow: wf ? { run_id: wf.runId ?? null, agent_index: wf.agentIndex ?? null } : null,
118    parent_agent_id: e.parentAgentId ?? null,
119  }
120}
121
122// A step row is a TUPLE in STEP_FIELDS order (~67 B, against ~236 B as an object): the ledger is
123// read by analysis code, never by eye, and the file names the fields once, as `step_fields`. The
124// token counts are null when the request reported no usage; `usage_model` (the model the API
125// reported) is null when it equals `model` or no usage came back.
126export const STEP_FIELDS = ['index', 'model', 'effort', 'message_count', 'stop_reason', 'input_tokens',
127  'output_tokens', 'cache_read_input_tokens', 'cache_creation_input_tokens', 'usage_model']
128
129export function stepRow(e, r) {
130  const res = r && typeof r === 'object' ? r : null
131  const u = usageRow(res ? res.usage : null)
132  const model = e.model ?? null
133  return [
134    e.index ?? null,
135    model,
136    e.effort ?? null,
137    e.messageCount ?? null,
138    res ? (res.stopReason ?? null) : null,
139    u ? u.input_tokens : null,
140    u ? u.output_tokens : null,
141    u ? u.cache_read_input_tokens : null,
142    u ? u.cache_creation_input_tokens : null,
143    u && u.model !== model ? u.model : null,
144  ]
145}
146
147const present = x => x !== null && x !== undefined
148
149// The ledger's bookkeeping, engine-free. `spawn` / `step` / `complete` take the event and the result
150// `next` resolved to; `complete` answers the loop's row (empty steps when nothing was open).
151export function createLedger({ maxOpen = MAX_OPEN, maxSpawns = MAX_SPAWNS, maxSteps = MAX_STEPS } = {}) {
152  const spawns = new Map()   // agentId -> spawn row (kept across that agent's runs)
153  const open = new Map()     // `${agentId|main}\u0000${turnId}` -> { steps, dropped, models, efforts }
154  let evicted = 0
155  const cap = (map, max, onDrop) => { while (map.size > max) { map.delete(map.keys().next().value); onDrop?.() } }
156  const key = (agentId, turnId) => `${agentId ?? MAIN}\u0000${turnId}`
157  return {
158    spawn(e, r) {
159      const agentId = r && typeof r === 'object' ? r.agentId : undefined
160      if (!agentId) return
161      spawns.set(agentId, spawnRow(e, r))
162      cap(spawns, maxSpawns, () => { evicted++ })
163    },
164    step(e, r) {
165      const k = key(e.agentId, e.turnId)
166      if (!open.has(k)) { open.set(k, { steps: [], dropped: 0, models: new Set(), efforts: new Set() }); cap(open, maxOpen) }
167      const loop = open.get(k)
168      if (!loop) return
169      const row = stepRow(e, r)
170      if (present(row[1])) loop.models.add(row[1])
171      if (present(row[2])) loop.efforts.add(row[2])
172      if (loop.steps.length < maxSteps) loop.steps.push(row)
173      else loop.dropped++
174    },
175    complete(e, { iso }) {
176      const k = key(e.agentId, e.turnId)
177      const loop = open.get(k) ?? { steps: [], dropped: 0, models: new Set(), efforts: new Set() }
178      open.delete(k)
179      const spawn = e.agentId ? (spawns.get(e.agentId) ?? null) : null
180      return {
181        ts: iso,
182        loop: !e.agentId ? MAIN : !spawn ? 'unspawned' : spawn.workflow ? 'workflow' : 'agent-tool',
183        agent_id: e.agentId ?? null,
184        turn_id: e.turnId ?? null,
185        spawn,
186        spawn_rows_evicted: evicted,
187        steps: loop.steps,
188        steps_dropped: loop.dropped,
189        models: [...loop.models],
190        efforts: [...loop.efforts],
191        complete: {
192          reason: e.reason ?? null,
193          aborted: !!e.isAborted,
194          duration_ms: e.durationMs ?? null,
195          usage: usageRow(e.usage),
196        },
197      }
198    },
199    sizes: () => ({ spawns: spawns.size, open: open.size, evicted }),
200  }
201}
202
203// One session's loop rows, kept serialized and capped; `render` is the whole file's text.
204export function createSessionLog({ maxLoops = MAX_LOOPS, maxBytes = MAX_FILE_BYTES } = {}) {
205  const loops = []   // { text, bytes }
206  let bytes = 0
207  let dropped = 0
208  return {
209    add(row) {
210      const text = JSON.stringify(row)
211      const b = utf8Bound(text) + 1   // + its comma
212      loops.push({ text, bytes: b })
213      bytes += b
214      while (loops.length > maxLoops || (loops.length > 0 && bytes + HEADER_RESERVE > maxBytes)) {
215        bytes -= loops.shift().bytes
216        dropped++
217      }
218    },
219    // `header` is the file's own fields; `loops` goes last, spliced in as the rows' kept text.
220    render(header) {
221      const h = JSON.stringify({ ...header, loops_dropped: dropped, loops: [] })
222      return `${h.slice(0, -3)}[${loops.map(l => l.text).join(',')}]}\n`
223    },
224    sizes: () => ({ loops: loops.length, bytes, dropped }),
225  }
226}
227
228// --- the engine half ---------------------------------------------------------------------------
229// Declared at the top of the file: the engine follows `$` only into such functions.
230
231// The project root to write under, or null where the root holds no `.harness.json`. `adopted`
232// caches the answer per root (a `/cd` or worktree move changes the root mid-session).
233async function target($, adopted) {
234  const root = await $.session.root()
235  if (!adopted.has(root)) adopted.set(root, await $.fs.exists(`${root}/.harness.json`))
236  return adopted.get(root) ? root : null
237}
238
239async function statOf($, path) {
240  try {
241    const st = await $.fs.stat(path)
242    return { mtimeMs: st.mtimeMs, size: st.size }
243  } catch { return null }
244}
245
246// The start every file this session renders begins with (`render` puts the header keys first).
247const headerOf = id => `{"schema":${SCHEMA},"session_id":${JSON.stringify(id)},`
248
249// A 32-bit FNV-1a over the UTF-16 code units: the mark of the text this session last wrote whole.
250export function fingerprint(text) {
251  let h = 0x811c9dc5
252  for (let i = 0; i < text.length; i++) h = Math.imul(h ^ text.charCodeAt(i), 0x01000193)
253  return h >>> 0
254}
255
256// One rewrite of a session's file. Runs on the session's chain, never two at once for one session.
257async function flush($, s, root, iso) {
258  const dir = `${root}/${LEDGER_DIR}`
259  if (s.root !== root) { s.root = root; s.slot = null; s.seen = null; s.mark = null }
260  if (s.slot !== null && s.seen) {
261    const now = await statOf($, `${dir}/${slotName(s.slot)}`)
262    // moved since this session's last write: another session claimed it. A stat that fails (a missing
263    // slot among the causes) proves nothing either way, so the read-back below decides.
264    if (!now) s.seen = null
265    else if (now.mtimeMs !== s.seen.mtimeMs || now.size !== s.seen.size) s.slot = null
266  }
267  if (s.slot !== null && !s.seen) {
268    // the last write or a stat failed, so there is no stat to compare: the file itself says whose
269    // it is. A missing or empty file stays ours; so does the very text this session last wrote whole
270    // (its fingerprint, which covers a stat that keeps failing), and a file starting with our own
271    // header (a partial write of ours keeps it). A file that exists but cannot be read proves
272    // nothing, and neither does a null session id's header (every id-less session writes the same
273    // one), so both are given up.
274    const path = `${dir}/${slotName(s.slot)}`
275    let text = ''
276    try { text = await $.fs.read(path) } catch {
277      let there = true
278      try { there = await $.fs.exists(path) } catch { /* unknown: not provably ours */ }
279      if (there) text = null
280    }
281    const ours = text === '' || (text !== null && ((s.mark !== null && fingerprint(text) === s.mark) ||
282      (s.id !== null && text.startsWith(headerOf(s.id)))))
283    if (!ours) s.slot = null
284  }
285  if (s.slot === null) {
286    let entries = []
287    try { entries = await $.fs.list(dir) } catch { /* no directory yet: every slot is missing */ }
288    s.slot = pickSlot(entries)
289    if (s.slot === null) return
290  }
291  const path = `${dir}/${slotName(s.slot)}`
292  const text = s.log.render({ schema: SCHEMA, session_id: s.id, slot: s.slot, updated: iso, step_fields: STEP_FIELDS })
293  s.seen = null
294  s.mark = null
295  await $.fs.write(path, text)
296  s.mark = fingerprint(text)
297  s.seen = await statOf($, path)
298}
299
300export function register(on) {
301  const ledger = createLedger()
302  const adopted = new Map()    // project root -> whether it holds .harness.json
303  const sessions = new Map()   // session id -> { id, log, root, slot, seen, mark, chain }
304  const sessionOf = id => {
305    let s = sessions.get(id)
306    if (s) { sessions.delete(id); sessions.set(id, s) }   // most recently used last: eviction takes the idlest
307    else {
308      s = { id, log: createSessionLog(), root: null, slot: null, seen: null, mark: null, chain: Promise.resolve() }
309      sessions.set(id, s)
310      while (sessions.size > MAX_SESSIONS) sessions.delete(sessions.keys().next().value)
311    }
312    return s
313  }
314
315  on('agent.spawn', async ($, e, next) => {
316    const r = await next(e)
317    try { ledger.spawn(e, r) } catch { /* observe-only: never fail the spawn */ }
318    return r
319  })
320
321  on('turn.step', async function* ($, e, next) {
322    const r = yield* next(e)
323    try { ledger.step(e, r) } catch { /* observe-only */ }
324    return r
325  })
326
327  on('turn.complete', async ($, e, next) => {
328    const r = await next(e)
329    try {
330      // The loop leaves memory first: a failing engine call below costs this one record, never a leak.
331      const row = ledger.complete(e, { iso: null })
332      const id = (await $.session.id()) ?? null
333      row.ts = new Date(await $.clock.now()).toISOString()
334      const root = await target($, adopted)
335      if (root) {
336        const s = sessionOf(id)
337        s.log.add(row)
338        const run = s.chain.then(() => flush($, s, root, row.ts))
339        s.chain = run.catch(() => {})
340        await run
341      }
342    } catch { /* observe-only: a ledger write never reaches the turn */ }
343    return r
344  })
345}
346