Claude Code platform guards, the adversarial code-review standard, and a docs-as-code scaffold — portable across repos and stacks

Portable Claude Code platform guards plus the adversarial code-review standard. Everything here is coupled to Claude Code's own hook payloads and runtime behavior — not to any repo's tech stack, directory layout, or conventions. That is why it ships as a plugin: the knowledge is identical in every repo, so it should be maintained in one place and consumed, not re-derived.
Defects in anything here are fixed in this repo, never patched in a consuming repo — see docs/adr/0010-harness-defects-fixed-in-canon.md. The sanctioned escape hatch for urgency is claude plugin disable, which is reversible and visible; a local fork is neither.
Managed settings register the ywrlabs marketplace on every org machine (ADR 0010). Installing stays each member's choice:
/plugin install ywr-harness@ywrlabs
Plugin components resolve as ywr-harness:<name>. A bare name does not resolve:
/ywr-harness:verify /ywr-harness:slice-close /ywr-harness:harness-init /ywr-harness:feedback /ywr-harness:artifact-publish
Workflow({name: 'ywr-harness:adversarial-review', args: {...}})
This is enforced — manifest-gate.ps1 fails on a shipped instruction that names a component bare, in either the Workflow({name: ...}) or the backticked-slash form. The gate exists because a live session found exactly that defect in shipped text while every selftest passed: the suite calls the scripts directly and never goes through the host's component registry.
| Hook | Event | Contract | ||
|---|---|---|---|---|
agent-model-warn.mjs | PreToolUse (matcher Agent) | Warn-only (ADR 0086). Speaks exactly when an Agent-tool spawn passes an explicit model naming the opus or fable family (opus, fable, or a full id such as claude-opus-5-5): the member sees a Korean banner, the model gets the org-guide rule (ADR 0109) — implementation and research workers run sonnet in every session, so the Agent-tool route is ywr-harness:worker with no per-call model; mechanical work runs haiku; opus · low is for review fan-out on an Opus session; any other opus/fable worker needs a demonstrable reason in the spawn's description; a per-call model overrides a pinned agent's frontmatter; ultracode does not lift the pins (ADR 0108). Never returns a permission decision, never blocks, and is silent on an omitted model — it names per-call overrides only; an unpinned type such as general-purpose spawned without a model inherits the session model, which the org guide's explicit-model rule governs, not this hook. No subagent_type is exempt: no plugin agent pins opus. The hook cannot see the session model. | ||
config-change-audit.mjs | ConfigChange | Visibility only. Surfaces mid-session permission/hook self-modification. Never blocks. | ||
directory-added-guard.mjs | DirectoryAdded | Visibility only by construction — the event carries no decision control and fires after the permission refresh. Its banner never reaches you: an /add-dir message is next-turn context for CLAUDE, an SDK register_repo_root one goes to the debug log. So it is English and tells the model that the added tree is editable and not to be assumed covered by this project's gates, which surfaces it contributes (.claude/skills, .claude/commands, .claude/agents, the two settings keys), and whether its instruction files (CLAUDE.md, .claude/CLAUDE.md, .claude/rules, CLAUDE.local.md) are in context — and asks it to tell you in one sentence (ADR 0097). It speaks on every mid-session add whose payload parses; an unusable directory gets a SCHEMA DRIFT banner instead. | ||
session-start-githooks-nudge.mjs | SessionStart | Suggest-only (ADR 0029). Speaks exactly when the work tree carries .githooks/ and this clone's core.hooksPath is unset — the state where no git hook runs and nothing else says so until slice close or CI. Names the one-line fix; never sets it. A wired clone, a repo without .githooks/, and a deliberate foreign hooksPath are all silent. | ||
session-start-model-route-guard.mjs | SessionStart (matcher `startup\ | resume\ | fork`) | Warn-only (ADR 0107, ADR 0130). Speaks exactly when a setting can move a model alias to an older model, or a CLAUDE_CODE_EFFORT_LEVEL overrides every worker effort pin: an ANTHROPIC_DEFAULT_*_MODEL, a CLAUDE_CODE_SUBAGENT_MODEL (any value), a full id in ANTHROPIC_MODEL / ANTHROPIC_DEFAULT_MODEL / a settings model, a non-empty modelOverrides, or a CLAUDE_CODE_USE_* cloud provider; or a non-empty CLAUDE_CODE_EFFORT_LEVEL, which beats an agent's frontmatter effort, a per-call Agent effort (both per the docs) and a Workflow agent() effort (measured on 2.1.295) — read from the environment and the user, project and local settings files. One Korean line naming each setting and where it is set; never model context, never a change. Silent otherwise. node session-start-model-route-guard.mjs --preflight is the environment-only half the eval runner uses to refuse a paid run (one line per finding, exit 1; preflight: clear, exit 0); it never reads the effort variable, which the runner's wrapper sets itself. |
session-start-node-check.ps1 | SessionStart (matcher `startup\ | resume\ | fork`) | Suggest-only. Speaks exactly when no node is found on PATH — the state in which the eight Node hooks cannot start and the host shows a hook error at session start, on a config change or added directory, on every Agent call and on every subagent stop, which reads like a defect in your repo. It names that cause, the fix (install Node.js LTS and restart Claude Code; winget / brew / the distro package), and the installed-but-invisible case (nvm, or Claude Code launched from a GUI/Dock that did not inherit the shell's PATH). Probes PATH entries by file, as the host launches an exec-form node: on Windows only node.exe counts (a node.cmd/.bat/.ps1 shim cannot be launched and is named in the notice); elsewhere an executable node file (a non-executable one is named). Found: byte-silent. Stays PowerShell — it detects node's absence — and calls no cmdlet. Never blocks. |
session-start-scaffold-refresh-nudge.mjs | SessionStart | Suggest-only (ADR 0033). Speaks exactly when the work tree carries a ywr-harness scaffold whose TOOLCHAIN placements differ from the installed plugin's templates — the stale-vendor state ADR 0014 recorded as undetected. Byte comparison, EOL-insensitive (the seed .gitattributes makes CRLF/LF checkout variance legitimate); the placement map is read from init.ps1's own literals by a small lexer (anything it cannot read as a bare string constant or a literal table is EXTRACTION DRIFT), so no second copy exists to drift. Comparison is exact apart from line endings — letter case, a BOM, U+00AD and NUL all count. Names the count, the files (capped list, cap stated), and the remedy; writes nothing. Seeds are never compared; a marker-less post-commit is skipped exactly as the scaffold refuses it. The advice is DIRECTION-AWARE (ADR 0042) via the repo's .harness-version stamp (written by harness-init on every successful run): repo ahead of this install → update the plugin, harness-init forbidden; repo behind → refresh, direction stated as measured; same version → hand-edit named; no readable stamp → the direction-blind caveat on the human banner AND the model context (a canon working tree mid-slice, or a multi-writer repo refreshed by a newer plugin, is newer than the installed copy and a re-run would revert it). When the running copy is itself a superseded cache install (a session that outlived a plugin update — ADR 0039), the advice flips: same file list, but the basis is named STALE, /reload-plugins (or a restart) is instructed, and harness-init is forbidden from that session — the loaded skill would place the old templates. The registry probe is best-effort; any failure returns the normal nudge. | ||
session-start-version-announce.mjs | SessionStart | Announce-once-per-version (ADR 0030). At the first session that loads a new plugin version it says so once — old → new, up to three bullets from CHANGELOG.md (the member release-notes canon, Korean), and the onboarding artifact's release-notes tab — then records the version in ~/.claude/ywr-harness/announced-version, the plugin's only user-scope write. A machine's very first run gets a one-time link-only welcome instead (ADR 0031) — "업데이트됨" is claimed only when a previous version was recorded. Steady state and downgrades are silent; manifest-gate.ps1 refuses a release whose top CHANGELOG entry does not match plugin.json. | ||
subagent-telemetry.mjs | SubagentStop | Appends a per-agent JSONL ledger to <project>/.claude/telemetry/ — documented SubagentStop fields only (no model: the event carries none). Fail-open. |
hooks/delegation-ledger.mjs is a function-hooks module, named in hooks.json under modules (ADR 0117, ADR 0118). The host runs it in its own environment — not on node — from Claude Code 2.1.287. It hooks agent.spawn, turn.step and turn.complete and is observe-only: it passes every event on unchanged, never rewrites a model or effort, never denies and never draws. Each finished model loop (a main-loop turn, an Agent-tool subagent, a Workflow agent() worker) becomes one row in its session's file, one of a ring of twenty under <project>/.claude/telemetry/delegations/ (slot-00.json..slot-19.json): the model, effort and usage of every request, the spawn's type and per-call model when the Agent tool started it, and the loop's duration. A new session takes a free slot or the oldest one; each file keeps its session's last 500 loops within 360 KiB, so the slot files never pass ~7 MiB per repo. A workflow worker shows as loop: "unspawned"; its steps show the model it inherited. No prompt, task description or answer is written. It writes only in a repo whose root holds .harness.json. Hosts below 2.1.287, disableAllHooks, --safe-mode and --bare drop it, and the nine hooks above keep running.
workflows/adversarial-review.js — lens finders → semantic dedupe → severity-gated skeptic verification. Invoke it as a workflow, or by name from the skill listing (a workflow's meta.whenToUse is what surfaces there; there is no separate skill file).
Workflow({ name: 'ywr-harness:adversarial-review', args: { scope: {
files: [<the emitter's file list>], context: '<what the slice does>',
invariants: [<from REVIEW.md>], gates_passed: '<the gate commands and results, verbatim>' } } })
scope is an object. A string scope still runs, but it loses sharding, the out-of-scope bucket, and skeptic #2's check against gates_passed, without saying so.
Lens defaults are deliberately repo-agnostic. Two knobs keep them that way:
| arg | Effect |
|---|---|
lensExtra | A house-specific angle appended to every lens prompt. Use this for repo vocabulary — how tenancy isolation is implemented, naming rules, a particular determinism boundary. |
lenses | Full override, [{key, prompt}]. An array with no valid entry throws: zero lenses is not a review. |
Prefer lensExtra. Redefining the lens set to add one angle means later improvements to the canonical defaults never reach that repo — the same outcome as a local fork.
root is optional. When omitted the "repo root" line is dropped from the prompts entirely and agents use the session's working directory; baking one repo's absolute path in as a default would be knowledge that is false everywhere else.
The scope block cites the consuming repo's own root REVIEW.md, seeded by /ywr-harness:harness-init (ADR 0054). The plugin's REVIEW.md is the canon repo's own review invariants, not a shared canon.
/ywr-harness:harness-init places two hooks into a consuming repo and wires core.hooksPath conditionally — set when unset, left alone (byte-identical) when it already names .githooks in any spelling — the literal, ./, an absolute path, the MSYS /c/… form, a ~/ value, or a link's real target — refused when it points anywhere else, and a worktree-scoped value that outranks the local one is reported as such (ADR 0015).
| Hook | Scope | Contract |
|---|---|---|
.githooks/pre-commit | staged files | Runs the emitter's file-scoped gates. Whole-program gates (tests, typecheck) are deferred to CI and the deferral is reported. A failing gate blocks the commit. |
.githooks/pre-push | added lines of the pushed range | Regex secret scan. Pre-existing secrets outside the range are not re-flagged; a false positive is exempted per line with harness:allow-secret. |
.githooks/post-commit | the commit just made | Slice retro gate (ADR 0017) — seven deterministic docs-drift checks, zero tokens, silent when clean. Advisory: it never blocks. |
post-commit is placed under a third mode, GUARDED: written when absent, refreshed when the existing file carries the ywr-harness:post-commit marker, refused otherwise. It is the one hook filename repos commonly already use — this repo's own post-commit republishes the docs artifact — so TOOLCHAIN would destroy working automation while SEED would silently deny the retro to every repo that has one.
The retro's checks (DEP · MIGRATION · SPEC · BUILD · FEAT · UNMAPPED · DEADMAP) read their scope from the retro block in .harness.json. An empty list disables the checks it drives, and --coverage says so — silence must mean clean, never "not configured".
python scripts/harness/harness_retro.py main~3..HEAD # whole slice — absorbs mid-slice splits
python scripts/harness/harness_retro.py --coverage # unowned files + dead implements_in
SLICE_RETRO=0 git commit ... # skip once
core.hooksPath lives in .git/config, which is per-clone and never committed — so the scaffold wires exactly the machine it ran on. That gap is closed by reporting, not by wiring harder: harness_gates.py prints a hooks: line on every /ywr-harness:slice-close and CI run, so an unwired clone can never look identical to a wired one.
The emitter also checks any declared claude.ai Artifact (ADR 0032): the artifacts section of .harness.json names each url and title, and the emitter verifies the README carries the link and the title starts with the repo name (convention <repo> · <purpose>). The emitter only reports (artifact: ok | VIOLATION | none declared); the vendored CI is the enforcement point — it fails on artifact: VIOLATION. The gate cannot see claude.ai, so an undeclared Artifact is unchecked, and that state is reported rather than silent.
Both hooks degrade the way everything else here does. A missing python, awk, or emitter is a reported skip, never a silent pass — an unparsed emitter would otherwise read as "no gates matched", which is the hardest kind of gate failure to notice.
pwsh -NoProfile -File "${CLAUDE_PLUGIN_ROOT}/statusline/install.ps1" # add -DryRun to preview
Renders location · model · effort · ctx N%/SIZE · 5h N% · 7d N% · ywr-harness vX.Y.Z, dropping the model's (1M context) suffix — a session constant costing width on every render — in favour of how much of that window is actually gone.
The trailing segment is the plugin version installed on this machine, read from Claude Code's install registry (~/.claude/plugins/installed_plugins.json). It is deliberately the on-disk value, never the repo's or the marketplace's: with marketplace auto-update on (ADR 0026), disk moves ahead of a live session, and that visible mismatch is the "restart to apply" signal. A machine without the plugin installed renders no segment. The org-guide version segment this replaces is retired (ADR 0027) — guide drift detection is spec 0003 §7's job, where it always really lived.
A plugin cannot contribute the main status line: a plugin's settings.json supports only agent and subagentStatusLine. So this ships as canon plus an installer that writes the user's own ~/.claude/ (ADR 0016). Run it once per machine; it is not applied automatically, and nothing detects a machine that never ran it.
The installer separates its two effects because they carry different risk. The script is TOOLCHAIN, overwritten on every run. The statusLine setting follows ADR 0015's rule — written when absent, left alone when already ours, refused when it points anywhere else. settings.json is parsed and re-emitted so unrelated keys survive; a file that does not parse is a hard refusal that prints the snippet to add by hand.
A key the payload does not carry removes its segment — never 0%. Unmeasured and zero are different states, and rate-limit keys are legitimately absent on a session's first render, before the first API response. Context and quota use different threshold curves: 50% context is ordinary working state, 50% of a rate limit is already worth watching.
/ywr-harness:feedback <description> — send a defect or request to the canon (ADR 0064). The canon repo is private, so the report is filed as an issue on the PUBLIC dist repo orcait-co/ywr-harness, and the canon's session-start inbox lists every open dist issue (ADR 0085 — no label is set, because GitHub drops one set by a filer without push access). The skill drafts the body (running + registered plugin versions, claude --version, OS, owner/repo + .harness-version, the refresh nudge's verdict and init.ps1 -DryRun output quoted verbatim, git log --oneline -3 of every file a re-run would change, a dedupe fingerprint), shows it, and files it only after ONE confirmation — the reviewed file is what gets filed. It never includes file contents or diffs (the tracker is public; the canon asks in-thread), never runs harness-init, and without gh keeps the body and prints the by-hand URL (NOT FILED, exit 2).
/ywr-harness:artifact-publish — republish a declared claude.ai Artifact safely (ADR 0068). A repo that commits a GENERATED Artifact source declares source + check on its artifacts.items[] entry; the emitter enforces that the source exists and rides the drift check through CI, and this skill closes the loop: list the declared items, run each check, ONE confirmation, then the Artifact-tool publish with the enforced read-before-republish sequence and a byte-level lockstep proof against the committed file. Never headless (headless sessions have no Artifact tool — measured), never from a hook, never past a VIOLATION or a failed check; an ownership refusal is reported, not worked around.
/ywr-harness:update — for the member who does not want to wait for the background auto-update (ADR 0026 converges every machine by its next session start). It reads the plugin's marketplace version and scope from claude plugin list --json, runs the two CLI commands with --json (claude plugin marketplace update, then claude plugin update -s <scope>), reports the on-disk old → new from the result line, falls back to the text forms on a host without --json, never accepts a marketplace-declared command (-y, --accept-command), and hands off the two steps no skill can perform: /reload-plugins (a REPL built-in — the remedy for the update CLI's own "restart required to apply"), and /ywr-harness:harness-init only when the ADR 0033 nudge reports drifted toolchain files. It never runs the scaffold refresh itself — the byte comparison alone cannot prove direction (the .harness-version stamp orients the nudge's advice since ADR 0042, but a stampless repo stays ambiguous), and an automatic re-run could revert a working tree deliberately newer than the installed plugin (ADR 0034).
scripts/workflow-gates.mjs (parse + behavioral gate over a workflow corpus; --dir selects the corpus, default .claude/workflows) and scripts/resolve-base.sh (CI diff-range base resolution — since ADR 0043 vendored by the scaffold as scripts/harness/resolve-base.sh and invoked by the vendored workflow to resolve the PR base once, MODE=blocking). lib/selftest-lib.ps1 is the shared assertion core the selftests dot-source — dot-source it, do not run it.
node — runs eight of the nine hooks (hooks/*.mjs, with the shared hooks/hook-lib.mjs; Node costs ~70 ms per hook call against ~780 ms for pwsh, ADR 0116) and the gates scripts/workflow-gates.mjs and the workflow corpus gate. Absent node: those hooks cannot start on a member machine — the host reports a hook error and never a blocked call, since no hook can block — and session-start-node-check (a PowerShell hook, so it can run without node) says so once at each session start, with the fix; the suites report a skip locally and a hard failure on CI, where they must actually run.pwsh (PowerShell 7+) on PATH — session-start-node-check.ps1, the skills' scripts (harness-init, feedback) and every selftest. Without it the Node hooks still run, but the node notice itself shows a hook error naming pwsh at each session start.python 3.9+ (stdlib only) — every gate script under scripts/ requires it: harness_gates.py (invoked by the pre-commit hook), harness_retro.py (by the post-commit hook), and verify_map.py (by ywr-harness:verify).Hooks run wherever the session runs; the gates and selftests are also exercised on Linux pwsh, so nothing here may assume Windows.
selftest.ps1 is the single entry point — the manifest/wiring gate, the workflow corpus gate, then every shipped PowerShell selftest:
pwsh -NoProfile -File ./selftest.ps1
The two gates run first; the suites then run side by side (-Jobs N, default the CPU count capped at 8; -Jobs 1 runs them one at a time). Each finished suite prints one done line, then the results come back in discovery order: a failed suite in full, a passing suite as one line plus any SKIP lines. -Full prints every suite's whole output. To iterate on one component, run its own *.selftest.ps1 directly.
manifest-gate.ps1 is deterministic and CLI-free. It fails on the defects with distribution blast radius: version missing (omitted version falls back to the git commit SHA, so every commit ships as a new version to every consumer), a hook path that does not resolve, a regression from exec form to shell form where a path placeholder is used, and a release-notes canon that lies (top CHANGELOG.md entry ≠ plugin.json version, or the artifact link diverging between the CHANGELOG and the announce hook — ADR 0030). manifest-gate.selftest.ps1 proves it can fail — one or more mutations per check class (the suite prints the count) plus an unmutated control, because a suite that only ever passes and a gate that fails on everything score identically without the control. The control has earned its place: it caught a broken identity gate while the negative suite was reporting every mutation caught.
evals/ is a claude plugin eval suite (Claude Code ≥ 2.1.269; ADR 0076, spec 0014). Where the selftests above are hermetic — they call the scripts and never the host — each eval case runs a fresh, isolated claude -p child with ONLY this plugin loaded and grades what came out: a SessionStart hook's additionalContext reaching the model, ywr-harness:verify triggering on natural phrasing, the five disable-model-invocation skills staying uninvoked, and ywr-harness:reviewer resolving by its namespaced name with a SubagentStop ledger line behind it. Every case pins model: to one worker model alias (opus, ADR 0103/0107), and manifest-gate.ps1 refuses a case that does not (or
hooks/delegation-ledger.mjs 346 lines1// Hooks module (Claude Mods, Claude Code 2.1.287+; ADR 0117, ADR 0118) — the OBSERVE-ONLY per-request
2// delegation ledger. `hooks.json` names this file under `modules`; the settings hooks beside it are
3// unaffected.
4//
5// What it records, per finished model loop (the main loop's turn, or one run of a subagent's loop):
6// every request the loop made (`turn.step`: the model the engine resolved, the effort it sends, the
7// usage the API reported), the loop's end (`turn.complete`: reason, duration, summed usage) and, for
8// a spawned agent, the spawn itself (`agent.spawn`: type, the per-call model as given, the parent
9// model, the resolved model). An agent-team teammate raises `agent.spawn` too from Claude Code 2.1.289
10// (`e.isTeammate`); its loop keeps the `agent-tool` kind and its spawn row says `teammate: true`, which
11// is what the reader splits on (schema 2 unchanged — the field is additive). A Workflow `agent()` worker
12// raises `agent.spawn` from Claude Code 2.1.292 with `e.workflow` ({runId, agentIndex}); its spawn row
13// carries `workflow: {run_id, agent_index}` and its loop is kind `workflow` (ADR 0126), and a row with
14// no `model_param` proves the call gave no model, so the worker inherited the session's. On 2.1.287 to
15// 2.1.291 such a worker raises no `agent.spawn` (fact 89): its loop is written with `spawn: null` and
16// `loop: 'unspawned'`, whose steps still name the model and effort it ran on. On 2.1.292+ `unspawned`
17// means an evicted spawn row.
18//
19// Observe-only, by contract: every hook passes the event on unchanged (`next(e)`, `yield* next(e)`),
20// never rewrites `model`/`effort`, never denies, never draws, and swallows its own failures — the
21// recording code runs after `next` resolved and inside try/catch, because a hook that fails
22// mid-stream is "left where it stood". It never carries a gate: hosts < 2.1.287, `disableAllHooks`,
23// `--safe-mode` and a managed `allowManagedModsOnly` all drop it (ADR 0117).
24//
25// What it never writes: a prompt, a task description, an answer or any other message text (the
26// secret-adjacent rule `subagent-telemetry` keeps). Where it writes (ADR 0118): a RING of twenty
27// session files, `<project root>/.claude/telemetry/delegations/slot-00.json`..`slot-19.json` —
28// gitignored by the scaffold and the canon — and only where the root holds `.harness.json` (a repo
29// that adopted the harness), so a session in a home directory or a stranger's checkout leaves nothing
30// behind. `$.fs` has no delete, so the bound is by construction: a session claims a missing slot or
31// the oldest by `mtimeMs`, and its file keeps its last 500 loops within 360 KiB. `$.fs.write` is not
32// atomic and replaces the whole file, so a session's writes are serialized, and a slot another
33// session took (its `mtimeMs`/`size` moved since this session's last write) is given up for a new one.
34//
35// The pure half (`createLedger`, `createSessionLog`, `pickSlot`, `stepRow`, `spawnRow`, `fingerprint`)
36// is exported so `delegation-ledger.selftest.ps1` drives it under plain node with a fake engine; the
37// engine imports only `register`.
38
39export const SCHEMA = 2
40export const LEDGER_DIR = '.claude/telemetry/delegations'
41export const RING = 20
42// One session file: its last MAX_LOOPS loops within MAX_FILE_BYTES (UTF-8, counted as an upper
43// bound). The ring's worst case is RING x MAX_FILE_BYTES, ~7 MiB per adopted repo.
44export const MAX_LOOPS = 500
45export const MAX_FILE_BYTES = 360 * 1024
46// Room left in MAX_FILE_BYTES for the file's own fields around `loops` (a session id is ~40 B).
47const HEADER_RESERVE = 1024
48// A loop keeps its first MAX_STEPS step rows; the rest are counted. Bounds one row (~67 B a step)
49// and the memory of an open loop.
50export const MAX_STEPS = 200
51// Open loops are held in module memory until their loop ends. A loop that never ends (killed agent,
52// crash) would otherwise leak; past this many, the oldest is dropped unwritten.
53export const MAX_OPEN = 256
54// Spawn rows outlive their loop (a resumed agent runs again under the same id), so they are capped
55// separately and higher (a row is ~300 B). An evicted row would make that agent's later run read as
56// `unspawned` — the workflow signal ADR 0117 Decision 6 counts on hosts before 2.1.292, and a plain
57// eviction after — so each record carries `spawn_rows_evicted`, the session's running count, and a
58// reader discounts `unspawned` when it is > 0.
59export const MAX_SPAWNS = 4096
60// Session loop lists held at once (a `/clear` starts a new session in the same module).
61export const MAX_SESSIONS = 4
62
63const MAIN = 'main'
64
65export function slotName(i) {
66 return `slot-${String(i).padStart(2, '0')}.json`
67}
68
69// The slot a new session claims from a `$.fs.list` of the ledger directory: the lowest-numbered slot
70// with no entry, else the regular slot file with the oldest `mtimeMs` (ties: the lowest number), else
71// null. A slot name that is a directory, a link or anything but a file is never claimed.
72export function pickSlot(entries, ring = RING) {
73 const byName = new Map()
74 for (const en of Array.isArray(entries) ? entries : []) if (en && typeof en.name === 'string') byName.set(en.name, en)
75 let oldest = null
76 for (let i = 0; i < ring; i++) {
77 const en = byName.get(slotName(i))
78 if (!en) return i
79 if (en.kind !== 'file' || en.isLink || typeof en.mtimeMs !== 'number') continue
80 if (oldest === null || en.mtimeMs < oldest.mtimeMs) oldest = { i, mtimeMs: en.mtimeMs }
81 }
82 return oldest ? oldest.i : null
83}
84
85// An upper bound on the UTF-8 length of a string: a code unit above U+007F counts 3 bytes (a
86// surrogate pair counts 6 for its 4).
87export function utf8Bound(s) {
88 let n = s.length
89 for (let i = 0; i < s.length; i++) if (s.charCodeAt(i) > 0x7f) n += 2
90 return n
91}
92
93function usageRow(u) {
94 if (!u || typeof u !== 'object') return null
95 return {
96 model: u.model ?? null,
97 input_tokens: u.input_tokens ?? 0,
98 output_tokens: u.output_tokens ?? 0,
99 cache_read_input_tokens: u.cache_read_input_tokens ?? 0,
100 cache_creation_input_tokens: u.cache_creation_input_tokens ?? 0,
101 }
102}
103
104export function spawnRow(e, r) {
105 const wf = e.workflow && typeof e.workflow === 'object' ? e.workflow : null
106 return {
107 tool_use_id: e.tool_use_id ?? null,
108 subagent_type: e.subagentType ?? null,
109 provider: e.provider ? `${e.provider.plugin}/${e.provider.tier}` : null,
110 model_param: e.model ?? null,
111 parent_model: e.parentModel ?? null,
112 model: r && typeof r === 'object' && !('deny' in r && r.deny) ? (r.model ?? null) : null,
113 denied: !!(r && typeof r === 'object' && r.deny),
114 fork: !!e.fork,
115 background: !!e.background,
116 teammate: !!e.isTeammate,
117 workflow: wf ? { run_id: wf.runId ?? null, agent_index: wf.agentIndex ?? null } : null,
118 parent_agent_id: e.parentAgentId ?? null,
119 }
120}
121
122// A step row is a TUPLE in STEP_FIELDS order (~67 B, against ~236 B as an object): the ledger is
123// read by analysis code, never by eye, and the file names the fields once, as `step_fields`. The
124// token counts are null when the request reported no usage; `usage_model` (the model the API
125// reported) is null when it equals `model` or no usage came back.
126export const STEP_FIELDS = ['index', 'model', 'effort', 'message_count', 'stop_reason', 'input_tokens',
127 'output_tokens', 'cache_read_input_tokens', 'cache_creation_input_tokens', 'usage_model']
128
129export function stepRow(e, r) {
130 const res = r && typeof r === 'object' ? r : null
131 const u = usageRow(res ? res.usage : null)
132 const model = e.model ?? null
133 return [
134 e.index ?? null,
135 model,
136 e.effort ?? null,
137 e.messageCount ?? null,
138 res ? (res.stopReason ?? null) : null,
139 u ? u.input_tokens : null,
140 u ? u.output_tokens : null,
141 u ? u.cache_read_input_tokens : null,
142 u ? u.cache_creation_input_tokens : null,
143 u && u.model !== model ? u.model : null,
144 ]
145}
146
147const present = x => x !== null && x !== undefined
148
149// The ledger's bookkeeping, engine-free. `spawn` / `step` / `complete` take the event and the result
150// `next` resolved to; `complete` answers the loop's row (empty steps when nothing was open).
151export function createLedger({ maxOpen = MAX_OPEN, maxSpawns = MAX_SPAWNS, maxSteps = MAX_STEPS } = {}) {
152 const spawns = new Map() // agentId -> spawn row (kept across that agent's runs)
153 const open = new Map() // `${agentId|main}\u0000${turnId}` -> { steps, dropped, models, efforts }
154 let evicted = 0
155 const cap = (map, max, onDrop) => { while (map.size > max) { map.delete(map.keys().next().value); onDrop?.() } }
156 const key = (agentId, turnId) => `${agentId ?? MAIN}\u0000${turnId}`
157 return {
158 spawn(e, r) {
159 const agentId = r && typeof r === 'object' ? r.agentId : undefined
160 if (!agentId) return
161 spawns.set(agentId, spawnRow(e, r))
162 cap(spawns, maxSpawns, () => { evicted++ })
163 },
164 step(e, r) {
165 const k = key(e.agentId, e.turnId)
166 if (!open.has(k)) { open.set(k, { steps: [], dropped: 0, models: new Set(), efforts: new Set() }); cap(open, maxOpen) }
167 const loop = open.get(k)
168 if (!loop) return
169 const row = stepRow(e, r)
170 if (present(row[1])) loop.models.add(row[1])
171 if (present(row[2])) loop.efforts.add(row[2])
172 if (loop.steps.length < maxSteps) loop.steps.push(row)
173 else loop.dropped++
174 },
175 complete(e, { iso }) {
176 const k = key(e.agentId, e.turnId)
177 const loop = open.get(k) ?? { steps: [], dropped: 0, models: new Set(), efforts: new Set() }
178 open.delete(k)
179 const spawn = e.agentId ? (spawns.get(e.agentId) ?? null) : null
180 return {
181 ts: iso,
182 loop: !e.agentId ? MAIN : !spawn ? 'unspawned' : spawn.workflow ? 'workflow' : 'agent-tool',
183 agent_id: e.agentId ?? null,
184 turn_id: e.turnId ?? null,
185 spawn,
186 spawn_rows_evicted: evicted,
187 steps: loop.steps,
188 steps_dropped: loop.dropped,
189 models: [...loop.models],
190 efforts: [...loop.efforts],
191 complete: {
192 reason: e.reason ?? null,
193 aborted: !!e.isAborted,
194 duration_ms: e.durationMs ?? null,
195 usage: usageRow(e.usage),
196 },
197 }
198 },
199 sizes: () => ({ spawns: spawns.size, open: open.size, evicted }),
200 }
201}
202
203// One session's loop rows, kept serialized and capped; `render` is the whole file's text.
204export function createSessionLog({ maxLoops = MAX_LOOPS, maxBytes = MAX_FILE_BYTES } = {}) {
205 const loops = [] // { text, bytes }
206 let bytes = 0
207 let dropped = 0
208 return {
209 add(row) {
210 const text = JSON.stringify(row)
211 const b = utf8Bound(text) + 1 // + its comma
212 loops.push({ text, bytes: b })
213 bytes += b
214 while (loops.length > maxLoops || (loops.length > 0 && bytes + HEADER_RESERVE > maxBytes)) {
215 bytes -= loops.shift().bytes
216 dropped++
217 }
218 },
219 // `header` is the file's own fields; `loops` goes last, spliced in as the rows' kept text.
220 render(header) {
221 const h = JSON.stringify({ ...header, loops_dropped: dropped, loops: [] })
222 return `${h.slice(0, -3)}[${loops.map(l => l.text).join(',')}]}\n`
223 },
224 sizes: () => ({ loops: loops.length, bytes, dropped }),
225 }
226}
227
228// --- the engine half ---------------------------------------------------------------------------
229// Declared at the top of the file: the engine follows `$` only into such functions.
230
231// The project root to write under, or null where the root holds no `.harness.json`. `adopted`
232// caches the answer per root (a `/cd` or worktree move changes the root mid-session).
233async function target($, adopted) {
234 const root = await $.session.root()
235 if (!adopted.has(root)) adopted.set(root, await $.fs.exists(`${root}/.harness.json`))
236 return adopted.get(root) ? root : null
237}
238
239async function statOf($, path) {
240 try {
241 const st = await $.fs.stat(path)
242 return { mtimeMs: st.mtimeMs, size: st.size }
243 } catch { return null }
244}
245
246// The start every file this session renders begins with (`render` puts the header keys first).
247const headerOf = id => `{"schema":${SCHEMA},"session_id":${JSON.stringify(id)},`
248
249// A 32-bit FNV-1a over the UTF-16 code units: the mark of the text this session last wrote whole.
250export function fingerprint(text) {
251 let h = 0x811c9dc5
252 for (let i = 0; i < text.length; i++) h = Math.imul(h ^ text.charCodeAt(i), 0x01000193)
253 return h >>> 0
254}
255
256// One rewrite of a session's file. Runs on the session's chain, never two at once for one session.
257async function flush($, s, root, iso) {
258 const dir = `${root}/${LEDGER_DIR}`
259 if (s.root !== root) { s.root = root; s.slot = null; s.seen = null; s.mark = null }
260 if (s.slot !== null && s.seen) {
261 const now = await statOf($, `${dir}/${slotName(s.slot)}`)
262 // moved since this session's last write: another session claimed it. A stat that fails (a missing
263 // slot among the causes) proves nothing either way, so the read-back below decides.
264 if (!now) s.seen = null
265 else if (now.mtimeMs !== s.seen.mtimeMs || now.size !== s.seen.size) s.slot = null
266 }
267 if (s.slot !== null && !s.seen) {
268 // the last write or a stat failed, so there is no stat to compare: the file itself says whose
269 // it is. A missing or empty file stays ours; so does the very text this session last wrote whole
270 // (its fingerprint, which covers a stat that keeps failing), and a file starting with our own
271 // header (a partial write of ours keeps it). A file that exists but cannot be read proves
272 // nothing, and neither does a null session id's header (every id-less session writes the same
273 // one), so both are given up.
274 const path = `${dir}/${slotName(s.slot)}`
275 let text = ''
276 try { text = await $.fs.read(path) } catch {
277 let there = true
278 try { there = await $.fs.exists(path) } catch { /* unknown: not provably ours */ }
279 if (there) text = null
280 }
281 const ours = text === '' || (text !== null && ((s.mark !== null && fingerprint(text) === s.mark) ||
282 (s.id !== null && text.startsWith(headerOf(s.id)))))
283 if (!ours) s.slot = null
284 }
285 if (s.slot === null) {
286 let entries = []
287 try { entries = await $.fs.list(dir) } catch { /* no directory yet: every slot is missing */ }
288 s.slot = pickSlot(entries)
289 if (s.slot === null) return
290 }
291 const path = `${dir}/${slotName(s.slot)}`
292 const text = s.log.render({ schema: SCHEMA, session_id: s.id, slot: s.slot, updated: iso, step_fields: STEP_FIELDS })
293 s.seen = null
294 s.mark = null
295 await $.fs.write(path, text)
296 s.mark = fingerprint(text)
297 s.seen = await statOf($, path)
298}
299
300export function register(on) {
301 const ledger = createLedger()
302 const adopted = new Map() // project root -> whether it holds .harness.json
303 const sessions = new Map() // session id -> { id, log, root, slot, seen, mark, chain }
304 const sessionOf = id => {
305 let s = sessions.get(id)
306 if (s) { sessions.delete(id); sessions.set(id, s) } // most recently used last: eviction takes the idlest
307 else {
308 s = { id, log: createSessionLog(), root: null, slot: null, seen: null, mark: null, chain: Promise.resolve() }
309 sessions.set(id, s)
310 while (sessions.size > MAX_SESSIONS) sessions.delete(sessions.keys().next().value)
311 }
312 return s
313 }
314
315 on('agent.spawn', async ($, e, next) => {
316 const r = await next(e)
317 try { ledger.spawn(e, r) } catch { /* observe-only: never fail the spawn */ }
318 return r
319 })
320
321 on('turn.step', async function* ($, e, next) {
322 const r = yield* next(e)
323 try { ledger.step(e, r) } catch { /* observe-only */ }
324 return r
325 })
326
327 on('turn.complete', async ($, e, next) => {
328 const r = await next(e)
329 try {
330 // The loop leaves memory first: a failing engine call below costs this one record, never a leak.
331 const row = ledger.complete(e, { iso: null })
332 const id = (await $.session.id()) ?? null
333 row.ts = new Date(await $.clock.now()).toISOString()
334 const root = await target($, adopted)
335 if (root) {
336 const s = sessionOf(id)
337 s.log.add(row)
338 const run = s.chain.then(() => flush($, s, root, row.ts))
339 s.chain = run.catch(() => {})
340 await run
341 }
342 } catch { /* observe-only: a ledger write never reaches the turn */ }
343 return r
344 })
345}
346