SLOPSHOPPER

harness

Draws the recorded gate verdict, the open track and the context in use above the prompt, and adds /gates. It draws and relays; every judgement is made by a…

newpanebandcommandtoastprocess
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · harness
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /gates ⎿ harness: GATES NOT RUN: scripts/harness/panel.js could not be read (JSON Parse error: Unexpected identifier "dev"), so harness facts not read: JSON Parse error: Unexpected identifier "dev" context 49% ⟨Claude Code's own drawing⟩ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
harness facts not read: JSON Parse error: Unexpected identifier "dev" context 49% ⟨Claude Code's own drawing⟩
README

Agent Harness

CI Licence: MIT Dependencies: 0

Your coding agent's rules, enforced by code instead of by the agent's memory.

You wrote the rules down. The agent followed them, until the session it didn't: it deleted a failing test, quoted last week's test count, read a credential file into the transcript, or called half-finished work done. Nothing told you, because a rule in a prompt cannot fail.

Agent Harness is scripts and hooks you copy into your repo. Every rule that matters becomes a check that fails loudly with an exit code. Anything that is still just a written rule is labelled as one.

Delete a test and an ordinary suite stays green. Here the run is red, and says why:

GATES RED — 2 violation(s) against scripts\harness\gates.baseline.json:
  x test: pass 1 < baseline 2
  x test: 1 checks in the total, baseline had 2 — a check DISAPPEARED (pass-by-absence)

Node 20 or newer, built-ins only, zero dependencies, no service, no model calls. This repo is checked by the harness it ships.

What you rely on today, and what replaces it

TodayHow it failsHere
A rule in CLAUDE.md or AGENTS.mdFollowed until it isn't, and nothing reports the lapseA hook denies the action before it runs, or a gate fails the run
"All tests pass" in the agent's summaryThe count is stale, or a test was deleted or skippedEach total is compared with a committed baseline; a missing check is red
"Done" typed into a status fileNobody checked it against the tree as it standsA phase closes only through a script that sees current evidence
.gitignore for secretsIt stops the commit, not the read into the transcriptThe read itself is denied
A long recap pasted into the next sessionEvery session re-learns what the last one knewOne handoff file per track, and the session starts there
An instructions file that keeps growingIt loads on every turn and nobody sees the sizeA byte ceiling on always-loaded files, failed in CI

The hooks deny only under a coding agent they are wired for; everywhere else the same rules run as gates in CI. Your stack says which is which, and What it does not do is the honest list of limits.

Why this is different

Most agent harness repos give you prompts, agent definitions or an orchestration framework. This one gives you none of those as its main product. It gives you the checks around them:

  • A missing check is a failure. Each gate's total is compared with a committed baseline, so a test that was deleted or silently skipped turns the run red.
  • Bad actions are blocked before they run, by hooks, not asked for politely in a prompt.
  • "Done" is a script. A phase closes only when the script sees current evidence.
  • It says what it cannot do. See What it does not do.
  • Nothing to install. No dependencies, no build, no service, no model calls.

Try it before you copy anything

git clone https://github.com/arunpaul-H/coding-agent-harness.git
cd coding-agent-harness
node bin/harness.js selftest    # every guard's decisions, driven with literals
node bin/harness.js doctor      # every piece: OK / MISSING / DRIFTED, each with its fix

No npm install: there is nothing to install. When you want it in your own repo, go to Install.

Who this is for

Engineers and developers who build with agents: agentic developers, not prompt engineers. If your answer to an agent breaking a rule is a better-worded prompt, this is not that. Here the answer is a gate, a hook or a script that fails the work, so the rule holds whether or not the agent remembers it.

Not a developer? Start here

You do not need to read code to use this. You tell an AI what to build, and the harness checks the AI's work for you.

flowchart LR
    A[You describe what you want] --> B[The AI writes it]
    B --> C{The harness checks it}
    C -->|Green| D[Keep going]
    C -->|Red| E[The AI is told why and fixes it]
    E --> C

What the colours mean

ResultIn plain words
🟢 GreenEvery check the project has was run, and none failed or went missing.
🔴 RedA check failed or disappeared. Nothing is broken for good: the AI is told exactly what to fix.
⚪ SkippedA check did not run. It proved nothing, so ask why before you rely on the result.

Your first hour

  1. Open your project in Claude Code, the AI coding tool this guide is built for.
  2. Ask it to copy this harness into your project. It shows you the list of files first.
  3. Ask it to check that every piece is in place.
  4. Type /guide. It walks you through the rest in plain English, one step at a time.

What green does not mean. Green does not mean the idea is good, that customers want it, or that the product is secure. It means the checks your project has all ran and passed. A project with few checks goes green easily.

Everything below this line is written for developers.

Who it is for

People on a flat-rate subscription seat of an agentic coding tool: a plan with a usage limit, not a bill per token. What rations a seat is quota and context, not dollars. So the harness counts what a seat runs out of, the bytes loaded on every turn and the work a session has to redo, and it accounts for nothing in money.

  • You run a coding agent on a real repo, across many sessions.
  • You have seen an agent skip a rule, quote a stale test count or call unfinished work done.
  • You want that caught by CI and hooks instead of by review.
  • You are not a tech expert and you drive a coding agent yourself: start here.

Who it is not for

  • You pay per token on an API key and want cost in dollars tracked. Nothing here measures it.
  • You are buying for an organisation and need SSO, a policy server or compliance reporting.
  • You want a framework for building or orchestrating agents. This calls no model.
  • You want a prompt or skill library to drop in.
  • You cannot run Node 20 on your machine and in CI.
  • You want blocking hooks on a coding agent the harness does not wire yet (see Your stack).

When to use it

Use it once the same rule has been broken silently more than once, or when work spans sessions and each new session re-learns what the last one knew. For a one-off script or a weekend prototype it is more process than you need.

Install

There is no published package. Installing is copying. You need Node >= 20, git and a POSIX shell (on Windows, the one Git provides).

git clone https://github.com/arunpaul-H/coding-agent-harness.git agent-harness

node agent-harness/bin/harness.js init --into <your-repo> --dry-run   # the plan; writes nothing
node agent-harness/bin/harness.js init --into <your-repo>

Cloning a project you will run several agents in? Clone it as git clone <url> <project>/main: every workspace then sits beside main/ in one folder. A plain clone works too (step 5 below).

init never overwrites a file without --force, and it copies no numbers measured on another repo: your baseline starts empty. It writes no .gitattributes; pin your line endings yourself.

To try it here first:

node bin/harness.js selftest    # every --self-test in the repo
npm test                        # node --test
npm run gates                   # every gate against the committed baseline
node bin/harness.js doctor      # every piece: OK / MISSING / DRIFTED, each with its fix

On a fresh clone or a CI runner the docs gate reports one skip and says why. An undeclared skip is red on purpose, so allow it by name:

node scripts/harness/gates.js --check --allow-skip reproducible-build=1,docs=1

Quickstart

  1. Init, as above. It detects your package manager, test command, docs folder and default branch, and lists the placeholders it left for you to fill in AGENTS.md.
  2. Read harness.config.json. It is the only place the harness learns about your project. Add your own checks to its gates array. With no config file, every script stops; there is no default suite.
  3. Run the gates. git add first, because the context gate reads tracked files only.
   $ node scripts/harness/gates.js --check --allow-skip reproducible-build=1
   GATES RED — 1 violation(s) against scripts\harness\gates.baseline.json:
     x docs: skipped=1 and no --allow-skip docs=n was given

The first run is red, and should be. Read the reason, allow the skip, record the baseline:

   $ node scripts/harness/gates.js --accept --allow-skip reproducible-build=1,docs=1
   $ node scripts/harness/gates.js --check  --allow-skip reproducible-build=1,docs=1
   GATES GREEN — test 1/1 . docs 147/148 (1 skipped, declared) . context 6/6 . hygiene 7/7 . skills 6/6 . reproducible-build 0/1 (1 skipped, declared)
  1. Commit (a commit that changes the baseline needs a subject starting gates:), run node agent-harness/bin/harness.js doctor, and put the same --check command in CI.
  1. More than one agent? One folder each. Two sessions in one folder trip over each other's unfinished files. A workspace is a folder on its own branch over the same history:
   $ node bin/harness.js worktree new fix-export --doing "mend the CSV export"
   $ node bin/harness.js worktree list            # who is doing what; files two branches changed
   $ node bin/harness.js worktree land fix-export # clean, green, merged? prints the merge command
   $ node bin/harness.js worktree prune           # stale entries; removes only with --yes

Start the agent in the new folder. Nothing here merges or pushes: a person runs the merge. harness doctor says which layout you have and how to adopt the parent-folder one.

What the baseline buys you: delete a test later and nothing fails, yet the run is red.

GATES RED — 2 violation(s) against scripts\harness\gates.baseline.json:
  x test: pass 1 < baseline 2
  x test: 1 checks in the total, baseline had 2 — a check DISAPPEARED (pass-by-absence)

Full walk-through, including how to confirm the hooks fire: docs/methods/11-adopting-this-in-your-repo.md.

What you get

PieceWhat it stopsWhere
Gates and a baselineStale totals in docs; a check that vanished while the suite stayed greenscripts/harness/gates.js
Reproducible-build checkA source edit committed without its rebuilt bundlescripts/harness/gates.js
Secret-file guardA credential file read into a transcriptscripts/harness/guards/env-guard.cjs
Phase guardA status flipped by hand; a gate figure typed into a docscripts/harness/guards/phase-guard.cjs
Commit hygieneStray files, pasted tokens and root clutter in a commitscripts/harness/guards/commit-hygiene.cjs
Spawn guardA subagent of yours spawned with no OUTCOME, READ-SET, RETURN or BOUNDSscripts/harness/guards/spawn-guard.cjs
Stop reportA turn that ends on edits no gate run has seen (it tells you; stop.block forces a turn)scripts/harness/close-phase.js
Doc contractDocs over their line cap; an index that drifted from the docsscripts/harness/guards/docs-lint.cjs
Context budgetAlways-loaded instruction files growing unnoticedscripts/harness/measure-context.js
Tracks and handoffsA new session re-deriving what the last one knewtemplates/track/
Phase lifecycleA phase closed on evidence from a different treescripts/harness/close-phase.js
Skill lockA skill changing with no diff anyone read: a third-party one, or one the harness shippedscripts/harness/guards/skill-verify.cjs
Branch syncA long-lived branch running an old copy of the harnessscripts/harness/branch-sync.js
Installer and doctorAn install nobody verifiedbin/harness.js
Panel mod (Claude Code)A gate verdict that scrolled away; a model turn spent to run the gate or close a phase.claude/plugins/harness/
Guard mod (Claude Code)A spawn no Agent call made, or a file attached with @, that no settings hook sees.claude/plugins/harness-guard/

Each piece has a page under docs/methods/: the failure, the mechanism, its cost and its limits.

Reference docs for your own project

The harness writes no doc about your code. init installs a template and nothing else, because a doc generated from code repeats what a search already finds. When a subsystem has rules that live in someone's head, ask for one doc, by name:

node agent-harness/bin/harness.js reference new billing-export \
  --read-when "changing what the nightly export writes" --paths "src/billing/**,tests/billing/**"
node scripts/harness/docs-index.js

That writes docs/reference/billing-export/README.md with the frontmatter filled and the body left for you. It never overwrites, and both flags are required: --read-when is the one line an agent sees in the index, and --paths is what the doc gate compares against history.

  • Run it when the same question about a subsystem keeps being re-answered from the code, or the answer is not in the code at all.
  • Put in it invariants and why they hold, decisions and what was turned down, what runs where, and limits set outside the repo. Leave out anything a search returns.
  • Afterwards the doc gate warns, and never fails, when a commit touches those paths after the doc's last_verified day. Re-check the doc, then change the date. Changing only the date defeats the warning.

What we measured, once: the same read-only question put to two fresh agents on the same model, one tree with a hand-filled doc and one without. Both answered every part correctly. The run with the doc used 65083 tokens against 78068, the same 18 tool uses, and one fewer file; it still read the code, because the question asked for line numbers. One pair is an observation, not proof, and it does not count the time spent writing the doc. Expect a modest saving at best, and write a doc for what the code cannot say, not to save tokens. The write-up is in git history: git show b2183f0:docs/tracks/project-reference/research/2026-10-08-one-task-with-and-without-a-reference-doc.md.

Your stack

  • The harness is Node. Your project need not be, but Node >= 20 must be on the machine and in CI.
  • Any language can be gated. A gate is any shell command. If its last stdout line is GATE <name> pass=<n> fail=<n> skipped=<n> the counts are read exactly; otherwise the exit-code adapter grades it as one pass or one fail. init detects Node, Go, Rust and Python test commands.
  • The blocking hooks are wired for few tools. harness init --tools <ids> names the coding agents to wire. A hook file is written for Claude Code and, for shell calls and file reads only, for Cursor. Every other tool below gets the policy, the skills and a stated gap. The agent files and the status line are Claude Code's only.
  • Which tools, and on what evidence. ran means the tool was run here and the row records the command. docs means the row was read off the vendor's pages and the tool was never run. Printed by node scripts/harness/tools.js --list, 2026-10-07:

<!-- tools-list:start -->

  id              tier        evidence  enabled
  claude-code     full        ran       yes
  codex-cli       full        docs
  gemini-cli      full        docs
  antigravity     guard-only  docs
  cursor          full        docs
  kimi-code       guard-only  docs
  deepseek        inherits    docs
  github-copilot  backstop    docs
  windsurf        backstop    docs
  cline           backstop    docs
  aider           backstop    docs
  opencode        backstop    docs
  qwen-code       backstop    docs

<!-- tools-list:end -->

  • What enforces each rule under your tools. node scripts/harness/tools.js --matrix prints it per enabled tool: a deny at the tool call, a report, a gate, a git-hook backstop, or prose.
  • The policy is not tied to one tool. It lives in AGENTS.md, which most coding agents load; CLAUDE.md is a one-line pointer to it.
  • What works anywhere today: gates.js --check in CI, and the opt-in git hooks: harness init --git-hooks writes a pre-commit that runs commit-hygiene.cjs --report and a pre-push that runs the gate suite. A backstop after the work, not a deny.

Claude Code mods

A mod is a plugin whose code runs inside Claude Code (terminal, v2.1.287 and later). This repo carries two, and neither holds a decision: each runs a script under scripts/harness/ and draws or relays what it printed.

  • harness draws a band above the prompt (the recorded gate verdict, the open track, the context in use) and adds /gates, /track, /handoff and /close-phase, none of which takes a model turn. It never refuses a call.
  • harness-guard hands each tool call the settings hook also sees, each agent spawn and each @ file mention to scripts/harness/hook.js and relays the verdict. A commit the hygiene guard refuses is held and shown with its findings: you have it checked again, or cancel. Each of its hooks refuses when it fails. It draws nothing.

No settings file enables either. Load them for one session, from the repo root:

claude --plugin-dir .claude/plugins/harness --plugin-dir .claude/plugins/harness-guard

What is and is not established:

  • Graded without a session. claude plugin validate --strict and claude plugin test run as the optional mods gate. Where the claude program is absent, that gate is a declared skip: the harness depends on no agent.
  • Run headless only. harness-guard was loaded in claude -p sessions, where it passed a harmless call and refused a read of a guarded secret file. No interactive session has loaded either mod, so the band, the commands and every dialog are held by tests alone.
  • Not final. A mod loaded ahead of harness-guard can answer a call before it is asked, and another mod can approve a call a settings hook refused. Only a hook in managed settings is final. The settings hooks stay; the mods are an addition for Claude Code.
  • Not installed by init yet. They live in this repo; copy the two folders to use them elsewhere.

The measurements are in docs/tracks/claude-code-uptake/research/.

What it does not do

  • A green gate does not prove the code is correct. It proves the declared checks ran, none failed and none vanished.
  • Nobody has measured that it improves output. There is no run with the harness off to compare against.
  • The guards match patterns. A different spelling of the same action can get through. They stop accidents, not a determined agent or an attacker.
  • A mod does not make a guard final. A user-installed mod can approve a call a project hook refused; see Claude Code mods.
  • The skill lock cannot cover what lives outside the repo: plugin caches, user-level skills and MCP servers. It reports them as uncovered; it does not secure them.
  • The context budget counts bytes of the files you declare. Tool schemas and the system prompt are usually larger and are not counted.
  • The lints check shape, not quality. An unfilled handoff template passes.
  • It is not an agent framework. It orchestrates nothing and calls no model.

Repo layout

bin/                     harness.js: the installer and CLI (init, doctor, gates, docs, host, selftest)
scripts/harness/         the gate runner, config loader, lifecycle and doc scripts, baselines
scripts/harness/guards/  the guards, the doc linter, the skill verifier
.claude/                 hook wiring, four agent briefs, four skills, the status line
.claude/plugins/         two Claude Code mods: harness (panel and commands), harness-guard (guards)
schema/                  the JSON schema for harness.config.json
templates/               starting files for AGENTS.md, CLAUDE.md, the routing doc and a track
docs/methods/            one page per mechanism
docs/reference/          the enforced shapes: doc standard, routing, skill safety, branches
docs/tracks/             open work on this repo
tests/                   node:test; every script run as a real process in a throwaway repo
.github/                 the CI workflow, Dependabot and scanner configs

harness.config.json holds everything project-specific. AGENTS.md is this repo's own policy and the starting point for yours.

Contributing

See CONTRIBUTING.md. Changes are listed in CHANGELOG.md.

Licence

MIT. See LICENSE.

Source 1 files
hooks/register.js 273 lines
1/* ============================================================
2   The harness panel: a band above the prompt, /gates, and the track lifecycle as commands
3   (/track, /handoff, /close-phase).
4
5   THIS FILE HOLDS NO DECISION. It runs scripts/harness/panel.js, which judges, and draws what
6   that script printed; each command runs the argv that script named and relays what it said.
7   A receipt, a `status:` and a handoff are written by close-phase.js and by a person, never here.
8   A mod has no Node built-ins and no `require`, so anything decided here would be out of reach
9   of `node --test`. Colours, wording and where a line is cut are all this file owns.
10
11   IT NEVER GATES A CALL: no tool.call, tool.check or prompt.submit hook, and no $.fs.write
12   (it reads one file, the handoff /track shows),
13   $.http.fetch or $.model.complete. `claude plugin validate` prints the proof on its `hooks:`
14   and `calls:` lines.
15
16   WHERE NOTHING DRAWS (a headless run) the band's facts are not read at all. The commands still
17   work, except /close-phase: nobody is there to confirm it, so it closes nothing.
18   ============================================================ */
19
20// The one script this mod reads, relative to the folder the session runs in
21const PANEL = ['node', 'scripts/harness/panel.js']
22// The longest $.process.run allows. A suite that needs longer is reported as NOT RUN.
23const TEN_MINUTES = 600000
24const TOAST_MAX = 100
25// close-phase.js lints, checks hygiene and reads the tree: longer than the default wait
26const TWO_MINUTES = 120000
27// The pane /track opens
28const PANE = 'harness-track'
29const CLOSE_IT = 'Close it'
30
31// What panel.js printed, or null while it has not been read
32let facts = null
33// Why it could not be read, when it could not
34let unread = 'not read yet'
35// Context in use, in percent, or null when the session has not said
36let percent = null
37// True while /gates is running the suite
38let running = false
39// True when this session has somewhere to draw
40let draws = false
41// True from a compaction until the next turn ends
42let compacted = false
43// The handoff the pane shows: { id, text }, or null before /track
44let shown = null
45
46const firstLine = (error) => String((error && error.message) || error || 'no reason given').split('\n')[0]
47
48// Run panel.js and keep what it printed. Never throws: a failure is something to draw.
49async function refresh($) {
50  try {
51    const r = await $.process.run(PANEL)
52    const read = JSON.parse(r.stdout)
53    if (!read || typeof read !== 'object' || read.error) throw new Error((read && read.error) || 'panel.js printed no facts')
54    facts = read
55    unread = ''
56  } catch (error) {
57    facts = null
58    unread = firstLine(error)
59  }
60  $.ui.invalidate('ui.render')
61}
62
63// Ask the session how full the context is
64async function measure($) {
65  try {
66    const usage = await $.session.usage()
67    percent = usage && usage.context && typeof usage.context.percent === 'number' ? Math.round(usage.context.percent) : null
68  } catch {
69    percent = null
70  }
71  $.ui.invalidate('ui.render')
72}
73
74// The lines gates.js prints under its rule: the verdict, then each reason
75function closing(stdout) {
76  const lines = String(stdout || '').split('\n').map((line) => line.trimEnd())
77  const rule = lines.lastIndexOf('-'.repeat(56))
78  const tail = (rule >= 0 ? lines.slice(rule + 1) : lines).filter(Boolean)
79  return rule >= 0 ? tail.slice(0, 20) : tail.slice(-8)
80}
81
82// The open track a command means: the one named after it, or the only one there is
83function trackFor(args) {
84  const tracks = facts ? facts.tracks : []
85  const id = String(args || '').trim()
86  if (id) return tracks.find((track) => track.id === id) || null
87  return tracks.length === 1 ? tracks[0] : null
88}
89
90// What to say when a command has no track to act on
91function noTrack(args) {
92  if (!facts) return 'NOT RUN: scripts/harness/panel.js could not be read (' + unread + '), so no track is known.'
93  const open = facts.tracks.map((track) => track.id)
94  if (!open.length) return 'No open track.'
95  const id = String(args || '').trim()
96  return (id ? 'No open track named `' + id + '`.' : 'More than one track is open, so name one.') + ' Open: ' + open.join(', ')
97}
98
99// Run one harness command and hand back everything it printed
100async function relay($, argv) {
101  try {
102    const r = await $.process.run(argv, { timeoutMs: TWO_MINUTES })
103    const said = (String(r.stdout || '') + '\n' + String(r.stderr || '')).split('\n').map((line) => line.trimEnd()).filter(Boolean)
104    return said.length ? said.join('\n') : '`' + argv.join(' ') + '` printed nothing (exit ' + r.exitCode + ').'
105  } catch (error) {
106    return 'NOT RUN: `' + argv.join(' ') + '` did not finish (' + firstLine(error) + ').'
107  }
108}
109
110// One line per open track, for the transcript
111function listing() {
112  return facts.tracks
113    .map((track) => track.id + (track.phaseNext === null ? '' : '  phase ' + track.phaseNext) + (track.doing ? '  ' + track.doing : ''))
114    .join('\n')
115}
116
117const COLOURS = { green: 'green', red: 'red', stale: 'yellow', 'not-run': 'yellow' }
118
119export function register(on) {
120  // Runs before the first prompt, and again after a reload
121  on('session.start', async ($, e, next) => {
122    draws = e.isInteractive === true
123    if (draws) {
124      // On a timer, so the session does not wait for the facts
125      $.clock.after(0, () => refresh($))
126      $.clock.after(0, () => measure($))
127    }
128    // Last, because a refused name throws and the rest of the hook would be skipped
129    try {
130      await $.command.register({ name: 'gates', description: 'Run the gate suite and show its verdict, without a model turn' })
131      await $.command.register({ name: 'track', description: 'List the open tracks and show one handoff in a pane', argumentHint: '[track]' })
132      await $.command.register({ name: 'handoff', description: 'Lint the handoff of a track and show each finding', argumentHint: '[track]' })
133      await $.command.register({ name: 'close-phase', description: 'Close the phase of a track through close-phase.js, after you confirm', argumentHint: '[track]' })
134    } catch (error) {
135      $.ui.log('a command was not added: ' + firstLine(error))
136    }
137    return next(e)
138  })
139
140  // A compacted session has lost the handoff it was working from. This event can refuse a
141  // compaction and this hook never does: it notes that one happened, and if it fails the
142  // compaction goes ahead.
143  on('session.compact', async ($, e, next) => {
144    compacted = true
145    return next(e)
146  }).catch(async ($, e, next) => next(e))
147
148  // A turn may have edited a file, which makes a recorded run stale
149  on('turn.complete', async ($, e, next) => {
150    if (draws) $.clock.after(0, () => refresh($))
151    const result = await next(e)
152    if (!compacted) return result
153    compacted = false
154    const track = facts ? facts.tracks.find((open) => open.handoff) : null
155    if (!track) return result
156    // The line under the answer
157    const line = 'Context was compacted. The open track resumes from ' + track.handoff
158    return { ...result, text: result && result.text ? result.text + '\n' + line : line }
159  })
160
161  // Fires after each turn, and when a plan limit moves
162  on('session.measure', async ($, e, next) => {
163    if (draws) await measure($)
164    return next(e)
165  })
166
167  // Runs when the person types /gates
168  on('command.run', { command: 'gates' }, async ($) => {
169    if (running) return { text: 'The gate suite is already running.' }
170    if (!facts) await refresh($)
171    if (!facts) return { text: 'GATES NOT RUN: scripts/harness/panel.js could not be read (' + unread + '), so there is no gate command to run.' }
172    const command = facts.gates.command
173    running = true
174    $.ui.invalidate('ui.render')
175    let text
176    try {
177      const r = await $.process.run(facts.gates.argv, { timeoutMs: TEN_MINUTES })
178      const lines = closing(r.stdout)
179      text = lines.length ? lines.join('\n') : 'GATES NOT RUN: `' + command + '` printed nothing (exit ' + r.exitCode + ').'
180    } catch (error) {
181      text = 'GATES NOT RUN: `' + command + '` did not finish (' + firstLine(error) + '). A mod may run a command for ten minutes at most; run it in a shell.'
182    }
183    running = false
184    await refresh($)
185    $.ui.toast(text.split('\n')[0].slice(0, TOAST_MAX))
186    return { text }
187  })
188
189  // Runs when the person types /track
190  on('command.run', { command: 'track' }, async ($, e) => {
191    if (!facts) await refresh($)
192    const track = trackFor(e.args)
193    if (!track) return { text: noTrack(e.args) }
194    if (!track.handoff) return { text: listing() + '\n' + track.id + ' has no handoff to show.' }
195    let text
196    try {
197      text = await $.fs.read(track.handoff)
198    } catch (error) {
199      return { text: listing() + '\n' + track.handoff + ' could not be read (' + firstLine(error) + ').' }
200    }
201    shown = { id: track.id, text }
202    await $.ui.open({ id: PANE, title: 'handoff: ' + track.id, focus: true, closeOnEscape: true })
203    $.ui.invalidate('ui.render')
204    return { text: listing() }
205  })
206
207  // Runs when the person types /handoff
208  on('command.run', { command: 'handoff' }, async ($, e) => {
209    if (!facts) await refresh($)
210    const track = trackFor(e.args)
211    if (!track) return { text: noTrack(e.args) }
212    return { text: await relay($, track.lint) }
213  })
214
215  // Runs when the person types /close-phase
216  on('command.run', { command: 'close-phase' }, async ($, e) => {
217    if (!facts) await refresh($)
218    const track = trackFor(e.args)
219    if (!track) return { text: noTrack(e.args) }
220    const what = (track.closes === null ? 'the phase' : 'phase ' + track.closes) + ' of ' + track.id
221    let answer
222    try {
223      answer = await $.ui.ask('Close ' + what + '?', [CLOSE_IT, 'Cancel'])
224    } catch {
225      return { text: 'NOT CLOSED: nobody confirmed closing ' + what + '.' }
226    }
227    if (answer !== CLOSE_IT) return { text: 'NOT CLOSED: ' + what + ' was left open.' }
228    const text = await relay($, track.close)
229    await refresh($)
230    return { text }
231  })
232
233  // Runs each time a pane is drawn
234  on('ui.render', { component: 'Pane' }, async ($, e, next) => {
235    // Leave other mods' panes alone
236    if (e.requestId !== PANE || !shown) return next(e)
237    const { Box, Markdown } = $.ui.resolve(e)
238    return Box({ flexDirection: 'column', children: [Markdown({ key: 'handoff', text: shown.text })] })
239  })
240
241  // Runs each time the band above the prompt is drawn
242  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
243    const { Box, Text } = $.ui.resolve(e)
244    // What the mods after this one draw there, kept under this one's line
245    const theirs = await next(e)
246    const parts = [Text({ bold: true, children: ['harness'] })]
247
248    if (running) {
249      parts.push(Text({ color: 'yellow', children: ['gates: running'] }))
250    } else if (!facts) {
251      parts.push(Text({ color: 'yellow', wrap: 'truncate-end', children: ['facts not read: ' + unread] }))
252    } else {
253      const g = facts.gates
254      parts.push(Text({ color: COLOURS[g.state], wrap: 'truncate-end', children: ['gates: ' + g.state + (g.why ? ' (' + g.why + ')' : '')] }))
255    }
256
257    if (facts) {
258      const track = facts.tracks[0]
259      const more = facts.tracks.length > 1 ? ' +' + (facts.tracks.length - 1) + ' more' : ''
260      parts.push(Text({
261        dimColor: !track,
262        children: [track ? 'track: ' + track.id + (track.phaseNext === null ? '' : ' phase ' + track.phaseNext) + more : 'no open track'],
263      }))
264      if (facts.findings.length) parts.push(Text({ children: ['findings: ' + facts.findings.length] }))
265      if (facts.problems.length) parts.push(Text({ color: 'red', children: ['unread: ' + facts.problems.length] }))
266    }
267    if (percent !== null) parts.push(Text({ dimColor: true, children: ['context ' + percent + '%'] }))
268
269    const mine = Box({ flexDirection: 'row', columnGap: 2, children: parts })
270    return theirs ? Box({ flexDirection: 'column', children: [mine, theirs] }) : mine
271  })
272}
273