Hands each tool call, agent spawn and file mention to scripts/harness/hook.js and relays the verdict. It decides nothing and draws nothing, and it is not…

Your coding agent's rules, enforced by code instead of by the agent's memory.
You wrote the rules down. The agent followed them, until the session it didn't: it deleted a failing test, quoted last week's test count, read a credential file into the transcript, or called half-finished work done. Nothing told you, because a rule in a prompt cannot fail.
Agent Harness is scripts and hooks you copy into your repo. Every rule that matters becomes a check that fails loudly with an exit code. Anything that is still just a written rule is labelled as one.
Delete a test and an ordinary suite stays green. Here the run is red, and says why:
GATES RED — 2 violation(s) against scripts\harness\gates.baseline.json:
x test: pass 1 < baseline 2
x test: 1 checks in the total, baseline had 2 — a check DISAPPEARED (pass-by-absence)
Node 20 or newer, built-ins only, zero dependencies, no service, no model calls. This repo is checked by the harness it ships.
| Today | How it fails | Here |
|---|---|---|
A rule in CLAUDE.md or AGENTS.md | Followed until it isn't, and nothing reports the lapse | A hook denies the action before it runs, or a gate fails the run |
| "All tests pass" in the agent's summary | The count is stale, or a test was deleted or skipped | Each total is compared with a committed baseline; a missing check is red |
| "Done" typed into a status file | Nobody checked it against the tree as it stands | A phase closes only through a script that sees current evidence |
.gitignore for secrets | It stops the commit, not the read into the transcript | The read itself is denied |
| A long recap pasted into the next session | Every session re-learns what the last one knew | One handoff file per track, and the session starts there |
| An instructions file that keeps growing | It loads on every turn and nobody sees the size | A byte ceiling on always-loaded files, failed in CI |
The hooks deny only under a coding agent they are wired for; everywhere else the same rules run as gates in CI. Your stack says which is which, and What it does not do is the honest list of limits.
Most agent harness repos give you prompts, agent definitions or an orchestration framework. This one gives you none of those as its main product. It gives you the checks around them:
git clone https://github.com/arunpaul-H/coding-agent-harness.git
cd coding-agent-harness
node bin/harness.js selftest # every guard's decisions, driven with literals
node bin/harness.js doctor # every piece: OK / MISSING / DRIFTED, each with its fix
No npm install: there is nothing to install. When you want it in your own repo, go to Install.
Engineers and developers who build with agents: agentic developers, not prompt engineers. If your answer to an agent breaking a rule is a better-worded prompt, this is not that. Here the answer is a gate, a hook or a script that fails the work, so the rule holds whether or not the agent remembers it.
You do not need to read code to use this. You tell an AI what to build, and the harness checks the AI's work for you.
flowchart LR
A[You describe what you want] --> B[The AI writes it]
B --> C{The harness checks it}
C -->|Green| D[Keep going]
C -->|Red| E[The AI is told why and fixes it]
E --> C
What the colours mean
| Result | In plain words |
|---|---|
| 🟢 Green | Every check the project has was run, and none failed or went missing. |
| 🔴 Red | A check failed or disappeared. Nothing is broken for good: the AI is told exactly what to fix. |
| ⚪ Skipped | A check did not run. It proved nothing, so ask why before you rely on the result. |
Your first hour
/guide. It walks you through the rest in plain English, one step at a time.What green does not mean. Green does not mean the idea is good, that customers want it, or that the product is secure. It means the checks your project has all ran and passed. A project with few checks goes green easily.
Everything below this line is written for developers.
People on a flat-rate subscription seat of an agentic coding tool: a plan with a usage limit, not a bill per token. What rations a seat is quota and context, not dollars. So the harness counts what a seat runs out of, the bytes loaded on every turn and the work a session has to redo, and it accounts for nothing in money.
Use it once the same rule has been broken silently more than once, or when work spans sessions and each new session re-learns what the last one knew. For a one-off script or a weekend prototype it is more process than you need.
There is no published package. Installing is copying. You need Node >= 20, git and a POSIX shell (on Windows, the one Git provides).
git clone https://github.com/arunpaul-H/coding-agent-harness.git agent-harness
node agent-harness/bin/harness.js init --into <your-repo> --dry-run # the plan; writes nothing
node agent-harness/bin/harness.js init --into <your-repo>
Cloning a project you will run several agents in? Clone it as git clone <url> <project>/main: every workspace then sits beside main/ in one folder. A plain clone works too (step 5 below).
init never overwrites a file without --force, and it copies no numbers measured on another repo: your baseline starts empty. It writes no .gitattributes; pin your line endings yourself.
To try it here first:
node bin/harness.js selftest # every --self-test in the repo
npm test # node --test
npm run gates # every gate against the committed baseline
node bin/harness.js doctor # every piece: OK / MISSING / DRIFTED, each with its fix
On a fresh clone or a CI runner the docs gate reports one skip and says why. An undeclared skip is red on purpose, so allow it by name:
node scripts/harness/gates.js --check --allow-skip reproducible-build=1,docs=1
AGENTS.md.harness.config.json. It is the only place the harness learns about your project. Add your own checks to its gates array. With no config file, every script stops; there is no default suite.git add first, because the context gate reads tracked files only. $ node scripts/harness/gates.js --check --allow-skip reproducible-build=1
GATES RED — 1 violation(s) against scripts\harness\gates.baseline.json:
x docs: skipped=1 and no --allow-skip docs=n was given
The first run is red, and should be. Read the reason, allow the skip, record the baseline:
$ node scripts/harness/gates.js --accept --allow-skip reproducible-build=1,docs=1
$ node scripts/harness/gates.js --check --allow-skip reproducible-build=1,docs=1
GATES GREEN — test 1/1 . docs 147/148 (1 skipped, declared) . context 6/6 . hygiene 7/7 . skills 6/6 . reproducible-build 0/1 (1 skipped, declared)
gates:), run node agent-harness/bin/harness.js doctor, and put the same --check command in CI. $ node bin/harness.js worktree new fix-export --doing "mend the CSV export"
$ node bin/harness.js worktree list # who is doing what; files two branches changed
$ node bin/harness.js worktree land fix-export # clean, green, merged? prints the merge command
$ node bin/harness.js worktree prune # stale entries; removes only with --yes
Start the agent in the new folder. Nothing here merges or pushes: a person runs the merge. harness doctor says which layout you have and how to adopt the parent-folder one.
What the baseline buys you: delete a test later and nothing fails, yet the run is red.
GATES RED — 2 violation(s) against scripts\harness\gates.baseline.json:
x test: pass 1 < baseline 2
x test: 1 checks in the total, baseline had 2 — a check DISAPPEARED (pass-by-absence)
Full walk-through, including how to confirm the hooks fire: docs/methods/11-adopting-this-in-your-repo.md.
| Piece | What it stops | Where |
|---|---|---|
| Gates and a baseline | Stale totals in docs; a check that vanished while the suite stayed green | scripts/harness/gates.js |
| Reproducible-build check | A source edit committed without its rebuilt bundle | scripts/harness/gates.js |
| Secret-file guard | A credential file read into a transcript | scripts/harness/guards/env-guard.cjs |
| Phase guard | A status flipped by hand; a gate figure typed into a doc | scripts/harness/guards/phase-guard.cjs |
| Commit hygiene | Stray files, pasted tokens and root clutter in a commit | scripts/harness/guards/commit-hygiene.cjs |
| Spawn guard | A subagent of yours spawned with no OUTCOME, READ-SET, RETURN or BOUNDS | scripts/harness/guards/spawn-guard.cjs |
| Stop report | A turn that ends on edits no gate run has seen (it tells you; stop.block forces a turn) | scripts/harness/close-phase.js |
| Doc contract | Docs over their line cap; an index that drifted from the docs | scripts/harness/guards/docs-lint.cjs |
| Context budget | Always-loaded instruction files growing unnoticed | scripts/harness/measure-context.js |
| Tracks and handoffs | A new session re-deriving what the last one knew | templates/track/ |
| Phase lifecycle | A phase closed on evidence from a different tree | scripts/harness/close-phase.js |
| Skill lock | A skill changing with no diff anyone read: a third-party one, or one the harness shipped | scripts/harness/guards/skill-verify.cjs |
| Branch sync | A long-lived branch running an old copy of the harness | scripts/harness/branch-sync.js |
| Installer and doctor | An install nobody verified | bin/harness.js |
| Panel mod (Claude Code) | A gate verdict that scrolled away; a model turn spent to run the gate or close a phase | .claude/plugins/harness/ |
| Guard mod (Claude Code) | A spawn no Agent call made, or a file attached with @, that no settings hook sees | .claude/plugins/harness-guard/ |
Each piece has a page under docs/methods/: the failure, the mechanism, its cost and its limits.
The harness writes no doc about your code. init installs a template and nothing else, because a doc generated from code repeats what a search already finds. When a subsystem has rules that live in someone's head, ask for one doc, by name:
node agent-harness/bin/harness.js reference new billing-export \
--read-when "changing what the nightly export writes" --paths "src/billing/**,tests/billing/**"
node scripts/harness/docs-index.js
That writes docs/reference/billing-export/README.md with the frontmatter filled and the body left for you. It never overwrites, and both flags are required: --read-when is the one line an agent sees in the index, and --paths is what the doc gate compares against history.
last_verified day. Re-check the doc, then change the date. Changing only the date defeats the warning.What we measured, once: the same read-only question put to two fresh agents on the same model, one tree with a hand-filled doc and one without. Both answered every part correctly. The run with the doc used 65083 tokens against 78068, the same 18 tool uses, and one fewer file; it still read the code, because the question asked for line numbers. One pair is an observation, not proof, and it does not count the time spent writing the doc. Expect a modest saving at best, and write a doc for what the code cannot say, not to save tokens. The write-up is in git history: git show b2183f0:docs/tracks/project-reference/research/2026-10-08-one-task-with-and-without-a-reference-doc.md.
GATE <name> pass=<n> fail=<n> skipped=<n> the counts are read exactly; otherwise the exit-code adapter grades it as one pass or one fail. init detects Node, Go, Rust and Python test commands.harness init --tools <ids> names the coding agents to wire. A hook file is written for Claude Code and, for shell calls and file reads only, for Cursor. Every other tool below gets the policy, the skills and a stated gap. The agent files and the status line are Claude Code's only.ran means the tool was run here and the row records the command. docs means the row was read off the vendor's pages and the tool was never run. Printed by node scripts/harness/tools.js --list, 2026-10-07:<!-- tools-list:start -->
id tier evidence enabled
claude-code full ran yes
codex-cli full docs
gemini-cli full docs
antigravity guard-only docs
cursor full docs
kimi-code guard-only docs
deepseek inherits docs
github-copilot backstop docs
windsurf backstop docs
cline backstop docs
aider backstop docs
opencode backstop docs
qwen-code backstop docs
<!-- tools-list:end -->
node scripts/harness/tools.js --matrix prints it per enabled tool: a deny at the tool call, a report, a gate, a git-hook backstop, or prose.AGENTS.md, which most coding agents load; CLAUDE.md is a one-line pointer to it.gates.js --check in CI, and the opt-in git hooks: harness init --git-hooks writes a pre-commit that runs commit-hygiene.cjs --report and a pre-push that runs the gate suite. A backstop after the work, not a deny.A mod is a plugin whose code runs inside Claude Code (terminal, v2.1.287 and later). This repo carries two, and neither holds a decision: each runs a script under scripts/harness/ and draws or relays what it printed.
harness draws a band above the prompt (the recorded gate verdict, the open track, the context in use) and adds /gates, /track, /handoff and /close-phase, none of which takes a model turn. It never refuses a call.harness-guard hands each tool call the settings hook also sees, each agent spawn and each @ file mention to scripts/harness/hook.js and relays the verdict. A commit the hygiene guard refuses is held and shown with its findings: you have it checked again, or cancel. Each of its hooks refuses when it fails. It draws nothing.No settings file enables either. Load them for one session, from the repo root:
claude --plugin-dir .claude/plugins/harness --plugin-dir .claude/plugins/harness-guard
What is and is not established:
claude plugin validate --strict and claude plugin test run as the optional mods gate. Where the claude program is absent, that gate is a declared skip: the harness depends on no agent.harness-guard was loaded in claude -p sessions, where it passed a harmless call and refused a read of a guarded secret file. No interactive session has loaded either mod, so the band, the commands and every dialog are held by tests alone.harness-guard can answer a call before it is asked, and another mod can approve a call a settings hook refused. Only a hook in managed settings is final. The settings hooks stay; the mods are an addition for Claude Code.init yet. They live in this repo; copy the two folders to use them elsewhere.The measurements are in docs/tracks/claude-code-uptake/research/.
bin/ harness.js: the installer and CLI (init, doctor, gates, docs, host, selftest)
scripts/harness/ the gate runner, config loader, lifecycle and doc scripts, baselines
scripts/harness/guards/ the guards, the doc linter, the skill verifier
.claude/ hook wiring, four agent briefs, four skills, the status line
.claude/plugins/ two Claude Code mods: harness (panel and commands), harness-guard (guards)
schema/ the JSON schema for harness.config.json
templates/ starting files for AGENTS.md, CLAUDE.md, the routing doc and a track
docs/methods/ one page per mechanism
docs/reference/ the enforced shapes: doc standard, routing, skill safety, branches
docs/tracks/ open work on this repo
tests/ node:test; every script run as a real process in a throwaway repo
.github/ the CI workflow, Dependabot and scanner configs
harness.config.json holds everything project-specific. AGENTS.md is this repo's own policy and the starting point for yours.
See CONTRIBUTING.md. Changes are listed in CHANGELOG.md.
MIT. See LICENSE.
hooks/register.js 137 lines1/* ============================================================
2 The harness guards, in-process: a tool call, an agent spawn and a file mention are each handed
3 to scripts/harness/hook.js, and what it answered is relayed.
4
5 THIS FILE HOLDS NO DECISION. Every call is spelled as the pre-tool payload the settings hook
6 receives and piped to the same entry, so the verdict is the one the guards under
7 scripts/harness/guards/ give, each with its own --self-test. A mod has no Node built-ins and no
8 `require`, so anything decided here would be out of reach of `node --test`. What this file
9 owns: the payload's spelling, the wording around a relayed reason, and the question it asks
10 while a refused commit is held.
11
12 IT IS NOT FINAL. It runs in the user tier: a mod loaded ahead of it can answer a call before
13 it is asked, and a mod's `tool.check` can approve what it refused. Only a hook in managed
14 settings is final. The settings hooks stay; for a tool call this is a second hearing of what
15 they also judge. What only this file reaches is a spawn no `Agent` call made (a workflow
16 agent, a teammate) and a file attached by a mention.
17
18 IT FAILS CLOSED. Each hook has a `.catch` that refuses, and an entry that cannot be run or
19 read is a refusal. So a session started anywhere but the repo root, where the relative path
20 below does not resolve, has every guarded call refused: start it at the root.
21
22 IT DRAWS NOTHING: no ui.render hook. `claude plugin validate` prints the proof on its
23 `hooks:` and `calls:` lines.
24 ============================================================ */
25
26// The one script this mod runs, relative to the folder the session runs in
27const HOOK = ['node', 'scripts/harness/hook.js', '--tool', 'claude-code', '--event', 'pre-tool']
28// The wait the settings registration gives the same entry
29const THIRTY_SECONDS = 30000
30// How the hygiene guard opens its line for the person. Chooses the dialog, never the verdict.
31const HYGIENE = 'Commit hygiene gate:'
32const CHECK_AGAIN = 'Check again'
33const CANCEL = 'Cancel the commit'
34const ALLOW_IT = 'Allow it'
35const REFUSE_IT = 'Refuse it'
36// How many times a held commit is checked again before it is refused
37const HELD_CHECKS = 5
38const QUESTION_MAX = 1500
39
40// Where the session runs, once it has said
41let cwd = null
42
43// Ask the entry about one call. Returns { deny, ask, reason, message }; throws if it cannot be run.
44async function judge($, tool, input) {
45 const payload = { hook_event_name: 'PreToolUse', tool_name: tool, tool_input: input }
46 if (cwd) payload.cwd = cwd
47 const r = await $.process.run(HOOK, { stdin: JSON.stringify(payload), timeoutMs: THIRTY_SECONDS })
48 const out = String(r.stdout || '').trim()
49 let said = null
50 if (out) {
51 try {
52 said = JSON.parse(out)
53 } catch {
54 said = undefined
55 }
56 }
57 const spoken = (said && said.hookSpecificOutput) || {}
58 const reason = String(spoken.permissionDecisionReason || r.stderr || '').trim()
59 const message = String((said && said.systemMessage) || '')
60 // Exit 2 is the entry's refusal. Any other failure is an entry that did not look, which is not a pass.
61 if (r.exitCode !== 0) {
62 return { deny: true, ask: false, message, reason: reason || 'scripts/harness/hook.js left with exit ' + r.exitCode + ' and no reason, so no guard looked at this call.' }
63 }
64 if (said === undefined || spoken.permissionDecision === 'deny') {
65 return { deny: true, ask: false, message, reason: reason || 'scripts/harness/hook.js printed an answer that could not be read, so this call was not run.' }
66 }
67 return { deny: false, ask: spoken.permissionDecision === 'ask', reason, message }
68}
69
70// Put a question to the person. Nobody there, or a dismissed question, is the safe answer.
71async function ask($, question, options, safe) {
72 try {
73 return await $.ui.ask(question.slice(0, QUESTION_MAX), options)
74 } catch {
75 return safe
76 }
77}
78
79// An ask from the guards where no permission prompt follows: a person decides, or it is refused
80async function decided($, verdict, what) {
81 if (!verdict.ask) return verdict
82 const answer = await ask($, 'The harness guards could not read ' + what + ', so you decide.\n' + verdict.reason, [ALLOW_IT, REFUSE_IT], REFUSE_IT)
83 if (answer === ALLOW_IT) return { ...verdict, ask: false }
84 return { ...verdict, deny: true, ask: false, reason: verdict.reason + '\nNobody allowed it, so it was refused.' }
85}
86
87// What a `.catch` answers with: the refusal, unless the hook had already passed the call on
88function refusal(e, next, what) {
89 if (next.called) return next(e)
90 const kind = (next.error && next.error.kind) || 'failed'
91 return { deny: 'harness-guard could not ask scripts/harness/hook.js about ' + what + ' (' + kind + '), so it was refused. Start the session at the repo root, or run `node scripts/harness/hook.js --self-test`.' }
92}
93// One per gating hook: the validator reads a `.catch` handler only by its own name
94async function callFailed($, e, next) { return refusal(e, next, 'this tool call') }
95async function spawnFailed($, e, next) { return refusal(e, next, 'this spawn') }
96async function mentionFailed($, e, next) { return refusal(e, next, 'this mention') }
97
98export function register(on) {
99 // Runs before the first prompt, and again after a reload
100 on('session.start', async ($, e, next) => {
101 cwd = typeof e.cwd === 'string' && e.cwd ? e.cwd : null
102 return next(e)
103 })
104
105 // Every tool the settings matcher names, for the lead and for every agent. The names are
106 // written out so `claude plugin validate` can print them, and tests/harness-guard.test.js
107 // holds them to the dialect's list.
108 on('tool.call', { tool: /^(Bash|PowerShell|Read|Grep|Glob|Edit|Write|Monitor|NotebookEdit|Agent|mcp__.*)$/ }, async ($, e, next) => {
109 const { tool, tool_use_id, agentId, ...input } = e
110 let verdict = await judge($, tool, input)
111 // A refused commit is held: the person fixes the tree and has it checked again, or cancels
112 for (let checks = 0; verdict.deny && verdict.message.startsWith(HYGIENE) && checks < HELD_CHECKS; checks++) {
113 const answer = await ask($, verdict.reason, [CHECK_AGAIN, CANCEL], CANCEL)
114 if (answer !== CHECK_AGAIN) return { deny: verdict.reason + '\nThe commit was cancelled. Do not retry it until the findings above are fixed.' }
115 verdict = await judge($, tool, input)
116 }
117 if (verdict.deny) return { deny: verdict.reason }
118 // An ask is left to the permission check under this hook, which the settings hook answers too
119 return next(e)
120 }).catch(callFailed)
121
122 // Every spawn, including the ones no `Agent` call made: a workflow agent and a teammate
123 on('agent.spawn', async ($, e, next) => {
124 const input = { subagent_type: e.subagentType, prompt: e.prompt, description: e.description }
125 const verdict = await decided($, await judge($, 'Agent', input), 'this spawn')
126 if (verdict.deny) return { deny: verdict.reason }
127 return next(e)
128 }).catch(spawnFailed)
129
130 // A file named with @ is attached without a Read call, so no settings hook sees it
131 on('prompt.mention', async ($, e, next) => {
132 const verdict = await decided($, await judge($, 'Read', { file_path: e.path }), 'this mention')
133 if (verdict.deny) return { deny: verdict.reason }
134 return next(e)
135 }).catch(mentionFailed)
136}
137