Turns a pull request into a published review agenda — what changed, the judgments a reviewer has to make with the exact lines that settle each one, the order…

<h1 align="center"> <img src="examples/assets/logo.png" alt="accountable-review" width="560"> </h1>
AI-assisted code review for teams that want to move faster with coding agents without losing control of their codebase.
accountable-review turns a pull request into a Review Map: a published HTML page that guides a reviewer through what changed, how the new behaviour works, what existing code is affected, and what to understand before merge.
It ships two skills: review-map, which produces the page for a Rails or a Rust codebase — a Rust backend or Rust in general — with or without a separate client such as Next.js, with Elixir/Phoenix and React (a Next.js app or React in general) support still in development; and setup-ci, which arranges for one to be produced automatically on every review-ready pull request.
Built by WyeWorks.
AI coding tools can produce large changes much faster than teams can absorb them.
Tests can pass. AI reviewers can find issues. Coding agents can fix those issues.
And still:
Do you understand how the system works now?
That is the problem accountable-review focuses on.
When code enters the codebase faster than the team understands it, short-term speed can become long-term loss of control. That gap has a name — comprehension debt — and it comes due the first time someone has to change the code nobody read.
The answer is not less AI. It is more reviewer understanding alongside it — more productivity and more comprehension, rather than one bought with the other:
Move faster with AI. Stay in control.
AI code review is great at asking:
accountable-review helps the human reviewer ask:
The plugin does not approve the PR for you.
It helps a competent reviewer understand the change well enough to decide.
No approval verdict. No severity score. No fake confidence badge.
Human judgment stays final. A page good enough to approve from without reading the code would be a failure — the reviewer would be holding a verdict instead of a mental model.
Diffs are excellent for reviewing code line by line.
Is this condition correct?
Is this query efficient?
Did this line introduce a regression?
But modern AI-assisted development also pushes review to a higher level:
What behaviour changed?
How does this work across the system?
What must the team understand to maintain it?
A Review Map adds that layer.
A diff is organized by files.
A Rails, Rust, Phoenix or Next.js feature is not.
One behaviour may cross:
routes
↓
controller
↓
model / service
↓
policy
↓
job
↓
serializer
↓
frontend contract
On a Rust backend the hops are different — router and its layers, extractor, handler, error mapping, query, migration — and on Phoenix different again — router, controller or LiveView, context, changeset, Repo, worker, template — and on Next.js different again — page, server component, client component, server action, cache tag, route handler — and the point is the same: no directory contains the behaviour.
accountable-review traces the PR around the behaviour being implemented, not the order of files in the diff — and then publishes the part of that work you have to act on: the handful of judgments the change asks of you, and where to look to make each one.
A PR can change the meaning of code without changing its lines.
Examples:
A normal diff cannot show those relationships. A Review Map can.
A Review Map is a review agenda: the smallest set of things you have to judge before you can approve a change, with the code that settles each one attached.
Five parts, and a shut evidence block at the foot:
What changed one paragraph: what is now true that was not
Context when earned: the pieces of this repository the judgments rely
on, for a reader who knows the stack but not the codebase
What needs your attention 3-5 checkpoints per independent change the PR makes, 7 at the
outside; each one judgment, framed as a question, and at
most one small figure where its shape is clearer than prose:
a chain, paths converging on one invariant, a lifecycle, or
an entity and its relationships
Read the code in this order 3-7 stops, in the order that builds understanding
Impact outside the diff 1-3 chains, from changed code into code it gives new meaning to,
each node linked to the file it lives in
▸ Evidence & diff coverage every changed path, the searches run, the rest of what was found
Claims about unchanged code come with the code attached: a collapsed excerpt of the real source, quoted by a script rather than retyped, that you open when you are ready to check that particular claim. The page reads completely with every excerpt closed — the sentence carries the consequence, the excerpt carries the proof.
The Review Map does not replace the diff.
It helps the diff make sense.
📖 Full anatomy of the page — every section, the checkpoint, the staging behaviour, the excerpt and framework-anchor rules — is in docs/review-map.md.
Add the WyeWorks marketplace and install, from inside Claude Code:
/plugin marketplace add wyeworks/claude-plugins
/plugin install accountable-review@wyeworks
or from a terminal:
claude plugin marketplace add wyeworks/claude-plugins
claude plugin install accountable-review@wyeworks
Restart, and the skills are available in every project as /accountable-review:review-map and /accountable-review:setup-ci. claude plugin update accountable-review@wyeworks takes the next release; you only receive one when the version in the plugin's manifest moves, so an update is something you ask for.
Add --scope project to both commands to declare the plugin in a repository's own settings, so everyone working on it gets the same one.
To run it from a checkout instead — which is what you want if you are changing the plugin, since /reload-plugins then picks up your edits without restarting:
git clone https://github.com/wyeworks/accountable-review.git
cd /path/to/your/app
claude --plugin-dir /path/to/accountable-review
Nothing is installed that way, so git pull in the checkout is how you update, and the skills exist only for sessions started with that flag.
This repository is its own Codex marketplace. From a terminal:
codex plugin marketplace add wyeworks/accountable-review
codex plugin add accountable-review@accountable-review
Restart Codex, and the skill is available in every project as $accountable-review:review-map — the same name as in Claude Code, behind Codex's $ instead of /. codex plugin marketplace upgrade accountable-review fetches the latest catalogue, and as with Claude Code you only receive a release when the version in the plugin's manifest moves. codex plugin remove accountable-review@accountable-review uninstalls it.
Open the repository you want to review in Codex, then invoke:
$accountable-review:review-map
The result is a local HTML file, linked from the response, using the same template and completeness gate as Claude. To choose its destination, add --output /absolute/path/outside-the-repo; the page is written as index.html there. No hosting service or publishing plugin is required.
High effort is the default and requires Codex subagent tools for the independent falsification pass. If those tools are unavailable, the skill reports the limitation; use --effort low explicitly to run without independent readers. The page never presents that pass as an approval.
Codex CI execution is not supported yet. setup-ci and the CI runner still use Claude Code and Anthropic credentials, so the Codex plugin ships review-map alone.
To run it from a checkout instead — which is what you want if you are changing the skill, since edits apply without reinstalling — link it into Codex's skills directory:
mkdir -p ~/.agents/skills
ln -s /path/to/accountable-review/skills/review-map ~/.agents/skills/review-map
Linked that way it is a standalone skill invoked as $review-map, and git pull in the checkout is how you update. Use the marketplace or the link, not both, or Codex lists the skill twice.
See Codex setup and verification for removal and the boundaries of this integration.
The skills live under skills/ as ordinary Agent Skills, so the cross-agent skills CLI installs them into Claude Code and Codex without cloning anything:
npx skills add wyeworks/accountable-review
It asks which skills and which agents; answer in flags instead with, for example:
npx skills add wyeworks/accountable-review -g --skill review-map -a claude-code -a codex
Prefer -g. Without it the skills are copied into the current project (.agents/skills/, .claude/skills/ and so on), which puts files in the repository you are about to review. npx skills update takes new commits, and npx skills remove review-map uninstalls.
Installed this way the skill is a standalone one rather than part of a plugin, so in Claude Code it is invoked as /review-map instead of /accountable-review:review-map, and in Codex as $review-map instead of $accountable-review:review-map. --effort high still works in Claude Code: the independent readers are launched as general-purpose subagents handed the same bundled mandate, rather than as the plugin's registered claim-falsifier agent. Install through the plugin or the skills CLI, not both, or Claude Code lists the skill twice.
The CLI can place the skills in any other agent it supports, Pi for example. Outside Claude Code, review-map follows one set of instructions whichever agent it is: the page is a local HTML file, and --effort high needs an independent reader — the agent's own subagents, or else a read-only background run of its own CLI. An agent with neither says so and offers --effort low. Only Claude Code and Codex are tested.
setup-ci is offered too — /setup-ci in Claude Code, the only host it works in. The workflow it writes clones the plugin itself at a pinned tag, so how you installed the skill locally does not matter to CI.
Run it from inside the repository you are reviewing:
/accountable-review:review-map # current branch against its base
/accountable-review:review-map 412 # a PR number
/accountable-review:review-map https://github.com/org/repo/pull/412
/accountable-review:review-map feature/some-branch
In Claude Code you get back a URL. The page is a private Claude Artifact until you share it, and re-running for the same PR republishes to the same URL — so the Review Map tracks the PR across pushes instead of scattering links.
It publishes early and fills in as parts complete: open it at minute two, watch it arrive, start reading the moment the part you need lands. The checkpoints arrive one at a time, and their questions arrive first — so you know what the change is asking you to judge well before the explanations land. While it is unfinished it says so in a banner, and every part still coming is marked pending, so a half-written page can never be mistaken for a finished one.
While it runs, Claude Code's status line shows what it is doing — tracing consumers, cutting excerpts, writing the page — with the elapsed time and how many checkpoints are written. Every field is read from the run rather than estimated. The one exception is a bar, labelled ~. It measures against earlier runs of the same kind in that repository, so it appears once one such run has finished with its coverage gate passed.
To have one generated for every pull request instead of by hand, see CI integration below.
This generates the default Review Map for PR #412. The skill works best for PRs that are about to be merged, but it also works well with already-merged PRs when you want to catch up on changes.
/accountable-review:review-map 412
What makes it onto the page is the part you actually need to act on. For a small or medium-sized PR, that usually means around 700 to 1,500 words.
/accountable-review:review-map 412 # --effort high, the default
/accountable-review:review-map 412 --effort low # opt out of the falsification pass
Everything the skill writes rests on claims it checked itself — and the context that wrote a claim is the one least able to see what it assumed. --effort high, the default, adds a second reader that does not share that context. The run writes its analysis down first, one note per behaviour, and sends a read-only claim-falsifier subagent at each note before any of it reaches the page, with a single mandate: assume this is wrong in ways that matter, and find evidence in the repository that contradicts it. It rewrites nothing; it returns challenges, each anchored in a line it opened. The run then opens the cited file itself and corrects, downgrades or drops the claim. A challenge it cannot confirm is dropped, exactly as an unconfirmed finding is.
It is on by default because it is very nearly free and it changes what the page finds: on a 28-file pull request the whole pass cost 23 seconds of waiting, under 1% of the run, and the same change reviewed without it missed five things the falsified page carried. --effort low turns it off, which is worth doing when the diff is small enough that a second reader has nothing to find.
A third axis, and the only one that puts anything on the page:
/accountable-review:review-map 412 --mentor # add framework primers
/accountable-review:review-map 412 --mentor rails # the same, naming the stack
A Review Map normally assumes you know the framework and are meeting this change for the first time. --mentor is for the other case — a reviewer new to Rails, to Rust, to Phoenix or to React, on their first pull requests in an unfamiliar codebase. Where a judgment turns on a framework rule they may not know, the page stops linking to the manual and states the rule: a short primer inside that checkpoint, with the API named, the behaviour explained, a worked example on a generic class, the line in your repository that made it relevant, and the pinned documentation link it came from.
Phoenix and React today: a primer is gated on the documentation link it escalates from, and the Elixir and React catalogues ship closed until a verification run has opened every row in them. So
--mentoron a Phoenix, React or Next.js project currently produces no primers and says so. That is the fail-closed rule doing its job — a page with no primer is narrower, a page with an invented link is wrong. The Rust catalogue shipped closed the same way and has since been opened, so a Rust page can carry primers, framed in graphite rather than Rails red.
/accountable-review:review-map 412 --update # re-read only the new commits
A Review Map describes one revision, and a pull request that keeps gaining commits after review opens ends up with a map that quietly describes an older one. --update is the cheap way to keep it current: it reads the commits since the page it is updating, re-analyses the judgments those commits reach, and edits that page in place at the same URL. What the new commits did not touch is carried, which is where the saving comes from — tracing consumers across the diff is most of a run's cost, and an update traces only what moved.
You are trading re-reading for speed, and the page says so. A carried judgment is one this run did not verify again. So an updated page names both revisions in its masthead and carries one sentence saying which parts still describe the earlier one. If you want the fuller read, run it again without --update — a map generated from scratch describes one revision throughout, and that is the right call before a final review pass on a branch that has moved a lot.
It refuses rather than guessing. A force-push or rebase, a base branch that moved underneath, a delta covering more than half the diff, a lock file bump, or a recorded search that now finds the changed code — each of those means the previous page is not a safe thing to build on, so the run regenerates from scratch and tells you which one it hit.
The tool separates what is known from what is inferred.
A statement may come from:
A claim the diff shows directly carries no label. Everything else is marked — from unchanged code, inferred from tests, inferred, uncertain — so an inference can never pass as a fact. Where the tool cannot establish intent, it says so:
It is unclear whether existing time entries stay editable after archival. No test covers it.
Every claim also carries a file:line into your repository, and by default a second reader attacks those claims before the page is finished (see Usage → effort).
Nor does the page claim to have found everything — see What a Review Map cannot do below. It never reads as a clean bill of health, and it never carries a boilerplate disclaimer saying so either: the limits are the same on every Review Map, so they are written down once, here, rather than reprinted under a heading you have already read a dozen times.
The goal is not to sound confident. The goal is to help the reviewer investigate the change.
A Review Map is written by a model reading your repository. That is what lets it trace a consequence into code the diff never opened, and it is also the honest limit on the page.
file:line, so open the citation for anything you would act on.And it depends on the model behind it. We develop and test with Claude Opus 5 most of the time, and that is what the page's depth is calibrated against. Other models will trade cost for reach differently — try a few against your own codebase and keep the one whose maps you actually trust.
The plugin is optimized and most heavily tested for Rails applications. That matters because Rails behaviour often emerges from several pieces working together:
The plugin is designed to help reviewers understand those relationships as a system, and it knows where the framework's own rules bite: update_all at a call site the diff never opened skips the validation this PR adds, a uniqueness validation is not a unique index, --sandbox rolls back so after_commit never fires there.
It also reads the app for what your team already decided. A value object in app/models is ordinary; a value object in app/models when four of its kind live in app/services is a choice someone made, and the page will ask whether it was deliberate — with the four siblings cited, because without them there is nothing to ask. That question is never a recommendation, never appears more than once, and never takes a slot from a judgment about behaviour.
Rust is also supported and tested, and it comes in two lenses, because what a reviewer looks for in a crate that serves requests and in one that does not overlaps less than the language suggests. Rust in general — a library, a CLI, a workspace that serves nothing — is reviewed against the things the compiler was told not to check or cannot see: the _ => arm a new enum variant falls into silently, a pub item a downstream crate depends on, a feature combination CI never builds, the module whose safe code keeps an unsafe block sound. A Rust backend — axum, actix-web, tonic and the rest — gets that l
hooks/progress.ts 237 lines1// progress.ts — the e2e progress board (skills/review-map/evals/e2e/progress.rb), as one status
2// line in Claude Code while /accountable-review:review-map runs interactively.
3//
4// 🌔 review map ▕███████████▍ ▏ 47% 11:02 / ~23:30 🔎 tracing consumers 🔧 38 🥊 4 📄 2/4
5//
6// Every field is READ, never guessed, and that is the rule this file keeps: a progress line that
7// invents progress is the verification badge of the terminal, and the page itself is forbidden
8// that badge. Three sources, each one the engine or the run already produces:
9//
10// * engine events — asking for the skill starts the run, each main-loop tool call names the
11// activity through hooks/activity.json (the table progress.rb reads too), each claim-falsifier
12// spawn is counted, and the coverage gate passing then the turn answering ends the run;
13// * the page the run is staging — checkpoints written against pending ones;
14// * the clock, against the median of this repository's earlier generations of the same kind,
15// which this file records itself. With none recorded there is no bar, only elapsed time.
16//
17// The activity is the latest tool call by purpose, never a step number: steps interleave and
18// cannot be read off a run (evals/README.md § Profiling one run), and this does not pretend
19// otherwise. It observes and nothing else — every hook passes its event on unchanged, and every
20// one has a .catch that still does, so a bug in the line can never stall a generation.
21
22import type { EngineInterface, Register, Timer } from 'claude-code'
23
24import {
25 type Activity,
26 type RunKind,
27 type Vars,
28 assignedW,
29 checkpoints,
30 label,
31 line,
32 median,
33 outputPage,
34 pagePath,
35 parseActivity,
36 repoKey,
37 runKind,
38 runsGate,
39 summary,
40} from './board'
41
42const REVIEW_MAP = /(^|:)review-map$/
43const FALSIFIER = /(^|:)claim-falsifier$/
44const KEEP = 20 // generations remembered per repository and kind
45
46type Run = {
47 started: number
48 paused: number // milliseconds spent waiting between turns, which are not generation
49 waitingSince?: number
50 ended?: 'done' | 'stopped' | 'interrupted'
51 endedAt?: number
52 activity?: string
53 tools: number
54 falsifiers: number
55 page?: string
56 vars: Vars
57 checkpoints?: string
58 sawGate: boolean
59 kind: RunKind
60 eta?: number
61 repo?: string
62 tick: number
63 timer?: Timer
64}
65
66// The module's own variables start over on a reload, as register runs again.
67let table: Activity = []
68let run: Run | undefined
69
70function elapsedMs(r: Run, now: number) {
71 return (r.endedAt ?? r.waitingSince ?? now) - r.started - r.paused
72}
73
74async function draw($: EngineInterface) {
75 if (!run) return
76 const now = await $.clock.now()
77 $.ui.status(line({ ...run, waiting: run.waitingSince !== undefined, elapsed: elapsedMs(run, now) / 1000 }))
78}
79
80async function recount($: EngineInterface) {
81 if (!run?.page) return
82 try {
83 run.checkpoints = checkpoints(await $.fs.read(run.page))
84 } catch {
85 // a page mid-Edit or not written yet: the next page-touching call will have it
86 }
87}
88
89function tick($: EngineInterface, r: Run) {
90 r.timer?.cancel()
91 r.timer = $.clock.every(1000, () => {
92 r.tick += 1
93 void draw($)
94 })
95}
96
97function historyKey(repo: string, kind: RunKind) {
98 return `generations:${kind}:${repo}`
99}
100
101// The command as typed, or as the engine records a typed skill command; a sentence that merely
102// mentions review-map is not a request for one.
103function isCommand(text: string) {
104 return /(^\s*|<command-name>)\/(accountable-review:)?review-map(?=\s|<|$)/.test(text)
105}
106
107async function start($: EngineInterface, command: string) {
108 run?.timer?.cancel()
109 if (table.length === 0) {
110 table = parseActivity(await $.fs.read(`${$.plugin.root}/hooks/activity.json`))
111 }
112 const kind = runKind(command)
113 const repo = await $.session.repo()
114 const key = repo ? repoKey(repo.remote, repo.root) : undefined
115 const past = key ? await $.store.get(historyKey(key, kind)) : undefined
116 const mine: Run = {
117 started: await $.clock.now(),
118 paused: 0,
119 tools: 0,
120 falsifiers: 0,
121 sawGate: false,
122 kind,
123 page: outputPage(command),
124 vars: { TMPDIR: await $.env.get('TMPDIR') },
125 tick: 0,
126 repo: key,
127 eta: Array.isArray(past) ? median(past.filter((n): n is number => typeof n === 'number')) : undefined,
128 }
129 run = mine
130 tick($, mine)
131 await draw($)
132}
133
134export const register: Register = on => {
135 table = []
136 run = undefined
137
138 // A run starts where the skill is asked for, in one of the two ways a person or the model can ask:
139 // the command typed as a prompt, or a Skill tool call naming it. `skill.prompt` would be the
140 // obvious event and is not one this mod can use: the engine's own security module sends it
141 // beneath the user tier, so a plugin's hook on it never runs (the debug log says "skill.prompt
142 // bypassed by cc-plugin-sec-default").
143 //
144 // A prompt typed over a running turn (`turnId`) is a note to that turn, not a new run. Any new
145 // prompt takes down the line a finished run left standing.
146 on('prompt.submit', async ($, e, next) => {
147 if (run?.ended) {
148 run = undefined
149 $.ui.status(undefined)
150 }
151 if (e.turnId === undefined && isCommand(e.text)) await start($, e.text)
152 return next(e)
153 }).catch(($, e, next) => next(e))
154
155 // The main loop only: progress.rb reads the parent transcript, and a falsifier's own reads are
156 // its work, not the run's activity. Its spawn is what gets counted, below.
157 on('tool.call', async ($, e, next) => {
158 if (e.agentId !== undefined) return next(e)
159 const args = e as unknown as Record<string, unknown>
160 const asked = e.tool === 'Skill' && REVIEW_MAP.test(String(args.skill ?? ''))
161 if (asked && (!run || run.ended)) {
162 await start($, `/review-map ${String(args.args ?? '')}`)
163 return next(e)
164 }
165 if (!run || run.ended) return next(e)
166 const mine = run
167 const text = summary(args)
168 mine.tools += 1
169 mine.activity = label(table, text) ?? mine.activity
170 if (typeof args.command === 'string') mine.vars.W = assignedW(args.command, mine.vars) ?? mine.vars.W
171 mine.page = pagePath(args, mine.vars) ?? mine.page
172 await draw($)
173 const result = await next(e)
174 // Passing, not mentioning: the gate exits non-zero when the inventory and the diff disagree.
175 if (runsGate(args) && result.deny === undefined && result.isError !== true) mine.sawGate = true
176 if (mine.page && /page-skeleton\.sh|\.html\b/.test(text)) {
177 await recount($)
178 await draw($)
179 }
180 return result
181 }).catch(($, e, next) => next(e))
182
183 on('agent.spawn', async ($, e, next) => {
184 if (run && !run.ended && FALSIFIER.test(e.subagentType)) {
185 run.falsifiers += 1
186 await draw($)
187 }
188 return next(e)
189 }).catch(($, e, next) => next(e))
190
191 // A run can span turns: the skill stops to ask (two stacks, say), or the main loop ends its turn
192 // while background falsifiers read and is woken by their notifications. So a turn that answers
193 // before the gate has passed pauses the run rather than ending it, and the next main-loop turn
194 // resumes it; the time between is the person's or the notification's, not generation's.
195 on('turn.start', async ($, e, next) => {
196 if (run && !run.ended && run.waitingSince !== undefined) {
197 run.paused += (await $.clock.now()) - run.waitingSince
198 run.waitingSince = undefined
199 tick($, run)
200 await draw($)
201 }
202 return next(e)
203 }).catch(($, e, next) => next(e))
204
205 // Its length feeds the next run's ETA only when the gate passed and the turn answered, so a run
206 // that stopped early never shortens the bar.
207 on('turn.complete', async ($, e, next) => {
208 if (run && !run.ended && e.agentId === undefined) {
209 const mine = run
210 mine.timer?.cancel()
211 const now = await $.clock.now()
212 if (e.reason === 'answer' && !mine.sawGate) {
213 mine.waitingSince = now
214 } else {
215 mine.endedAt = now
216 mine.ended = e.reason === 'answer' ? 'done' : e.isAborted ? 'interrupted' : 'stopped'
217 }
218 await recount($)
219 await draw($)
220 if (mine.ended === 'done' && mine.repo) {
221 const key = historyKey(mine.repo, mine.kind)
222 const past = await $.store.get(key)
223 const kept = Array.isArray(past) ? past.filter(n => typeof n === 'number') : []
224 await $.store.set(key, [...kept, elapsedMs(mine, now) / 1000].slice(-KEEP))
225 }
226 }
227 return next(e)
228 }).catch(($, e, next) => next(e))
229
230 on('session.end', async ($, e, next) => {
231 run?.timer?.cancel()
232 run = undefined
233 $.ui.status(undefined)
234 return next(e)
235 }).catch(($, e, next) => next(e))
236}
237hooks/board.ts 162 lines1// board.ts — the pure half of the progress line: what a field says, given what was read. Nothing
2// here reads anything; progress.ts does the reading, and this file turns readings into one line.
3// The rule both keep is progress.rb's: every field is READ, never guessed.
4
5export type Activity = readonly (readonly [RegExp, string])[]
6
7export const MOON = ['🌑', '🌒', '🌓', '🌔', '🌕', '🌖', '🌗', '🌘'] as const
8const BLOCKS = ['▏', '▎', '▍', '▌', '▋', '▊', '▉'] as const
9
10// hooks/activity.json is the one copy of the table, read by this mod and by progress.rb alike.
11export function parseActivity(json: string): Activity {
12 const rows = JSON.parse(json) as [string, string][]
13 return rows.map(([pattern, label]) => [new RegExp(pattern), label] as const)
14}
15
16// The same summary progress.rb#calls builds from a transcript's tool_use block, so the one table
17// matches the same text in both. A tool call's arguments arrive flat on the event.
18export function summary(e: Record<string, unknown>): string {
19 const s = (k: string) => (typeof e[k] === 'string' ? (e[k] as string) : '')
20 return `${s('tool')} ${s('subagent_type')} ${s('command')} ${s('file_path')} ${s('pattern')}`
21}
22
23// First match wins, so specific beats general; no match leaves the previous activity standing.
24export function label(table: Activity, text: string): string | undefined {
25 return table.find(([re]) => re.test(text))?.[1]
26}
27
28// Which kind of run the command asked for. An --update re-run reads only the delta and a low-effort
29// run sends no falsifiers, so each finishes on its own clock and keeps its own history: a median
30// over all three would set a full run's bar by a three-minute update.
31export type RunKind = 'update' | 'low' | 'high'
32
33export function runKind(command: string): RunKind {
34 if (/(^|\s)--update(?=\s|$)/.test(command)) return 'update'
35 if (/(^|\s)--effort[= ]low(?=\s|$)/.test(command)) return 'low'
36 return 'high'
37}
38
39// Where --output puts the page: <dir>/index.html, named in the command and nowhere else.
40export function outputPage(command: string): string | undefined {
41 const m = command.match(/(?:^|\s)--output[= ]+("([^"]+)"|'([^']+)'|(\S+))/)
42 const dir = m && (m[2] ?? m[3] ?? m[4])
43 return dir ? `${dir.replace(/\/+$/, '')}/index.html` : undefined
44}
45
46// The shell variables SKILL.md writes the page path with. Step 1 derives
47// W="${TMPDIR:-/tmp}/review-map/<repo>-pr-<N>" and step 9 runs `page-skeleton.sh --out
48// "$W/page.html"`, so the path is read by expanding what the run itself assigned. Whatever is
49// still a variable after that is unknown, and an unknown path is no path.
50export type Vars = { W?: string; TMPDIR?: string }
51
52export function expand(text: string, vars: Vars): string | undefined {
53 const out = text
54 .replace(/\$\{TMPDIR:-([^}]*)\}/g, (_, fallback: string) => vars.TMPDIR || fallback)
55 .replace(/\$\{?(W|TMPDIR)\}?(?![A-Za-z0-9_])/g, (whole, name: 'W' | 'TMPDIR') => vars[name] ?? whole)
56 .replace(/\/{2,}/g, '/')
57 return out.includes('$') ? undefined : out
58}
59
60// The W a Bash command assigns, expanded, or undefined when it assigns none.
61export function assignedW(command: string, vars: Vars): string | undefined {
62 const m = command.match(/(?:^|[\s;&|])W=("([^"]*)"|'([^']*)'|(\S+))/)
63 const raw = m && (m[2] ?? m[3] ?? m[4])
64 return raw === null || raw === undefined ? undefined : expand(raw, vars)
65}
66
67// The page a run is staging: `page-skeleton.sh --out <path>` names it first, and an Edit or a
68// Write of the derived page names it again.
69export function pagePath(e: Record<string, unknown>, vars: Vars = {}): string | undefined {
70 if (typeof e.command === 'string') {
71 const m = e.command.match(/page-skeleton\.sh\b.*?--out[= ]+("([^"]+)"|'([^']+)'|(\S+))/)
72 const path = m && (m[2] ?? m[3] ?? m[4])
73 if (path) return expand(path, vars)
74 }
75 if (typeof e.file_path === 'string' && /\/review-map\/[^/]+\/page\.html$/.test(e.file_path)) {
76 return e.file_path
77 }
78 return undefined
79}
80
81// The coverage gate run, as distinct from read: the script is the command word — first, after a
82// separator, or handed to a shell — rather than the argument of cat, sed or grep.
83export function runsGate(e: Record<string, unknown>): boolean {
84 return (
85 e.tool === 'Bash' &&
86 typeof e.command === 'string' &&
87 /(^|[;&|(]\s*|\b(?:ba)?sh\s+)[^\s;&|]*coverage-gate\.sh(?=\s|$)/m.test(e.command)
88 )
89}
90
91// A pending checkpoint is a <section class="cp"> whose <h3> carries span.pending, so each
92// checkpoint is read up to its own close. Comments are stripped first: the template's own
93// comments describe these sections in the same words.
94export function checkpoints(html: string): string {
95 const bare = html.replace(/<!--[\s\S]*?-->/g, '')
96 const cps = bare.match(/<section class="cp\b[\s\S]*?<\/section>/g) ?? []
97 const written = cps.filter(c => !c.includes('class="pending"')).length
98 return cps.length === 0 ? '📄 page started' : `📄 ${written}/${cps.length}`
99}
100
101// What a repository's generation times are kept under: its remote, so every clone and worktree of
102// one repository shares a history, with any credential an HTTPS remote carries taken out first —
103// the store is a file on disk. A repository with no remote is keyed by its root.
104export function repoKey(remote: string | null, root: string): string {
105 if (!remote) return root
106 return remote.replace(/^([a-z][a-z0-9+.-]*:\/\/)[^/@]*@/i, '$1')
107}
108
109export function median(secs: readonly number[]): number | undefined {
110 if (secs.length === 0) return undefined
111 const sorted = [...secs].sort((a, b) => a - b)
112 return sorted[Math.floor(sorted.length / 2)]
113}
114
115export function clock(secs: number): string {
116 const s = Math.max(0, Math.floor(secs))
117 return `${Math.floor(s / 60)}:${String(s % 60).padStart(2, '0')}`
118}
119
120export function bar(frac: number, cells = 20): string {
121 const full = Math.floor(frac * cells)
122 const part = Math.floor((frac * cells - full) * BLOCKS.length)
123 let s = '█'.repeat(full)
124 if (full < cells && part > 0) s += BLOCKS[part]
125 return `▕${s.padEnd(cells)}▏`
126}
127
128export type Reading = {
129 tick: number
130 elapsed: number // seconds of generation: frozen once the run ends, paused while it waits
131 eta?: number // the median of this repository's earlier generations; absent with none recorded
132 ended?: 'done' | 'stopped' | 'interrupted'
133 waiting?: boolean // the turn ended with the run unfinished; the next turn resumes it
134 activity?: string
135 tools: number
136 falsifiers: number
137 checkpoints?: string
138}
139
140// The bar is the only estimate on the line and it is labelled as one (~), held below 100% until
141// the run ends. With no earlier generation recorded there is no bar at all: the eval harness can
142// fall back to a documented order of magnitude, but here that would be progress invented.
143export function line(r: Reading): string {
144 const icon =
145 r.ended === 'done' ? '✅' : r.ended ? '⏹' : r.waiting ? '⏸' : MOON[r.tick % MOON.length]
146 // The icon carries the state's glyph, so the words beside it carry none.
147 const status = r.ended ?? (r.waiting ? 'waiting for the next turn' : (r.activity ?? '🤔 thinking'))
148 const bits = [`${icon} review map`]
149 if (r.eta !== undefined && r.eta > 0) {
150 const frac = r.ended === 'done' ? 1 : Math.min(r.elapsed / r.eta, 0.99)
151 bits.push(bar(frac), `${String(Math.floor(frac * 100)).padStart(3)}%`,
152 `${clock(r.elapsed)} / ~${clock(r.eta)}`)
153 } else {
154 bits.push(clock(r.elapsed))
155 }
156 bits.push(status)
157 if (r.tools > 0) bits.push(`🔧 ${r.tools}`)
158 if (r.falsifiers > 0) bits.push(`🥊 ${r.falsifiers}`)
159 if (r.checkpoints) bits.push(r.checkpoints)
160 return bits.join(' ')
161}
162