SLOPSHOPPER

Accountable Review

Turns a pull request into a published review agenda — what changed, the judgments a reviewer has to make with the exact lines that settle each one, the order…

newguardstatusprompttimeragents
★ 41v1.3.0MITupdated 2026-10-09wyeworks/accountable-review
A shopper browsing a rack in a slop shop
README

<h1 align="center"> <img src="examples/assets/logo.png" alt="accountable-review" width="560"> </h1>

AI-assisted code review for teams that want to move faster with coding agents without losing control of their codebase.

accountable-review turns a pull request into a Review Map: a published HTML page that guides a reviewer through what changed, how the new behaviour works, what existing code is affected, and what to understand before merge.

It ships two skills: review-map, which produces the page for a Rails or a Rust codebase — a Rust backend or Rust in general — with or without a separate client such as Next.js, with Elixir/Phoenix and React (a Next.js app or React in general) support still in development; and setup-ci, which arranges for one to be produced automatically on every review-ready pull request.

Built by WyeWorks.


Contents 📑


Motivation ⚡

Comprehension debt

AI coding tools can produce large changes much faster than teams can absorb them.

Tests can pass. AI reviewers can find issues. Coding agents can fix those issues.

And still:

Do you understand how the system works now?

That is the problem accountable-review focuses on.

When code enters the codebase faster than the team understands it, short-term speed can become long-term loss of control. That gap has a name — comprehension debt — and it comes due the first time someone has to change the code nobody read.

The answer is not less AI. It is more reviewer understanding alongside it — more productivity and more comprehension, rather than one bought with the other:

Move faster with AI. Stay in control.

AI code review ≠ AI-assisted code review

AI code review is great at asking:

  • Is there a bug here?
  • Is this implementation suspicious?
  • What should be fixed?

accountable-review helps the human reviewer ask:

  • What behaviour now exists?
  • Why was it introduced?
  • How do the pieces work together?
  • What existing assumptions now matter?
  • What should I remember when maintaining this later?

The plugin does not approve the PR for you.

It helps a competent reviewer understand the change well enough to decide.

No approval verdict. No severity score. No fake confidence badge.

Human judgment stays final. A page good enough to approve from without reading the code would be a failure — the reviewer would be holding a verdict instead of a mental model.

Diffs are necessary. They are no longer sufficient.

Diffs are excellent for reviewing code line by line.

Is this condition correct?
Is this query efficient?
Did this line introduce a regression?

But modern AI-assisted development also pushes review to a higher level:

What behaviour changed?
How does this work across the system?
What must the team understand to maintain it?

A Review Map adds that layer.

Review behaviours, not file lists

A diff is organized by files.

A Rails, Rust, Phoenix or Next.js feature is not.

One behaviour may cross:

routes
  ↓
controller
  ↓
model / service
  ↓
policy
  ↓
job
  ↓
serializer
  ↓
frontend contract

On a Rust backend the hops are different — router and its layers, extractor, handler, error mapping, query, migration — and on Phoenix different again — router, controller or LiveView, context, changeset, Repo, worker, template — and on Next.js different again — page, server component, client component, server action, cache tag, route handler — and the point is the same: no directory contains the behaviour.

accountable-review traces the PR around the behaviour being implemented, not the order of files in the diff — and then publishes the part of that work you have to act on: the handful of judgments the change asks of you, and where to look to make each one.

The most important code may not have changed

A PR can change the meaning of code without changing its lines.

Examples:

  • an existing policy now controls a new flow,
  • an existing model invariant becomes load-bearing,
  • an unchanged serializer gains new meaning,
  • an existing query gets a new caller,
  • a TypeScript type now participates in a changed backend contract.

A normal diff cannot show those relationships. A Review Map can.


What is a Review Map? 🧭

A Review Map is a review agenda: the smallest set of things you have to judge before you can approve a change, with the code that settles each one attached.

Five parts, and a shut evidence block at the foot:

What changed                    one paragraph: what is now true that was not
Context                         when earned: the pieces of this repository the judgments rely
                                on, for a reader who knows the stack but not the codebase
What needs your attention       3-5 checkpoints per independent change the PR makes, 7 at the
                                outside; each one judgment, framed as a question, and at
                                most one small figure where its shape is clearer than prose:
                                a chain, paths converging on one invariant, a lifecycle, or
                                an entity and its relationships
Read the code in this order     3-7 stops, in the order that builds understanding
Impact outside the diff         1-3 chains, from changed code into code it gives new meaning to,
                                each node linked to the file it lives in
▸ Evidence & diff coverage      every changed path, the searches run, the rest of what was found

Claims about unchanged code come with the code attached: a collapsed excerpt of the real source, quoted by a script rather than retyped, that you open when you are ready to check that particular claim. The page reads completely with every excerpt closed — the sentence carries the consequence, the excerpt carries the proof.

The Review Map does not replace the diff.

It helps the diff make sense.

📖 Full anatomy of the page — every section, the checkpoint, the staging behaviour, the excerpt and framework-anchor rules — is in docs/review-map.md.


Installation ⚙️

Claude Code

Add the WyeWorks marketplace and install, from inside Claude Code:

/plugin marketplace add wyeworks/claude-plugins
/plugin install accountable-review@wyeworks

or from a terminal:

claude plugin marketplace add wyeworks/claude-plugins
claude plugin install accountable-review@wyeworks

Restart, and the skills are available in every project as /accountable-review:review-map and /accountable-review:setup-ci. claude plugin update accountable-review@wyeworks takes the next release; you only receive one when the version in the plugin's manifest moves, so an update is something you ask for.

Add --scope project to both commands to declare the plugin in a repository's own settings, so everyone working on it gets the same one.

To run it from a checkout instead — which is what you want if you are changing the plugin, since /reload-plugins then picks up your edits without restarting:

git clone https://github.com/wyeworks/accountable-review.git
cd /path/to/your/app
claude --plugin-dir /path/to/accountable-review

Nothing is installed that way, so git pull in the checkout is how you update, and the skills exist only for sessions started with that flag.

Codex

This repository is its own Codex marketplace. From a terminal:

codex plugin marketplace add wyeworks/accountable-review
codex plugin add accountable-review@accountable-review

Restart Codex, and the skill is available in every project as $accountable-review:review-map — the same name as in Claude Code, behind Codex's $ instead of /. codex plugin marketplace upgrade accountable-review fetches the latest catalogue, and as with Claude Code you only receive a release when the version in the plugin's manifest moves. codex plugin remove accountable-review@accountable-review uninstalls it.

Open the repository you want to review in Codex, then invoke:

$accountable-review:review-map

The result is a local HTML file, linked from the response, using the same template and completeness gate as Claude. To choose its destination, add --output /absolute/path/outside-the-repo; the page is written as index.html there. No hosting service or publishing plugin is required.

High effort is the default and requires Codex subagent tools for the independent falsification pass. If those tools are unavailable, the skill reports the limitation; use --effort low explicitly to run without independent readers. The page never presents that pass as an approval.

Codex CI execution is not supported yet. setup-ci and the CI runner still use Claude Code and Anthropic credentials, so the Codex plugin ships review-map alone.

To run it from a checkout instead — which is what you want if you are changing the skill, since edits apply without reinstalling — link it into Codex's skills directory:

mkdir -p ~/.agents/skills
ln -s /path/to/accountable-review/skills/review-map ~/.agents/skills/review-map

Linked that way it is a standalone skill invoked as $review-map, and git pull in the checkout is how you update. Use the marketplace or the link, not both, or Codex lists the skill twice.

See Codex setup and verification for removal and the boundaries of this integration.

Skills CLI

The skills live under skills/ as ordinary Agent Skills, so the cross-agent skills CLI installs them into Claude Code and Codex without cloning anything:

npx skills add wyeworks/accountable-review

It asks which skills and which agents; answer in flags instead with, for example:

npx skills add wyeworks/accountable-review -g --skill review-map -a claude-code -a codex

Prefer -g. Without it the skills are copied into the current project (.agents/skills/, .claude/skills/ and so on), which puts files in the repository you are about to review. npx skills update takes new commits, and npx skills remove review-map uninstalls.

Installed this way the skill is a standalone one rather than part of a plugin, so in Claude Code it is invoked as /review-map instead of /accountable-review:review-map, and in Codex as $review-map instead of $accountable-review:review-map. --effort high still works in Claude Code: the independent readers are launched as general-purpose subagents handed the same bundled mandate, rather than as the plugin's registered claim-falsifier agent. Install through the plugin or the skills CLI, not both, or Claude Code lists the skill twice.

The CLI can place the skills in any other agent it supports, Pi for example. Outside Claude Code, review-map follows one set of instructions whichever agent it is: the page is a local HTML file, and --effort high needs an independent reader — the agent's own subagents, or else a read-only background run of its own CLI. An agent with neither says so and offers --effort low. Only Claude Code and Codex are tested.

setup-ci is offered too — /setup-ci in Claude Code, the only host it works in. The workflow it writes clones the plugin itself at a pinned tag, so how you installed the skill locally does not matter to CI.


Usage 🚀

Run it from inside the repository you are reviewing:

/accountable-review:review-map              # current branch against its base
/accountable-review:review-map 412          # a PR number
/accountable-review:review-map https://github.com/org/repo/pull/412
/accountable-review:review-map feature/some-branch

In Claude Code you get back a URL. The page is a private Claude Artifact until you share it, and re-running for the same PR republishes to the same URL — so the Review Map tracks the PR across pushes instead of scattering links.

It publishes early and fills in as parts complete: open it at minute two, watch it arrive, start reading the moment the part you need lands. The checkpoints arrive one at a time, and their questions arrive first — so you know what the change is asking you to judge well before the explanations land. While it is unfinished it says so in a banner, and every part still coming is marked pending, so a half-written page can never be mistaken for a finished one.

While it runs, Claude Code's status line shows what it is doing — tracing consumers, cutting excerpts, writing the page — with the elapsed time and how many checkpoints are written. Every field is read from the run rather than estimated. The one exception is a bar, labelled ~. It measures against earlier runs of the same kind in that repository, so it appears once one such run has finished with its coverage gate passed.

To have one generated for every pull request instead of by hand, see CI integration below.

Basic usage

This generates the default Review Map for PR #412. The skill works best for PRs that are about to be merged, but it also works well with already-merged PRs when you want to catch up on changes.

/accountable-review:review-map 412

What makes it onto the page is the part you actually need to act on. For a small or medium-sized PR, that usually means around 700 to 1,500 words.

How hard it works

/accountable-review:review-map 412                 # --effort high, the default
/accountable-review:review-map 412 --effort low    # opt out of the falsification pass

Everything the skill writes rests on claims it checked itself — and the context that wrote a claim is the one least able to see what it assumed. --effort high, the default, adds a second reader that does not share that context. The run writes its analysis down first, one note per behaviour, and sends a read-only claim-falsifier subagent at each note before any of it reaches the page, with a single mandate: assume this is wrong in ways that matter, and find evidence in the repository that contradicts it. It rewrites nothing; it returns challenges, each anchored in a line it opened. The run then opens the cited file itself and corrects, downgrades or drops the claim. A challenge it cannot confirm is dropped, exactly as an unconfirmed finding is.

It is on by default because it is very nearly free and it changes what the page finds: on a 28-file pull request the whole pass cost 23 seconds of waiting, under 1% of the run, and the same change reviewed without it missed five things the falsified page carried. --effort low turns it off, which is worth doing when the diff is small enough that a second reader has nothing to find.

Onboarding a reviewer into the stack

A third axis, and the only one that puts anything on the page:

/accountable-review:review-map 412 --mentor          # add framework primers
/accountable-review:review-map 412 --mentor rails    # the same, naming the stack

A Review Map normally assumes you know the framework and are meeting this change for the first time. --mentor is for the other case — a reviewer new to Rails, to Rust, to Phoenix or to React, on their first pull requests in an unfamiliar codebase. Where a judgment turns on a framework rule they may not know, the page stops linking to the manual and states the rule: a short primer inside that checkpoint, with the API named, the behaviour explained, a worked example on a generic class, the line in your repository that made it relevant, and the pinned documentation link it came from.

Phoenix and React today: a primer is gated on the documentation link it escalates from, and the Elixir and React catalogues ship closed until a verification run has opened every row in them. So --mentor on a Phoenix, React or Next.js project currently produces no primers and says so. That is the fail-closed rule doing its job — a page with no primer is narrower, a page with an invented link is wrong. The Rust catalogue shipped closed the same way and has since been opened, so a Rust page can carry primers, framed in graphite rather than Rails red.

When the branch keeps moving

/accountable-review:review-map 412 --update    # re-read only the new commits

A Review Map describes one revision, and a pull request that keeps gaining commits after review opens ends up with a map that quietly describes an older one. --update is the cheap way to keep it current: it reads the commits since the page it is updating, re-analyses the judgments those commits reach, and edits that page in place at the same URL. What the new commits did not touch is carried, which is where the saving comes from — tracing consumers across the diff is most of a run's cost, and an update traces only what moved.

You are trading re-reading for speed, and the page says so. A carried judgment is one this run did not verify again. So an updated page names both revisions in its masthead and carries one sentence saying which parts still describe the earlier one. If you want the fuller read, run it again without --update — a map generated from scratch describes one revision throughout, and that is the right call before a final review pass on a branch that has moved a lot.

It refuses rather than guessing. A force-push or rebase, a base branch that moved underneath, a delta covering more than half the diff, a lock file bump, or a recorded search that now finds the changed code — each of those means the previous page is not a safe thing to build on, so the run regenerates from scratch and tells you which one it hit.


What to expect from a map 🔬

Evidence over confidence

The tool separates what is known from what is inferred.

A statement may come from:

  • changed code,
  • unchanged repository code,
  • tests,
  • inference,
  • or uncertain intent.

A claim the diff shows directly carries no label. Everything else is marked — from unchanged code, inferred from tests, inferred, uncertain — so an inference can never pass as a fact. Where the tool cannot establish intent, it says so:

It is unclear whether existing time entries stay editable after archival. No test covers it.

Every claim also carries a file:line into your repository, and by default a second reader attacks those claims before the page is finished (see Usage → effort).

Nor does the page claim to have found everything — see What a Review Map cannot do below. It never reads as a clean bill of health, and it never carries a boilerplate disclaimer saying so either: the limits are the same on every Review Map, so they are written down once, here, rather than reprinted under a heading you have already read a dozen times.

The goal is not to sound confident. The goal is to help the reviewer investigate the change.

What a Review Map cannot do

A Review Map is written by a model reading your repository. That is what lets it trace a consequence into code the diff never opened, and it is also the honest limit on the page.

  • It is a pass, not an audit. Three passes over the same 109-file diff produced eight headline findings between them, only one of which appeared in all three. Explanation is reproducible; defect discovery is sampling.
  • It has blind spots, and they are not random. Behaviour living in configuration, in data, in a queue, in another service or in the gap between two deploys is harder to reach from a diff than behaviour living in a method — so those are the regions a map is quietest about, and quiet is not the same as clear. On a very large diff it also runs out of room before it runs out of diff, and says which region it skimmed.
  • Some of it can simply be wrong — a misread method, a framework default that does not hold for your version, a consequence prevented somewhere the run never looked. Every claim is labelled by how it is known and carries a file:line, so open the citation for anything you would act on.
  • It never decides anything. A map that says nothing about a file is not telling you the file is fine, only that this pass surfaced no judgment there.

And it depends on the model behind it. We develop and test with Claude Opus 5 most of the time, and that is what the page's depth is calibrated against. Other models will trade cost for reach differently — try a few against your own codebase and keep the one whose maps you actually trust.

Ruby on Rails first, then Rust, Phoenix and React

The plugin is optimized and most heavily tested for Rails applications. That matters because Rails behaviour often emerges from several pieces working together:

  • routes
  • controllers
  • Active Record models
  • validations and callbacks
  • policies
  • jobs
  • serializers
  • service objects
  • tests
  • frontend clients

The plugin is designed to help reviewers understand those relationships as a system, and it knows where the framework's own rules bite: update_all at a call site the diff never opened skips the validation this PR adds, a uniqueness validation is not a unique index, --sandbox rolls back so after_commit never fires there.

It also reads the app for what your team already decided. A value object in app/models is ordinary; a value object in app/models when four of its kind live in app/services is a choice someone made, and the page will ask whether it was deliberate — with the four siblings cited, because without them there is nothing to ask. That question is never a recommendation, never appears more than once, and never takes a slot from a judgment about behaviour.

Rust is also supported and tested, and it comes in two lenses, because what a reviewer looks for in a crate that serves requests and in one that does not overlaps less than the language suggests. Rust in general — a library, a CLI, a workspace that serves nothing — is reviewed against the things the compiler was told not to check or cannot see: the _ => arm a new enum variant falls into silently, a pub item a downstream crate depends on, a feature combination CI never builds, the module whose safe code keeps an unsafe block sound. A Rust backend — axum, actix-web, tonic and the rest — gets that l

Source 2 files
hooks/progress.ts 237 lines
1// progress.ts — the e2e progress board (skills/review-map/evals/e2e/progress.rb), as one status
2// line in Claude Code while /accountable-review:review-map runs interactively.
3//
4//   🌔 review map  ▕███████████▍       ▏  47%  11:02 / ~23:30  🔎 tracing consumers  🔧 38  🥊 4  📄 2/4
5//
6// Every field is READ, never guessed, and that is the rule this file keeps: a progress line that
7// invents progress is the verification badge of the terminal, and the page itself is forbidden
8// that badge. Three sources, each one the engine or the run already produces:
9//
10//   * engine events — asking for the skill starts the run, each main-loop tool call names the
11//     activity through hooks/activity.json (the table progress.rb reads too), each claim-falsifier
12//     spawn is counted, and the coverage gate passing then the turn answering ends the run;
13//   * the page the run is staging — checkpoints written against pending ones;
14//   * the clock, against the median of this repository's earlier generations of the same kind,
15//     which this file records itself. With none recorded there is no bar, only elapsed time.
16//
17// The activity is the latest tool call by purpose, never a step number: steps interleave and
18// cannot be read off a run (evals/README.md § Profiling one run), and this does not pretend
19// otherwise. It observes and nothing else — every hook passes its event on unchanged, and every
20// one has a .catch that still does, so a bug in the line can never stall a generation.
21
22import type { EngineInterface, Register, Timer } from 'claude-code'
23
24import {
25  type Activity,
26  type RunKind,
27  type Vars,
28  assignedW,
29  checkpoints,
30  label,
31  line,
32  median,
33  outputPage,
34  pagePath,
35  parseActivity,
36  repoKey,
37  runKind,
38  runsGate,
39  summary,
40} from './board'
41
42const REVIEW_MAP = /(^|:)review-map$/
43const FALSIFIER = /(^|:)claim-falsifier$/
44const KEEP = 20 // generations remembered per repository and kind
45
46type Run = {
47  started: number
48  paused: number // milliseconds spent waiting between turns, which are not generation
49  waitingSince?: number
50  ended?: 'done' | 'stopped' | 'interrupted'
51  endedAt?: number
52  activity?: string
53  tools: number
54  falsifiers: number
55  page?: string
56  vars: Vars
57  checkpoints?: string
58  sawGate: boolean
59  kind: RunKind
60  eta?: number
61  repo?: string
62  tick: number
63  timer?: Timer
64}
65
66// The module's own variables start over on a reload, as register runs again.
67let table: Activity = []
68let run: Run | undefined
69
70function elapsedMs(r: Run, now: number) {
71  return (r.endedAt ?? r.waitingSince ?? now) - r.started - r.paused
72}
73
74async function draw($: EngineInterface) {
75  if (!run) return
76  const now = await $.clock.now()
77  $.ui.status(line({ ...run, waiting: run.waitingSince !== undefined, elapsed: elapsedMs(run, now) / 1000 }))
78}
79
80async function recount($: EngineInterface) {
81  if (!run?.page) return
82  try {
83    run.checkpoints = checkpoints(await $.fs.read(run.page))
84  } catch {
85    // a page mid-Edit or not written yet: the next page-touching call will have it
86  }
87}
88
89function tick($: EngineInterface, r: Run) {
90  r.timer?.cancel()
91  r.timer = $.clock.every(1000, () => {
92    r.tick += 1
93    void draw($)
94  })
95}
96
97function historyKey(repo: string, kind: RunKind) {
98  return `generations:${kind}:${repo}`
99}
100
101// The command as typed, or as the engine records a typed skill command; a sentence that merely
102// mentions review-map is not a request for one.
103function isCommand(text: string) {
104  return /(^\s*|<command-name>)\/(accountable-review:)?review-map(?=\s|<|$)/.test(text)
105}
106
107async function start($: EngineInterface, command: string) {
108  run?.timer?.cancel()
109  if (table.length === 0) {
110    table = parseActivity(await $.fs.read(`${$.plugin.root}/hooks/activity.json`))
111  }
112  const kind = runKind(command)
113  const repo = await $.session.repo()
114  const key = repo ? repoKey(repo.remote, repo.root) : undefined
115  const past = key ? await $.store.get(historyKey(key, kind)) : undefined
116  const mine: Run = {
117    started: await $.clock.now(),
118    paused: 0,
119    tools: 0,
120    falsifiers: 0,
121    sawGate: false,
122    kind,
123    page: outputPage(command),
124    vars: { TMPDIR: await $.env.get('TMPDIR') },
125    tick: 0,
126    repo: key,
127    eta: Array.isArray(past) ? median(past.filter((n): n is number => typeof n === 'number')) : undefined,
128  }
129  run = mine
130  tick($, mine)
131  await draw($)
132}
133
134export const register: Register = on => {
135  table = []
136  run = undefined
137
138  // A run starts where the skill is asked for, in one of the two ways a person or the model can ask:
139  // the command typed as a prompt, or a Skill tool call naming it. `skill.prompt` would be the
140  // obvious event and is not one this mod can use: the engine's own security module sends it
141  // beneath the user tier, so a plugin's hook on it never runs (the debug log says "skill.prompt
142  // bypassed by cc-plugin-sec-default").
143  //
144  // A prompt typed over a running turn (`turnId`) is a note to that turn, not a new run. Any new
145  // prompt takes down the line a finished run left standing.
146  on('prompt.submit', async ($, e, next) => {
147    if (run?.ended) {
148      run = undefined
149      $.ui.status(undefined)
150    }
151    if (e.turnId === undefined && isCommand(e.text)) await start($, e.text)
152    return next(e)
153  }).catch(($, e, next) => next(e))
154
155  // The main loop only: progress.rb reads the parent transcript, and a falsifier's own reads are
156  // its work, not the run's activity. Its spawn is what gets counted, below.
157  on('tool.call', async ($, e, next) => {
158    if (e.agentId !== undefined) return next(e)
159    const args = e as unknown as Record<string, unknown>
160    const asked = e.tool === 'Skill' && REVIEW_MAP.test(String(args.skill ?? ''))
161    if (asked && (!run || run.ended)) {
162      await start($, `/review-map ${String(args.args ?? '')}`)
163      return next(e)
164    }
165    if (!run || run.ended) return next(e)
166    const mine = run
167    const text = summary(args)
168    mine.tools += 1
169    mine.activity = label(table, text) ?? mine.activity
170    if (typeof args.command === 'string') mine.vars.W = assignedW(args.command, mine.vars) ?? mine.vars.W
171    mine.page = pagePath(args, mine.vars) ?? mine.page
172    await draw($)
173    const result = await next(e)
174    // Passing, not mentioning: the gate exits non-zero when the inventory and the diff disagree.
175    if (runsGate(args) && result.deny === undefined && result.isError !== true) mine.sawGate = true
176    if (mine.page && /page-skeleton\.sh|\.html\b/.test(text)) {
177      await recount($)
178      await draw($)
179    }
180    return result
181  }).catch(($, e, next) => next(e))
182
183  on('agent.spawn', async ($, e, next) => {
184    if (run && !run.ended && FALSIFIER.test(e.subagentType)) {
185      run.falsifiers += 1
186      await draw($)
187    }
188    return next(e)
189  }).catch(($, e, next) => next(e))
190
191  // A run can span turns: the skill stops to ask (two stacks, say), or the main loop ends its turn
192  // while background falsifiers read and is woken by their notifications. So a turn that answers
193  // before the gate has passed pauses the run rather than ending it, and the next main-loop turn
194  // resumes it; the time between is the person's or the notification's, not generation's.
195  on('turn.start', async ($, e, next) => {
196    if (run && !run.ended && run.waitingSince !== undefined) {
197      run.paused += (await $.clock.now()) - run.waitingSince
198      run.waitingSince = undefined
199      tick($, run)
200      await draw($)
201    }
202    return next(e)
203  }).catch(($, e, next) => next(e))
204
205  // Its length feeds the next run's ETA only when the gate passed and the turn answered, so a run
206  // that stopped early never shortens the bar.
207  on('turn.complete', async ($, e, next) => {
208    if (run && !run.ended && e.agentId === undefined) {
209      const mine = run
210      mine.timer?.cancel()
211      const now = await $.clock.now()
212      if (e.reason === 'answer' && !mine.sawGate) {
213        mine.waitingSince = now
214      } else {
215        mine.endedAt = now
216        mine.ended = e.reason === 'answer' ? 'done' : e.isAborted ? 'interrupted' : 'stopped'
217      }
218      await recount($)
219      await draw($)
220      if (mine.ended === 'done' && mine.repo) {
221        const key = historyKey(mine.repo, mine.kind)
222        const past = await $.store.get(key)
223        const kept = Array.isArray(past) ? past.filter(n => typeof n === 'number') : []
224        await $.store.set(key, [...kept, elapsedMs(mine, now) / 1000].slice(-KEEP))
225      }
226    }
227    return next(e)
228  }).catch(($, e, next) => next(e))
229
230  on('session.end', async ($, e, next) => {
231    run?.timer?.cancel()
232    run = undefined
233    $.ui.status(undefined)
234    return next(e)
235  }).catch(($, e, next) => next(e))
236}
237
hooks/board.ts 162 lines
1// board.ts — the pure half of the progress line: what a field says, given what was read. Nothing
2// here reads anything; progress.ts does the reading, and this file turns readings into one line.
3// The rule both keep is progress.rb's: every field is READ, never guessed.
4
5export type Activity = readonly (readonly [RegExp, string])[]
6
7export const MOON = ['🌑', '🌒', '🌓', '🌔', '🌕', '🌖', '🌗', '🌘'] as const
8const BLOCKS = ['▏', '▎', '▍', '▌', '▋', '▊', '▉'] as const
9
10// hooks/activity.json is the one copy of the table, read by this mod and by progress.rb alike.
11export function parseActivity(json: string): Activity {
12  const rows = JSON.parse(json) as [string, string][]
13  return rows.map(([pattern, label]) => [new RegExp(pattern), label] as const)
14}
15
16// The same summary progress.rb#calls builds from a transcript's tool_use block, so the one table
17// matches the same text in both. A tool call's arguments arrive flat on the event.
18export function summary(e: Record<string, unknown>): string {
19  const s = (k: string) => (typeof e[k] === 'string' ? (e[k] as string) : '')
20  return `${s('tool')} ${s('subagent_type')} ${s('command')} ${s('file_path')} ${s('pattern')}`
21}
22
23// First match wins, so specific beats general; no match leaves the previous activity standing.
24export function label(table: Activity, text: string): string | undefined {
25  return table.find(([re]) => re.test(text))?.[1]
26}
27
28// Which kind of run the command asked for. An --update re-run reads only the delta and a low-effort
29// run sends no falsifiers, so each finishes on its own clock and keeps its own history: a median
30// over all three would set a full run's bar by a three-minute update.
31export type RunKind = 'update' | 'low' | 'high'
32
33export function runKind(command: string): RunKind {
34  if (/(^|\s)--update(?=\s|$)/.test(command)) return 'update'
35  if (/(^|\s)--effort[= ]low(?=\s|$)/.test(command)) return 'low'
36  return 'high'
37}
38
39// Where --output puts the page: <dir>/index.html, named in the command and nowhere else.
40export function outputPage(command: string): string | undefined {
41  const m = command.match(/(?:^|\s)--output[= ]+("([^"]+)"|'([^']+)'|(\S+))/)
42  const dir = m && (m[2] ?? m[3] ?? m[4])
43  return dir ? `${dir.replace(/\/+$/, '')}/index.html` : undefined
44}
45
46// The shell variables SKILL.md writes the page path with. Step 1 derives
47// W="${TMPDIR:-/tmp}/review-map/<repo>-pr-<N>" and step 9 runs `page-skeleton.sh --out
48// "$W/page.html"`, so the path is read by expanding what the run itself assigned. Whatever is
49// still a variable after that is unknown, and an unknown path is no path.
50export type Vars = { W?: string; TMPDIR?: string }
51
52export function expand(text: string, vars: Vars): string | undefined {
53  const out = text
54    .replace(/\$\{TMPDIR:-([^}]*)\}/g, (_, fallback: string) => vars.TMPDIR || fallback)
55    .replace(/\$\{?(W|TMPDIR)\}?(?![A-Za-z0-9_])/g, (whole, name: 'W' | 'TMPDIR') => vars[name] ?? whole)
56    .replace(/\/{2,}/g, '/')
57  return out.includes('$') ? undefined : out
58}
59
60// The W a Bash command assigns, expanded, or undefined when it assigns none.
61export function assignedW(command: string, vars: Vars): string | undefined {
62  const m = command.match(/(?:^|[\s;&|])W=("([^"]*)"|'([^']*)'|(\S+))/)
63  const raw = m && (m[2] ?? m[3] ?? m[4])
64  return raw === null || raw === undefined ? undefined : expand(raw, vars)
65}
66
67// The page a run is staging: `page-skeleton.sh --out <path>` names it first, and an Edit or a
68// Write of the derived page names it again.
69export function pagePath(e: Record<string, unknown>, vars: Vars = {}): string | undefined {
70  if (typeof e.command === 'string') {
71    const m = e.command.match(/page-skeleton\.sh\b.*?--out[= ]+("([^"]+)"|'([^']+)'|(\S+))/)
72    const path = m && (m[2] ?? m[3] ?? m[4])
73    if (path) return expand(path, vars)
74  }
75  if (typeof e.file_path === 'string' && /\/review-map\/[^/]+\/page\.html$/.test(e.file_path)) {
76    return e.file_path
77  }
78  return undefined
79}
80
81// The coverage gate run, as distinct from read: the script is the command word — first, after a
82// separator, or handed to a shell — rather than the argument of cat, sed or grep.
83export function runsGate(e: Record<string, unknown>): boolean {
84  return (
85    e.tool === 'Bash' &&
86    typeof e.command === 'string' &&
87    /(^|[;&|(]\s*|\b(?:ba)?sh\s+)[^\s;&|]*coverage-gate\.sh(?=\s|$)/m.test(e.command)
88  )
89}
90
91// A pending checkpoint is a <section class="cp"> whose <h3> carries span.pending, so each
92// checkpoint is read up to its own close. Comments are stripped first: the template's own
93// comments describe these sections in the same words.
94export function checkpoints(html: string): string {
95  const bare = html.replace(/<!--[\s\S]*?-->/g, '')
96  const cps = bare.match(/<section class="cp\b[\s\S]*?<\/section>/g) ?? []
97  const written = cps.filter(c => !c.includes('class="pending"')).length
98  return cps.length === 0 ? '📄 page started' : `📄 ${written}/${cps.length}`
99}
100
101// What a repository's generation times are kept under: its remote, so every clone and worktree of
102// one repository shares a history, with any credential an HTTPS remote carries taken out first —
103// the store is a file on disk. A repository with no remote is keyed by its root.
104export function repoKey(remote: string | null, root: string): string {
105  if (!remote) return root
106  return remote.replace(/^([a-z][a-z0-9+.-]*:\/\/)[^/@]*@/i, '$1')
107}
108
109export function median(secs: readonly number[]): number | undefined {
110  if (secs.length === 0) return undefined
111  const sorted = [...secs].sort((a, b) => a - b)
112  return sorted[Math.floor(sorted.length / 2)]
113}
114
115export function clock(secs: number): string {
116  const s = Math.max(0, Math.floor(secs))
117  return `${Math.floor(s / 60)}:${String(s % 60).padStart(2, '0')}`
118}
119
120export function bar(frac: number, cells = 20): string {
121  const full = Math.floor(frac * cells)
122  const part = Math.floor((frac * cells - full) * BLOCKS.length)
123  let s = '█'.repeat(full)
124  if (full < cells && part > 0) s += BLOCKS[part]
125  return `▕${s.padEnd(cells)}▏`
126}
127
128export type Reading = {
129  tick: number
130  elapsed: number // seconds of generation: frozen once the run ends, paused while it waits
131  eta?: number // the median of this repository's earlier generations; absent with none recorded
132  ended?: 'done' | 'stopped' | 'interrupted'
133  waiting?: boolean // the turn ended with the run unfinished; the next turn resumes it
134  activity?: string
135  tools: number
136  falsifiers: number
137  checkpoints?: string
138}
139
140// The bar is the only estimate on the line and it is labelled as one (~), held below 100% until
141// the run ends. With no earlier generation recorded there is no bar at all: the eval harness can
142// fall back to a documented order of magnitude, but here that would be progress invented.
143export function line(r: Reading): string {
144  const icon =
145    r.ended === 'done' ? '✅' : r.ended ? '⏹' : r.waiting ? '⏸' : MOON[r.tick % MOON.length]
146  // The icon carries the state's glyph, so the words beside it carry none.
147  const status = r.ended ?? (r.waiting ? 'waiting for the next turn' : (r.activity ?? '🤔 thinking'))
148  const bits = [`${icon} review map`]
149  if (r.eta !== undefined && r.eta > 0) {
150    const frac = r.ended === 'done' ? 1 : Math.min(r.elapsed / r.eta, 0.99)
151    bits.push(bar(frac), `${String(Math.floor(frac * 100)).padStart(3)}%`,
152      `${clock(r.elapsed)} / ~${clock(r.eta)}`)
153  } else {
154    bits.push(clock(r.elapsed))
155  }
156  bits.push(status)
157  if (r.tools > 0) bits.push(`🔧 ${r.tools}`)
158  if (r.falsifiers > 0) bits.push(`🥊 ${r.falsifiers}`)
159  if (r.checkpoints) bits.push(r.checkpoints)
160  return bits.join('  ')
161}
162