SLOPSHOPPER

orca

Autonomous multi-agent development: init/doctor once, then orca:feature from idea to integration branch, and orca:review to walk the deliverable in your own…

newbandguardpromptprocesstimer
★ 1v1.4.0MITupdated 2026-10-07miguelbacalhau/orca
A shopper browsing a rack in a slop shop
README

<img src=".github/logo.png" width="280" alt="Orca logo">

<h1 align="center">orca</h1>

A Claude Code plugin for autonomous, multi-agent development. Set the repository up once, then /orca:feature takes a feature from idea to a committed integration branch: an adversarial interview captures intent as a durable brief; a deterministic workflow plans, implements, independently reviews, fixes, commits, merges, and verifies. The user interacts exactly once after the interview (twice, by opt-in).

/orca:feature <idea>     # triage → interview/brief → spec → work loop → report
                         # (interactive until one confirmation, autonomous after)
/orca:review             # review a deliverable in your editor; your comments round-trip
                         # (addressed on consent, resolutions rendered inline next review)
/orca:pr                 # land a delivered run through a GitHub pull request: the report
                         # becomes an external-facing description, previewed, opened as a draft
/orca:retry              # finish a finished run's unmet items in the same run: audit against
                         # git, resolve the blocked decisions with you, relaunch the work loop
/orca:followup           # turn a finished run's optional follow-ups into the next brief
                         # for /orca:feature (new intent, not recovery)
/orca:iterate            # new work on a delivered, still-unlanded feature branch, whatever
                         # the size — same run, same branch, specced and reviewed like a run
/orca:status             # read-only dashboard: .orca state joined with git ground truth,
                         # grouped by next action, each state naming its skill
/orca:archive            # retire finished runs whose branches provably landed, so triage
                         # stops carrying the whole history, removing their branches + worktrees
/orca:init               # one-time repository layout setup      (interactive, consent per step)
/orca:doctor             # machine tooling + repo readiness setup (interactive, consent per step)
/orca:config             # optional per-repo reviewer & model/effort tuning

The deliverable of a run is a feature/<slug> branch on an integration worktree, which you review with /orca:review and land yourself with git merge --no-ff <branch>, or publish as a GitHub pull request with /orca:pr. The runs never touch your own worktree, and no commit they produce mentions Claude, AI, agents, or orca.

Table of contents

How it works

Three design choices carry the whole system:

Double isolation. Every stage — spec, plan, implement, review, fix, commit, merge, integrate — runs in a dedicated subagent with its own context window, and every unit of parallel work (a feature's work item) gets its own git worktree off a shared bare repository. Heavy context — codebase exploration, diffs, test output — lives and dies inside subagents; the main conversation only reads artifact files and the workflow's structured result. Parallel items can never corrupt each other's files: overlap surfaces as an explicit merge conflict, resolved by a merge agent holding both items' plans.

Independent review. Claude implements; a separate reviewer attacks the result before anything is committed. The featured default is cross-model: Codex reviews, run non-interactively by the plugin's own CLI (orca.sh codex → codex exec) against the global codex binary, driven adversarially over each item's diff by a dedicated courier agent — an independent second opinion from a different model family, one that does not share the implementer's blind spots. The review explicitly attacks the tests, since the same model family wrote the code and the tests. With reviewer=claude — the detected default wherever codex isn't installed, or an explicit pin via /orca:config — a dedicated Claude review agent performs the same adversarial review itself: it keeps fresh-context independence (a separate agent, only the artifacts and the diff), but it is same-model, so it may share the implementer's blind spots. That trade-off is stated wherever the choice is made; cross-model stays the stronger design.

A deterministic work loop. The long autonomous middle of a run is one bundled script executed through Claude Code's Workflow tool, not conversational orchestration. Scheduling, retry bounds, review throttling, merge serialization, and the commit-attribution check are code, so the guarantees are structural rather than model discipline that decays over a long context. Judgment calls (plan reconciliation, escalation, review verdicts) stay in agents, but as schema'd calls whose reasons land in the run's artifacts. Every agent call is journaled, so an interrupted run resumes where it stopped. The workflow sandbox has no shell, so its git plumbing has to leave the sandbox to run: orca's verb hook, a function hook the plugin ships, runs each orca.sh verb itself and hands back its exact output — no model in the path. That hook is the only way out: a run requires function hooks (mods) loaded and a bypassPermissions session, and without either it fails at its first verb with a message naming the cause — there is no fallback.

State lives in files, never in conversation memory: the brief, the spec with its work breakdown and decision log, one plan per item, the raw review findings per round, and a final report — all under a per-run directory in .orca/.

Requirements

RequirementDetail
Claude CodeA harness with the Workflow tool (the work loop runs through it). /orca:feature checks and refuses without it.
Codex CLIRequired only when the reviewer is codex (the default wherever it is installed): the global codex binary on PATH, version ≥ 0.142.5, authenticated via codex login, and still carrying the codex exec flags orca drives (--output-schema, --output-last-message, --sandbox, --cd) — the pre-flight asks the binary directly, because a version number does not tell you the interface has not moved. Never install codex via npm — use brew install codex or the official release binaries. With reviewer=claude, the codex rows don't apply.
Repository layoutBare-repo-with-worktrees (.bare/ + peer worktrees). /orca:init sets this up, including converting an existing conventional checkout in place.
git≥ 2.31 (rev-parse --path-format — every orca script's repository resolution; older gits get a typed OLD_GIT failure naming the upgrade, never a misdiagnosis); ≥ 2.42 for worktree add --orphan when /orca:init creates a brand-new repository.
BASH_MAX_TIMEOUT_MSCodex-only, like the Codex CLI row: set to 1200000 (~20 min) in a Claude Code settings env block, so Codex reviews are not killed at the default Bash tool timeout. A plugin cannot ship session env, so /orca:doctor writes it for you.
Permission modeRuns need bypassPermissions for the session — see Permissions and autonomy. The verb hook enforces it: in any other mode it refuses, and the run fails at its first verb.
Function hooks (mods)Runs need orca's hooks module loaded in the session: a Claude Code build with function hooks (2.1.290+ was probed), the feature enabled (CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 in a settings env block where the server-side flag is off), not started with --bare, and no disableAllHooks in any settings file, managed ones included. The work loop reaches orca.sh only through the verb hook; without it a run fails at its first verb, naming the fix. No shell check can see the hook, so /orca:doctor and the pre-flight state this requirement rather than test it.

Everything else — the fifteen stage agents and the CLI that drives codex — ships inside the plugin itself; there is nothing to install per repository beyond the layout.

Installation

This repository hosts its own plugin marketplace (.claude-plugin/marketplace.json), so installing is two commands inside Claude Code:

/plugin marketplace add miguelbacalhau/orca
/plugin install orca@orca

The install persists across sessions. Updates are manual by default for third-party marketplaces — pull new versions with /plugin marketplace update orca, or toggle auto-update in the /plugin → Marketplaces UI. Removal is symmetric: /plugin marketplace remove orca uninstalls the marketplace and the plugin in one step.

Teams can make the install declarative instead: commit this to the project's .claude/settings.json, and Claude Code prompts each teammate to install the marketplace and pre-enables the plugin when they trust the workspace:

{
  "extraKnownMarketplaces": {
    "orca": { "source": { "source": "github", "repo": "miguelbacalhau/orca" } }
  },
  "enabledPlugins": { "orca@orca": true }
}

For local development on the plugin itself, load a checkout directly for a single session:

claude --plugin-dir /path/to/this/repo

Plugins load at session start, so after installing or enabling one, start a fresh session before running. There is nothing session-scoped left to verify for the reviewer: codex is reached by running the binary, and /orca:feature's pre-flight checks it from a shell.

Quick start

# 1a. One-time per repo: the bare-repo-with-worktrees layout. Interactive,
#     consent per step. Converting an existing checkout preserves untracked
#     files (.env, caches) but changes every path — see /orca:init.
/orca:init

# 1b. One-time per machine, only if the pre-flight flags it: Codex CLI
#     install/auth guidance and the review timeout settings write. Skippable
#     outright if you'll run with the Claude reviewer.
/orca:doctor

# 2. The feature. An interview captures intent as a brief file — outcome,
#    features, non-goals, constraints, doubt rule, as many rounds as the
#    idea needs — then, after ONE confirmation, the run drives autonomously
#    to a final report and a feature/<slug> branch. Decline the run prompt
#    to leave the brief queued for a later /orca:feature instead.
/orca:feature add rate limiting to the public API

# 3. Walk the deliverable's diff in your own editor (nvim, via a tmux
#    window in your session), then land it from your own worktree —
#    or publish it as a GitHub pull request instead with /orca:pr.
/orca:review
git merge --no-ff feature/<slug>
/orca:pr                             # the PR path: report → description, draft PR via gh

# 4. If the report left blocked items: resolve their recorded decisions
#    and finish them inside the same run, on the same branch.
/orca:retry

# 5. If it left optional follow-ups you want: select them into the next
#    run's brief.
/orca:followup

Briefs can be queued ahead of time — everything at the top level of .orca/feat-briefs/ is ready and unconsumed; a later /orca:feature finds them, offers to run one, and a run consumes exactly one.

Commands

/orca:feature <idea>

The one feature verb — a dispatcher over .orca state. Triage looks at what is on disk and offers the first match, never forcing it:

  1. An interrupted run — a run directory whose spec.md records a workflow runId with no report.md beside it. Offers to resume from the journal; completed work replays instantly, only in-flight and remaining work runs live.
  2. A queued brief — anything at the top level of .orca/feat-briefs/. Offers to run one, or to interview a new idea instead.
  3. Nothing waiting — interviews the idea into a brief, then asks once: run it now, or leave it queued?

The interview sharpens a rough idea into a durable brief at .orca/feat-briefs/<timestamp>-<slug>.md. Before asking anything substantive it researches the subsystems the idea touches through the dedicated orca:research agent (the report stays behind the scenes; the main context never explores deeply), then interviews from that picture — opening with a reflection of the idea against the system as it exists, so tensions between the two surface in conversation, where changing course is cheap. The brief is the entire intent the run acts on — the run asks nothing beyond one confirmation — so the interview is deliberately adversarial about scope: it pushes back, hunts for unstated non-goals, and makes you resolve ambiguities now rather than leaving them for an autonomous run to guess at. Direction decisions it settles land in the brief with their rationale and bind the spec like constraints; the decomposition itself still belongs to the spec stage.

A brief records: outcome, features, non-goals, inputs & outputs, constraints, plus two run-controlling choices:

  • Doubt rule — when the run hits an ambiguity, does it prefer the smaller interpretation and cut scope (prefer-smaller-scope, the default) or the more complete one (prefer-complete)?
  • Breakdown checkpoint — review the spec and work breakdown once before any code (review-once), or run straight through (straight-through, the default)?

The brief is what and why, never how: no work breakdown, no interfaces, no file ownership — those come from the run's spec stage, grounded in real codebase exploration. Location is status: top-level briefs are ready; park unfinished ones in .orca/feat-briefs/drafts/. One brief is one run's scope — two independent efforts get two briefs, and declining the run prompt leaves a brief queued for any later /orca:feature.

The run. See Anatomy of a run for the full lifecycle. Interaction surface, in total:

  1. One confirmation of the restated brief (plus trunk-branch confirmation) — this authorizes everything that follows. It runs in full even when the brief was written seconds earlier in the same session: the file, not the conversation, is the authorized intent.
  2. One optional checkpoint — only if the brief opted in: review the spec and work breakdown before any code.

After that, nothing asks you anything. Ambiguities resolve against the spec and the doubt rule; what cannot be resolved that way becomes a blocked item in the final report, with the options you must choose between recorded. The outcome lands in report.md.

/orca:review [branch]

The human half of review — the runs' adversarial review stage is automated and lives inside them; this opens the finished deliverable in your editor before you land it. It discovers unmerged feature/<slug> deliverable branches (one → opens it; several → asks; a gone integration worktree → offers to add it back) and opens the deliverable's integration worktree in your editor, running an orca review session — file list, native side-by-side merge-base diffs, your LSP and colors. Two editors are supported:

  • nvim — a new tmux window in your current session running orca.nvim's :OrcaReview. Quit nvim (:qa) and the window closes, dropping you back where you invoked it, ready to git merge --no-ff. tmux is the mechanism, not a convenience — the harness has no TTY to run an interactive editor in, so a new window in the session $TMUX points at is the one way a skill puts a live editor in front of you.
  • vscode — orca.vscode's review session, opened via code --open-url and the extension's URI handler, which routes it to a window on the worktree (opening one if needed). A detached GUI launch — no tmux involved; close the review window when done.

Without tmux (on the nvim path), or without either companion installed (/orca:doctor prescribes both), it prints the exact command to run instead (cd <worktree> && nvim "+OrcaReview", cd <worktree> && code . plus the palette command, or a git difftool fallback) and stops.

Review comments round-trip. Comments you leave in the review with :OrcaComment persist to .orca/review-notes/<key>.json (<key> is the deliverable branch, sanitized — orca.nvim's rule). On the nvim path the skill leaves a background waiter on the tmux window: quit nvim and it wakes, reads the file, and lays out every open comment alongside what it intends to do about it — fix (with a one-line approach), answer, decline with reasoning, or defer as too big for review — then asks before doing anything (window death means "nvim closed," not "review finished"; you can address now, resume the review with comments intact, or leave them for later). Consent is to that stated plan, not a blank check: correct a misread interpretation right there, before any code changes. On consent, a dedicated agent converts each open comment into a fix or an answer in the integration worktree — change requests get fixed with tests and committed under the same no-attribution rules as every run commit; questions get answered from the code — and writes status/resolution back into the same file, where the next :OrcaReview renders them inline under their anchors. Each round is also appended to the run's report.md under ## Review rounds — commit, per-comment resolutions, verification result — so /orca:pr describes the branch as it now stands; a round whose verification fails downgrades the report's deliverable state to unverified. Editing a resolution's comment there re-opens it for the next round; that loop is the convergence mechanism, so there is deliberately no machine re-review after addressing. The file is the source of truth and the window-exit only the wake signal: comments left from the vscode tier, a print-only session, or a review whose session died are picked up by the next /orca:review, which surfaces "N unaddressed comments" in triage. The workflow is one-writer-at-a-time by design, and both sides fail loud on a notes schema version they don't speak (/orca:doctor's nvim check diagnoses the skew).

Two config keys govern it, set via /orca:config: editor (nvim|vscode|none) and terminal (tmux|none, nvim path only), with the same semantics as reviewer — absent detects (nvim first, then vscode), a pin fails loudly instead of silently falling back, none opts out to the printed command.

/orca:pr [run]

The PR path for landing — the report template's Landing section ends at a local git merge --no-ff, which is wrong for a repo that lands work through GitHub pull requests. This skill takes a delivered-but-unlanded run (a finished run whose deliverable branch exists and is unmerged, found via triage snapshot — newest by default, or the one the argument names) and publishes it with the gh CLI. It is report-driven by design: every fact in the PR body traces to a section of the run's report.md, nothing is re-derived from the diff — so a branch no run produced is out of scope (plain gh pr create already covers it).

Two guards gate it, because even a draft PR asserts "the work is finished": the report must say Deliverable state: verified, and ## Blocked must be "None" — anything else is refused with a pointer at the owning skill (/orca:feature's resume, /orca:retry). Draft status is no fallback there: it is the unconditional default, not an escape hatch, and it still pushes an unverified branch under a description claiming work the run never verified.

New PRs are always drafts, because two readinesses are in play and the skill can only vouch for one. The guards settle run readiness — the run finished and verified its own work. Social readiness is yours: at publish time no human has read the diff, and a ready PR announces the opposite, auto-requesting CODEOWNERS reviewers and releasing whatever CI and merge automation keys on non-draft. So it goes up as a draft and you mark it ready on GitHub after your own pass — no flag and no config key, since the wrong default this way costs one click and the other way costs a notification you cannot recall. Draft is a creation-time choice only: refreshing an existing PR never moves its status in either direction, so a PR you already promoted stays promoted.

Composition is translation into a fixed shape. The reader has never heard of orca, so run vocabulary is dropped wholesale — item IDs, hashes, counts, run-dir paths, follow-ups, every /orca:* pointer — and so is the run's item-by-item shape, which is what made descriptions long: work items are the run's unit of parallelism, not the reader's unit of understanding, so bullets group by outcome. The body fills the repo's own pull_request_template.md when it has one; otherwise it follows a three-section skeleton — a plain-language paragraph on what was wrong or missing (from the run's brief.md) and what the branch does about it, What changed, Testing, plus How it works and Notes only when they earn their place — held to 200–350 words. The no-attribution rule that governs run commits extends verbatim to the PR title and body, enforced by the same deterministic marker scan — explicitly overriding the harness habit of appending a "Generated with Claude Code" footer. You see the exact title, body, base, and head before anything happens; one confirmation gates both the push and the PR creation (re-running an existing PR refreshes it). It never merges, never edits the report, never touches the integration worktree.

/orca:retry [run]

Finishes a finished run's unmet work items inside the same run — no new run directory, no new spec, no brief. Recovery never creates a run; only new intent does. It picks a finished feature run with leftovers (newest by default, or the one the argument names), audits it through the same read-only agent /orca:followup uses — reconciling the report's claims against the spec's work breakdown and git ground truth — then resolves every blocked item's recorded decision with you. That interview is the entry fee, on principle: the run's escalation agents already spent its bounded machine retries deciding each block could not be resolved within the spec, so a retry without new human input would reproduce it. You can also resurrect items the run cut autonomously (a machine-made scope reduction you never directly approved — overruling it restores scope the original brief contained), and a resolution may authorize the retried item to amend code an earlier item already merged.

Your resolutions are appended to the run's own spec.md as binding Decisions bullets; the superseded plans, review findings, and report are archived in place (plans/<ID>.round<N>.md, reviews/prev<N>/, report.round<N>.md) so the failure evidence survives the next round; and the work loop relaunches over only the unmet items, on the same integration branch — each item carrying a retry note that points at its recorded failure evidence, and each surviving item branch (salvaged WIP commit included) resumed rather than restarted. The rewritten report keeps the whole-run picture, with a Prior rounds section naming the archived reports. An interrupted run (no report yet) is redirected to /orca:feature's resume; a clean run has nothing to retry — its follow-ups are /orca:followup

Source 5 files
hooks/index.ts 14 lines
1// orca's hooks module: the run band (./run-band.tsx), drawn above the prompt,
2// and the verb hook (./relay.ts), which answers the work loop's orca.sh verb
3// calls without a model. hooks.json names one module, so this one registers both.
4
5import type { Register } from 'claude-code'
6
7import { register as relay } from './relay'
8import { register as runBand } from './run-band'
9
10export const register: Register = (on, options) => {
11  runBand(on, options)
12  relay(on, options)
13}
14
hooks/relay.ts 177 lines
1// The verb hook: answers the work loop's orca.sh verb calls without a model.
2//
3// The work loop (scripts/work-loop.workflow.js) asks for a verb with
4// `agent('', { model: 'orca:{"v":1,"verb":…,"args":[…]}' })`. That request
5// reaches `turn.step` before any API call; this hook answers it itself by
6// running `bash <this plugin>/scripts/orca.sh <verb> <args…>` and returning
7// the result as the agent's text. It never calls `next` for an `orca:` model:
8// that would send the request (and the arguments in its model name) to the API.
9//
10// What it will run, and when:
11//   - only an `orca:` request from a subagent (a workflow agent), protocol
12//     version 1, whose verb (and subcommand) is one the work loop calls;
13//   - only while the session's permission mode is known to be
14//     bypassPermissions. The mode is read off the classic hook events that
15//     carry it and kept in $.state; no mode seen yet means refuse.
16//   - always through bash + orca.sh, never git directly: `$.process.run` runs
17//     git with repo hooks off, a script that bash starts does not.
18//
19// The answer is `b64:` + base64(UTF-8 JSON): `{ v, rc, stdout, stderr }`, or
20// `{ v, refused }` when a guard says no, or `{ v, error, timedOut? }` when the
21// command could not be run to completion. Encoded whole because the harness
22// rewrites subagent results that look instruction-shaped (a verb's output
23// naming the session's flags, say); base64 passes untouched.
24//
25// This hook is the work loop's only way to orca.sh: there is no fallback.
26// Where it is not loaded, an `orca:` request goes to the API, which has no
27// such model, and the work loop fails loudly at that call (its first is
28// `triage claim`), naming the cause.
29
30import { atom, read, update } from 'claude-code'
31import type { EngineInterface, Register } from 'claude-code'
32
33export const RELAY_VERSION = 1
34export const MODEL_PREFIX = 'orca:'
35// The longest `$.process.run` allows. setup's own cap (540 s, verbs/setup.sh)
36// sits under it, so a slow worktree setup ends in setup's own soft timeout;
37// a verb that still runs past ten minutes fails here, as a timeout naming it.
38export const RUN_TIMEOUT_MS = 600_000
39
40// Exactly the verbs work-loop.workflow.js calls through orca(), with the
41// subcommands it uses where a verb has several; null means any arguments.
42export const ALLOWED: Readonly<Record<string, readonly string[] | null>> = {
43  'self-test': null,
44  'triage': ['claim', 'release'],
45  'spec': ['mark', 'changed'],
46  'secrets': ['place', 'remove'],
47  'archive': ['attempt'],
48  'notes': ['write'],
49  'worktree-item': null,
50  'provision': null,
51  'commit-verify': null,
52  'merge-prepare': null,
53  'merge-finalize': null,
54  'commit-prepare': null,
55  'salvage': null,
56}
57
58export type RelayRequest = { v: number; verb: string; args: string[] }
59export type RelayAnswer =
60  | { v: number; rc: number; stdout: string; stderr: string }
61  | { v: number; refused: string }
62  | { v: number; error: string; timedOut?: true }
63
64const modeAtom = atom({ plugin: 'orca', key: 'permissionMode' } as const, null)
65
66// Parses and checks the request the model name carries. A string answer is
67// the refusal reason.
68export function parseRequest(model: string): RelayRequest | string {
69  if (!model.startsWith(MODEL_PREFIX)) return 'not an orca: request'
70  let req: unknown
71  try { req = JSON.parse(model.slice(MODEL_PREFIX.length)) } catch { return 'request is not JSON' }
72  if (typeof req !== 'object' || req === null || Array.isArray(req)) return 'request is not an object'
73  const { v, verb, args } = req as Record<string, unknown>
74  if (v !== RELAY_VERSION) return `unknown protocol version ${JSON.stringify(v)} (this hook speaks ${RELAY_VERSION})`
75  if (typeof verb !== 'string' || !Object.prototype.hasOwnProperty.call(ALLOWED, verb))
76    return `verb ${JSON.stringify(verb)} is not on the work loop's allow-list`
77  if (!Array.isArray(args) || args.some(a => typeof a !== 'string')) return 'args must be an array of strings'
78  const subcommands = ALLOWED[verb]
79  if (subcommands && !subcommands.includes(args[0] as string))
80    return `${verb} ${JSON.stringify(args[0] ?? null)} is not on the work loop's allow-list (${subcommands.join(', ')})`
81  return { v, verb, args: args as string[] }
82}
83
84export function argvFor(root: string, req: RelayRequest): string[] {
85  return ['bash', `${root.replace(/\/+$/, '')}/scripts/orca.sh`, req.verb, ...req.args]
86}
87
88export function encodeAnswer(answer: RelayAnswer): string {
89  const bytes = new TextEncoder().encode(JSON.stringify(answer))
90  let bin = ''
91  for (let i = 0; i < bytes.length; i += 0x8000)
92    bin += String.fromCharCode(...bytes.subarray(i, i + 0x8000))
93  return `b64:${btoa(bin)}`
94}
95
96// Main-thread events set the mode either way. An event from inside a
97// subagent may only close the gate: a subagent's own mode is not the
98// session's, and must never be what opens it.
99export function nextMode(current: string | null, seen: unknown, fromSubagent: boolean): string | null {
100  if (typeof seen !== 'string' || !seen) return current
101  if (!fromSubagent) return seen
102  return seen === 'bypassPermissions' ? current : seen
103}
104
105async function noteMode($: EngineInterface, e: { permission_mode?: string; agent_id?: string }) {
106  const current = await read($, modeAtom)
107  const mode = nextMode(current, e.permission_mode, e.agent_id !== undefined)
108  if (mode !== current) await update($, modeAtom, () => mode)
109}
110
111// A rejected $.process.run: the command could not start, or was still
112// running at the cap and was killed. The two are told apart by the time it
113// took, or by the message (2.1.291 rejects a timeout with "<plugin>:
114// $.process.run(<argv0>) aborted: still running after <ms>ms"), so the work
115// loop can name a timeout.
116export function runFailure(verb: string, message: string, elapsedMs: number): RelayAnswer {
117  if (elapsedMs >= RUN_TIMEOUT_MS - 1000 || /still running after|timed? ?out/i.test(message))
118    return { v: RELAY_VERSION, error: `orca.sh ${verb} was still running after ${RUN_TIMEOUT_MS / 60_000} minutes, the longest $.process.run allows, and was stopped: ${message}`, timedOut: true }
119  return { v: RELAY_VERSION, error: `orca.sh ${verb} did not run to completion: ${message}` }
120}
121
122async function answer($: EngineInterface, model: string, agentId: string | undefined): Promise<RelayAnswer> {
123  if (!agentId) return { v: RELAY_VERSION, refused: 'orca: requests are answered for workflow agents only' }
124  const req = parseRequest(model)
125  if (typeof req === 'string') return { v: RELAY_VERSION, refused: req }
126  const mode = await read($, modeAtom)
127  if (mode !== 'bypassPermissions')
128    return { v: RELAY_VERSION, refused: mode === null
129      ? 'permission mode not known yet (none seen since this session or hook loaded); orca runs verbs here only in bypassPermissions sessions'
130      : `permission mode is ${mode}; orca runs verbs here only in bypassPermissions sessions` }
131  let run
132  const started = Date.now()
133  try { run = await $.process.run(argvFor($.plugin.root, req), { timeoutMs: RUN_TIMEOUT_MS }) }
134  catch (err) { return runFailure(req.verb, String((err as Error)?.message ?? err), Date.now() - started) }
135  if (run.isStdoutTruncated) return { v: RELAY_VERSION, error: `orca.sh ${req.verb} wrote more than 4 MiB to stdout` }
136  return { v: RELAY_VERSION, rc: run.exitCode, stdout: run.stdout, stderr: run.stderr }
137}
138
139export const register: Register = on => {
140  // The events Phase 3 saw carry permission_mode (2.1.290/291); SubagentStart
141  // and SessionStart do not.
142  on('classic.UserPromptSubmit', async ($, e, next) => {
143    try { await noteMode($, e) } catch { /* the gate stays as it was */ }
144    return next(e)
145  }).catch(($, e, next) => next(e))
146  on('classic.PostToolUse', async ($, e, next) => {
147    try { await noteMode($, e) } catch { /* the gate stays as it was */ }
148    return next(e)
149  }).catch(($, e, next) => next(e))
150  on('classic.SubagentStop', async ($, e, next) => {
151    try { await noteMode($, e) } catch { /* the gate stays as it was */ }
152    return next(e)
153  }).catch(($, e, next) => next(e))
154  on('classic.Stop', async ($, e, next) => {
155    try { await noteMode($, e) } catch { /* the gate stays as it was */ }
156    return next(e)
157  }).catch(($, e, next) => next(e))
158
159  on('turn.step', async function* ($, e, next) {
160    if (!e.model.startsWith(MODEL_PREFIX)) return yield* next(e)
161    let reply: RelayAnswer
162    try { reply = await answer($, e.model, e.agentId) }
163    catch (err) { reply = { v: RELAY_VERSION, error: `orca hook failed: ${String((err as Error)?.message ?? err)}` } }
164    const text = encodeAnswer(reply)
165    yield { kind: 'text', index: 0, text }
166    yield { kind: 'stop', stopReason: 'end_turn', usage: null }
167    return { turnId: e.turnId, index: e.index, answer: text, toolUses: [], stopReason: 'end_turn', usage: null }
168  }).catch(async function* ($, e, next) {
169    // An orca: request never goes to the API, even when this hook failed.
170    if (next.called || !e.model.startsWith(MODEL_PREFIX)) return yield* next(e)
171    const text = encodeAnswer({ v: RELAY_VERSION, error: `orca hook failed: ${next.error.message ?? next.error.kind}` })
172    yield { kind: 'text', index: 0, text }
173    yield { kind: 'stop', stopReason: 'end_turn', usage: null }
174    return { turnId: e.turnId, index: e.index, answer: text, toolUses: [], stopReason: 'end_turn', usage: null }
175  })
176}
177
hooks/run-band.tsx 281 lines
1// The run band: a live view of the orca run this session launched, drawn
2// above the prompt. A read-only observer — it watches this session's own
3// Workflow launches of spec.workflow.js / work-loop.workflow.js, then reads
4// the run directory and the workflow's journal on a short tick while that run
5// is tracked. It writes nothing anywhere but its own $.state, and draws
6// nothing where the AbovePrompt band is not raised (VS Code, `claude -p`).
7//
8// The derivation itself is pure and lives in ./derive.ts.
9
10import { atom, read, update } from 'claude-code'
11import type { EngineInterface, Register, Timer } from 'claude-code'
12
13import type { OrcaBand, OrcaRun, OrcaTone } from '../types'
14import { derive, headerRule, itemsFromArgs, layoutRows, parseTaskNotifications, projectDirName, slugOf } from './derive'
15import type { JournalRead, RuleTone, RunFiles } from './derive'
16
17const runAtom = atom({ plugin: 'orca', key: 'run' } as const, null)
18const bandAtom = atom({ plugin: 'orca', key: 'band' } as const, null)
19
20const TICK_MS = 1500
21// $.fs.read refuses anything larger.
22const MAX_READ = 4 * 1024 * 1024
23// A spec that finished and then sat through this many of the person's own
24// prompts with no work-loop launch was abandoned at the checkpoint.
25const SPEC_IDLE_PROMPTS = 3
26
27// Module variables start over on a hot reload; the run itself is in $.state.
28// A timer keeps the `$` of the hook that started it, which stays usable after
29// that dispatch ends (checked under `claude -p` on 2.1.290).
30let ticker: Timer | null = null
31let ticking = false
32const textCache = new Map<string, { size: number; mtimeMs: number; text: string }>()
33
34function kindOf(scriptPath: unknown): OrcaRun['kind'] | null {
35  if (typeof scriptPath !== 'string') return null
36  if (/(^|\/)spec\.workflow\.js$/.test(scriptPath)) return 'spec'
37  if (/(^|\/)work-loop\.workflow\.js$/.test(scriptPath)) return 'build'
38  return null
39}
40
41function record(v: unknown): Record<string, unknown> | null {
42  if (typeof v === 'string') { try { v = JSON.parse(v) } catch { return null } }
43  return typeof v === 'object' && v !== null && !Array.isArray(v) ? (v as Record<string, unknown>) : null
44}
45
46// Reads a file through a size+mtime cache: an unchanged journal is not read
47// again every tick. 'too-large' past what $.fs.read takes.
48async function readCached($: EngineInterface, path: string): Promise<JournalRead> {
49  let stat
50  try { stat = await $.fs.stat(path) } catch { return { status: 'missing' } }
51  if (stat.size > MAX_READ) return { status: 'too-large' }
52  const hit = textCache.get(path)
53  if (hit && hit.size === stat.size && hit.mtimeMs === stat.mtimeMs) return { status: 'ok', text: hit.text }
54  try {
55    const text = await $.fs.read(path)
56    textCache.set(path, { size: stat.size, mtimeMs: stat.mtimeMs, text })
57    return { status: 'ok', text }
58  } catch { return { status: 'unreadable' } }
59}
60
61async function listNames($: EngineInterface, dir: string): Promise<{ name: string; mtimeMs: number }[]> {
62  try { return (await $.fs.list(dir)).map(f => ({ name: f.name, mtimeMs: f.mtimeMs })) } catch { return [] }
63}
64
65async function readRunFiles($: EngineInterface, run: OrcaRun): Promise<RunFiles> {
66  const files: RunFiles = { runDirExists: true, spec: null, reviews: [], plans: [], merged: null, reportMtimeMs: null, agents: [] }
67  let entries
68  try { entries = await $.fs.list(run.runDir) } catch {
69    files.runDirExists = await $.fs.exists(run.runDir).catch(() => true)
70    return files
71  }
72  const named = new Map(entries.map(f => [f.name, f]))
73  const report = named.get('report.md')
74  files.reportMtimeMs = report ? report.mtimeMs : null
75  if (run.amendPath) {
76    const st = await $.fs.stat(run.amendPath).catch(() => null)
77    files.spec = st ? { mtimeMs: st.mtimeMs, text: null } : null
78  } else {
79    const spec = named.get('spec.md')
80    // spec.md's text is needed only when the launch's own item list was unreadable.
81    const needText = run.kind === 'build' && run.items.length === 0
82    const text = spec && needText ? await readCached($, `${run.runDir}/spec.md`) : null
83    files.spec = spec ? { mtimeMs: spec.mtimeMs, text: text && text.status === 'ok' ? text.text : null } : null
84  }
85  if (named.has('reviews')) files.reviews = await listNames($, `${run.runDir}/reviews`)
86  if (named.has('plans')) files.plans = (await listNames($, `${run.runDir}/plans`)).map(f => f.name)
87  if (run.kind === 'build' && named.has('merged.tsv')) {
88    const merged = await readCached($, `${run.runDir}/merged.tsv`)
89    files.merged = merged.status === 'ok' ? merged.text : null
90  }
91  if (run.kind === 'spec' && run.journal) files.agents = await listNames($, run.journal.replace(/\/[^/]*$/, ''))
92  return files
93}
94
95// Writes only a change: every write redraws the band.
96async function setBand($: EngineInterface, band: OrcaBand | null) {
97  if (JSON.stringify(await read($, bandAtom)) === JSON.stringify(band)) return
98  await update($, bandAtom, () => band)
99}
100
101async function forget($: EngineInterface) {
102  if (ticker) ticker.cancel()
103  ticker = null
104  await update($, runAtom, () => null)
105  await setBand($, null)
106}
107
108async function tick($: EngineInterface) {
109  if (ticking) return
110  ticking = true
111  try {
112    const run = await read($, runAtom)
113    if (!run) {
114      if (ticker) ticker.cancel()
115      ticker = null
116      await setBand($, null)
117      return
118    }
119    // Nothing draws the band here (a `claude -p` run, VS Code): read nothing.
120    const surfaces = await $.session.surfaces()
121    if (!surfaces.some(s => s === 'terminal' || s === 'desktop')) return
122    const files = await readRunFiles($, run)
123    const journal: JournalRead = run.journal ? await readCached($, run.journal) : { status: 'missing' }
124    const out = derive(run, files, journal, await $.clock.now())
125    // The run may have ended (a notification) or been replaced (a new launch)
126    // while this tick read files; a stale answer must not redraw the band.
127    if ((await read($, runAtom))?.runId !== run.runId) return
128    if (out.finished) await forget($)
129    else await setBand($, out.band)
130  } catch {
131    // Never crash the session over a status line; the next tick tries again.
132  } finally {
133    ticking = false
134  }
135}
136
137function follow($: EngineInterface) {
138  if (!ticker) ticker = $.clock.every(TICK_MS, () => { void tick($) })
139  void tick($)
140}
141
142// The launch result names the workflow's transcript directory, which holds
143// journal.jsonl. Without it, the same path is rebuilt from the session.
144async function journalPath($: EngineInterface, transcriptDir: unknown, runId: string): Promise<string | null> {
145  if (typeof transcriptDir === 'string' && transcriptDir) return `${transcriptDir.replace(/\/+$/, '')}/journal.jsonl`
146  const config = (await $.env.get('CLAUDE_CONFIG_DIR')) ?? `${(await $.env.get('HOME')) ?? ''}/.claude`
147  if (!config.startsWith('/')) return null
148  const cwd = await $.session.cwd()
149  const session = await $.session.id()
150  return `${config}/projects/${projectDirName(cwd)}/${session}/subagents/workflows/${runId}/journal.jsonl`
151}
152
153async function track($: EngineInterface, kind: OrcaRun['kind'], input: Record<string, unknown>, launched: unknown, since: number) {
154  const result = record(launched)
155  if (!result || result.status !== 'async_launched') return
156  const runId = result.runId
157  const taskId = result.taskId
158  if (typeof runId !== 'string' || typeof taskId !== 'string') return
159  const args = record(input.args)
160  const runDir = args && typeof args.runDir === 'string' ? args.runDir.replace(/\/+$/, '') : null
161  if (!args || !runDir) return
162  const run: OrcaRun = {
163    kind,
164    runDir,
165    slug: typeof args.slug === 'string' && args.slug ? args.slug : slugOf(runDir),
166    runId,
167    taskId,
168    journal: await journalPath($, result.transcriptDir, runId),
169    since,
170    items: kind === 'build' ? itemsFromArgs(args) : [],
171    amendPath: kind === 'spec' && typeof args.amendPath === 'string' ? args.amendPath : null,
172    reviewer: kind === 'spec' && typeof args.reviewer === 'string' && args.reviewer ? args.reviewer : null,
173    specEnded: false,
174    idlePrompts: 0,
175  }
176  await update($, runAtom, () => run)
177  follow($)
178}
179
180async function onPrompt($: EngineInterface, origin: string, text: string) {
181  const run = await read($, runAtom)
182  if (!run) return
183  if (origin === 'task-notification') {
184    const ended = parseTaskNotifications(text).find(n => n.taskId === run.taskId)
185    if (!ended) return
186    // The work loop ending ends the run's live phase. A spec that completed
187    // stays on screen as ready (the checkpoint, or the launch about to come);
188    // one that failed or was stopped does not.
189    if (run.kind === 'build' || ended.status !== 'completed') await forget($)
190    else await update($, runAtom, r => (r ? { ...r, specEnded: true } : r))
191    return
192  }
193  if ((origin === 'composer' || origin === 'bridge') && run.kind === 'spec') {
194    const band = await read($, bandAtom)
195    if (!run.specEnded && band?.detail !== 'ready') return
196    if (run.idlePrompts + 1 >= SPEC_IDLE_PROMPTS) await forget($)
197    else await update($, runAtom, r => (r ? { ...r, idlePrompts: r.idlePrompts + 1 } : r))
198  }
199}
200
201type Style = { color?: string; dimColor?: boolean; bold?: boolean }
202
203const TONE: Record<OrcaTone, Style> = {
204  done: { color: 'success' },
205  live: { color: 'suggestion' },
206  wait: { dimColor: true },
207  queued: { dimColor: true },
208  blocked: { color: 'error' },
209  cut: { dimColor: true },
210}
211
212// The top border: the rule and the name in Claude's accent, the facts set in it.
213const RULE: Record<RuleTone, Style> = {
214  rule: { color: 'claude', dimColor: true },
215  name: { color: 'claude', bold: true },
216  slug: { bold: true },
217  phase: { color: 'suggestion' },
218  detail: {},
219  sep: { dimColor: true },
220  note: { color: 'warning' },
221}
222
223export const register: Register = on => {
224  on('session.start', async ($, e, next) => {
225    // A hot reload drops the timer; the tracked run is still in $.state.
226    if (await read($, runAtom)) follow($)
227    return next(e)
228  })
229
230  on('tool.call', { tool: 'Workflow' }, async ($, e, next) => {
231    // Only the main conversation's own launches: the skills launch from there.
232    if (e.tool !== 'Workflow' || e.agentId !== undefined) return next(e)
233    const kind = kindOf(e.scriptPath)
234    if (!kind) return next(e)
235    const since = await $.clock.now()
236    const ran = await next(e)
237    if (!ran.isError && 'result' in ran) {
238      try { await track($, kind, { args: e.args }, ran.result, since) } catch { /* observe only */ }
239    }
240    return ran
241  }).catch(($, e, next) => next(e)) // an observer never stands in a call's way
242
243  on('prompt.submit', async ($, e, next) => {
244    try { await onPrompt($, e.origin.kind, e.text) } catch { /* observe only */ }
245    return next(e)
246  }).catch(($, e, next) => next(e))
247
248  on('session.end', async ($, e, next) => {
249    try { await forget($) } catch { /* observe only */ }
250    return next(e)
251  })
252
253  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
254    if (e.props.hasSurvey) return next(e)
255    const band = await read($, bandAtom)
256    if (!band || !(await read($, runAtom))) return next(e)
257    const { Box, Text } = $.ui.resolve(e)
258    const columns = Math.max(24, e.props.bodyColumns || 80)
259
260    // A blank row above the rule sets the band off from the transcript's last
261    // line, and one below it sets the rule off from the rows.
262    return (
263      <Box flexDirection="column" marginTop={1}>
264        <Box marginBottom={band.rows.length ? 1 : 0}>
265          {headerRule(band, columns).map(piece => (
266            <Text {...RULE[piece.tone]} wrap="truncate-end">{piece.text}</Text>
267          ))}
268        </Box>
269        {layoutRows(band.rows, columns).map(line => (
270          <Box>
271            <Text>{'  '}</Text>
272            <Text bold>{line.id}</Text>
273            <Text>{`  ${line.title}  `}</Text>
274            <Text {...TONE[line.row.tone]} wrap="truncate-end">{line.tail}</Text>
275          </Box>
276        ))}
277      </Box>
278    )
279  })
280}
281
hooks/derive.ts 567 lines
1// Pure derivation for the run band: run files and the workflow journal in,
2// the rows the band draws out. No `$` here — run-band.tsx gathers the inputs
3// (and owns every engine call), so everything below is testable as plain data.
4
5import type { OrcaBand, OrcaItem, OrcaRow, OrcaRun } from '../types'
6
7// What run-band.tsx reads from the run directory on one tick.
8export type RunFiles = {
9  runDirExists: boolean
10  // spec.md (or the amendment file on an iterate spec launch): its mtime,
11  // and its text when the caller needed it.
12  spec: { mtimeMs: number; text: string | null } | null
13  // reviews/ entries, by name.
14  reviews: { name: string; mtimeMs: number }[]
15  // plans/ entry names.
16  plans: string[]
17  // merged.tsv's text, null when absent.
18  merged: string | null
19  // report.md's mtime, null when absent.
20  reportMtimeMs: number | null
21  // The workflow's transcript directory (the journal's own), by name: a spec
22  // launch times its stages from the agent files there.
23  agents: { name: string; mtimeMs: number }[]
24}
25
26export type JournalRead =
27  | { status: 'ok'; text: string }
28  | { status: 'missing' | 'unreadable' | 'too-large' }
29
30export type Derived = { finished: boolean; band: OrcaBand | null }
31
32// One agent() call the journal recorded.
33export type JournalStage = {
34  seq: number
35  label: string
36  agentId: string
37  running: boolean
38  failed: boolean
39  result: unknown
40}
41
42// ---------- journal ----------
43
44// One JSON object per line: {type:"started",key,agentId,label,phase} and
45// later {type:"result"|"failed",key,agentId,...}. A line that does not parse
46// (the torn tail of a write in progress, a partial flush) is skipped — the
47// next tick reads it whole. A started entry with no label predates labels and
48// says nothing the band can use.
49export function parseJournal(text: string): JournalStage[] {
50  const stages: JournalStage[] = []
51  const open = new Map<string, JournalStage[]>()
52  for (const line of text.split('\n')) {
53    if (!line.trim()) continue
54    let row: unknown
55    try { row = JSON.parse(line) } catch { continue }
56    if (!isRecord(row)) continue
57    const key = typeof row.key === 'string' ? row.key : null
58    if (row.type === 'started' && key !== null) {
59      const stage: JournalStage = {
60        seq: stages.length, label: typeof row.label === 'string' ? row.label : '',
61        agentId: typeof row.agentId === 'string' ? row.agentId : '',
62        running: true, failed: false, result: undefined,
63      }
64      stages.push(stage)
65      const list = open.get(key) ?? []
66      list.push(stage)
67      open.set(key, list)
68    } else if ((row.type === 'result' || row.type === 'failed') && key !== null) {
69      const list = open.get(key)
70      const stage = list?.shift()
71      if (!stage) continue
72      stage.running = false
73      stage.failed = row.type === 'failed'
74      stage.result = row.result
75    }
76  }
77  return stages.filter(s => s.label !== '')
78}
79
80// What a label in work-loop.workflow.js / spec.workflow.js means for the band.
81export type LabelMeaning =
82  | { kind: 'item'; ids: string[]; stage: string }
83  | { kind: 'escalate'; ids: string[] }
84  | { kind: 'salvage'; id: string }
85  | { kind: 'integration'; stage: string }
86  | { kind: 'spec'; state: 'drafting' | 'reviewing' | 'revising' }
87  | { kind: 'wrap-up' }
88  | null
89
90export function classifyLabel(label: string): LabelMeaning {
91  // Retries keep their stage: review:W1#0~retry.
92  const bare = label.replace(/~retry$/, '')
93  if (bare === 'spec') return { kind: 'spec', state: 'drafting' }
94  if (bare === 'spec-review') return { kind: 'spec', state: 'reviewing' }
95  if (bare === 'spec-revise') return { kind: 'spec', state: 'revising' }
96  if (bare === 'run-lease-release') return { kind: 'wrap-up' }
97  if (bare === 'integration-verify') return { kind: 'integration', stage: 'verify' }
98  if (bare === 'integration-review-clean') return { kind: 'integration', stage: 'review' }
99  if (bare === 'integration-status') return { kind: 'integration', stage: 'commit' }
100  if (bare === 'integration-fix-notes') return { kind: 'integration', stage: 'fix' }
101  const colon = bare.indexOf(':')
102  if (colon < 0) return null
103  // reconcile#2 and reconcile~serial are reconcile; plan:W3#2 is a replan.
104  const head = bare.slice(0, colon).replace(/[#~].*$/, '')
105  const tail = bare.slice(colon + 1)
106  const hash = tail.indexOf('#')
107  const subject = hash < 0 ? tail : tail.slice(0, hash)
108  const suffix = hash < 0 ? '' : tail.slice(hash + 1)
109  const round = /^\d+$/.test(suffix) ? Number(suffix) : null
110  const ids = subject.split('+').filter(Boolean)
111  if (!ids.length) return null
112
113  let stage: string | null
114  switch (head) {
115    case 'plan': stage = suffix ? 'replan' : 'plan'; break
116    case 'plan-archive':
117    case 'review-archive': stage = 'replan'; break
118    case 'reconcile': stage = 'reconcile'; break
119    case 'escalate': return { kind: 'escalate', ids }
120    case 'salvage': return { kind: 'salvage', id: ids[0] ?? subject }
121    case 'worktree': stage = 'worktree'; break
122    case 'provision': stage = 'provision'; break
123    case 'implement': stage = 'implement'; break
124    // Rounds are 0-based in labels: review:W3#0 is the first review.
125    case 'secrets-remove':
126    case 'review': stage = round === null ? 'review' : `review r${round + 1}`; break
127    case 'secrets-place':
128    case 'fix': stage = round === null ? 'fix' : `fix r${round}`; break
129    case 'secrets-restore': stage = 'review'; break
130    case 'commit':
131    case 'commit-verify': stage = 'commit'; break
132    case 'merge-reset':
133    case 'merge':
134    case 'merge-finalize': stage = 'merge'; break
135    default: stage = null // spec-hash, run-lease and other plumbing
136  }
137  if (stage === null) return null
138  if (subject === 'integration') return { kind: 'integration', stage }
139  return { kind: 'item', ids, stage }
140}
141
142// ---------- spec phase ----------
143
144export type SpecState = 'drafting' | 'reviewing' | 'revising' | 'ready'
145
146// Files decide by default: no fresh spec is drafting, a fresh spec with no
147// fresh review artifact is reviewing, both is ready (which covers the
148// checkpoint wait). "Fresh" is "written since this launch": a checkpoint
149// re-spawn or an iterate amend round runs over a directory that already holds
150// a spec and a review. A running spec-stage agent in the journal overrides
151// the files, and a completed spec workflow is ready.
152export function specState(run: OrcaRun, files: RunFiles, stages: JournalStage[] | null): SpecState {
153  if (run.specEnded) return 'ready'
154  const running = (stages ?? []).filter(s => s.running).map(s => classifyLabel(s.label))
155  const live = running.filter((m): m is Extract<LabelMeaning, { kind: 'spec' }> => m?.kind === 'spec').pop()
156  if (live) return live.state
157  if (!files.spec || files.spec.mtimeMs < run.since) return 'drafting'
158  const artifact = run.amendPath ? /^spec-amend-[a-z]+\.json$/ : /^spec-[a-z]+\.json$/
159  const reviewed = files.reviews.some(r => artifact.test(r.name) && r.mtimeMs >= run.since)
160  return reviewed ? 'ready' : 'reviewing'
161}
162
163// An agent's transcript directory holds agent-<id>.meta.json, written once at
164// its spawn, and agent-<id>.jsonl, appended to until it returns: the spawn and
165// the last word. Whole minutes, so the band redraws once a minute at most.
166export function elapsed(stage: JournalStage, agents: { name: string; mtimeMs: number }[], now: number): string {
167  const meta = agents.find(a => a.name === `agent-${stage.agentId}.meta.json`)
168  const end = stage.running ? now : agents.find(a => a.name === `agent-${stage.agentId}.jsonl`)?.mtimeMs
169  if (!stage.agentId || !meta || end === undefined) return ''
170  const minutes = Math.floor(Math.max(0, end - meta.mtimeMs) / 60_000)
171  return minutes < 1 ? '<1m' : `${minutes}m`
172}
173
174// One review attempt as the spec workflow judges it: a written artifact with
175// possible counts passes; anything else is a failed attempt and its reason.
176type ReviewVerdict = { ok: true; total: number; criticalHigh: number } | { ok: false; reason: string }
177
178function reviewVerdict(stage: JournalStage): ReviewVerdict {
179  let r = stage.result
180  if (typeof r === 'string') { try { r = JSON.parse(r) } catch { r = null } }
181  if (stage.failed || !isRecord(r)) return { ok: false, reason: 'review agent was skipped or died' }
182  if (r.written !== true) return { ok: false, reason: typeof r.reason === 'string' && r.reason ? r.reason : 'no artifact written' }
183  const total = typeof r.total === 'number' ? r.total : 0
184  const criticalHigh = typeof r.criticalHigh === 'number' ? r.criticalHigh : 0
185  if (criticalHigh > total) return { ok: false, reason: `impossible count (${criticalHigh} Critical/High of ${total})` }
186  return { ok: true, total, criticalHigh }
187}
188
189const plural = (n: number, word: string) => `${n} ${word}${n === 1 ? '' : 's'}`
190
191// The gate's three stages as rows, from the journal: spec.workflow.js labels
192// them spec, spec-review (spec-review~retry for its second attempt) and
193// spec-revise. A stage that returned nothing (skipped, died) is "died".
194function specRowsFromJournal(run: OrcaRun, stages: JournalStage[], agents: RunFiles['agents'], now: number): OrcaRow[] {
195  const bare = (s: JournalStage) => s.label.replace(/~retry$/, '')
196  const last = (label: string) => stages.filter(s => bare(s) === label).pop()
197  const died = (s: JournalStage) => s.failed || s.result === null || s.result === undefined
198  const time = (s: JournalStage) => { const t = elapsed(s, agents, now); return t ? ` · ${t}` : '' }
199  const [specRow, reviewRow, reviseRow] = specRowShells(run)
200
201  const spec = last('spec')
202  const specDied = spec !== undefined && !spec.running && died(spec)
203  if (!spec) specRow.set('●', 'starting', 'live')
204  else if (spec.running) specRow.set('●', `drafting${time(spec)}`, 'live')
205  else if (specDied) specRow.set('✕', 'died', 'blocked')
206  else specRow.set('✓', `done${time(spec)}`, 'done')
207
208  const review = last('spec-review')
209  const retry = review?.label.endsWith('~retry') ?? false
210  const verdict = review && !review.running ? reviewVerdict(review) : null
211  const failedOpen = verdict !== null && !verdict.ok && retry
212  if (specDied) reviewRow.set('–', 'skipped', 'cut')
213  else if (!review) reviewRow.set('○', 'queued', 'queued')
214  else if (review.running) reviewRow.set('●', `reviewing${retry ? ' (retry)' : ''}${time(review)}`, 'live')
215  else if (verdict?.ok) {
216    const found = verdict.total === 0 ? 'no findings'
217      : `${plural(verdict.total, 'finding')}, ${verdict.criticalHigh} Critical/High`
218    reviewRow.set('✓', `${found}${time(review)}`, 'done')
219  } else if (failedOpen) reviewRow.set('✕', `failed open: ${verdict?.ok === false ? verdict.reason : ''}`, 'blocked')
220  else reviewRow.set('●', 'retrying', 'live')
221
222  const revise = last('spec-revise')
223  const criticalHigh = verdict?.ok ? verdict.criticalHigh : 0
224  if (revise?.running) reviseRow.set('●', `revising${time(revise)}`, 'live')
225  else if (revise && died(revise)) reviseRow.set('✕', 'died: the findings stand', 'blocked')
226  else if (revise) reviseRow.set('✓', `revised${time(revise)}`, 'done')
227  else if (specDied || failedOpen) reviseRow.set('–', 'skipped', 'cut')
228  else if (verdict?.ok && criticalHigh === 0) reviseRow.set('–', 'not needed', 'cut')
229  else if (verdict?.ok) reviseRow.set('●', 'starting', 'live')
230  else reviseRow.set('○', 'on Critical/High only', 'queued')
231
232  return [specRow.row, reviewRow.row, reviseRow.row]
233}
234
235// Without a readable journal only the files speak: what was written, never
236// what a review found or whether a revise ran.
237function specRowsFromFiles(run: OrcaRun, state: SpecState): OrcaRow[] {
238  const [specRow, reviewRow, reviseRow] = specRowShells(run)
239  if (state === 'drafting') specRow.set('●', 'drafting', 'live')
240  else specRow.set('✓', 'done', 'done')
241  if (state === 'drafting') reviewRow.set('○', 'queued', 'queued')
242  else if (state === 'ready') reviewRow.set('✓', 'done', 'done')
243  else reviewRow.set('●', 'reviewing', 'live')
244  if (state === 'ready') reviseRow.set('◌', 'unknown', 'wait')
245  else reviseRow.set('○', 'on Critical/High only', 'queued')
246  return [specRow.row, reviewRow.row, reviseRow.row]
247}
248
249function specRowShells(run: OrcaRun) {
250  const shell = (id: string, title: string) => {
251    const row: OrcaRow = { id, title, mark: '○', status: '', tone: 'queued' }
252    return { row, set: (mark: string, status: string, tone: OrcaRow['tone']) => Object.assign(row, { mark, status, tone }) }
253  }
254  return [
255    shell('spec', run.amendPath ? 'author the amendment' : 'author spec.md from the brief'),
256    shell('review', `${run.reviewer ?? 'independent'} review against brief + codebase`),
257    shell('revise', 'fold Critical/High findings in'),
258  ] as const
259}
260
261export function deriveSpec(run: OrcaRun, files: RunFiles, journal: JournalRead, now: number): OrcaBand {
262  const stages = journal.status === 'ok' ? parseJournal(journal.text) : null
263  const state = specState(run, files, stages)
264  const rows = stages ? specRowsFromJournal(run, stages, files.agents, now) : specRowsFromFiles(run, state)
265  return { slug: run.slug, phase: 'spec', detail: state, rows, source: stages ? 'journal' : 'files' }
266}
267
268// ---------- build phase ----------
269
270// The last `**Workflow args:**` line of spec.md: the item list a resume
271// replays. Used only when the launch's own args could not be read.
272export function itemsFromSpec(text: string | null): OrcaItem[] {
273  if (!text) return []
274  const lines = text.split('\n').filter(l => l.startsWith('**Workflow args:** '))
275  const last = lines[lines.length - 1]
276  if (!last) return []
277  try { return itemsFromArgs(JSON.parse(last.slice('**Workflow args:** '.length))) } catch { return [] }
278}
279
280export function itemsFromArgs(args: unknown): OrcaItem[] {
281  let parsed = args
282  if (typeof parsed === 'string') { try { parsed = JSON.parse(parsed) } catch { return [] } }
283  if (!isRecord(parsed)) return []
284  let items = parsed.items
285  if (typeof items === 'string') { try { items = JSON.parse(items) } catch { return [] } }
286  if (!Array.isArray(items)) return []
287  return items.filter(isRecord).filter(i => typeof i.id === 'string').map(i => ({
288    id: i.id as string,
289    title: typeof i.title === 'string' ? i.title : '',
290    deps: Array.isArray(i.deps) ? i.deps.filter((d): d is string => typeof d === 'string') : [],
291  }))
292}
293
294export function mergedIds(text: string | null): Set<string> {
295  const ids = new Set<string>()
296  for (const line of (text ?? '').split('\n')) {
297    const [id, sha] = line.split('\t')
298    if (id && sha && sha.trim()) ids.add(id)
299  }
300  return ids
301}
302
303type ItemTrack = {
304  item: OrcaItem
305  blocked: string | null
306  cut: string | null
307  salvaged: boolean
308  live: string | null
309  last: string | null
310}
311
312// Escalation results follow ESCALATE ({replan, cut, blocked, addDeps}) for a
313// wave and BUILD_ESCALATE ({action, reason}) for one item mid-build.
314function applyEscalation(track: Map<string, ItemTrack>, items: OrcaItem[], ids: string[], result: unknown, merged: Set<string>) {
315  let r = result
316  if (typeof r === 'string') { try { r = JSON.parse(r) } catch { return } }
317  if (!isRecord(r)) return
318  if (typeof r.action === 'string') {
319    const t = track.get(ids[0] ?? '')
320    const reason = typeof r.reason === 'string' ? r.reason : ''
321    if (t && r.action === 'block') t.blocked = reason
322    if (t && r.action === 'cut') t.cut = reason
323    return
324  }
325  for (const b of Array.isArray(r.blocked) ? r.blocked : []) {
326    if (!isRecord(b) || typeof b.id !== 'string' || !ids.includes(b.id)) continue
327    const t = track.get(b.id)
328    if (t) t.blocked = typeof b.reason === 'string' ? b.reason : ''
329  }
330  for (const c of Array.isArray(r.cut) ? r.cut : []) {
331    if (!isRecord(c) || typeof c.id !== 'string' || !ids.includes(c.id)) continue
332    const t = track.get(c.id)
333    if (t) t.cut = typeof c.reason === 'string' ? c.reason : ''
334  }
335  // The scheduler applies addDeps to items not yet building, never as a
336  // cycle; an item already merged keeps its deps as they were.
337  for (const a of Array.isArray(r.addDeps) ? r.addDeps : []) {
338    if (!isRecord(a) || typeof a.id !== 'string' || !Array.isArray(a.dependsOn)) continue
339    const it = items.find(i => i.id === a.id)
340    if (!it || merged.has(it.id)) continue
341    for (const dep of a.dependsOn) {
342      if (typeof dep !== 'string' || dep === it.id || it.deps.includes(dep)) continue
343      if (!items.some(i => i.id === dep) || reaches(items, dep, it.id)) continue
344      it.deps.push(dep)
345    }
346  }
347}
348
349function reaches(items: OrcaItem[], from: string, to: string): boolean {
350  const seen = new Set<string>()
351  const stack = [from]
352  while (stack.length) {
353    const cur = stack.pop() as string
354    if (cur === to) return true
355    if (seen.has(cur)) continue
356    seen.add(cur)
357    const it = items.find(i => i.id === cur)
358    if (it) stack.push(...it.deps)
359  }
360  return false
361}
362
363// A review artifact on disk names the last completed round: <ID>-<reviewer>.json
364// is the latest, <ID>-<reviewer>.round<N>.json one per round.
365function reviewRound(id: string, reviews: { name: string }[]): number {
366  let round = 0
367  const esc = id.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')
368  const re = new RegExp(`^${esc}-[a-z]+(?:\\.round(\\d+))?\\.json$`)
369  for (const r of reviews) {
370    const m = re.exec(r.name)
371    if (!m) continue
372    round = Math.max(round, m[1] === undefined ? 1 : Number(m[1]) + 1)
373  }
374  return round
375}
376
377const isCut = (t: ItemTrack | undefined) => t !== undefined && t.cut !== null
378const isBlocked = (t: ItemTrack | undefined) =>
379  t !== undefined && !isCut(t) && (t.blocked !== null || t.salvaged)
380
381// Dependencies that have not merged; a cut one counts as satisfied, as it
382// does for the scheduler. A blocked one is named as such: the item waits on
383// a decision, not on work in progress.
384function waitingOn(item: OrcaItem, merged: Set<string>, track: Map<string, ItemTrack>): string[] {
385  return item.deps
386    .filter(d => !merged.has(d) && !isCut(track.get(d)))
387    .map(d => (isBlocked(track.get(d)) ? `${d} (blocked)` : d))
388}
389
390export function deriveBuild(run: OrcaRun, files: RunFiles, journal: JournalRead): OrcaBand {
391  const items = (run.items.length ? run.items : itemsFromSpec(files.spec?.text ?? null))
392    .map(i => ({ ...i, deps: [...i.deps] }))
393  const merged = mergedIds(files.merged)
394  const track = new Map<string, ItemTrack>(items.map(item =>
395    [item.id, { item, blocked: null, cut: null, salvaged: false, live: null, last: null }]))
396  const stages = journal.status === 'ok' ? parseJournal(journal.text) : null
397
398  const integration: { live: string | null; seen: boolean } = { live: null, seen: false }
399  let wrapping = false
400  for (const s of stages ?? []) {
401    const m = classifyLabel(s.label)
402    if (!m) continue
403    if (m.kind === 'item') {
404      for (const id of m.ids) {
405        const t = track.get(id)
406        if (!t) continue
407        t.last = m.stage
408        if (s.running) t.live = m.stage
409      }
410    } else if (m.kind === 'escalate') {
411      for (const id of m.ids) {
412        const t = track.get(id)
413        if (!t) continue
414        t.last = 'escalate'
415        if (s.running) t.live = 'escalate'
416      }
417      if (!s.running && !s.failed) applyEscalation(track, items, m.ids, s.result, merged)
418    } else if (m.kind === 'salvage') {
419      const t = track.get(m.id)
420      if (t) t.salvaged = true
421    } else if (m.kind === 'integration') {
422      integration.seen = true
423      if (s.running) integration.live = m.stage
424    } else if (m.kind === 'wrap-up') {
425      wrapping = true
426    }
427  }
428
429  const rows: OrcaRow[] = items.map(item => {
430    const t = track.get(item.id) as ItemTrack
431    const row = (mark: string, status: string, tone: OrcaRow['tone']): OrcaRow =>
432      ({ id: item.id, title: item.title, mark, status, tone })
433    if (merged.has(item.id)) return row('✓', 'done', 'done')
434    if (isCut(t)) return row('–', t.cut ? `cut: ${t.cut}` : 'cut', 'cut')
435    if (isBlocked(t)) {
436      const why = t.blocked || (t.last ? `after ${t.last}` : '')
437      return row('✕', why ? `blocked: ${why}` : 'blocked', 'blocked')
438    }
439    if (t.live) return row('●', t.live, 'live')
440    const waits = waitingOn(item, merged, track)
441    if (waits.length) return row('◌', `waiting on ${waits.join(', ')}`, 'wait')
442    if (stages) return t.last ? row('●', t.last, 'live') : row('○', 'queued', 'queued')
443    // File-only fallback: coarse, from what the stages leave on disk.
444    const round = reviewRound(item.id, files.reviews)
445    if (round > 0) return row('●', `reviewed r${round}`, 'live')
446    if (files.plans.includes(`${item.id}.md`)) return row('●', 'planning done', 'live')
447    return row('○', 'queued', 'queued')
448  })
449
450  const done = items.filter(i => merged.has(i.id)).length
451  const phase = wrapping ? 'wrapping up'
452    : integration.seen ? (integration.live ? `integrate · ${integration.live}` : 'integrate')
453      : 'build'
454  return { slug: run.slug, phase, detail: `${done}/${items.length} merged`, rows,
455    source: stages ? 'journal' : 'files' }
456}
457
458// ---------- the whole band ----------
459
460export function derive(run: OrcaRun, files: RunFiles, journal: JournalRead, now: number): Derived {
461  // A run directory that went away (archived, removed) takes the band with it.
462  if (!files.runDirExists) return { finished: true, band: null }
463  // The feature skill writes report.md after the work loop returns: the run
464  // is over. Only a report written since this launch counts — a retry or an
465  // iterate round relaunches over a directory whose old report was archived.
466  if (run.kind === 'build' && files.reportMtimeMs !== null && files.reportMtimeMs >= run.since)
467    return { finished: true, band: null }
468  if (run.kind === 'spec') return { finished: false, band: deriveSpec(run, files, journal, now) }
469  return { finished: false, band: deriveBuild(run, files, journal) }
470}
471
472// <task-notification><task-id>…</task-id>…<status>…</status>: the message the
473// session receives when a background task (a workflow included) ends. One
474// delivery may carry several.
475export function parseTaskNotifications(text: string): { taskId: string; status: string }[] {
476  const out: { taskId: string; status: string }[] = []
477  for (const block of text.split('<task-notification>').slice(1)) {
478    const id = /<task-id>([^<]+)<\/task-id>/.exec(block)
479    if (!id || !id[1]) continue
480    const status = /<status>([^<]+)<\/status>/.exec(block)
481    out.push({ taskId: id[1].trim(), status: status?.[1]?.trim() ?? '' })
482  }
483  return out
484}
485
486// .orca/<timestamp>-feat-<slug>/ → <slug>
487export function slugOf(runDir: string): string {
488  const base = runDir.replace(/\/+$/, '').split('/').pop() ?? runDir
489  return base.replace(/^\d{8}-\d{6}-(feat|bug|proto)-/, '')
490}
491
492// Claude Code's project folder name for a directory: every character that is
493// not a letter or a digit becomes '-'.
494export function projectDirName(cwd: string): string {
495  return cwd.replace(/[^a-zA-Z0-9]/g, '-')
496}
497
498// ---------- layout ----------
499
500export function truncate(text: string, width: number): string {
501  const flat = text.replace(/\s+/g, ' ').trim()
502  if (width <= 0) return ''
503  return flat.length <= width ? flat : `${flat.slice(0, Math.max(0, width - 1))}…`
504}
505
506// The band's top border: a rule across the width with the header set in it,
507// `── orca ─ slug · phase · detail ─────`. Each piece keeps its own tone; the
508// facts are cut from the end to fit and the rule takes whatever is left.
509export type RuleTone = 'rule' | 'name' | 'slug' | 'phase' | 'detail' | 'sep' | 'note'
510export type RulePiece = { text: string; tone: RuleTone }
511
512export function headerRule(band: OrcaBand, columns: number): RulePiece[] {
513  const facts: RulePiece[] = [
514    { text: band.slug, tone: 'slug' },
515    { text: band.phase, tone: 'phase' },
516    { text: band.detail, tone: 'detail' },
517  ]
518  if (band.source === 'files') facts.push({ text: 'no journal', tone: 'note' })
519  const out: RulePiece[] = [{ text: '── ', tone: 'rule' }, { text: 'orca', tone: 'name' }, { text: ' ─ ', tone: 'rule' }]
520  let used = out.reduce((n, p) => n + p.text.length, 0)
521  // A space and at least two rule cells close the line after the facts.
522  let room = columns - used - 3
523  let first = true
524  for (const fact of facts) {
525    const text = fact.text.replace(/\s+/g, ' ').trim()
526    if (!text) continue
527    if (!first) {
528      // A separator with nothing after it says nothing: stop instead.
529      if (room < 5) break
530      out.push({ text: ' · ', tone: 'sep' })
531      room -= 3
532    }
533    first = false
534    if (room <= 0) break
535    const cut = truncate(text, room)
536    out.push({ text: cut, tone: fact.tone })
537    room -= cut.length
538    if (cut !== text) break
539  }
540  used = out.reduce((n, p) => n + p.text.length, 0)
541  out.push({ text: ` ${'─'.repeat(Math.max(2, columns - used - 1))}`, tone: 'rule' })
542  return out
543}
544
545// Fits the rows to the band's width: ids padded to one column, titles cut to
546// what is left beside the widest status (statuses capped so a long blocked
547// reason never squeezes every title away), statuses cut to the remainder.
548// `lead` is `  <id>  <title>  `, with `id` and `title` its two padded cells.
549export function layoutRows(rows: OrcaRow[], columns: number): { lead: string; id: string; title: string; tail: string; row: OrcaRow }[] {
550  const idW = Math.max(0, ...rows.map(r => r.id.length))
551  const statusW = Math.min(30, Math.max(0, ...rows.map(r => r.status.length + 2)))
552  const lead = 2 + idW + 2
553  const titleMax = Math.max(0, ...rows.map(r => r.title.replace(/\s+/g, ' ').trim().length))
554  const titleW = Math.min(titleMax, Math.max(8, columns - lead - 2 - statusW))
555  return rows.map(row => {
556    const id = row.id.padEnd(idW)
557    const title = truncate(row.title, titleW).padEnd(titleW)
558    const head = `  ${id}  ${title}  `
559    const tail = truncate(`${row.mark} ${row.status}`, Math.max(1, columns - head.length))
560    return { lead: head, id, title, tail, row }
561  })
562}
563
564function isRecord(v: unknown): v is Record<string, unknown> {
565  return typeof v === 'object' && v !== null && !Array.isArray(v)
566}
567
types/index.d.ts 61 lines
1// orca's $.state contract: what the hooks module keeps in $.state so it
2// survives a hot reload — the run band's run and band (hooks/run-band.tsx),
3// and the verb hook's permission gate (hooks/relay.ts).
4
5// One work item as the work-loop launch passed it.
6export type OrcaItem = { id: string; title: string; deps: string[] }
7
8// The one orca workflow this session launched and the band follows.
9export type OrcaRun = {
10  // `spec` for spec.workflow.js, `build` for work-loop.workflow.js.
11  kind: 'spec' | 'build'
12  runDir: string
13  slug: string
14  runId: string
15  taskId: string
16  // <transcriptDir>/journal.jsonl from the launch result; null when the
17  // result carried no transcript directory and none could be derived.
18  journal: string | null
19  // When the launch was seen (ms since the epoch): run files older than
20  // this belong to an earlier launch over the same run directory.
21  since: number
22  items: OrcaItem[]
23  // The iterate skill's amendment file, when the spec launch passed one.
24  amendPath: string | null
25  // The spec launch's reviewer (codex, claude); null for a build launch.
26  reviewer: string | null
27  // The spec workflow's task notification arrived with status completed.
28  specEnded: boolean
29  // Prompts the person sent since the spec became ready with no launch.
30  idlePrompts: number
31}
32
33export type OrcaTone = 'done' | 'live' | 'wait' | 'blocked' | 'cut' | 'queued'
34
35export type OrcaRow = { id: string; title: string; mark: string; status: string; tone: OrcaTone }
36
37// What the band draws; null draws nothing.
38export type OrcaBand = {
39  slug: string
40  // `spec`, `build`, `integrate` or `integrate · <stage>`, `wrapping up`.
41  phase: string
42  // The spec state (drafting, reviewing, revising, ready), or `<merged>/<total> merged`.
43  detail: string
44  rows: OrcaRow[]
45  // Whether item stages came from the workflow journal or from run files alone.
46  source: 'journal' | 'files'
47}
48
49declare module 'claude-code' {
50  interface PluginState {
51    orca: {
52      run: OrcaRun | null
53      band: OrcaBand | null
54      // The session's permission mode as the classic hook events last
55      // reported it; null until one has. The verb hook runs verbs only while
56      // it is bypassPermissions.
57      permissionMode: string | null
58    }
59  }
60}
61