Autonomous multi-agent development: init/doctor once, then orca:feature from idea to integration branch, and orca:review to walk the deliverable in your own…

<img src=".github/logo.png" width="280" alt="Orca logo">
<h1 align="center">orca</h1>
A Claude Code plugin for autonomous, multi-agent development. Set the repository up once, then /orca:feature takes a feature from idea to a committed integration branch: an adversarial interview captures intent as a durable brief; a deterministic workflow plans, implements, independently reviews, fixes, commits, merges, and verifies. The user interacts exactly once after the interview (twice, by opt-in).
/orca:feature <idea> # triage → interview/brief → spec → work loop → report
# (interactive until one confirmation, autonomous after)
/orca:review # review a deliverable in your editor; your comments round-trip
# (addressed on consent, resolutions rendered inline next review)
/orca:pr # land a delivered run through a GitHub pull request: the report
# becomes an external-facing description, previewed, opened as a draft
/orca:retry # finish a finished run's unmet items in the same run: audit against
# git, resolve the blocked decisions with you, relaunch the work loop
/orca:followup # turn a finished run's optional follow-ups into the next brief
# for /orca:feature (new intent, not recovery)
/orca:iterate # new work on a delivered, still-unlanded feature branch, whatever
# the size — same run, same branch, specced and reviewed like a run
/orca:status # read-only dashboard: .orca state joined with git ground truth,
# grouped by next action, each state naming its skill
/orca:archive # retire finished runs whose branches provably landed, so triage
# stops carrying the whole history, removing their branches + worktrees
/orca:init # one-time repository layout setup (interactive, consent per step)
/orca:doctor # machine tooling + repo readiness setup (interactive, consent per step)
/orca:config # optional per-repo reviewer & model/effort tuning
The deliverable of a run is a feature/<slug> branch on an integration worktree, which you review with /orca:review and land yourself with git merge --no-ff <branch>, or publish as a GitHub pull request with /orca:pr. The runs never touch your own worktree, and no commit they produce mentions Claude, AI, agents, or orca.
Three design choices carry the whole system:
Double isolation. Every stage — spec, plan, implement, review, fix, commit, merge, integrate — runs in a dedicated subagent with its own context window, and every unit of parallel work (a feature's work item) gets its own git worktree off a shared bare repository. Heavy context — codebase exploration, diffs, test output — lives and dies inside subagents; the main conversation only reads artifact files and the workflow's structured result. Parallel items can never corrupt each other's files: overlap surfaces as an explicit merge conflict, resolved by a merge agent holding both items' plans.
Independent review. Claude implements; a separate reviewer attacks the result before anything is committed. The featured default is cross-model: Codex reviews, run non-interactively by the plugin's own CLI (orca.sh codex → codex exec) against the global codex binary, driven adversarially over each item's diff by a dedicated courier agent — an independent second opinion from a different model family, one that does not share the implementer's blind spots. The review explicitly attacks the tests, since the same model family wrote the code and the tests. With reviewer=claude — the detected default wherever codex isn't installed, or an explicit pin via /orca:config — a dedicated Claude review agent performs the same adversarial review itself: it keeps fresh-context independence (a separate agent, only the artifacts and the diff), but it is same-model, so it may share the implementer's blind spots. That trade-off is stated wherever the choice is made; cross-model stays the stronger design.
A deterministic work loop. The long autonomous middle of a run is one bundled script executed through Claude Code's Workflow tool, not conversational orchestration. Scheduling, retry bounds, review throttling, merge serialization, and the commit-attribution check are code, so the guarantees are structural rather than model discipline that decays over a long context. Judgment calls (plan reconciliation, escalation, review verdicts) stay in agents, but as schema'd calls whose reasons land in the run's artifacts. Every agent call is journaled, so an interrupted run resumes where it stopped. The workflow sandbox has no shell, so its git plumbing has to leave the sandbox to run: orca's verb hook, a function hook the plugin ships, runs each orca.sh verb itself and hands back its exact output — no model in the path. That hook is the only way out: a run requires function hooks (mods) loaded and a bypassPermissions session, and without either it fails at its first verb with a message naming the cause — there is no fallback.
State lives in files, never in conversation memory: the brief, the spec with its work breakdown and decision log, one plan per item, the raw review findings per round, and a final report — all under a per-run directory in .orca/.
| Requirement | Detail |
|---|---|
| Claude Code | A harness with the Workflow tool (the work loop runs through it). /orca:feature checks and refuses without it. |
| Codex CLI | Required only when the reviewer is codex (the default wherever it is installed): the global codex binary on PATH, version ≥ 0.142.5, authenticated via codex login, and still carrying the codex exec flags orca drives (--output-schema, --output-last-message, --sandbox, --cd) — the pre-flight asks the binary directly, because a version number does not tell you the interface has not moved. Never install codex via npm — use brew install codex or the official release binaries. With reviewer=claude, the codex rows don't apply. |
| Repository layout | Bare-repo-with-worktrees (.bare/ + peer worktrees). /orca:init sets this up, including converting an existing conventional checkout in place. |
| git | ≥ 2.31 (rev-parse --path-format — every orca script's repository resolution; older gits get a typed OLD_GIT failure naming the upgrade, never a misdiagnosis); ≥ 2.42 for worktree add --orphan when /orca:init creates a brand-new repository. |
BASH_MAX_TIMEOUT_MS | Codex-only, like the Codex CLI row: set to 1200000 (~20 min) in a Claude Code settings env block, so Codex reviews are not killed at the default Bash tool timeout. A plugin cannot ship session env, so /orca:doctor writes it for you. |
| Permission mode | Runs need bypassPermissions for the session — see Permissions and autonomy. The verb hook enforces it: in any other mode it refuses, and the run fails at its first verb. |
| Function hooks (mods) | Runs need orca's hooks module loaded in the session: a Claude Code build with function hooks (2.1.290+ was probed), the feature enabled (CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 in a settings env block where the server-side flag is off), not started with --bare, and no disableAllHooks in any settings file, managed ones included. The work loop reaches orca.sh only through the verb hook; without it a run fails at its first verb, naming the fix. No shell check can see the hook, so /orca:doctor and the pre-flight state this requirement rather than test it. |
Everything else — the fifteen stage agents and the CLI that drives codex — ships inside the plugin itself; there is nothing to install per repository beyond the layout.
This repository hosts its own plugin marketplace (.claude-plugin/marketplace.json), so installing is two commands inside Claude Code:
/plugin marketplace add miguelbacalhau/orca
/plugin install orca@orca
The install persists across sessions. Updates are manual by default for third-party marketplaces — pull new versions with /plugin marketplace update orca, or toggle auto-update in the /plugin → Marketplaces UI. Removal is symmetric: /plugin marketplace remove orca uninstalls the marketplace and the plugin in one step.
Teams can make the install declarative instead: commit this to the project's .claude/settings.json, and Claude Code prompts each teammate to install the marketplace and pre-enables the plugin when they trust the workspace:
{
"extraKnownMarketplaces": {
"orca": { "source": { "source": "github", "repo": "miguelbacalhau/orca" } }
},
"enabledPlugins": { "orca@orca": true }
}
For local development on the plugin itself, load a checkout directly for a single session:
claude --plugin-dir /path/to/this/repo
Plugins load at session start, so after installing or enabling one, start a fresh session before running. There is nothing session-scoped left to verify for the reviewer: codex is reached by running the binary, and /orca:feature's pre-flight checks it from a shell.
# 1a. One-time per repo: the bare-repo-with-worktrees layout. Interactive,
# consent per step. Converting an existing checkout preserves untracked
# files (.env, caches) but changes every path — see /orca:init.
/orca:init
# 1b. One-time per machine, only if the pre-flight flags it: Codex CLI
# install/auth guidance and the review timeout settings write. Skippable
# outright if you'll run with the Claude reviewer.
/orca:doctor
# 2. The feature. An interview captures intent as a brief file — outcome,
# features, non-goals, constraints, doubt rule, as many rounds as the
# idea needs — then, after ONE confirmation, the run drives autonomously
# to a final report and a feature/<slug> branch. Decline the run prompt
# to leave the brief queued for a later /orca:feature instead.
/orca:feature add rate limiting to the public API
# 3. Walk the deliverable's diff in your own editor (nvim, via a tmux
# window in your session), then land it from your own worktree —
# or publish it as a GitHub pull request instead with /orca:pr.
/orca:review
git merge --no-ff feature/<slug>
/orca:pr # the PR path: report → description, draft PR via gh
# 4. If the report left blocked items: resolve their recorded decisions
# and finish them inside the same run, on the same branch.
/orca:retry
# 5. If it left optional follow-ups you want: select them into the next
# run's brief.
/orca:followup
Briefs can be queued ahead of time — everything at the top level of .orca/feat-briefs/ is ready and unconsumed; a later /orca:feature finds them, offers to run one, and a run consumes exactly one.
/orca:feature <idea>The one feature verb — a dispatcher over .orca state. Triage looks at what is on disk and offers the first match, never forcing it:
spec.md records a workflow runId with no report.md beside it. Offers to resume from the journal; completed work replays instantly, only in-flight and remaining work runs live..orca/feat-briefs/. Offers to run one, or to interview a new idea instead.The interview sharpens a rough idea into a durable brief at .orca/feat-briefs/<timestamp>-<slug>.md. Before asking anything substantive it researches the subsystems the idea touches through the dedicated orca:research agent (the report stays behind the scenes; the main context never explores deeply), then interviews from that picture — opening with a reflection of the idea against the system as it exists, so tensions between the two surface in conversation, where changing course is cheap. The brief is the entire intent the run acts on — the run asks nothing beyond one confirmation — so the interview is deliberately adversarial about scope: it pushes back, hunts for unstated non-goals, and makes you resolve ambiguities now rather than leaving them for an autonomous run to guess at. Direction decisions it settles land in the brief with their rationale and bind the spec like constraints; the decomposition itself still belongs to the spec stage.
A brief records: outcome, features, non-goals, inputs & outputs, constraints, plus two run-controlling choices:
prefer-smaller-scope, the default) or the more complete one (prefer-complete)?review-once), or run straight through (straight-through, the default)?The brief is what and why, never how: no work breakdown, no interfaces, no file ownership — those come from the run's spec stage, grounded in real codebase exploration. Location is status: top-level briefs are ready; park unfinished ones in .orca/feat-briefs/drafts/. One brief is one run's scope — two independent efforts get two briefs, and declining the run prompt leaves a brief queued for any later /orca:feature.
The run. See Anatomy of a run for the full lifecycle. Interaction surface, in total:
After that, nothing asks you anything. Ambiguities resolve against the spec and the doubt rule; what cannot be resolved that way becomes a blocked item in the final report, with the options you must choose between recorded. The outcome lands in report.md.
/orca:review [branch]The human half of review — the runs' adversarial review stage is automated and lives inside them; this opens the finished deliverable in your editor before you land it. It discovers unmerged feature/<slug> deliverable branches (one → opens it; several → asks; a gone integration worktree → offers to add it back) and opens the deliverable's integration worktree in your editor, running an orca review session — file list, native side-by-side merge-base diffs, your LSP and colors. Two editors are supported:
:OrcaReview. Quit nvim (:qa) and the window closes, dropping you back where you invoked it, ready to git merge --no-ff. tmux is the mechanism, not a convenience — the harness has no TTY to run an interactive editor in, so a new window in the session $TMUX points at is the one way a skill puts a live editor in front of you.code --open-url and the extension's URI handler, which routes it to a window on the worktree (opening one if needed). A detached GUI launch — no tmux involved; close the review window when done.Without tmux (on the nvim path), or without either companion installed (/orca:doctor prescribes both), it prints the exact command to run instead (cd <worktree> && nvim "+OrcaReview", cd <worktree> && code . plus the palette command, or a git difftool fallback) and stops.
Review comments round-trip. Comments you leave in the review with :OrcaComment persist to .orca/review-notes/<key>.json (<key> is the deliverable branch, sanitized — orca.nvim's rule). On the nvim path the skill leaves a background waiter on the tmux window: quit nvim and it wakes, reads the file, and lays out every open comment alongside what it intends to do about it — fix (with a one-line approach), answer, decline with reasoning, or defer as too big for review — then asks before doing anything (window death means "nvim closed," not "review finished"; you can address now, resume the review with comments intact, or leave them for later). Consent is to that stated plan, not a blank check: correct a misread interpretation right there, before any code changes. On consent, a dedicated agent converts each open comment into a fix or an answer in the integration worktree — change requests get fixed with tests and committed under the same no-attribution rules as every run commit; questions get answered from the code — and writes status/resolution back into the same file, where the next :OrcaReview renders them inline under their anchors. Each round is also appended to the run's report.md under ## Review rounds — commit, per-comment resolutions, verification result — so /orca:pr describes the branch as it now stands; a round whose verification fails downgrades the report's deliverable state to unverified. Editing a resolution's comment there re-opens it for the next round; that loop is the convergence mechanism, so there is deliberately no machine re-review after addressing. The file is the source of truth and the window-exit only the wake signal: comments left from the vscode tier, a print-only session, or a review whose session died are picked up by the next /orca:review, which surfaces "N unaddressed comments" in triage. The workflow is one-writer-at-a-time by design, and both sides fail loud on a notes schema version they don't speak (/orca:doctor's nvim check diagnoses the skew).
Two config keys govern it, set via /orca:config: editor (nvim|vscode|none) and terminal (tmux|none, nvim path only), with the same semantics as reviewer — absent detects (nvim first, then vscode), a pin fails loudly instead of silently falling back, none opts out to the printed command.
/orca:pr [run]The PR path for landing — the report template's Landing section ends at a local git merge --no-ff, which is wrong for a repo that lands work through GitHub pull requests. This skill takes a delivered-but-unlanded run (a finished run whose deliverable branch exists and is unmerged, found via triage snapshot — newest by default, or the one the argument names) and publishes it with the gh CLI. It is report-driven by design: every fact in the PR body traces to a section of the run's report.md, nothing is re-derived from the diff — so a branch no run produced is out of scope (plain gh pr create already covers it).
Two guards gate it, because even a draft PR asserts "the work is finished": the report must say Deliverable state: verified, and ## Blocked must be "None" — anything else is refused with a pointer at the owning skill (/orca:feature's resume, /orca:retry). Draft status is no fallback there: it is the unconditional default, not an escape hatch, and it still pushes an unverified branch under a description claiming work the run never verified.
New PRs are always drafts, because two readinesses are in play and the skill can only vouch for one. The guards settle run readiness — the run finished and verified its own work. Social readiness is yours: at publish time no human has read the diff, and a ready PR announces the opposite, auto-requesting CODEOWNERS reviewers and releasing whatever CI and merge automation keys on non-draft. So it goes up as a draft and you mark it ready on GitHub after your own pass — no flag and no config key, since the wrong default this way costs one click and the other way costs a notification you cannot recall. Draft is a creation-time choice only: refreshing an existing PR never moves its status in either direction, so a PR you already promoted stays promoted.
Composition is translation into a fixed shape. The reader has never heard of orca, so run vocabulary is dropped wholesale — item IDs, hashes, counts, run-dir paths, follow-ups, every /orca:* pointer — and so is the run's item-by-item shape, which is what made descriptions long: work items are the run's unit of parallelism, not the reader's unit of understanding, so bullets group by outcome. The body fills the repo's own pull_request_template.md when it has one; otherwise it follows a three-section skeleton — a plain-language paragraph on what was wrong or missing (from the run's brief.md) and what the branch does about it, What changed, Testing, plus How it works and Notes only when they earn their place — held to 200–350 words. The no-attribution rule that governs run commits extends verbatim to the PR title and body, enforced by the same deterministic marker scan — explicitly overriding the harness habit of appending a "Generated with Claude Code" footer. You see the exact title, body, base, and head before anything happens; one confirmation gates both the push and the PR creation (re-running an existing PR refreshes it). It never merges, never edits the report, never touches the integration worktree.
/orca:retry [run]Finishes a finished run's unmet work items inside the same run — no new run directory, no new spec, no brief. Recovery never creates a run; only new intent does. It picks a finished feature run with leftovers (newest by default, or the one the argument names), audits it through the same read-only agent /orca:followup uses — reconciling the report's claims against the spec's work breakdown and git ground truth — then resolves every blocked item's recorded decision with you. That interview is the entry fee, on principle: the run's escalation agents already spent its bounded machine retries deciding each block could not be resolved within the spec, so a retry without new human input would reproduce it. You can also resurrect items the run cut autonomously (a machine-made scope reduction you never directly approved — overruling it restores scope the original brief contained), and a resolution may authorize the retried item to amend code an earlier item already merged.
Your resolutions are appended to the run's own spec.md as binding Decisions bullets; the superseded plans, review findings, and report are archived in place (plans/<ID>.round<N>.md, reviews/prev<N>/, report.round<N>.md) so the failure evidence survives the next round; and the work loop relaunches over only the unmet items, on the same integration branch — each item carrying a retry note that points at its recorded failure evidence, and each surviving item branch (salvaged WIP commit included) resumed rather than restarted. The rewritten report keeps the whole-run picture, with a Prior rounds section naming the archived reports. An interrupted run (no report yet) is redirected to /orca:feature's resume; a clean run has nothing to retry — its follow-ups are /orca:followup
hooks/index.ts 14 lines1// orca's hooks module: the run band (./run-band.tsx), drawn above the prompt,
2// and the verb hook (./relay.ts), which answers the work loop's orca.sh verb
3// calls without a model. hooks.json names one module, so this one registers both.
4
5import type { Register } from 'claude-code'
6
7import { register as relay } from './relay'
8import { register as runBand } from './run-band'
9
10export const register: Register = (on, options) => {
11 runBand(on, options)
12 relay(on, options)
13}
14hooks/relay.ts 177 lines1// The verb hook: answers the work loop's orca.sh verb calls without a model.
2//
3// The work loop (scripts/work-loop.workflow.js) asks for a verb with
4// `agent('', { model: 'orca:{"v":1,"verb":…,"args":[…]}' })`. That request
5// reaches `turn.step` before any API call; this hook answers it itself by
6// running `bash <this plugin>/scripts/orca.sh <verb> <args…>` and returning
7// the result as the agent's text. It never calls `next` for an `orca:` model:
8// that would send the request (and the arguments in its model name) to the API.
9//
10// What it will run, and when:
11// - only an `orca:` request from a subagent (a workflow agent), protocol
12// version 1, whose verb (and subcommand) is one the work loop calls;
13// - only while the session's permission mode is known to be
14// bypassPermissions. The mode is read off the classic hook events that
15// carry it and kept in $.state; no mode seen yet means refuse.
16// - always through bash + orca.sh, never git directly: `$.process.run` runs
17// git with repo hooks off, a script that bash starts does not.
18//
19// The answer is `b64:` + base64(UTF-8 JSON): `{ v, rc, stdout, stderr }`, or
20// `{ v, refused }` when a guard says no, or `{ v, error, timedOut? }` when the
21// command could not be run to completion. Encoded whole because the harness
22// rewrites subagent results that look instruction-shaped (a verb's output
23// naming the session's flags, say); base64 passes untouched.
24//
25// This hook is the work loop's only way to orca.sh: there is no fallback.
26// Where it is not loaded, an `orca:` request goes to the API, which has no
27// such model, and the work loop fails loudly at that call (its first is
28// `triage claim`), naming the cause.
29
30import { atom, read, update } from 'claude-code'
31import type { EngineInterface, Register } from 'claude-code'
32
33export const RELAY_VERSION = 1
34export const MODEL_PREFIX = 'orca:'
35// The longest `$.process.run` allows. setup's own cap (540 s, verbs/setup.sh)
36// sits under it, so a slow worktree setup ends in setup's own soft timeout;
37// a verb that still runs past ten minutes fails here, as a timeout naming it.
38export const RUN_TIMEOUT_MS = 600_000
39
40// Exactly the verbs work-loop.workflow.js calls through orca(), with the
41// subcommands it uses where a verb has several; null means any arguments.
42export const ALLOWED: Readonly<Record<string, readonly string[] | null>> = {
43 'self-test': null,
44 'triage': ['claim', 'release'],
45 'spec': ['mark', 'changed'],
46 'secrets': ['place', 'remove'],
47 'archive': ['attempt'],
48 'notes': ['write'],
49 'worktree-item': null,
50 'provision': null,
51 'commit-verify': null,
52 'merge-prepare': null,
53 'merge-finalize': null,
54 'commit-prepare': null,
55 'salvage': null,
56}
57
58export type RelayRequest = { v: number; verb: string; args: string[] }
59export type RelayAnswer =
60 | { v: number; rc: number; stdout: string; stderr: string }
61 | { v: number; refused: string }
62 | { v: number; error: string; timedOut?: true }
63
64const modeAtom = atom({ plugin: 'orca', key: 'permissionMode' } as const, null)
65
66// Parses and checks the request the model name carries. A string answer is
67// the refusal reason.
68export function parseRequest(model: string): RelayRequest | string {
69 if (!model.startsWith(MODEL_PREFIX)) return 'not an orca: request'
70 let req: unknown
71 try { req = JSON.parse(model.slice(MODEL_PREFIX.length)) } catch { return 'request is not JSON' }
72 if (typeof req !== 'object' || req === null || Array.isArray(req)) return 'request is not an object'
73 const { v, verb, args } = req as Record<string, unknown>
74 if (v !== RELAY_VERSION) return `unknown protocol version ${JSON.stringify(v)} (this hook speaks ${RELAY_VERSION})`
75 if (typeof verb !== 'string' || !Object.prototype.hasOwnProperty.call(ALLOWED, verb))
76 return `verb ${JSON.stringify(verb)} is not on the work loop's allow-list`
77 if (!Array.isArray(args) || args.some(a => typeof a !== 'string')) return 'args must be an array of strings'
78 const subcommands = ALLOWED[verb]
79 if (subcommands && !subcommands.includes(args[0] as string))
80 return `${verb} ${JSON.stringify(args[0] ?? null)} is not on the work loop's allow-list (${subcommands.join(', ')})`
81 return { v, verb, args: args as string[] }
82}
83
84export function argvFor(root: string, req: RelayRequest): string[] {
85 return ['bash', `${root.replace(/\/+$/, '')}/scripts/orca.sh`, req.verb, ...req.args]
86}
87
88export function encodeAnswer(answer: RelayAnswer): string {
89 const bytes = new TextEncoder().encode(JSON.stringify(answer))
90 let bin = ''
91 for (let i = 0; i < bytes.length; i += 0x8000)
92 bin += String.fromCharCode(...bytes.subarray(i, i + 0x8000))
93 return `b64:${btoa(bin)}`
94}
95
96// Main-thread events set the mode either way. An event from inside a
97// subagent may only close the gate: a subagent's own mode is not the
98// session's, and must never be what opens it.
99export function nextMode(current: string | null, seen: unknown, fromSubagent: boolean): string | null {
100 if (typeof seen !== 'string' || !seen) return current
101 if (!fromSubagent) return seen
102 return seen === 'bypassPermissions' ? current : seen
103}
104
105async function noteMode($: EngineInterface, e: { permission_mode?: string; agent_id?: string }) {
106 const current = await read($, modeAtom)
107 const mode = nextMode(current, e.permission_mode, e.agent_id !== undefined)
108 if (mode !== current) await update($, modeAtom, () => mode)
109}
110
111// A rejected $.process.run: the command could not start, or was still
112// running at the cap and was killed. The two are told apart by the time it
113// took, or by the message (2.1.291 rejects a timeout with "<plugin>:
114// $.process.run(<argv0>) aborted: still running after <ms>ms"), so the work
115// loop can name a timeout.
116export function runFailure(verb: string, message: string, elapsedMs: number): RelayAnswer {
117 if (elapsedMs >= RUN_TIMEOUT_MS - 1000 || /still running after|timed? ?out/i.test(message))
118 return { v: RELAY_VERSION, error: `orca.sh ${verb} was still running after ${RUN_TIMEOUT_MS / 60_000} minutes, the longest $.process.run allows, and was stopped: ${message}`, timedOut: true }
119 return { v: RELAY_VERSION, error: `orca.sh ${verb} did not run to completion: ${message}` }
120}
121
122async function answer($: EngineInterface, model: string, agentId: string | undefined): Promise<RelayAnswer> {
123 if (!agentId) return { v: RELAY_VERSION, refused: 'orca: requests are answered for workflow agents only' }
124 const req = parseRequest(model)
125 if (typeof req === 'string') return { v: RELAY_VERSION, refused: req }
126 const mode = await read($, modeAtom)
127 if (mode !== 'bypassPermissions')
128 return { v: RELAY_VERSION, refused: mode === null
129 ? 'permission mode not known yet (none seen since this session or hook loaded); orca runs verbs here only in bypassPermissions sessions'
130 : `permission mode is ${mode}; orca runs verbs here only in bypassPermissions sessions` }
131 let run
132 const started = Date.now()
133 try { run = await $.process.run(argvFor($.plugin.root, req), { timeoutMs: RUN_TIMEOUT_MS }) }
134 catch (err) { return runFailure(req.verb, String((err as Error)?.message ?? err), Date.now() - started) }
135 if (run.isStdoutTruncated) return { v: RELAY_VERSION, error: `orca.sh ${req.verb} wrote more than 4 MiB to stdout` }
136 return { v: RELAY_VERSION, rc: run.exitCode, stdout: run.stdout, stderr: run.stderr }
137}
138
139export const register: Register = on => {
140 // The events Phase 3 saw carry permission_mode (2.1.290/291); SubagentStart
141 // and SessionStart do not.
142 on('classic.UserPromptSubmit', async ($, e, next) => {
143 try { await noteMode($, e) } catch { /* the gate stays as it was */ }
144 return next(e)
145 }).catch(($, e, next) => next(e))
146 on('classic.PostToolUse', async ($, e, next) => {
147 try { await noteMode($, e) } catch { /* the gate stays as it was */ }
148 return next(e)
149 }).catch(($, e, next) => next(e))
150 on('classic.SubagentStop', async ($, e, next) => {
151 try { await noteMode($, e) } catch { /* the gate stays as it was */ }
152 return next(e)
153 }).catch(($, e, next) => next(e))
154 on('classic.Stop', async ($, e, next) => {
155 try { await noteMode($, e) } catch { /* the gate stays as it was */ }
156 return next(e)
157 }).catch(($, e, next) => next(e))
158
159 on('turn.step', async function* ($, e, next) {
160 if (!e.model.startsWith(MODEL_PREFIX)) return yield* next(e)
161 let reply: RelayAnswer
162 try { reply = await answer($, e.model, e.agentId) }
163 catch (err) { reply = { v: RELAY_VERSION, error: `orca hook failed: ${String((err as Error)?.message ?? err)}` } }
164 const text = encodeAnswer(reply)
165 yield { kind: 'text', index: 0, text }
166 yield { kind: 'stop', stopReason: 'end_turn', usage: null }
167 return { turnId: e.turnId, index: e.index, answer: text, toolUses: [], stopReason: 'end_turn', usage: null }
168 }).catch(async function* ($, e, next) {
169 // An orca: request never goes to the API, even when this hook failed.
170 if (next.called || !e.model.startsWith(MODEL_PREFIX)) return yield* next(e)
171 const text = encodeAnswer({ v: RELAY_VERSION, error: `orca hook failed: ${next.error.message ?? next.error.kind}` })
172 yield { kind: 'text', index: 0, text }
173 yield { kind: 'stop', stopReason: 'end_turn', usage: null }
174 return { turnId: e.turnId, index: e.index, answer: text, toolUses: [], stopReason: 'end_turn', usage: null }
175 })
176}
177hooks/run-band.tsx 281 lines1// The run band: a live view of the orca run this session launched, drawn
2// above the prompt. A read-only observer — it watches this session's own
3// Workflow launches of spec.workflow.js / work-loop.workflow.js, then reads
4// the run directory and the workflow's journal on a short tick while that run
5// is tracked. It writes nothing anywhere but its own $.state, and draws
6// nothing where the AbovePrompt band is not raised (VS Code, `claude -p`).
7//
8// The derivation itself is pure and lives in ./derive.ts.
9
10import { atom, read, update } from 'claude-code'
11import type { EngineInterface, Register, Timer } from 'claude-code'
12
13import type { OrcaBand, OrcaRun, OrcaTone } from '../types'
14import { derive, headerRule, itemsFromArgs, layoutRows, parseTaskNotifications, projectDirName, slugOf } from './derive'
15import type { JournalRead, RuleTone, RunFiles } from './derive'
16
17const runAtom = atom({ plugin: 'orca', key: 'run' } as const, null)
18const bandAtom = atom({ plugin: 'orca', key: 'band' } as const, null)
19
20const TICK_MS = 1500
21// $.fs.read refuses anything larger.
22const MAX_READ = 4 * 1024 * 1024
23// A spec that finished and then sat through this many of the person's own
24// prompts with no work-loop launch was abandoned at the checkpoint.
25const SPEC_IDLE_PROMPTS = 3
26
27// Module variables start over on a hot reload; the run itself is in $.state.
28// A timer keeps the `$` of the hook that started it, which stays usable after
29// that dispatch ends (checked under `claude -p` on 2.1.290).
30let ticker: Timer | null = null
31let ticking = false
32const textCache = new Map<string, { size: number; mtimeMs: number; text: string }>()
33
34function kindOf(scriptPath: unknown): OrcaRun['kind'] | null {
35 if (typeof scriptPath !== 'string') return null
36 if (/(^|\/)spec\.workflow\.js$/.test(scriptPath)) return 'spec'
37 if (/(^|\/)work-loop\.workflow\.js$/.test(scriptPath)) return 'build'
38 return null
39}
40
41function record(v: unknown): Record<string, unknown> | null {
42 if (typeof v === 'string') { try { v = JSON.parse(v) } catch { return null } }
43 return typeof v === 'object' && v !== null && !Array.isArray(v) ? (v as Record<string, unknown>) : null
44}
45
46// Reads a file through a size+mtime cache: an unchanged journal is not read
47// again every tick. 'too-large' past what $.fs.read takes.
48async function readCached($: EngineInterface, path: string): Promise<JournalRead> {
49 let stat
50 try { stat = await $.fs.stat(path) } catch { return { status: 'missing' } }
51 if (stat.size > MAX_READ) return { status: 'too-large' }
52 const hit = textCache.get(path)
53 if (hit && hit.size === stat.size && hit.mtimeMs === stat.mtimeMs) return { status: 'ok', text: hit.text }
54 try {
55 const text = await $.fs.read(path)
56 textCache.set(path, { size: stat.size, mtimeMs: stat.mtimeMs, text })
57 return { status: 'ok', text }
58 } catch { return { status: 'unreadable' } }
59}
60
61async function listNames($: EngineInterface, dir: string): Promise<{ name: string; mtimeMs: number }[]> {
62 try { return (await $.fs.list(dir)).map(f => ({ name: f.name, mtimeMs: f.mtimeMs })) } catch { return [] }
63}
64
65async function readRunFiles($: EngineInterface, run: OrcaRun): Promise<RunFiles> {
66 const files: RunFiles = { runDirExists: true, spec: null, reviews: [], plans: [], merged: null, reportMtimeMs: null, agents: [] }
67 let entries
68 try { entries = await $.fs.list(run.runDir) } catch {
69 files.runDirExists = await $.fs.exists(run.runDir).catch(() => true)
70 return files
71 }
72 const named = new Map(entries.map(f => [f.name, f]))
73 const report = named.get('report.md')
74 files.reportMtimeMs = report ? report.mtimeMs : null
75 if (run.amendPath) {
76 const st = await $.fs.stat(run.amendPath).catch(() => null)
77 files.spec = st ? { mtimeMs: st.mtimeMs, text: null } : null
78 } else {
79 const spec = named.get('spec.md')
80 // spec.md's text is needed only when the launch's own item list was unreadable.
81 const needText = run.kind === 'build' && run.items.length === 0
82 const text = spec && needText ? await readCached($, `${run.runDir}/spec.md`) : null
83 files.spec = spec ? { mtimeMs: spec.mtimeMs, text: text && text.status === 'ok' ? text.text : null } : null
84 }
85 if (named.has('reviews')) files.reviews = await listNames($, `${run.runDir}/reviews`)
86 if (named.has('plans')) files.plans = (await listNames($, `${run.runDir}/plans`)).map(f => f.name)
87 if (run.kind === 'build' && named.has('merged.tsv')) {
88 const merged = await readCached($, `${run.runDir}/merged.tsv`)
89 files.merged = merged.status === 'ok' ? merged.text : null
90 }
91 if (run.kind === 'spec' && run.journal) files.agents = await listNames($, run.journal.replace(/\/[^/]*$/, ''))
92 return files
93}
94
95// Writes only a change: every write redraws the band.
96async function setBand($: EngineInterface, band: OrcaBand | null) {
97 if (JSON.stringify(await read($, bandAtom)) === JSON.stringify(band)) return
98 await update($, bandAtom, () => band)
99}
100
101async function forget($: EngineInterface) {
102 if (ticker) ticker.cancel()
103 ticker = null
104 await update($, runAtom, () => null)
105 await setBand($, null)
106}
107
108async function tick($: EngineInterface) {
109 if (ticking) return
110 ticking = true
111 try {
112 const run = await read($, runAtom)
113 if (!run) {
114 if (ticker) ticker.cancel()
115 ticker = null
116 await setBand($, null)
117 return
118 }
119 // Nothing draws the band here (a `claude -p` run, VS Code): read nothing.
120 const surfaces = await $.session.surfaces()
121 if (!surfaces.some(s => s === 'terminal' || s === 'desktop')) return
122 const files = await readRunFiles($, run)
123 const journal: JournalRead = run.journal ? await readCached($, run.journal) : { status: 'missing' }
124 const out = derive(run, files, journal, await $.clock.now())
125 // The run may have ended (a notification) or been replaced (a new launch)
126 // while this tick read files; a stale answer must not redraw the band.
127 if ((await read($, runAtom))?.runId !== run.runId) return
128 if (out.finished) await forget($)
129 else await setBand($, out.band)
130 } catch {
131 // Never crash the session over a status line; the next tick tries again.
132 } finally {
133 ticking = false
134 }
135}
136
137function follow($: EngineInterface) {
138 if (!ticker) ticker = $.clock.every(TICK_MS, () => { void tick($) })
139 void tick($)
140}
141
142// The launch result names the workflow's transcript directory, which holds
143// journal.jsonl. Without it, the same path is rebuilt from the session.
144async function journalPath($: EngineInterface, transcriptDir: unknown, runId: string): Promise<string | null> {
145 if (typeof transcriptDir === 'string' && transcriptDir) return `${transcriptDir.replace(/\/+$/, '')}/journal.jsonl`
146 const config = (await $.env.get('CLAUDE_CONFIG_DIR')) ?? `${(await $.env.get('HOME')) ?? ''}/.claude`
147 if (!config.startsWith('/')) return null
148 const cwd = await $.session.cwd()
149 const session = await $.session.id()
150 return `${config}/projects/${projectDirName(cwd)}/${session}/subagents/workflows/${runId}/journal.jsonl`
151}
152
153async function track($: EngineInterface, kind: OrcaRun['kind'], input: Record<string, unknown>, launched: unknown, since: number) {
154 const result = record(launched)
155 if (!result || result.status !== 'async_launched') return
156 const runId = result.runId
157 const taskId = result.taskId
158 if (typeof runId !== 'string' || typeof taskId !== 'string') return
159 const args = record(input.args)
160 const runDir = args && typeof args.runDir === 'string' ? args.runDir.replace(/\/+$/, '') : null
161 if (!args || !runDir) return
162 const run: OrcaRun = {
163 kind,
164 runDir,
165 slug: typeof args.slug === 'string' && args.slug ? args.slug : slugOf(runDir),
166 runId,
167 taskId,
168 journal: await journalPath($, result.transcriptDir, runId),
169 since,
170 items: kind === 'build' ? itemsFromArgs(args) : [],
171 amendPath: kind === 'spec' && typeof args.amendPath === 'string' ? args.amendPath : null,
172 reviewer: kind === 'spec' && typeof args.reviewer === 'string' && args.reviewer ? args.reviewer : null,
173 specEnded: false,
174 idlePrompts: 0,
175 }
176 await update($, runAtom, () => run)
177 follow($)
178}
179
180async function onPrompt($: EngineInterface, origin: string, text: string) {
181 const run = await read($, runAtom)
182 if (!run) return
183 if (origin === 'task-notification') {
184 const ended = parseTaskNotifications(text).find(n => n.taskId === run.taskId)
185 if (!ended) return
186 // The work loop ending ends the run's live phase. A spec that completed
187 // stays on screen as ready (the checkpoint, or the launch about to come);
188 // one that failed or was stopped does not.
189 if (run.kind === 'build' || ended.status !== 'completed') await forget($)
190 else await update($, runAtom, r => (r ? { ...r, specEnded: true } : r))
191 return
192 }
193 if ((origin === 'composer' || origin === 'bridge') && run.kind === 'spec') {
194 const band = await read($, bandAtom)
195 if (!run.specEnded && band?.detail !== 'ready') return
196 if (run.idlePrompts + 1 >= SPEC_IDLE_PROMPTS) await forget($)
197 else await update($, runAtom, r => (r ? { ...r, idlePrompts: r.idlePrompts + 1 } : r))
198 }
199}
200
201type Style = { color?: string; dimColor?: boolean; bold?: boolean }
202
203const TONE: Record<OrcaTone, Style> = {
204 done: { color: 'success' },
205 live: { color: 'suggestion' },
206 wait: { dimColor: true },
207 queued: { dimColor: true },
208 blocked: { color: 'error' },
209 cut: { dimColor: true },
210}
211
212// The top border: the rule and the name in Claude's accent, the facts set in it.
213const RULE: Record<RuleTone, Style> = {
214 rule: { color: 'claude', dimColor: true },
215 name: { color: 'claude', bold: true },
216 slug: { bold: true },
217 phase: { color: 'suggestion' },
218 detail: {},
219 sep: { dimColor: true },
220 note: { color: 'warning' },
221}
222
223export const register: Register = on => {
224 on('session.start', async ($, e, next) => {
225 // A hot reload drops the timer; the tracked run is still in $.state.
226 if (await read($, runAtom)) follow($)
227 return next(e)
228 })
229
230 on('tool.call', { tool: 'Workflow' }, async ($, e, next) => {
231 // Only the main conversation's own launches: the skills launch from there.
232 if (e.tool !== 'Workflow' || e.agentId !== undefined) return next(e)
233 const kind = kindOf(e.scriptPath)
234 if (!kind) return next(e)
235 const since = await $.clock.now()
236 const ran = await next(e)
237 if (!ran.isError && 'result' in ran) {
238 try { await track($, kind, { args: e.args }, ran.result, since) } catch { /* observe only */ }
239 }
240 return ran
241 }).catch(($, e, next) => next(e)) // an observer never stands in a call's way
242
243 on('prompt.submit', async ($, e, next) => {
244 try { await onPrompt($, e.origin.kind, e.text) } catch { /* observe only */ }
245 return next(e)
246 }).catch(($, e, next) => next(e))
247
248 on('session.end', async ($, e, next) => {
249 try { await forget($) } catch { /* observe only */ }
250 return next(e)
251 })
252
253 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
254 if (e.props.hasSurvey) return next(e)
255 const band = await read($, bandAtom)
256 if (!band || !(await read($, runAtom))) return next(e)
257 const { Box, Text } = $.ui.resolve(e)
258 const columns = Math.max(24, e.props.bodyColumns || 80)
259
260 // A blank row above the rule sets the band off from the transcript's last
261 // line, and one below it sets the rule off from the rows.
262 return (
263 <Box flexDirection="column" marginTop={1}>
264 <Box marginBottom={band.rows.length ? 1 : 0}>
265 {headerRule(band, columns).map(piece => (
266 <Text {...RULE[piece.tone]} wrap="truncate-end">{piece.text}</Text>
267 ))}
268 </Box>
269 {layoutRows(band.rows, columns).map(line => (
270 <Box>
271 <Text>{' '}</Text>
272 <Text bold>{line.id}</Text>
273 <Text>{` ${line.title} `}</Text>
274 <Text {...TONE[line.row.tone]} wrap="truncate-end">{line.tail}</Text>
275 </Box>
276 ))}
277 </Box>
278 )
279 })
280}
281hooks/derive.ts 567 lines1// Pure derivation for the run band: run files and the workflow journal in,
2// the rows the band draws out. No `$` here — run-band.tsx gathers the inputs
3// (and owns every engine call), so everything below is testable as plain data.
4
5import type { OrcaBand, OrcaItem, OrcaRow, OrcaRun } from '../types'
6
7// What run-band.tsx reads from the run directory on one tick.
8export type RunFiles = {
9 runDirExists: boolean
10 // spec.md (or the amendment file on an iterate spec launch): its mtime,
11 // and its text when the caller needed it.
12 spec: { mtimeMs: number; text: string | null } | null
13 // reviews/ entries, by name.
14 reviews: { name: string; mtimeMs: number }[]
15 // plans/ entry names.
16 plans: string[]
17 // merged.tsv's text, null when absent.
18 merged: string | null
19 // report.md's mtime, null when absent.
20 reportMtimeMs: number | null
21 // The workflow's transcript directory (the journal's own), by name: a spec
22 // launch times its stages from the agent files there.
23 agents: { name: string; mtimeMs: number }[]
24}
25
26export type JournalRead =
27 | { status: 'ok'; text: string }
28 | { status: 'missing' | 'unreadable' | 'too-large' }
29
30export type Derived = { finished: boolean; band: OrcaBand | null }
31
32// One agent() call the journal recorded.
33export type JournalStage = {
34 seq: number
35 label: string
36 agentId: string
37 running: boolean
38 failed: boolean
39 result: unknown
40}
41
42// ---------- journal ----------
43
44// One JSON object per line: {type:"started",key,agentId,label,phase} and
45// later {type:"result"|"failed",key,agentId,...}. A line that does not parse
46// (the torn tail of a write in progress, a partial flush) is skipped — the
47// next tick reads it whole. A started entry with no label predates labels and
48// says nothing the band can use.
49export function parseJournal(text: string): JournalStage[] {
50 const stages: JournalStage[] = []
51 const open = new Map<string, JournalStage[]>()
52 for (const line of text.split('\n')) {
53 if (!line.trim()) continue
54 let row: unknown
55 try { row = JSON.parse(line) } catch { continue }
56 if (!isRecord(row)) continue
57 const key = typeof row.key === 'string' ? row.key : null
58 if (row.type === 'started' && key !== null) {
59 const stage: JournalStage = {
60 seq: stages.length, label: typeof row.label === 'string' ? row.label : '',
61 agentId: typeof row.agentId === 'string' ? row.agentId : '',
62 running: true, failed: false, result: undefined,
63 }
64 stages.push(stage)
65 const list = open.get(key) ?? []
66 list.push(stage)
67 open.set(key, list)
68 } else if ((row.type === 'result' || row.type === 'failed') && key !== null) {
69 const list = open.get(key)
70 const stage = list?.shift()
71 if (!stage) continue
72 stage.running = false
73 stage.failed = row.type === 'failed'
74 stage.result = row.result
75 }
76 }
77 return stages.filter(s => s.label !== '')
78}
79
80// What a label in work-loop.workflow.js / spec.workflow.js means for the band.
81export type LabelMeaning =
82 | { kind: 'item'; ids: string[]; stage: string }
83 | { kind: 'escalate'; ids: string[] }
84 | { kind: 'salvage'; id: string }
85 | { kind: 'integration'; stage: string }
86 | { kind: 'spec'; state: 'drafting' | 'reviewing' | 'revising' }
87 | { kind: 'wrap-up' }
88 | null
89
90export function classifyLabel(label: string): LabelMeaning {
91 // Retries keep their stage: review:W1#0~retry.
92 const bare = label.replace(/~retry$/, '')
93 if (bare === 'spec') return { kind: 'spec', state: 'drafting' }
94 if (bare === 'spec-review') return { kind: 'spec', state: 'reviewing' }
95 if (bare === 'spec-revise') return { kind: 'spec', state: 'revising' }
96 if (bare === 'run-lease-release') return { kind: 'wrap-up' }
97 if (bare === 'integration-verify') return { kind: 'integration', stage: 'verify' }
98 if (bare === 'integration-review-clean') return { kind: 'integration', stage: 'review' }
99 if (bare === 'integration-status') return { kind: 'integration', stage: 'commit' }
100 if (bare === 'integration-fix-notes') return { kind: 'integration', stage: 'fix' }
101 const colon = bare.indexOf(':')
102 if (colon < 0) return null
103 // reconcile#2 and reconcile~serial are reconcile; plan:W3#2 is a replan.
104 const head = bare.slice(0, colon).replace(/[#~].*$/, '')
105 const tail = bare.slice(colon + 1)
106 const hash = tail.indexOf('#')
107 const subject = hash < 0 ? tail : tail.slice(0, hash)
108 const suffix = hash < 0 ? '' : tail.slice(hash + 1)
109 const round = /^\d+$/.test(suffix) ? Number(suffix) : null
110 const ids = subject.split('+').filter(Boolean)
111 if (!ids.length) return null
112
113 let stage: string | null
114 switch (head) {
115 case 'plan': stage = suffix ? 'replan' : 'plan'; break
116 case 'plan-archive':
117 case 'review-archive': stage = 'replan'; break
118 case 'reconcile': stage = 'reconcile'; break
119 case 'escalate': return { kind: 'escalate', ids }
120 case 'salvage': return { kind: 'salvage', id: ids[0] ?? subject }
121 case 'worktree': stage = 'worktree'; break
122 case 'provision': stage = 'provision'; break
123 case 'implement': stage = 'implement'; break
124 // Rounds are 0-based in labels: review:W3#0 is the first review.
125 case 'secrets-remove':
126 case 'review': stage = round === null ? 'review' : `review r${round + 1}`; break
127 case 'secrets-place':
128 case 'fix': stage = round === null ? 'fix' : `fix r${round}`; break
129 case 'secrets-restore': stage = 'review'; break
130 case 'commit':
131 case 'commit-verify': stage = 'commit'; break
132 case 'merge-reset':
133 case 'merge':
134 case 'merge-finalize': stage = 'merge'; break
135 default: stage = null // spec-hash, run-lease and other plumbing
136 }
137 if (stage === null) return null
138 if (subject === 'integration') return { kind: 'integration', stage }
139 return { kind: 'item', ids, stage }
140}
141
142// ---------- spec phase ----------
143
144export type SpecState = 'drafting' | 'reviewing' | 'revising' | 'ready'
145
146// Files decide by default: no fresh spec is drafting, a fresh spec with no
147// fresh review artifact is reviewing, both is ready (which covers the
148// checkpoint wait). "Fresh" is "written since this launch": a checkpoint
149// re-spawn or an iterate amend round runs over a directory that already holds
150// a spec and a review. A running spec-stage agent in the journal overrides
151// the files, and a completed spec workflow is ready.
152export function specState(run: OrcaRun, files: RunFiles, stages: JournalStage[] | null): SpecState {
153 if (run.specEnded) return 'ready'
154 const running = (stages ?? []).filter(s => s.running).map(s => classifyLabel(s.label))
155 const live = running.filter((m): m is Extract<LabelMeaning, { kind: 'spec' }> => m?.kind === 'spec').pop()
156 if (live) return live.state
157 if (!files.spec || files.spec.mtimeMs < run.since) return 'drafting'
158 const artifact = run.amendPath ? /^spec-amend-[a-z]+\.json$/ : /^spec-[a-z]+\.json$/
159 const reviewed = files.reviews.some(r => artifact.test(r.name) && r.mtimeMs >= run.since)
160 return reviewed ? 'ready' : 'reviewing'
161}
162
163// An agent's transcript directory holds agent-<id>.meta.json, written once at
164// its spawn, and agent-<id>.jsonl, appended to until it returns: the spawn and
165// the last word. Whole minutes, so the band redraws once a minute at most.
166export function elapsed(stage: JournalStage, agents: { name: string; mtimeMs: number }[], now: number): string {
167 const meta = agents.find(a => a.name === `agent-${stage.agentId}.meta.json`)
168 const end = stage.running ? now : agents.find(a => a.name === `agent-${stage.agentId}.jsonl`)?.mtimeMs
169 if (!stage.agentId || !meta || end === undefined) return ''
170 const minutes = Math.floor(Math.max(0, end - meta.mtimeMs) / 60_000)
171 return minutes < 1 ? '<1m' : `${minutes}m`
172}
173
174// One review attempt as the spec workflow judges it: a written artifact with
175// possible counts passes; anything else is a failed attempt and its reason.
176type ReviewVerdict = { ok: true; total: number; criticalHigh: number } | { ok: false; reason: string }
177
178function reviewVerdict(stage: JournalStage): ReviewVerdict {
179 let r = stage.result
180 if (typeof r === 'string') { try { r = JSON.parse(r) } catch { r = null } }
181 if (stage.failed || !isRecord(r)) return { ok: false, reason: 'review agent was skipped or died' }
182 if (r.written !== true) return { ok: false, reason: typeof r.reason === 'string' && r.reason ? r.reason : 'no artifact written' }
183 const total = typeof r.total === 'number' ? r.total : 0
184 const criticalHigh = typeof r.criticalHigh === 'number' ? r.criticalHigh : 0
185 if (criticalHigh > total) return { ok: false, reason: `impossible count (${criticalHigh} Critical/High of ${total})` }
186 return { ok: true, total, criticalHigh }
187}
188
189const plural = (n: number, word: string) => `${n} ${word}${n === 1 ? '' : 's'}`
190
191// The gate's three stages as rows, from the journal: spec.workflow.js labels
192// them spec, spec-review (spec-review~retry for its second attempt) and
193// spec-revise. A stage that returned nothing (skipped, died) is "died".
194function specRowsFromJournal(run: OrcaRun, stages: JournalStage[], agents: RunFiles['agents'], now: number): OrcaRow[] {
195 const bare = (s: JournalStage) => s.label.replace(/~retry$/, '')
196 const last = (label: string) => stages.filter(s => bare(s) === label).pop()
197 const died = (s: JournalStage) => s.failed || s.result === null || s.result === undefined
198 const time = (s: JournalStage) => { const t = elapsed(s, agents, now); return t ? ` · ${t}` : '' }
199 const [specRow, reviewRow, reviseRow] = specRowShells(run)
200
201 const spec = last('spec')
202 const specDied = spec !== undefined && !spec.running && died(spec)
203 if (!spec) specRow.set('●', 'starting', 'live')
204 else if (spec.running) specRow.set('●', `drafting${time(spec)}`, 'live')
205 else if (specDied) specRow.set('✕', 'died', 'blocked')
206 else specRow.set('✓', `done${time(spec)}`, 'done')
207
208 const review = last('spec-review')
209 const retry = review?.label.endsWith('~retry') ?? false
210 const verdict = review && !review.running ? reviewVerdict(review) : null
211 const failedOpen = verdict !== null && !verdict.ok && retry
212 if (specDied) reviewRow.set('–', 'skipped', 'cut')
213 else if (!review) reviewRow.set('○', 'queued', 'queued')
214 else if (review.running) reviewRow.set('●', `reviewing${retry ? ' (retry)' : ''}${time(review)}`, 'live')
215 else if (verdict?.ok) {
216 const found = verdict.total === 0 ? 'no findings'
217 : `${plural(verdict.total, 'finding')}, ${verdict.criticalHigh} Critical/High`
218 reviewRow.set('✓', `${found}${time(review)}`, 'done')
219 } else if (failedOpen) reviewRow.set('✕', `failed open: ${verdict?.ok === false ? verdict.reason : ''}`, 'blocked')
220 else reviewRow.set('●', 'retrying', 'live')
221
222 const revise = last('spec-revise')
223 const criticalHigh = verdict?.ok ? verdict.criticalHigh : 0
224 if (revise?.running) reviseRow.set('●', `revising${time(revise)}`, 'live')
225 else if (revise && died(revise)) reviseRow.set('✕', 'died: the findings stand', 'blocked')
226 else if (revise) reviseRow.set('✓', `revised${time(revise)}`, 'done')
227 else if (specDied || failedOpen) reviseRow.set('–', 'skipped', 'cut')
228 else if (verdict?.ok && criticalHigh === 0) reviseRow.set('–', 'not needed', 'cut')
229 else if (verdict?.ok) reviseRow.set('●', 'starting', 'live')
230 else reviseRow.set('○', 'on Critical/High only', 'queued')
231
232 return [specRow.row, reviewRow.row, reviseRow.row]
233}
234
235// Without a readable journal only the files speak: what was written, never
236// what a review found or whether a revise ran.
237function specRowsFromFiles(run: OrcaRun, state: SpecState): OrcaRow[] {
238 const [specRow, reviewRow, reviseRow] = specRowShells(run)
239 if (state === 'drafting') specRow.set('●', 'drafting', 'live')
240 else specRow.set('✓', 'done', 'done')
241 if (state === 'drafting') reviewRow.set('○', 'queued', 'queued')
242 else if (state === 'ready') reviewRow.set('✓', 'done', 'done')
243 else reviewRow.set('●', 'reviewing', 'live')
244 if (state === 'ready') reviseRow.set('◌', 'unknown', 'wait')
245 else reviseRow.set('○', 'on Critical/High only', 'queued')
246 return [specRow.row, reviewRow.row, reviseRow.row]
247}
248
249function specRowShells(run: OrcaRun) {
250 const shell = (id: string, title: string) => {
251 const row: OrcaRow = { id, title, mark: '○', status: '', tone: 'queued' }
252 return { row, set: (mark: string, status: string, tone: OrcaRow['tone']) => Object.assign(row, { mark, status, tone }) }
253 }
254 return [
255 shell('spec', run.amendPath ? 'author the amendment' : 'author spec.md from the brief'),
256 shell('review', `${run.reviewer ?? 'independent'} review against brief + codebase`),
257 shell('revise', 'fold Critical/High findings in'),
258 ] as const
259}
260
261export function deriveSpec(run: OrcaRun, files: RunFiles, journal: JournalRead, now: number): OrcaBand {
262 const stages = journal.status === 'ok' ? parseJournal(journal.text) : null
263 const state = specState(run, files, stages)
264 const rows = stages ? specRowsFromJournal(run, stages, files.agents, now) : specRowsFromFiles(run, state)
265 return { slug: run.slug, phase: 'spec', detail: state, rows, source: stages ? 'journal' : 'files' }
266}
267
268// ---------- build phase ----------
269
270// The last `**Workflow args:**` line of spec.md: the item list a resume
271// replays. Used only when the launch's own args could not be read.
272export function itemsFromSpec(text: string | null): OrcaItem[] {
273 if (!text) return []
274 const lines = text.split('\n').filter(l => l.startsWith('**Workflow args:** '))
275 const last = lines[lines.length - 1]
276 if (!last) return []
277 try { return itemsFromArgs(JSON.parse(last.slice('**Workflow args:** '.length))) } catch { return [] }
278}
279
280export function itemsFromArgs(args: unknown): OrcaItem[] {
281 let parsed = args
282 if (typeof parsed === 'string') { try { parsed = JSON.parse(parsed) } catch { return [] } }
283 if (!isRecord(parsed)) return []
284 let items = parsed.items
285 if (typeof items === 'string') { try { items = JSON.parse(items) } catch { return [] } }
286 if (!Array.isArray(items)) return []
287 return items.filter(isRecord).filter(i => typeof i.id === 'string').map(i => ({
288 id: i.id as string,
289 title: typeof i.title === 'string' ? i.title : '',
290 deps: Array.isArray(i.deps) ? i.deps.filter((d): d is string => typeof d === 'string') : [],
291 }))
292}
293
294export function mergedIds(text: string | null): Set<string> {
295 const ids = new Set<string>()
296 for (const line of (text ?? '').split('\n')) {
297 const [id, sha] = line.split('\t')
298 if (id && sha && sha.trim()) ids.add(id)
299 }
300 return ids
301}
302
303type ItemTrack = {
304 item: OrcaItem
305 blocked: string | null
306 cut: string | null
307 salvaged: boolean
308 live: string | null
309 last: string | null
310}
311
312// Escalation results follow ESCALATE ({replan, cut, blocked, addDeps}) for a
313// wave and BUILD_ESCALATE ({action, reason}) for one item mid-build.
314function applyEscalation(track: Map<string, ItemTrack>, items: OrcaItem[], ids: string[], result: unknown, merged: Set<string>) {
315 let r = result
316 if (typeof r === 'string') { try { r = JSON.parse(r) } catch { return } }
317 if (!isRecord(r)) return
318 if (typeof r.action === 'string') {
319 const t = track.get(ids[0] ?? '')
320 const reason = typeof r.reason === 'string' ? r.reason : ''
321 if (t && r.action === 'block') t.blocked = reason
322 if (t && r.action === 'cut') t.cut = reason
323 return
324 }
325 for (const b of Array.isArray(r.blocked) ? r.blocked : []) {
326 if (!isRecord(b) || typeof b.id !== 'string' || !ids.includes(b.id)) continue
327 const t = track.get(b.id)
328 if (t) t.blocked = typeof b.reason === 'string' ? b.reason : ''
329 }
330 for (const c of Array.isArray(r.cut) ? r.cut : []) {
331 if (!isRecord(c) || typeof c.id !== 'string' || !ids.includes(c.id)) continue
332 const t = track.get(c.id)
333 if (t) t.cut = typeof c.reason === 'string' ? c.reason : ''
334 }
335 // The scheduler applies addDeps to items not yet building, never as a
336 // cycle; an item already merged keeps its deps as they were.
337 for (const a of Array.isArray(r.addDeps) ? r.addDeps : []) {
338 if (!isRecord(a) || typeof a.id !== 'string' || !Array.isArray(a.dependsOn)) continue
339 const it = items.find(i => i.id === a.id)
340 if (!it || merged.has(it.id)) continue
341 for (const dep of a.dependsOn) {
342 if (typeof dep !== 'string' || dep === it.id || it.deps.includes(dep)) continue
343 if (!items.some(i => i.id === dep) || reaches(items, dep, it.id)) continue
344 it.deps.push(dep)
345 }
346 }
347}
348
349function reaches(items: OrcaItem[], from: string, to: string): boolean {
350 const seen = new Set<string>()
351 const stack = [from]
352 while (stack.length) {
353 const cur = stack.pop() as string
354 if (cur === to) return true
355 if (seen.has(cur)) continue
356 seen.add(cur)
357 const it = items.find(i => i.id === cur)
358 if (it) stack.push(...it.deps)
359 }
360 return false
361}
362
363// A review artifact on disk names the last completed round: <ID>-<reviewer>.json
364// is the latest, <ID>-<reviewer>.round<N>.json one per round.
365function reviewRound(id: string, reviews: { name: string }[]): number {
366 let round = 0
367 const esc = id.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')
368 const re = new RegExp(`^${esc}-[a-z]+(?:\\.round(\\d+))?\\.json$`)
369 for (const r of reviews) {
370 const m = re.exec(r.name)
371 if (!m) continue
372 round = Math.max(round, m[1] === undefined ? 1 : Number(m[1]) + 1)
373 }
374 return round
375}
376
377const isCut = (t: ItemTrack | undefined) => t !== undefined && t.cut !== null
378const isBlocked = (t: ItemTrack | undefined) =>
379 t !== undefined && !isCut(t) && (t.blocked !== null || t.salvaged)
380
381// Dependencies that have not merged; a cut one counts as satisfied, as it
382// does for the scheduler. A blocked one is named as such: the item waits on
383// a decision, not on work in progress.
384function waitingOn(item: OrcaItem, merged: Set<string>, track: Map<string, ItemTrack>): string[] {
385 return item.deps
386 .filter(d => !merged.has(d) && !isCut(track.get(d)))
387 .map(d => (isBlocked(track.get(d)) ? `${d} (blocked)` : d))
388}
389
390export function deriveBuild(run: OrcaRun, files: RunFiles, journal: JournalRead): OrcaBand {
391 const items = (run.items.length ? run.items : itemsFromSpec(files.spec?.text ?? null))
392 .map(i => ({ ...i, deps: [...i.deps] }))
393 const merged = mergedIds(files.merged)
394 const track = new Map<string, ItemTrack>(items.map(item =>
395 [item.id, { item, blocked: null, cut: null, salvaged: false, live: null, last: null }]))
396 const stages = journal.status === 'ok' ? parseJournal(journal.text) : null
397
398 const integration: { live: string | null; seen: boolean } = { live: null, seen: false }
399 let wrapping = false
400 for (const s of stages ?? []) {
401 const m = classifyLabel(s.label)
402 if (!m) continue
403 if (m.kind === 'item') {
404 for (const id of m.ids) {
405 const t = track.get(id)
406 if (!t) continue
407 t.last = m.stage
408 if (s.running) t.live = m.stage
409 }
410 } else if (m.kind === 'escalate') {
411 for (const id of m.ids) {
412 const t = track.get(id)
413 if (!t) continue
414 t.last = 'escalate'
415 if (s.running) t.live = 'escalate'
416 }
417 if (!s.running && !s.failed) applyEscalation(track, items, m.ids, s.result, merged)
418 } else if (m.kind === 'salvage') {
419 const t = track.get(m.id)
420 if (t) t.salvaged = true
421 } else if (m.kind === 'integration') {
422 integration.seen = true
423 if (s.running) integration.live = m.stage
424 } else if (m.kind === 'wrap-up') {
425 wrapping = true
426 }
427 }
428
429 const rows: OrcaRow[] = items.map(item => {
430 const t = track.get(item.id) as ItemTrack
431 const row = (mark: string, status: string, tone: OrcaRow['tone']): OrcaRow =>
432 ({ id: item.id, title: item.title, mark, status, tone })
433 if (merged.has(item.id)) return row('✓', 'done', 'done')
434 if (isCut(t)) return row('–', t.cut ? `cut: ${t.cut}` : 'cut', 'cut')
435 if (isBlocked(t)) {
436 const why = t.blocked || (t.last ? `after ${t.last}` : '')
437 return row('✕', why ? `blocked: ${why}` : 'blocked', 'blocked')
438 }
439 if (t.live) return row('●', t.live, 'live')
440 const waits = waitingOn(item, merged, track)
441 if (waits.length) return row('◌', `waiting on ${waits.join(', ')}`, 'wait')
442 if (stages) return t.last ? row('●', t.last, 'live') : row('○', 'queued', 'queued')
443 // File-only fallback: coarse, from what the stages leave on disk.
444 const round = reviewRound(item.id, files.reviews)
445 if (round > 0) return row('●', `reviewed r${round}`, 'live')
446 if (files.plans.includes(`${item.id}.md`)) return row('●', 'planning done', 'live')
447 return row('○', 'queued', 'queued')
448 })
449
450 const done = items.filter(i => merged.has(i.id)).length
451 const phase = wrapping ? 'wrapping up'
452 : integration.seen ? (integration.live ? `integrate · ${integration.live}` : 'integrate')
453 : 'build'
454 return { slug: run.slug, phase, detail: `${done}/${items.length} merged`, rows,
455 source: stages ? 'journal' : 'files' }
456}
457
458// ---------- the whole band ----------
459
460export function derive(run: OrcaRun, files: RunFiles, journal: JournalRead, now: number): Derived {
461 // A run directory that went away (archived, removed) takes the band with it.
462 if (!files.runDirExists) return { finished: true, band: null }
463 // The feature skill writes report.md after the work loop returns: the run
464 // is over. Only a report written since this launch counts — a retry or an
465 // iterate round relaunches over a directory whose old report was archived.
466 if (run.kind === 'build' && files.reportMtimeMs !== null && files.reportMtimeMs >= run.since)
467 return { finished: true, band: null }
468 if (run.kind === 'spec') return { finished: false, band: deriveSpec(run, files, journal, now) }
469 return { finished: false, band: deriveBuild(run, files, journal) }
470}
471
472// <task-notification><task-id>…</task-id>…<status>…</status>: the message the
473// session receives when a background task (a workflow included) ends. One
474// delivery may carry several.
475export function parseTaskNotifications(text: string): { taskId: string; status: string }[] {
476 const out: { taskId: string; status: string }[] = []
477 for (const block of text.split('<task-notification>').slice(1)) {
478 const id = /<task-id>([^<]+)<\/task-id>/.exec(block)
479 if (!id || !id[1]) continue
480 const status = /<status>([^<]+)<\/status>/.exec(block)
481 out.push({ taskId: id[1].trim(), status: status?.[1]?.trim() ?? '' })
482 }
483 return out
484}
485
486// .orca/<timestamp>-feat-<slug>/ → <slug>
487export function slugOf(runDir: string): string {
488 const base = runDir.replace(/\/+$/, '').split('/').pop() ?? runDir
489 return base.replace(/^\d{8}-\d{6}-(feat|bug|proto)-/, '')
490}
491
492// Claude Code's project folder name for a directory: every character that is
493// not a letter or a digit becomes '-'.
494export function projectDirName(cwd: string): string {
495 return cwd.replace(/[^a-zA-Z0-9]/g, '-')
496}
497
498// ---------- layout ----------
499
500export function truncate(text: string, width: number): string {
501 const flat = text.replace(/\s+/g, ' ').trim()
502 if (width <= 0) return ''
503 return flat.length <= width ? flat : `${flat.slice(0, Math.max(0, width - 1))}…`
504}
505
506// The band's top border: a rule across the width with the header set in it,
507// `── orca ─ slug · phase · detail ─────`. Each piece keeps its own tone; the
508// facts are cut from the end to fit and the rule takes whatever is left.
509export type RuleTone = 'rule' | 'name' | 'slug' | 'phase' | 'detail' | 'sep' | 'note'
510export type RulePiece = { text: string; tone: RuleTone }
511
512export function headerRule(band: OrcaBand, columns: number): RulePiece[] {
513 const facts: RulePiece[] = [
514 { text: band.slug, tone: 'slug' },
515 { text: band.phase, tone: 'phase' },
516 { text: band.detail, tone: 'detail' },
517 ]
518 if (band.source === 'files') facts.push({ text: 'no journal', tone: 'note' })
519 const out: RulePiece[] = [{ text: '── ', tone: 'rule' }, { text: 'orca', tone: 'name' }, { text: ' ─ ', tone: 'rule' }]
520 let used = out.reduce((n, p) => n + p.text.length, 0)
521 // A space and at least two rule cells close the line after the facts.
522 let room = columns - used - 3
523 let first = true
524 for (const fact of facts) {
525 const text = fact.text.replace(/\s+/g, ' ').trim()
526 if (!text) continue
527 if (!first) {
528 // A separator with nothing after it says nothing: stop instead.
529 if (room < 5) break
530 out.push({ text: ' · ', tone: 'sep' })
531 room -= 3
532 }
533 first = false
534 if (room <= 0) break
535 const cut = truncate(text, room)
536 out.push({ text: cut, tone: fact.tone })
537 room -= cut.length
538 if (cut !== text) break
539 }
540 used = out.reduce((n, p) => n + p.text.length, 0)
541 out.push({ text: ` ${'─'.repeat(Math.max(2, columns - used - 1))}`, tone: 'rule' })
542 return out
543}
544
545// Fits the rows to the band's width: ids padded to one column, titles cut to
546// what is left beside the widest status (statuses capped so a long blocked
547// reason never squeezes every title away), statuses cut to the remainder.
548// `lead` is ` <id> <title> `, with `id` and `title` its two padded cells.
549export function layoutRows(rows: OrcaRow[], columns: number): { lead: string; id: string; title: string; tail: string; row: OrcaRow }[] {
550 const idW = Math.max(0, ...rows.map(r => r.id.length))
551 const statusW = Math.min(30, Math.max(0, ...rows.map(r => r.status.length + 2)))
552 const lead = 2 + idW + 2
553 const titleMax = Math.max(0, ...rows.map(r => r.title.replace(/\s+/g, ' ').trim().length))
554 const titleW = Math.min(titleMax, Math.max(8, columns - lead - 2 - statusW))
555 return rows.map(row => {
556 const id = row.id.padEnd(idW)
557 const title = truncate(row.title, titleW).padEnd(titleW)
558 const head = ` ${id} ${title} `
559 const tail = truncate(`${row.mark} ${row.status}`, Math.max(1, columns - head.length))
560 return { lead: head, id, title, tail, row }
561 })
562}
563
564function isRecord(v: unknown): v is Record<string, unknown> {
565 return typeof v === 'object' && v !== null && !Array.isArray(v)
566}
567types/index.d.ts 61 lines1// orca's $.state contract: what the hooks module keeps in $.state so it
2// survives a hot reload — the run band's run and band (hooks/run-band.tsx),
3// and the verb hook's permission gate (hooks/relay.ts).
4
5// One work item as the work-loop launch passed it.
6export type OrcaItem = { id: string; title: string; deps: string[] }
7
8// The one orca workflow this session launched and the band follows.
9export type OrcaRun = {
10 // `spec` for spec.workflow.js, `build` for work-loop.workflow.js.
11 kind: 'spec' | 'build'
12 runDir: string
13 slug: string
14 runId: string
15 taskId: string
16 // <transcriptDir>/journal.jsonl from the launch result; null when the
17 // result carried no transcript directory and none could be derived.
18 journal: string | null
19 // When the launch was seen (ms since the epoch): run files older than
20 // this belong to an earlier launch over the same run directory.
21 since: number
22 items: OrcaItem[]
23 // The iterate skill's amendment file, when the spec launch passed one.
24 amendPath: string | null
25 // The spec launch's reviewer (codex, claude); null for a build launch.
26 reviewer: string | null
27 // The spec workflow's task notification arrived with status completed.
28 specEnded: boolean
29 // Prompts the person sent since the spec became ready with no launch.
30 idlePrompts: number
31}
32
33export type OrcaTone = 'done' | 'live' | 'wait' | 'blocked' | 'cut' | 'queued'
34
35export type OrcaRow = { id: string; title: string; mark: string; status: string; tone: OrcaTone }
36
37// What the band draws; null draws nothing.
38export type OrcaBand = {
39 slug: string
40 // `spec`, `build`, `integrate` or `integrate · <stage>`, `wrapping up`.
41 phase: string
42 // The spec state (drafting, reviewing, revising, ready), or `<merged>/<total> merged`.
43 detail: string
44 rows: OrcaRow[]
45 // Whether item stages came from the workflow journal or from run files alone.
46 source: 'journal' | 'files'
47}
48
49declare module 'claude-code' {
50 interface PluginState {
51 orca: {
52 run: OrcaRun | null
53 band: OrcaBand | null
54 // The session's permission mode as the classic hook events last
55 // reported it; null until one has. The verb hook runs verbs only while
56 // it is bypassPermissions.
57 permissionMode: string | null
58 }
59 }
60}
61