Spend human attention like money. Intent before code, evidence over claims, checks with teeth, an audit trail, and a profile whose taste evolves with you.

Spend human attention like money. flow-stack is a Claude Code plugin that makes the agent define "done" before coding, prove its work with evidence it can't fake, stop fixating on its first idea, prefer deleting code to adding it, and keep a trail you can audit. It adapts to the tools you already use.
It is built from skills, bash + jq hooks, and existing tools. There is no new runtime. It was inspired by gstack, pstack, and mattpocock/skills.
# inside Claude Code
/plugin marketplace add nayanraj210401/flow-stack
/plugin install flow-stack@flow-stack
/flow-stack:setup
setup learns your existing toolchain, creates ~/.flow-stack/profile.md from a preset, and offers the helper tools you're missing. It asks before installing anything.
For local development: claude --plugin-dir ./plugins/flow-stack.
Requirements: jq, git, perl, curl, and shasum (standard on macOS and most Linux). Optional tools setup can install: rtk, serena, headroom, graphify, ccusage, Playwright MCP, and ntfy.
flowchart LR
YOU(["👤 You<br/>request<br/>+ 3 gates:<br/>intent · merge<br/>profile changes"])
subgraph L1["① Context · ~2k tokens"]
PROF["profile.md<br/>who · rules · taste<br/>toolchain · rigor"]
REPO["repo/.flow/<br/>taste · map · lessons<br/>gates · playbooks"]
FEAT["feature map<br/>.flow/features/<br/>features · entries · owns"]
YOURS["your tools<br/>pstack · gstack · rtk<br/>serena · Playwright"]
FORGE["🔨 forge (you run)<br/>setup · make-verifier<br/>make-gates · tend"]
end
subgraph L2["② Hooks · 0 tokens"]
H1["SessionStart<br/>anchor"]
H2["guard<br/>read-guard"]
H3["seal · fence<br/>slice gate"]
H4["trail<br/>circuit breaker"]
H5["claims check<br/>pre-compact"]
end
subgraph L3["③ Skills"]
FLOW["🧭 flow<br/>classify → playbook<br/>feature · bug · refactor<br/>optimize · spike"]
OUTER["outer loop<br/>understand → intent<br/>→ challenge → slice<br/>→ prove → present<br/>→ close"]
LOOP["🔁 loop per slice<br/>red → subtract → build<br/>→ verify → probe<br/>→ proof gate"]
SUPPORT["on demand<br/>brief · gate · budget<br/>handoff · debt · tdd<br/>diagnose · principles"]
end
subgraph L4["④ Workers"]
AG["subagents · model per role<br/>advocate · checker<br/>reviewer · worker"]
SC["⚙ scripts · 0 tokens<br/>task.sh · evidence.sh<br/>seal · probe · diffstat<br/>blind-run · trace-stats"]
end
subgraph L5["⑤ Task memory<br/>.flow/tasks/slug/"]
T1["INTENT · SLICES"]
T2["SEALS · blind/"]
T3["EVIDENCE<br/>DECISIONS"]
T4["trail.jsonl"]
T5["HANDOFF · TRACE"]
end
YOU ==> L1 ==> L2 ==> L3 ==> L4 ==> L5
Read it left to right. A request passes through five layers:
.flow/ files (including the feature map), and your Toolchain delegations. The forge skills, which you run yourself, generate the repo-specific parts.flow classifies the request and runs a playbook. Its outer loop goes understand → intent → challenge → slice → prove → present → close, and each slice runs the inner loop. The other skills load only when a step needs them.budget: model for its role. Bookkeeping goes to shell scripts, which cost 0 tokens..flow/tasks/<slug>/: intent, sealed checks, evidence, decisions, the tool-call trail, and the handoff and trace. A fresh session resumes from these files, not from chat history.The diagrams are Mermaid. GitHub renders them. VS Code's built-in preview doesn't, unless you install a Mermaid extension such as "Markdown Preview Mermaid Support".
Mostly, you just talk. You don't need to remember skill names:
flow, which picks a playbook and runs it. You get three decision points:Type a skill yourself when you want a specific thing: /flow-stack:challenge, /flow-stack:brief, /flow-stack:wrap.
| Situation | Use | You'll get |
|---|---|---|
| New feature, multi-file change | just ask, or /flow-stack:flow | intent → approach → slices → proof → a short review tour |
| "What does this app do, and how do we prove it?" | /flow-stack:feature-map | one file per user-facing feature: sub-feature IDs, every entry point, the code it owns, a scenario |
| "What does my change affect?" | features.sh impact (flow runs it for you) | the features your diff touches, plus changed code no feature owns |
| Bug with an unknown cause | just describe it, or /flow-stack:diagnose | a reproduction first, then the root cause, then a fix with a regression check |
| Refactor, rename, migration | /flow-stack:flow (refactor playbook) | behavior pinned first; kept only if the code gets easier to read |
| Make something faster, smaller, cheaper | /flow-stack:flow (optimize playbook) | one change, one measurement, keep or revert |
| Not sure what to build | /flow-stack:flow (spike playbook) | 2 or 3 throwaway prototypes and a recommendation |
| "Is there a better way?" / you suspect over-engineering | /flow-stack:challenge | an independent design plus the delete/reuse/skip option, as a table |
| "What exactly are we building?" | /flow-stack:intent | INTENT.md with runnable acceptance checks |
| "How does X work?" / "Why is it like this?" | how / why | a cited answer (file:line, commits, PRs) |
| Understand something deeply | /flow-stack:teach | an explanation at your level, then a question to check you got it |
| "What changed?" / catching up | /flow-stack:what | the shape of a diff, PR, or the agent's session in about 8 lines |
| Prove it works | /flow-stack:verify | evidence recorded from the real app, not "it compiles" |
| Before you review a diff | /flow-stack:tour | what to read, skim, or skip, and why |
| Reply too long | /flow-stack:brief | the last reply in 3 plain lines |
| Clean up a diff / prose | /flow-stack:deslop / /flow-stack:unslop | slop removed, graded against your taste |
| Context filling up / stopping for the day | /flow-stack:handoff | HANDOFF.md; the next session starts with /flow resume |
| "What did we decide about X?" | /flow-stack:recall | a short brief from past conversations and flow's records, your own words quoted |
| Many independent slices | /flow-stack:delegate | parallel worktree workers, capped by your review budget |
| Audit what the agent did | /flow-stack:trace | TRACE.md: timeline, who decided what, evidence, cost |
| End of day | /flow-stack:wrap | shipped, waiting on you, and your first step tomorrow |
| See everything at once | /flow-stack:board | one private artifact link, refreshed in place: what needs you, tasks, features, decisions, debt, spend; your layout, theme, and custom panels persist in the profile's # Board section |
| After a frustrating session | /flow-stack:reflect | proposed taste and lesson updates you approve |
| New repo | /flow-stack:setup → /flow-stack:feature-map → /flow-stack:make-verifier | a driver that proves this app works the way a user uses it |
| Installed a new plugin or hook | /flow-stack:setup adapt | flow-stack re-fits itself around your toolchain |
Invoked says who triggers the skill:
flow calls it at the right step.Agent means the skill spawns a subagent, the most expensive kind of step. Its model comes from the profile's budget: for that role.
| Skill | Invoked | What it does |
|---|---|---|
flow | auto · you | The operating mode and orchestrator (like pstack's poteto-mode). It holds a routing table from situation to skill, classifies the request, and runs a playbook (feature, bug, investigate, refactor, optimize, spike, multi-session, or your repo's own) with human gates. It scales ceremony by rigor, sets autonomy (interactive / autonomous / quick), and holds the subagent rules. |
loop | flow | The inner loop for one slice: red → subtract scan → build → verify → probe → proof-gated done, with a circuit breaker. |
| Skill | Invoked | What it does |
|---|---|---|
intent | auto · flow | A short grilling interview (it answers from the code whatever the code can answer) → INTENT.md with goal, non-goals, and runnable checks. Blind checks are written by the checker agent. Agent (blind checks only) |
challenge | auto · flow | Breaks fixation. The advocate agent designs from the goal alone, never seeing our approach, then attacks ours. The null option (delete, reuse, configure, don't build) is always weighed. Agent |
slice | flow | Splits the work into slices, each with one check, a file fence, and a line budget. The first slice is a thin end-to-end tracer. |
| Skill | Invoked | What it does |
|---|---|---|
how | auto | How code works: runtime flow, ownership, where a change belongs. Cites file:line. |
why | auto | Why it's built this way, from git log and blame, PRs, issues, and docs, with a confidence level. |
what | auto | A summary of a diff, PR, branch, module, or the agent's session. |
recall | auto · you | What happened in earlier conversations in this repo: flow's records first, then recall.sh searches past Claude Code transcripts (your prompts and Claude's replies only, redacted) and checks what it finds against git and gh. |
teach | auto · you | A deep explanation at your level (from the profile), ending with a check for understanding. |
map | auto · flow | Builds .flow/map.md once, so later sessions read it instead of re-exploring the repo. |
| Skill | Invoked | What it does |
|---|---|---|
verify | auto · flow | Runs checks on the real artifact through evidence.sh (the only way to write EVIDENCE.md), runs the scenarios of the features your diff touches, uses your repo's driver, and runs the blind checks. |
feature-map | auto · you | Builds and maintains .flow/features/: one file per user-facing feature, with sub-feature IDs, every entry point, the code it owns, a scenario, and a proof. features.sh does impact, run, stale, coverage, and check in bash. |
seal | flow | Hashes the approved checks. Editing one then needs your confirmation. |
probe | flow | Reverts the change and re-runs the check. If it still passes, the check proves nothing (TOOTHLESS). |
tdd | auto | Red → green → refactor for fast checks. The slice-done gate enforces the cadence. |
diagnose | auto · flow | Reproduce as a failing check → hypotheses → observations → root cause. No symptom patches. |
review | flow · you | The reviewer agent grades the diff against INTENT, EVIDENCE, and your taste, and asks whether fewer lines could do it. Agent |
| Skill | Invoked | What it does |
|---|---|---|
brief | always on · you | The reply contract: answer first, ≤ 5 lines. /brief restates the last reply plainly. |
tour | flow · you | A risk-ranked diff: READ / SKIM / SKIP with reasons, net lines, and estimated review minutes. |
gate | auto | Classifies decisions: reversible ones proceed and are logged; irreversible or taste ones are asked, batched into one message, and you're notified. |
claims | auto | Tags every statement ✓ verified (evidence) or ~ assumed. The Stop hook enforces it for "done". |
| Skill | Invoked | What it does |
|---|---|---|
deslop | auto · flow | Removes code slop from the diff (narrating comments, dead code, defensive clutter, one-caller wrappers, duplicate helpers). Behavior must stay green. |
unslop | auto · flow | Removes AI tells from replies, commits, PRs, and docs. |
debt | flow · you | A ledger of shortcuts taken, each with a repay trigger. |
principles | auto | 18 principles (design-it-twice, subtract-before-add, red-before-green, evidence-over-claims, …). Each loads on demand as one short file. |
| Skill | Invoked | What it does |
|---|---|---|
budget | auto · you | Context and cost rules (subagents for bulk reading, cheap models, symbol reads, the handoff threshold). /budget reports spend. |
handoff | auto · you | Writes HANDOFF.md for a cold resume. /handoff resume loads it. |
delegate | you · flow | Parallel worktree worker agents, capped by attention.max_parallel_agents. A result is accepted only with evidence and a review. Agent |
trace | flow · you | Writes TRACE.md from the trail, decisions, evidence, git, and cost, including estimate vs. actual. |
reflect | flow · you | Proposes taste, lesson, and calibration updates from your corrections. Prefers encoding a lesson as a lint rule over prose. |
dojo | flow · you | Keeps your skills sharp: leaves a TODO(you) piece, or asks an explain-back question. Off by default. |
| Skill | What it does |
|---|---|
setup | Learns your toolchain and usage, adapts to it, creates your profile, offers tools, and runs a forge census. /flow-stack:setup adapt re-fits after changes. |
profile | Shows, edits, or lints ~/.flow-stack/profile.md. |
wrap | End-of-day digest across your repos. |
make-verifier | Generates .claude/skills/verify-<repo>/, a driver that runs this app like a user (CLI, HTTP, browser) and proves itself on a real feature. |
make-runner | Generates .claude/skills/run-<repo>/ with verified install, start, seed, reset, and stop commands. |
make-playbook | Generates .flow/playbooks/<name>.md from past traces, for workflows this repo repeats. |
make-gates | Generates .flow/gates.md: this repo's irreversible commands as deny/ask rules the guard hook enforces. |
make-skill | Generates any other repo or personal skill, with a self-test. |
tend | Re-runs every generated skill's self-test and flags drift against the code. |
| Agent | Spawned by | Sees |
|---|---|---|
advocate | challenge | the problem only (design mode), then our approach (attack mode) |
checker | intent | INTENT; writes held-out blind checks and returns only a count |
reviewer | review | INTENT, the diff, EVIDENCE, taste; read-only |
worker | delegate | one slice, in its own worktree; never the blind checks |
flow-agent | any other delegated step | loads flow first, so fences, evidence, and gates hold in delegated work |
flow-stack writes files and makes decisions, and that costs something. These are the numbers, measured with claude plugin eval against the same tasks without the plugin.
| Tokens | When | |
|---|---|---|
| Skill list the model sees | ≈ 1.75k | every session (cached after the first turn) |
| SessionStart context (profile, rules, taste, active task) | ≈ 0.3–0.5k | every session |
| Intent anchor | ≈ 40 | every prompt, only while a task is active |
flow + a playbook + conventions | ≈ 3.6k | only when a non-trivial task starts |
On small eval tasks, where fixed overhead dominates, a session with the plugin costs 1.43× a plain session on average. That was 1.78× before the optimizations below.
| Kind of task | With / without |
|---|---|
| Quick question / feature routing | 1.1× |
| Intent writing | 1.0× |
| One-turn replies | 1.2–1.5× (about +$0.02 fixed) |
challenge (spawns an agent) | 2.8× |
The numbers above are the price side. The value side is what the evals show the plugin prevents:
.env secrets reaching the transcript,A rework loop or a wrong approach usually costs more than the overhead. Judge it on your own traces: trace records estimate vs. actual for every task.
rigor in your profile decides what runs:| Step | lean | standard (default) | strict |
|---|---|---|---|
| Blind checks (agent) | skip | when a contract changes | always |
challenge (agent) | on request, or when the circuit trips | multi-slice features and refactors | every task |
review (agent) | only for diffs over 300 lines | one reviewer | one per area, in parallel |
| Evidence, seals, probe, slice gate (bash) | on | on | on |
| Trace / reflect | on request | at close | at close |
The solo-hacker preset uses lean. Override per task by saying "quick" or "go strict".
budget:: design_model for review and design, build_model for building, subagent_model (default Haiku) for exploration. The agents pin no model, so a spawn that names none inherits the session's.map.md replaces re-exploring the repo each session. HANDOFF.md replaces a lossy compaction. The anchor re-injects 3 lines instead of re-reading INTENT.setup records your token savers (rtk, headroom, serena) and delegates to tools you already pay for, instead of duplicating them.Check your own spend with /flow-stack:budget (uses ccusage) and the estimate-vs-actual row in each TRACE.md.
| Hook | Blocks or does | |
|---|---|---|
| guard (Bash) | Denies rm -rf /, force-push to main, reading .env, and packages that don't exist on npm or PyPI (hallucinated dependencies). Asks you before reset --hard, destructive SQL, `curl \ | sh, and anything in your repo's .flow/gates.md`. |
| seal (Edit, Write, Bash) | Editing a sealed check, INTENT.md, or SEALS needs your confirmation. | |
| fence (Edit, Write) | Edits outside the current slice's file fence are denied until the fence is widened on purpose and logged. | |
| blind / read-guard | Blind checks and .env files can't be read by the builder. | |
| slice gate | status: done can only be set by task.sh slice <id> done, which requires red first, green after the last edit, probe TEETH, and the line budget. | |
| circuit (PostToolUse) | The same failing check 3 times → stop, attack the premise, challenge, and ask you. | |
| trail (PostToolUse) | Every tool call goes to trail.jsonl with secrets redacted. | |
| claims (Stop) | "Done", "works", or "fixed" with no passing evidence since the last edit → sent back once to verify or say ~ assumed. | |
| anchor (UserPromptSubmit) | A 3-line goal and slice reminder on every prompt while a task is active. | |
| session-start / pre-compact / notify | Profile and toolchain context; a handoff snapshot before compaction; a notification when the agent is blocked on you. |
Hooks fail open: if jq is missing or a script errors, the action is allowed and a warning goes to ~/.flow-stack/hooks.log. Turn any hook off in .flow/config.json (repo) or ~/.flow-stack/config.json (global): {"hooks": {"fence": false}}.
hooks/register.tsx is a mod: code that runs inside Claude Code (2.1.287+), next to the bash hooks. It changes no guard. It draws only in the terminal and the Desktop app; in VS Code chat and claude -p nothing draws, but model per role still applies.
| Mod | What you get | Reads |
|---|---|---|
| Flow band | A dim line above the prompt: flow · <task> · <slice> · n/m slices · tdd red · C1 PASS · edited since. Empty when no task is active. | status.sh |
| Spinner | While Claude works, the spinner shows the slice: Thinking · S2 · red… | status.sh |
/flow-pane | A pane with slices, open gates with Approve / Reject, and your other leads. It is a command, so it runs instantly with no Claude turn and no tokens. | status.sh --full |
| Model per role | A flow-stack agent started with no model gets one from your profile's budget:. advocate and reviewer get design_model; worker, checker and flow-agent get build_model. A model the call names wins. Built-ins such as Explore are left alone. A toast names the model the first time each role spawns; a budget value that isn't a model name is ignored with a warning. | ~/.flow-stack/profile.md |
╭ flow ─────────────────────────────────────────────╮
│ flow · rate-limit · S2 token bucket · 1/3 slices │ ← the band's line
│ │
│ S1 bucket · done · PASS │ ← slices: status · last verdict
│ S2 token bucket · doing · FAIL │
│ S3 headers · todo │
│ │
│ GATES │ ← open gates in GATES.md
│ merge PR #3 │
│ options: A) merge B) wait (recommend: A) │
│ [ Approve ] [ Reject ] hooks/register.tsx 388 lines1// flow-stack's mod: what bash hooks can't do.
2// - the flow band above the prompt, and the slice beside the spinner (state from skills/flow/scripts/status.sh)
3// - each flow-stack agent spawns on the model its role gets in the profile's budget:
4// - /flow-pane: the task's slices, open gates with Approve/Reject, the other leads, this session's
5// subagents, the evidence history, and context and cost gauges, animated while open
6import type { EngineInterface, Register, SessionUsage, Timer } from 'claude-code'
7
8type Status = {
9 task?: string
10 where?: 'main' | 'lead' | 'lane'
11 slice: { id: string; title: string } | null
12 done: number
13 total: number
14 tdd: string
15 evidence: { label: string; verdict: string; ts: string } | null
16 stale: boolean
17 // with status.sh --full, while the pane is open
18 slices?: { id: string; title: string; status: string; verdict: string }[]
19 gates?: { n: number; question: string; detail: string }[]
20 leads?: { id: string; repo: string; branch: string; task: string; slice: string }[]
21 runs?: string[]
22 est_usd?: number | null
23}
24
25// A subagent this session spawned: started at agent.spawn, counted on each tool.call it makes,
26// ended at its turn.complete.
27type Lane = { role: string; model: string; what: string; start: number; tools: number; end?: number; ok?: boolean }
28
29const PANE = 'flow-pane'
30
31type Budget = 'subagent_model' | 'build_model' | 'design_model'
32
33// A flow-stack agent the Agent call names without a model runs on the model its role gets.
34// Only flow-stack's own agents: built-ins like Explore keep Claude Code's choice.
35const ROLE: Record<string, Budget> = {
36 'flow-stack:advocate': 'design_model',
37 'flow-stack:reviewer': 'design_model',
38 'flow-stack:worker': 'build_model',
39 'flow-stack:checker': 'build_model',
40 'flow-stack:flow-agent': 'build_model',
41}
42
43// A model alias, or a full id from the API, Bedrock (us.anthropic.claude-…, ARNs), or Vertex (claude-…@date).
44const MODEL = /^(haiku|sonnet|opus|fable|inherit|[\w.:@\/-]*claude[\w.:@\/-]*)(\[1m\])?$/
45
46// The profile's frontmatter `budget:` block (" build_model: sonnet # comment"). A value that
47// isn't a model is dropped into `rejected`, so a typo leaves the spawn on Claude Code's choice.
48export function parseBudget(profile: string) {
49 const models: Partial<Record<Budget, string>> = {}
50 const rejected: string[] = []
51 const front = /^---\r?\n([\s\S]*?)\r?\n---/.exec(profile)?.[1] ?? ''
52 for (const m of front.matchAll(/^[ \t]+(subagent_model|build_model|design_model):[ \t]*([^\s#]+)/gm)) {
53 const [, key, value] = m as unknown as [string, Budget, string]
54 if (MODEL.test(value)) models[key] = value
55 else rejected.push(`${key}: ${value}`)
56 }
57 return { models, rejected }
58}
59
60export function bandText(s: Status): string {
61 const parts = [`flow · ${s.where === 'lane' ? 'lane of ' : ''}${s.task}`]
62 if (s.slice) parts.push(`${s.slice.id} ${s.slice.title}`)
63 if (s.total) parts.push(`${s.done}/${s.total} slices`)
64 if (s.tdd) parts.push(`tdd ${s.tdd}`)
65 if (s.evidence) {
66 parts.push(`${s.evidence.label} ${s.evidence.verdict}${s.stale ? ' · edited since' : ''}`)
67 }
68 return parts.join(' · ')
69}
70
71const SPIN = '⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏'
72const EIGHTHS = ' ▏▎▍▌▋▊▉'
73const FRAME_MS = 100
74
75// A progress bar `width` cells wide, filled to `fraction` in eighths of a cell.
76export function bar(fraction: number, width: number): string {
77 const eighths = Math.round(Math.min(1, Math.max(0, fraction)) * width * 8)
78 const full = Math.floor(eighths / 8)
79 const part = eighths % 8 ? EIGHTHS[eighths % 8] : ''
80 return '█'.repeat(full) + part + '░'.repeat(width - full - (part ? 1 : 0))
81}
82
83export const spin = (frame: number) => SPIN[frame % SPIN.length]
84
85// Consecutive equal verdicts as one run of ■, so the strip is a few Texts, not forty.
86export function strip(runs: string[]): { verdict: string; cells: string }[] {
87 const out: { verdict: string; cells: string }[] = []
88 for (const v of runs) {
89 const last = out.at(-1)
90 if (last?.verdict === v) last.cells += '■'
91 else out.push({ verdict: v, cells: '■' })
92 }
93 return out
94}
95
96export function laneText(l: Lane, now: number): string {
97 const secs = Math.max(0, Math.round(((l.end ?? now) - l.start) / 1000))
98 return `${l.role.padEnd(10)} ${l.model.padEnd(7)} ${l.what} ${secs}s · ${l.tools} tools`
99}
100
101const LANE_LINGER_MS = 8000
102
103// Drops lanes that ended more than LANE_LINGER_MS ago; called wherever the clock is read.
104function prune(t: number) {
105 now = t
106 for (const [id, l] of lanes) if (l.end !== undefined && t - l.end > LANE_LINGER_MS) lanes.delete(id)
107}
108
109// The pane's animation: a frame count, the bar easing toward done/total, rows fading in after open.
110let frame = 0
111let openedAt = 0
112let shown = 0
113let ticker: Timer | null = null
114let now = 0
115let usage: SessionUsage | null = null
116const lanes = new Map<string, Lane>()
117
118let status: Status | null = null
119let models: Partial<Record<Budget, string>> = {}
120let paneOpen = false
121const deciding = new Set<number>()
122let running = false
123let again = false
124const toasted = new Set<string>()
125
126// Re-read the task's state; calls that land while one runs fold into one more pass.
127async function refresh($: EngineInterface) {
128 if (running) {
129 again = true
130 return
131 }
132 running = true
133 try {
134 do {
135 again = false
136 try {
137 const args = [...(paneOpen ? ['--full'] : []), await $.session.cwd()]
138 const { stdout } = await $.process.run([`${$.plugin.root}/skills/flow/scripts/status.sh`, ...args])
139 const next = JSON.parse(stdout || '{}') as Status
140 status = next.task ? next : null
141 } catch {
142 status = null
143 }
144 $.ui.invalidate('ui.render')
145 } while (again)
146 } finally {
147 running = false
148 }
149}
150
151// A gate button: record the decision (only if gate n still reads as shown), then tell Claude.
152// The question stays out of the prompt: GATES.md is repo text, the decision is the human's.
153async function decide($: EngineInterface, n: number, question: string, verdict: 'approve' | 'reject') {
154 if (deciding.has(n)) return
155 deciding.add(n)
156 try {
157 const ran = await $.process.run(
158 [`${$.plugin.root}/skills/flow/scripts/task.sh`, 'gate', String(n), verdict, question],
159 { cwd: await $.session.cwd() },
160 )
161 if (ran.exitCode !== 0) {
162 $.ui.toast(`flow-stack: ${(ran.stderr || ran.stdout).trim()}`)
163 } else {
164 const past = verdict === 'approve' ? 'approved' : 'rejected'
165 await $.prompt.submit({ text: `The human ${past} gate ${n} in the /flow-pane; see GATES.md and DECISIONS.tsv.` })
166 }
167 } finally {
168 deciding.delete(n)
169 await refresh($)
170 }
171}
172
173export const register: Register = on => {
174 on('session.start', async ($, e, next) => {
175 const home = (await $.env.get('FLOW_STACK_HOME')) ?? `${await $.env.get('HOME')}/.flow-stack`
176 const budget = parseBudget(await $.fs.read(`${home}/profile.md`).catch(() => ''))
177 models = budget.models
178 if (budget.rejected.length) {
179 $.ui.toast(`flow-stack: ignoring budget ${budget.rejected.join(', ')} (not a model name)`, { timeoutMs: 10000 })
180 }
181 await $.command.register({ name: 'flow-pane', description: "Open flow-stack's pane: slices, gates to approve, leads" })
182 void refresh($)
183 return next(e)
184 })
185
186 on('command.run', { command: 'flow-pane' }, async $ => {
187 await $.ui.open({ id: PANE, title: 'flow' })
188 paneOpen = true
189 openedAt = frame
190 shown = 0
191 ticker ??= $.clock.every(FRAME_MS, () => {
192 frame++
193 const target = status?.total ? status.done / status.total : 0
194 shown = Math.abs(target - shown) < 0.005 ? target : shown + (target - shown) * 0.25
195 void $.clock.now().then(prune)
196 if (frame % 10 === 1) void $.session.usage().then(u => (usage = u), () => {})
197 $.ui.invalidate('ui.render')
198 })
199 void refresh($)
200 return {}
201 })
202
203 on('ui.close', { id: PANE }, async ($, e, next) => {
204 const closed = await next(e)
205 paneOpen = false
206 ticker?.cancel()
207 ticker = null
208 return closed
209 })
210
211 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
212 const { Box, Button, Text } = $.ui.resolve(e)
213 const { slices = [], gates = [], leads = [], runs = [], est_usd = null } = status ?? {}
214 const age = frame - openedAt
215 const pulse = Math.floor(frame / 5) % 2 === 0
216 // row i fades in on frame i after the pane opens
217 const faded = (i: number) => age <= i
218 const verdictColor = (v: string) => (v === 'PASS' ? 'success' : v === 'FAIL' ? 'error' : 'subtle')
219 const ctx = usage?.context.percent
220 const usd = usage?.cost?.usd
221 const agents = [...lanes.entries()]
222 const busy = agents.filter(([, l]) => l.end === undefined).length
223 return (
224 <Box flexDirection="column" gap={1}>
225 {status ? (
226 <Box gap={1}>
227 <Text color="claude">{spin(frame)}</Text>
228 <Text bold>{bandText(status)}</Text>
229 </Box>
230 ) : (
231 <Text dimColor>No active flow task. Start one with /flow-stack:flow.</Text>
232 )}
233 {status && status.total > 0 && (
234 <Box gap={1}>
235 <Text color="claude">{bar(shown, 24)}</Text>
236 <Text dimColor>
237 {status.done}/{status.total}
238 </Text>
239 </Box>
240 )}
241 <Box flexDirection="column">
242 {slices.map((s, i) => {
243 const active = s.id === status?.slice?.id && s.status !== 'done'
244 const icon = s.status === 'done' ? '✓' : active ? spin(frame + i) : '○'
245 const iconColor = faded(i) ? 'subtle' : s.status === 'done' ? 'success' : active ? 'claude' : 'subtle'
246 return (
247 <Box key={`slice-${s.id}`} gap={1}>
248 <Text color={iconColor}>{icon}</Text>
249 <Text dimColor={faded(i) || s.status === 'done'} bold={active && !faded(i)}>
250 {s.id} {s.title} · {s.status}
251 </Text>
252 {s.verdict !== '' && <Text color={faded(i) ? 'subtle' : verdictColor(s.verdict)}>{s.verdict}</Text>}
253 </Box>
254 )
255 })}
256 </Box>
257 {runs.length > 0 && (
258 <Box gap={1}>
259 <Box>
260 {strip(runs).map((g, i) => (
261 <Text key={`run-${i}`} color={verdictColor(g.verdict)}>
262 {g.cells}
263 </Text>
264 ))}
265 </Box>
266 <Text dimColor>
267 last {runs.length} checks · {runs.filter(v => v === 'PASS').length} pass
268 </Text>
269 </Box>
270 )}
271 {(ctx !== undefined || usd !== undefined) && (
272 <Box flexDirection="column">
273 {ctx !== undefined && (
274 <Box gap={1}>
275 <Text color={ctx >= 80 ? 'error' : ctx >= 60 ? 'warning' : 'success'}>{bar(ctx / 100, 24)}</Text>
276 <Text dimColor>context {Math.round(ctx)}%</Text>
277 </Box>
278 )}
279 {usd !== undefined && (
280 <Text dimColor>
281 ${usd.toFixed(2)} this session{est_usd ? ` · task est $${est_usd.toFixed(2)}` : ''}
282 </Text>
283 )}
284 </Box>
285 )}
286 {agents.length > 0 && (
287 <Box flexDirection="column">
288 <Text bold>AGENTS · {busy} running</Text>
289 {agents.map(([id, l], i) => (
290 <Box key={`agent-${id}`} gap={1}>
291 <Text color={l.end === undefined ? 'claude' : l.ok ? 'success' : 'error'}>
292 {l.end === undefined ? spin(frame + i * 2) : l.ok ? '✓' : '✗'}
293 </Text>
294 <Text dimColor={l.end !== undefined} wrap="truncate-end">
295 {laneText(l, now)}
296 </Text>
297 </Box>
298 ))}
299 </Box>
300 )}
301 {gates.length > 0 && (
302 <Box flexDirection="column">
303 <Text bold color={pulse ? 'warning' : 'subtle'}>
304 {pulse ? '◆' : '◇'} GATES · {gates.length} waiting on you
305 </Text>
306 {gates.map(g => (
307 <Box key={`gate-${g.n}`} flexDirection="column">
308 <Text>{g.question}</Text>
309 {g.detail !== '' && <Text dimColor>{g.detail}</Text>}
310 <Box gap={1}>
311 <Button key={`approve-${g.n}`} label="Approve" onPress={() => decide($, g.n, g.question, 'approve')} />
312 <Button key={`reject-${g.n}`} label="Reject" onPress={() => decide($, g.n, g.question, 'reject')} />
313 {deciding.has(g.n) && <Text color="claude">{spin(frame)} recording…</Text>}
314 </Box>
315 </Box>
316 ))}
317 </Box>
318 )}
319 {leads.length > 0 && (
320 <Box flexDirection="column">
321 <Text bold>LEADS</Text>
322 {leads.map((l, i) => (
323 <Box key={`lead-${l.id}`} gap={1}>
324 <Text color={Math.floor((frame + i * 3) / 4) % 2 === 0 ? 'success' : 'subtle'}>●</Text>
325 <Text dimColor>
326 {l.id} · {l.repo} · {l.branch}
327 {l.task ? ` · ${l.task}` : ''}
328 {l.slice ? ` ${l.slice}` : ''}
329 </Text>
330 </Box>
331 ))}
332 </Box>
333 )}
334 </Box>
335 )
336 })
337
338 // evidence.sh, task.sh, and edits all change what the band shows
339 on('tool.call', async ($, e, next) => {
340 const lane = e.agentId ? lanes.get(e.agentId) : undefined
341 if (lane) lane.tools++
342 const ran = await next(e)
343 if (/^(Bash|Edit|Write|MultiEdit)$/.test(e.tool)) void refresh($)
344 return ran
345 })
346
347 on('turn.complete', async ($, e, next) => {
348 const lane = e.agentId ? lanes.get(e.agentId) : undefined
349 if (lane) {
350 lane.end = await $.clock.now()
351 lane.ok = e.reason === 'answer'
352 prune(lane.end)
353 }
354 void refresh($)
355 return next(e)
356 })
357
358 on('agent.spawn', async ($, e, next) => {
359 const role = ROLE[e.subagentType]
360 const model = !e.model && role ? models[role] : undefined
361 if (model && !toasted.has(e.subagentType)) {
362 toasted.add(e.subagentType)
363 $.ui.toast(`${e.subagentType.replace('flow-stack:', '')} → ${model}`)
364 }
365 const started = await next(model ? { ...e, model } : e)
366 if (started.agentId) {
367 const what = e.description || e.subagentType
368 const start = await $.clock.now()
369 prune(start)
370 lanes.set(started.agentId, { role: e.subagentType.replace('flow-stack:', ''), model: started.model, what, start, tools: 0 })
371 }
372 return started
373 })
374
375 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
376 if (!status || e.props.hasSurvey) return next(e)
377 const { Text } = $.ui.resolve(e)
378 return <Text dimColor wrap="truncate-end">{bandText(status)}</Text>
379 })
380
381 on('ui.render', { component: 'Spinner' }, async ($, e, next) => {
382 const s = status?.slice
383 if (!s) return next(e)
384 const tdd = status?.tdd ? ` · ${status.tdd}` : ''
385 return next({ ...e, props: { ...e.props, suffix: ` · ${s.id}${tdd}…` } })
386 })
387}
388