A verification-first development pipeline for Claude Code: brainstorm to plan to phased implementation, with quality gates (verification-loop, code-review…

Custom skills and agents for Claude Code that enhance codebase research, context management, and implementation planning workflows.
New to this workflow? See the Workflow Overview for a step-by-step guide.
Installing? This repo ships as the
devflowClaude Code plugin — see Installation and Plugin Distribution.
This skill collection uses Claude Code's native Task tools for progress tracking.
Progress tracking is handled entirely through Claude Code's built-in Task system:
| Tool | Purpose |
|---|---|
TaskCreate | Create tasks for each phase with dependencies |
TaskUpdate | Mark tasks as in_progress or completed |
TaskList | View all tasks with status and blockers |
TaskGet | Get full task details including description |
Benefits:
Plans remain pure specification documents. See Progress Tracking for details.
This collection uses modern Claude Code skill features (v2.1.16+):
| Feature | Skills Using It | Purpose |
|---|---|---|
context: fork | codebase-research | Run in isolated subagent context. Only for skills that run to completion without the user and don't wait on agents of their own: agent reports arrive in the top-level session, so the phase leads (implement-phase, tt-implement-phase) run inline. |
agent: Explore/Plan | (none) | Subagent type for a forked skill. Both built-ins are read-only and lack the Agent tool, so any skill that writes files or spawns subagents must leave this unset. |
allowed-tools | code-review, verification-loop, security-review, adversarial-reviewer, codebase-research, strategic-compact | Restrict available tools (read-only enforcement) |
argument-hint | implement-plan, implement-phase, adr, e2e-testing, code-review, adversarial-reviewer, context-saver, prompt-generator | Show usage hints in autocomplete |
disable-model-invocation | context-saver, prompt-generator | User-only invocation (no auto-trigger) |
user-invocable: false | implement-phase | Hide from user menu (internal skill) |
Never add context: fork to an interactive skill. A forked skill runs in a subagent with no access to the conversation history, and since Claude Code v2.1.218 it is backgrounded by default — so a skill that asks the user questions and waits for answers (brainstorm) simply never reaches the user. agent: Explore compounds it: that agent type is read-only, so the skill cannot write its output document either. tests/test-interactive-skills.sh guards this.
Skills support the new argument syntax:
$0, $1, $2 - Positional arguments$ARGUMENTS - All arguments${CLAUDE_SESSION_ID} - Session trackingExample: /implement-plan docs/plans/my-feature.md passes the path as $0.
The core workflow for implementing features follows this hierarchy:
┌─────────────────┐
│ brainstorm │
│ (ideation) │
└────────┬────────┘
│
┌───────────┴───────────┐
▼ ▼
┌─────────────────┐ ┌──────────────────┐
│ adr │ │ user-story │
│ (decisions) │ │ (requirements) │
└────────┬────────┘ └──────────┬───────┘
└────────────┬──────────┘
│
▼
┌─────────────────┐
│ create-plan │
│ (spec → phases) │
└────────┬────────┘
│
▼
┌──────────────────────────────────────────────────────────────────────────────┐
│ implement-plan │
│ (orchestrates all phases) │
│ │
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │
│ │ Phase 1 │───▶│ Phase 2 │───▶│ Phase N │───▶ Complete │
│ └──────┬──────┘ └──────┬──────┘ └──────┬──────┘ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ┌─────────────────────────────────────────────────────────────────────┐ │
│ │ implement-phase (per phase) │ │
│ │ ┌────────────────────────────────────────────────────────────┐ │ │
│ │ │ Step 1: Implementation (subagents) [TDD mode: tests first] │ │ │
│ │ │ Step 2: verification-loop (6-phase exit conditions) ─────┐│ │ │
│ │ │ Step 3: Integration Testing (API/UI via Playwright) ││ │ │
│ │ │ Step 4: code-review ─────────────────────────────────────┼┼──┐ │ │
│ │ │ ├─► security-review (optional OWASP audit) ──────┼┼──┼─┐│ │
│ │ │ Step 5: ADR Compliance ──────────────────────────────────┼┼──┼─┼┤ │
│ │ │ Step 6: Plan Sync ││ │ ││ │
│ │ │ Step 7: Prompt Archival ││ │ ││ │
│ │ │ Step 8: Completion Report ││ │ ││ │
│ │ └──────────────────────────────────────────────────────────┼┼──┼─┼┘ │
│ └─────────────────────────────────────────────────────────────┼┼──┼─┼─────┘
│ ││ │ │ │
│ ┌───────────────────────────────────────────────────────┘│ │ │ │
│ │ ┌─────────────────────────────────────────┘ │ │ │
│ │ │ ┌──────────────────────────────┘ │ │
│ │ │ │ ┌─────────────────┘ │
│ ▼ ▼ ▼ ▼ │
│ ┌───────────┐ ┌───────────┐ ┌───────────┐ ┌─────────────────┐ │
│ │verification│ │code-review│ │security- │ │ adr │ │
│ │ -loop │ │ │ │ review │ │ │ │
│ └───────────┘ └───────────┘ └───────────┘ └─────────────────┘ │
│ (default) (optional) │
└───────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────┐
│ e2e-testing │
│ (validation) │
└────────┬────────┘
│
▼
┌──────────────────────────┐
│ continuous-learning │
│ (extract session patterns)│
└──────────────────────────┘
| Stage | Skill | Purpose |
|---|---|---|
| Ideation | brainstorm | Refine rough ideas through Socratic questioning |
| Requirements | user-story | Generate hierarchical user stories with acceptance criteria |
| Planning | create-plan | Create detailed, phased implementation plans |
| Iteration | iterate-plan | Update plans based on feedback |
| Execution | implement-plan | Orchestrate full plan execution |
| Phase Work | implement-phase | Execute single phase with quality gates |
| Quality | code-review | Verify code quality, patterns, ADR compliance |
| Adversarial Quality | adversarial-reviewer | Subagent-based hostile review (Saboteur, New Hire, Security Auditor) to break self-review blind spots |
| Comprehensive Audit | codebase-audit | Long-running full-codebase audit — partitions the repo, delegates to adversarial-reviewer per partition, synthesizes a written remediation report |
| Security | security-review | OWASP-aligned security audit (optional step) |
| Verification | verification-loop | 6-phase verification: build, type, lint, test, security, diff |
| Metrics Gate | code-quality-audit | Coverage, complexity, module size, deps, mutation — gate or on-demand |
| Decisions | adr | Document architectural decisions |
| Testing | e2e-testing | End-to-end validation with Playwright |
| Evaluation | eval-harness | Formal capability/regression testing with metrics |
| Learning | continuous-learning | Extract patterns from sessions for reuse |
Skills are invoked via the Skill tool or /skill-name shorthand.
| Skill | Trigger | Description |
|---|---|---|
| brainstorm | /brainstorm, "explore this idea" | Interactive idea refinement using Socratic questioning, written to docs/brainstorms/ for create-plan or tt-create-plan |
| user-story | /user-story, "create user stories" | Generate hierarchical user stories (epics/features/tasks) with Given/When/Then acceptance criteria |
| create-plan | /create-plan, "plan the implementation" | Creates detailed implementation plans through research |
| iterate-plan | "update the plan", "iterate on this plan" | Updates existing plans based on feedback |
| implement-plan | /implement-plan, "implement the plan" | Orchestrates execution of complete plans (subagent mode) |
| implement-phase | Called by implement-plan | Executes single phase with all quality gates |
| tt-create-plan | /tt-create-plan, "plan this in tasktracker" | TaskTracker-native planning: requirements with acceptance criteria and phase tasks from templates instead of a docs/plans/*.md file |
| tt-implement-plan | /tt-implement-plan, "execute the tasktracker plan" | Runs a TaskTracker plan phase by phase, delegating each phase to tt-implement-phase, with active-task time tracking and insight logging |
| tt-workflow-run | /tt-workflow-run, "drive the backlog to done" | Autonomous gated run over a TaskTracker backlog: next ready slice → tt-implement-phase → drift and defect gates → measured-time projection, until the backlog drains |
| tt-workflow-build | /tt-workflow-build, "parallel build (tasktracker)" | Parallel build of a TaskTracker project, on the Workflow tool when it is enabled and on parallel subagents otherwise, driven to done |
| tt-create-build-loop | /tt-create-build-loop, "set up an autonomous build loop" | Writes prompts/autonomous-build-loop.md, the prompt a /loop re-reads every iteration to plan, implement and verify TaskTracker phases unattended under a zero-error gate, with a human-only queue and a per-phase cost ledger. Detects the repo, stack, CI and TaskTracker phases, asks only what it can't, creates missing queue/ledger phases, optionally installs the bundled token-usage tooling, and installs a clean-worktree gate runner |
| workflow-guide | /workflow-guide, "which workflow should I use" | Routes work to a lane: file-based or TaskTracker pipeline, autonomous run, unattended loop, parallel build or audit, or issue-to-merge |
| Skill | Trigger | Description |
|---|---|---|
| code-review | /devflow:code-review (bare /code-review is Claude Code's built-in review), Step 4 of implement-phase and tt-implement-phase | Systematic review: SRP, patterns, ADR compliance |
| adversarial-reviewer | /adversarial-reviewer, "adversarial review", "critical review", "audit this repo" | Spawns three hostile-persona subagents (Saboteur, New Hire, Security Auditor) in parallel; each must find ≥1 issue; cross-persona findings get severity-promoted. Default mode reviews a diff; --codebase [path] reviews a whole repo/subtree with strategic per-persona deep-dives |
| grumpy-reviewer | /grumpy-reviewer, "structural review", "is this well factored" | One isolated-subagent reviewer that judges only maintainability (separation of concerns, helper extraction, small files, the rule of 7) and never learns how the code was produced |
| codebase-audit | /codebase-audit, "comprehensive codebase review", "thorough audit", "code due diligence" | Long-running full-coverage audit. Partitions the repo, delegates to /adversarial-reviewer --codebase per partition, synthesizes systemic findings, produces written remediation report. Resumable. Pairs with code-quality-audit for qualitative + quantitative picture |
| tt-workflow-audit | /tt-workflow-audit, "parallel audit (tasktracker)" | Read-only parallel audit of a TaskTracker project's repo, backlog or architecture: a ranked risk register, with fix tasks written back by the parent on approval. Resumable |
| adr | /adr, "document decision" | Creates Architecture Decision Records |
| e2e-testing | /e2e-testing, "test my webapp" | E2E testing with Playwright MCP |
| security-review | /devflow:security-review (bare /security-review is the built-in), auth/input code | 10-category OWASP-aligned security audit |
| verification-loop | /verification-loop, "verify implementation" | 6-phase verification: build, type, lint, test, security, diff |
| code-quality-audit | /code-quality-audit, "audit code quality", "run mutation testing" | Coverage + complexity + module size + dependency cycles + mutation score. Gate or on-demand modes |
| eval-harness | /eval-harness, "run evals" | Formal evaluation framework with pass@k metrics |
| Skill | Trigger | Description |
|---|---|---|
| codebase-research | "how does X work" | Parallel codebase research with sub-agents |
| context-saver | /context-saver, "save context" | Preserves session state for continuation |
| prompt-generator | /prompt, "generate prompt" | Creates implementation prompts for phases |
| Skill | Trigger | Description |
|---|---|---|
| continuous-learning | "save what we learned", /continuous-learning | Promotes reusable procedures from a session into learned skills at ~/.claude/skills/learned-<slug>/ (on demand; one-liners go to auto-memory) |
| strategic-compact | PreToolUse hook | Suggests /compact at logical boundaries, not arbitrary thresholds |
| skill-visualizer | /skill-visualizer, "visualize skills" | Generate interactive HTML visualizations of skills and codebase |
| Skill | Trigger | Description |
|---|---|---|
| agent-creator | "create agent", "build agent" | Creates composable AI agent systems in NestJS |
| design-language | "design my own design language", "build my design system", "I want my own look not the Claude default" | Guided design-brief → living-styleguide compiler. Reaction-first elicitation (web research + real references + live archetypes rendered in your own tokens → you react → targeted refinement) with a per-dimension 1–5 "safeness" gradient (conventional→experimental) across the entire system (color, type, spacing, shape, elevation, motion, imagery, components, patterns, voice). Resumable multi-phase flow; emits a self-contained, portable ./design-system/ (interactive dashboard + css/tokens.css single-source-of-truth + W3C design-tokens.json + DESIGN_LANGUAGE.md contract with an explicit anti-Claude reference + pre-ship checklist + a default-vs-yours compare.html) that future sessions read and build from. Different from frontend-design (which styles one UI in the moment). |
| Skill | Trigger | Description |
|---|---|---|
| video-explainer | /video-explainer <url-or-path>, "make an HTML explainer for this video", "turn this talk into visual notes" | Turns a video (YouTube, any yt-dlp URL, or a local file) into one self-contained HTML explainer page: headline + TL;DR, key takeaways, a section per part of the video, and diagrams built in pure HTML + CSS (17 components: flows, timelines, bar/line/dumbbell charts, threshold scales, comparisons, cycles, layers…). Watches the video via the watch plugin (frames + transcript), adds screenshots only when the video's own image is the information, renders the page headless at desktop and phone width to check its own layout, and links every claim to its timestamp. Requires the watch plugin: see the Video Explainer Guide for install and usage. |
| Skill | Trigger | Description |
|---|---|---|
| ship-issue | /ship-issue <issue-number-or-url> | One-command GitHub issue → merged PR pipeline across nine stages (preflight, plan, implement, review, ci, cloud_review, deploy, e2e, logs) with exactly two human gates — plan approval and merge confirmation. Model-tiered: Fable 5 plans, orchestrates, and runs the merge-gate review; Opus 4.8 implements (TDD); Sonnet 4.6 runs staging E2E and log checks. File-based run state gives lossless crash-resume; per-stage time tracking (work / gate-wait / crash-gap, with a per-model-tier rollup) is embedded in the Gate 2 merge brief. Pair with the single-file dashboard.py for a live view. |
Overview page: open
skills/ship-issue/pipeline.htmlin a browser for a one-page tour — the skill, its five agents, and the full stage-flow diagram with model-tier colour-coding (self-contained, no dependencies).
Agents are specialized subagents launched with the Agent tool. When the collection is installed as a plugin they are named devflow:<agent>.
| Agent | Purpose |
|---|---|
| implementer | sonnet, medium effort: writes code and tests for one briefed unit of work; used by tt-implement-phase for implementation and fix rounds |
| mechanic | haiku, low effort: runs build/type/lint/test checks and verification-loop, report-only, never edits |
| reviewer | opus, high effort (callers pass fable for security reviews): runs code-review / security-review on code it did not write |
| codebase-analyzer | Traces implementation with file:line references |
| codebase-locator | Finds files by topic/feature ("Super Grep/Glob") |
| codebase-pattern-finder | Finds concrete code examples and patterns |
| docs-analyzer | Extracts insights from docs, ADRs, design docs |
| docs-locator | Finds documentation and research notes |
| web-search-researcher | Web research for APIs, libraries, troubleshooting |
| browser-verification-agent | UI testing via Playwright MCP with screenshot evidence |
| design-researcher | claude-sonnet-4-6 — per-dimension inspiration research for the design-language skill (WebSearch → Playwright screenshots → cached manifest), steering away from the default Claude look |
Each agent pins its model in frontmatter — the tier is part of the contract, never switched mid-task (prompt caches are model-scoped; ADR-0007). A fix cycle is always a fresh task on the same tier, never a resumed task on a different model.
| Agent | Model | Purpose |
|---|---|---|
| issue-planner | claude-fable-5 (Fable 5) | Turns a GitHub issue + codebase into the plan presented at Gate 1. Outcome-prompted (ADR-0008). |
| merge-gate-reviewer | claude-fable-5 (Fable 5) | Last-line diff review at the merge gate; verdict contract APPROVE / FIX (itemized blockers). Outcome-prompted. |
| tdd-implementer | claude-opus-4-8 (Opus 4.8) | Tests-first implementation, UI components, API routes; receives reviewer/CI/E2E blockers verbatim on fix cycles. |
| staging-e2e-verifier | claude-sonnet-4-6 (Sonnet 4.6) | Runs the plan's E2E scenarios against the live staging URL via Playwright MCP; PASS/FAIL/FLAKY/BLOCKED with screenshot evidence. |
| staging-log-verifier | claude-sonnet-4-6 (Sonnet 4.6) | Scans staging service logs over the deploy window; CLEAN vs ERRORS_FOUND with cited lines. |
This repository is distributed exclusively as a Claude Code plugin named devflow, published through a marketplace named mhylle. Inside Claude Code:
/plugin marketplace add mhylle/claude-skills-collection
/plugin install devflow@mhylle
Then /reload-plugins (or restart) if the install summary asks for it.
Plugin skills are namespaced, so every skill is invoked as /devflow:<skill>:
/devflow:brainstorm
/devflow:create-plan
/devflow:implement-plan
Nothing is copied into ~/.claude/ — the plugin lives in Claude Code's plugin cache, carries a version, and updates in place:
/plugin marketplace update mhylle
To remove it: /plugin uninstall devflow.
See Plugin Distribution for the manifest layout, versioning, publishing, and local development.
Earlier versions shipped an install.sh that copied skills, agents, and hooks into ~/.claude/. It is gone (ADR-0011).
Those copies do not disappear on their own, and plugin skills do not override same-named personal skills — so until you remove them you will have both /brainstorm and /devflow:brainstorm pointing at two copies that drift apart, with Claude free to auto-invoke either. Remove the copies once, after installing the plugin:
# from a checkout of this repo
for s in $(ls skills); do rm -rf "$HOME/.claude/skills/$s"; done
for a in agents/*.md; do rm -f "$HOME/.claude/agents/$(basename "$a")"; done
~/.claude/hooks.json is your own user configuration, not a copy of this repo's file. The old script overwrote it (leaving hooks.json.backup); the plugin no longer touches it. Review it by hand and delete only the entries you recognise from this collection.
Skills, agents, and these hooks — all discovered automatically from the plugin's default component locations (skills/, agents/, hooks/hooks.json).
Hooks (automatic behaviors):
/devflow:strategic-compact is invoked on demand rather than by a hook — see ADR-0011.
| Hook Type | Name | Trigger | Purpose |
|---|---|---|---|
| PreToolUse | tmux-dev-block | npm run dev etc. | Block dev servers outside tmux |
| PreToolUse | tmux-reminder | Long-running commands | Suggest tmux for session persistence |
| PreToolUse | git-push-review | git push | Reminder to review before push |
| PreToolUse | doc-file-warn | .md/.txt creation | Warn about docs outside docs/ structure |
| PostToolUse | pr-url-logger | gh pr create | Log PR URL and review command |
| PostToolUse | prettier-format | JS/TS file edits | Auto-format with Prettier |
| PostToolUse | typescript-check | .ts/.tsx edits | Run tsc --noEmit and show errors |
| PostToolUse | console-log-warn | JS/TS file edits | Wa
hooks/register.tsx 649 lines1// devflow HUD (part of the devflow plugin): a progress band and a /hud pane for the devflow TaskTracker skills, and the functions
2// the skills lean on: the orchestrator guard, the in-flight watch, the time-tracking heartbeat and
3// pause, and the test-run alarms. See README.md.
4
5import { atom, read, update } from 'claude-code'
6import type { EngineInterface, Register, ToolCallResult } from 'claude-code'
7
8import type { HudAgent, HudCheck, HudGuard, HudPhase, HudRequirement, HudTask } from '../types'
9import {
10 formatDuration,
11 isGuardExempt,
12 linksTask,
13 parseBuildSummary,
14 parseCriteria,
15 parseRequirements,
16 parseSlnTestProjects,
17 parseTask,
18 parseTaskList,
19 parseTestRuns,
20 parseTestSummaries,
21 shapeOf,
22} from './parse'
23import { isRunning, PANE } from './state'
24import { band, hasContent, pane, type HudView } from './view'
25
26// Every function that takes `$`, and every state atom, lives in this file: the engine follows `$` and reads
27// `$.state` references only where they are declared in the file that uses them.
28
29// The session state the band and the pane draw from (declared in ../types, PluginState).
30const phaseAtom = atom({ plugin: 'devflow', key: 'phase' } as const, null as HudPhase | null)
31const activeTaskAtom = atom({ plugin: 'devflow', key: 'activeTaskId' } as const, null as string | null)
32const agentsAtom = atom({ plugin: 'devflow', key: 'agents' } as const, [] as HudAgent[])
33const checksAtom = atom({ plugin: 'devflow', key: 'checks' } as const, [] as HudCheck[])
34const guardAtom = atom({ plugin: 'devflow', key: 'guard' } as const, { isEnabled: true, isActive: false, skill: '' } as HudGuard)
35const bandHiddenAtom = atom({ plugin: 'devflow', key: 'isBandHidden' } as const, false)
36const tickAtom = atom({ plugin: 'devflow', key: 'tick' } as const, 0)
37
38/** How often the background tick polls the agents and runs a wanted refresh. */
39const TICK_MS = 2_000
40
41/** TaskTracker closes a time segment after ~5 minutes without a call; heartbeat inside that. */
42const HEARTBEAT_MS = 240_000
43
44/** TaskTracker tools after which the view is read again. */
45const WRITES = new Set([
46 'updateTaskStatus',
47 'batchUpdateStatus',
48 'createTask',
49 'batchCreateTasks',
50 'updateTask',
51 'batchUpdateTasks',
52 'archiveTask',
53 'deleteTask',
54 'reparentTask',
55 'completeWithCaveat',
56 'addAcceptanceCriterion',
57 'updateAcceptanceCriterion',
58 'deleteAcceptanceCriterion',
59])
60
61/** Writes that change which requirements a phase links (the cached links are read again). */
62const LINK_WRITES = new Set(['linkRequirementToTask', 'unlinkRequirementFromTask'])
63
64const PREFIX = 'mcp__tasktracker__tasktracker_'
65
66const USAGE = '/hud opens the pane. /hud refresh | guard on|off | band on|off | project <uuid>'
67
68/** Module state; a reload starts it over (what is drawn lives in $.state). */
69const hud = {
70 orchestrator: /implement-plan|implement-phase|tt-workflow-run/,
71 projectId: '',
72 /** After an automatic pause: no TaskTracker call of ours until the next turn, or it would book the wait. */
73 isQuiet: false,
74 isRefreshWanted: false,
75 isForcedRefresh: false,
76 isUserRefresh: false,
77 isTicking: false,
78 lastHeartbeat: 0,
79 lastRedraw: 0,
80 slnTests: new Map<string, string[]>(),
81}
82
83// TaskTracker, through the session's own MCP connection. Every call heartbeats the active task (TaskTracker
84// books time from calls), so none is made during a wait for the person (hud.isQuiet).
85
86/** One TaskTracker tool call; its text blocks joined. Rejects when the server reports an error. */
87async function call($: EngineInterface, tool: string, args: Record<string, unknown>): Promise<string> {
88 const result = await $.mcp.call('tasktracker', `tasktracker_${tool}`, args)
89 const text = result.content.map(block => block.text ?? '').join('\n')
90 if (result.isError) {
91 throw new Error(`${tool}: ${text.slice(0, 200)}`)
92 }
93
94 return text
95}
96
97/** The requirements linked to a phase, cached in $.store for an hour (links rarely change). */
98async function linkedRequirements($: EngineInterface, phaseId: string, isForced: boolean): Promise<{ id: string; slug: string }[]> {
99 const key = `links:${phaseId}`
100 const cached = (await $.store.get(key)) as { at: number; requirements: { id: string; slug: string }[] } | undefined
101 const now = await $.clock.now()
102 if (cached && !isForced && now - cached.at < 3_600_000) {
103 return cached.requirements
104 }
105
106 const all = parseRequirements(await call($, 'listRequirements', { projectId: hud.projectId, status: 'approved' }))
107 const linked: { id: string; slug: string }[] = []
108 for (const requirement of all) {
109 if (linksTask(await call($, 'listRequirementTaskLinks', { requirementId: requirement.id }), phaseId)) {
110 linked.push(requirement)
111 }
112 }
113
114 await $.store.set(key, { at: now, requirements: linked })
115 return linked
116}
117
118/** The phase of task `taskId` with its sub-tasks and (when the project is known) its requirements' ACs. */
119async function loadPhase($: EngineInterface, taskId: string, isForced: boolean): Promise<HudPhase> {
120 const task = parseTask(await call($, 'getTask', { taskId }))
121 if (task === null) {
122 throw new Error(`getTask ${taskId}: unexpected answer`)
123 }
124
125 let phase: HudTask = task
126 if (task.type !== 'phase') {
127 phase = parseTaskList(await call($, 'getTaskAncestors', { taskId })).find(t => t.type === 'phase') ?? task
128 }
129
130 const subtasks = parseTaskList(await call($, 'getChildTasks', { taskId: phase.id }))
131 const requirements: HudRequirement[] = []
132 if (hud.projectId !== '' && phase.type === 'phase') {
133 for (const requirement of await linkedRequirements($, phase.id, isForced)) {
134 const criteria = parseCriteria(await call($, 'listAcceptanceCriteria', { requirementId: requirement.id }))
135 requirements.push({ ...requirement, criteria })
136 }
137 }
138
139 return { phase, active: task.id === phase.id ? null : task, subtasks, requirements, refreshedAt: await $.clock.now(), isIdle: false }
140}
141
142function requestRefresh(isForced: boolean, isUser: boolean) {
143 hud.isRefreshWanted = true
144 hud.isForcedRefresh ||= isForced
145 hud.isUserRefresh ||= isUser
146}
147
148async function viewOf($: EngineInterface): Promise<HudView> {
149 return {
150 phase: await read($, phaseAtom),
151 agents: await read($, agentsAtom),
152 checks: await read($, checksAtom),
153 guard: await read($, guardAtom),
154 now: await $.clock.now(),
155 }
156}
157
158/** Reads the phase of the active task (or, on a user's refresh with none active, of the last phase seen). */
159async function refresh($: EngineInterface) {
160 const active = await read($, activeTaskAtom)
161 const shown = await read($, phaseAtom)
162 const target = active ?? (hud.isUserRefresh ? (shown?.phase.id ?? null) : null)
163 const wasQuiet = hud.isQuiet
164 const isForced = hud.isForcedRefresh
165 hud.isRefreshWanted = false
166 hud.isForcedRefresh = false
167 hud.isUserRefresh = false
168 if (target === null) {
169 return
170 }
171
172 try {
173 const phase: HudPhase = { ...(await loadPhase($, target, isForced)), isIdle: active === null }
174 await update($, phaseAtom, () => phase)
175 await $.store.set('lastPhase', phase)
176 } catch (error) {
177 const message = error instanceof Error ? error.message : String(error)
178 await update($, phaseAtom, p => (p === null ? p : { ...p, error: message }))
179 }
180
181 if (wasQuiet) {
182 // A refresh the person asked for during a wait: close the segment its calls opened.
183 await call($, 'pauseActiveTask', {}).catch(() => undefined)
184 }
185}
186
187/** Mirrors $.agent.list() into the state; a toast for each agent that finished. */
188async function pollAgents($: EngineInterface) {
189 const listed = await $.agent.list()
190 const known = await read($, agentsAtom)
191 const now = await $.clock.now()
192 const byId = new Map(known.map(agent => [agent.id, agent]))
193 const finished: HudAgent[] = []
194 let isChanged = false
195 for (const info of listed) {
196 const previous = byId.get(info.id)
197 if (previous === undefined) {
198 byId.set(info.id, {
199 id: info.id,
200 description: info.description,
201 type: info.type,
202 status: info.status,
203 startedAt: now,
204 ...(isRunning(info) ? {} : { endedAt: now }),
205 })
206 isChanged = true
207 } else if (previous.status !== info.status) {
208 const changed: HudAgent = { ...previous, status: info.status }
209 if (isRunning(previous) && !isRunning(info)) {
210 changed.endedAt = now
211 finished.push(changed)
212 }
213
214 byId.set(info.id, changed)
215 isChanged = true
216 }
217 }
218
219 for (const previous of known) {
220 if (isRunning(previous) && !listed.some(info => info.id === previous.id)) {
221 const ended: HudAgent = { ...previous, status: 'completed', endedAt: now }
222 byId.set(previous.id, ended)
223 finished.push(ended)
224 isChanged = true
225 }
226 }
227
228 if (isChanged) {
229 await update($, agentsAtom, () => [...byId.values()].sort((a, b) => a.startedAt - b.startedAt).slice(-30))
230 }
231
232 for (const agent of finished) {
233 const mark = agent.status === 'completed' ? '✔' : '✗'
234 $.ui.toast(`${mark} ${agent.type} "${agent.description}" ${agent.status} after ${formatDuration((agent.endedAt ?? now) - agent.startedAt)}`)
235 }
236}
237
238/** The background tick: agents, elapsed-time redraws, wanted refreshes and the heartbeat. */
239async function tick($: EngineInterface) {
240 if (hud.isTicking) {
241 return
242 }
243
244 hud.isTicking = true
245 try {
246 await pollAgents($)
247 const now = await $.clock.now()
248 const isAnyRunning = (await read($, agentsAtom)).some(isRunning)
249 if (isAnyRunning && now - hud.lastRedraw >= 30_000) {
250 hud.lastRedraw = now
251 await update($, tickAtom, n => n + 1)
252 }
253
254 if (hud.isRefreshWanted && (!hud.isQuiet || hud.isUserRefresh)) {
255 await refresh($)
256 }
257
258 const active = await read($, activeTaskAtom)
259 if (!hud.isQuiet && isAnyRunning && active !== null && now - hud.lastHeartbeat >= HEARTBEAT_MS) {
260 // Agents work while the main thread waits for them: keep the active task's segment open.
261 hud.lastHeartbeat = now
262 await call($, 'getCurrentTimer', { taskId: active }).catch(() => undefined)
263 }
264 } finally {
265 hud.isTicking = false
266 }
267}
268
269/** Pauses the active task's timer before a wait for the person, unless agents still work. */
270async function pauseIfIdle($: EngineInterface) {
271 if (hud.isQuiet || (await read($, activeTaskAtom)) === null) {
272 return
273 }
274
275 if ((await $.agent.list()).some(isRunning)) {
276 return
277 }
278
279 await call($, 'pauseActiveTask', {}).catch(() => undefined)
280 hud.isQuiet = true
281}
282
283async function agentLabel($: EngineInterface, agentId: string | undefined): Promise<string> {
284 if (agentId === undefined) return 'main'
285 const agent = (await read($, agentsAtom)).find(a => a.id === agentId)
286 return agent === undefined ? 'agent' : `${agent.type} "${agent.description}"`
287}
288
289async function testProjectsOf($: EngineInterface, sln: string): Promise<string[]> {
290 let projects = hud.slnTests.get(sln)
291 if (projects === undefined) {
292 projects = parseSlnTestProjects(await $.fs.read(sln))
293 hud.slnTests.set(sln, projects)
294 }
295
296 return projects
297}
298
299/** Reads build, test and verify results from a shell command's output; notes for the model when a project did not run. */
300async function readChecks($: EngineInterface, command: string, agentId: string | undefined, ran: ToolCallResult): Promise<ToolCallResult> {
301 if (ran.deny !== undefined) {
302 return ran
303 }
304
305 const text = typeof ran.text === 'string' ? ran.text : ''
306 const shape = shapeOf(command)
307 const summaries = parseTestSummaries(text)
308 const build = parseBuildSummary(text)
309 if (summaries.length === 0 && build === null && !shape.isVerify) {
310 return ran
311 }
312
313 const now = await $.clock.now()
314 const loop = await agentLabel($, agentId)
315 const checks: HudCheck[] = []
316 const notes: string[] = []
317 if (build !== null && summaries.length === 0 && !shape.isVerify) {
318 checks.push({
319 label: 'build',
320 isOk: build.errors === 0 && build.warnings === 0,
321 failed: 0,
322 passed: 0,
323 warnings: build.warnings,
324 errors: build.errors,
325 missing: [],
326 projects: 0,
327 at: now,
328 loop,
329 })
330 }
331
332 if (summaries.length > 0) {
333 let missing: string[] = []
334 const started = parseTestRuns(text)
335 if (shape.isDotnetTest && shape.sln !== null && (!shape.isPartial || started.length > 0)) {
336 const expected = await testProjectsOf($, shape.sln).catch(() => [] as string[])
337 const seen = new Set([...summaries.map(s => s.project), ...started])
338 missing = expected.filter(project => !seen.has(project))
339 }
340
341 const failed = summaries.reduce((n, s) => n + s.failed, 0)
342 checks.push({
343 label: 'tests',
344 isOk: failed === 0 && missing.length === 0,
345 failed,
346 passed: summaries.reduce((n, s) => n + s.passed, 0),
347 missing,
348 projects: summaries.length,
349 at: now,
350 loop,
351 })
352 if (missing.length > 0) {
353 notes.push(
354 `devflow-hud: ${missing.length} test project(s) of ${shape.sln} reported no result: ${missing.join(', ')}. ` +
355 'dotnet test can skip a project (for example while a test project it references fails): run each one separately before counting the totals.',
356 )
357 }
358 }
359
360 if (shape.isVerify) {
361 const summary = text.split('\n').reverse().find(line => /\b(passed|failed|FAIL)\b/i.test(line))?.trim()
362 checks.push({ label: 'verify', isOk: ran.isError !== true, failed: 0, passed: 0, missing: [], projects: 0, at: now, loop, summary })
363 }
364
365 await update($, checksAtom, list => [...list, ...checks].slice(-20))
366 const bad = checks.find(check => !check.isOk)
367 if (bad !== undefined) {
368 const detail = bad.missing.length > 0 ? `not run: ${bad.missing.join(', ')}` : bad.label === 'tests' ? `${bad.failed} failed` : bad.label
369 $.ui.toast(`✗ ${bad.label} (${loop}): ${detail}`)
370 }
371
372 return notes.length === 0 ? ran : { ...ran, context: [...(ran.context ?? []), ...notes] }
373}
374
375/** The guard's refusal of a main-thread write while an orchestrator skill leads, or undefined. */
376async function guardRefusal($: EngineInterface, path: string, agentId: string | undefined): Promise<string | undefined> {
377 if (agentId !== undefined || isGuardExempt(path)) {
378 return undefined
379 }
380
381 const guard = await read($, guardAtom)
382 if (!guard.isEnabled || !guard.isActive) {
383 return undefined
384 }
385
386 return (
387 `devflow-hud orchestrator guard: ${guard.skill} leads this session, and a plan orchestrator never writes code itself. ` +
388 'Dispatch this change to the implementer role agent (devflow:implementer), as /tt-implement-phase describes; do not work around this with shell commands. ' +
389 'Only the user can switch the guard off, with /hud guard off.'
390 )
391}
392
393/** Notifies when a task goes blocked or the shown phase completes. */
394async function notifyStatus($: EngineInterface, taskId: string, status: unknown) {
395 const phase = await read($, phaseAtom)
396 const task = [phase?.phase, phase?.active, ...(phase?.subtasks ?? [])].find(t => t?.id === taskId)
397 const title = task?.title ?? taskId
398 if (status === 'blocked') {
399 await $.ui.notify(`Blocked: ${title}`, { title: 'devflow' })
400 } else if (status === 'completed' && phase?.phase.id === taskId) {
401 await $.ui.notify(`Phase completed: ${title}`, { title: 'devflow' })
402 }
403}
404
405/** Reminder for the model of the agents still running, or undefined when none. */
406async function runningNote($: EngineInterface, lead: string): Promise<string | undefined> {
407 const running = (await $.agent.list()).filter(isRunning)
408 if (running.length === 0) {
409 return undefined
410 }
411
412 const names = running.map(agent => `${agent.type} "${agent.description}"`).join(', ')
413 return `devflow-hud: ${lead} ${running.length} agent(s) still run: ${names}. Their reports arrive in this session as messages when they finish; do not report on their work, predict it, or start work that depends on it before then.`
414}
415
416async function openPane($: EngineInterface) {
417 await $.ui.open({ id: PANE, title: 'devflow HUD' })
418}
419
420async function setGuardEnabled($: EngineInterface, isEnabled: boolean) {
421 await update($, guardAtom, guard => ({ ...guard, isEnabled }))
422 await $.store.set('guardEnabled', isEnabled)
423}
424
425async function toggleGuard($: EngineInterface) {
426 await setGuardEnabled($, !(await read($, guardAtom)).isEnabled)
427}
428
429async function toggleBand($: EngineInterface) {
430 await update($, bandHiddenAtom, isHidden => !isHidden)
431}
432
433/** Loads what the HUD keeps across sessions and starts the background tick. */
434async function start($: EngineInterface) {
435 await $.command.register({
436 name: 'hud',
437 description: 'devflow HUD: the TaskTracker phase, agents and checks',
438 argumentHint: '[refresh|guard on|off|band on|off|project <id>]',
439 })
440 const storedProject = await $.store.get('projectId')
441 if (hud.projectId === '' && typeof storedProject === 'string') {
442 hud.projectId = storedProject
443 }
444
445 const isGuardEnabled = await $.store.get('guardEnabled')
446 if (typeof isGuardEnabled === 'boolean') {
447 await update($, guardAtom, guard => ({ ...guard, isEnabled: isGuardEnabled }))
448 }
449
450 if ((await read($, phaseAtom)) === null) {
451 const last = (await $.store.get('lastPhase')) as HudPhase | undefined | null
452 if (last !== undefined && last !== null) {
453 await update($, phaseAtom, () => ({ ...last, isIdle: true }))
454 }
455 }
456
457 $.clock.every(TICK_MS, () => {
458 void tick($)
459 })
460}
461
462/** Follows a TaskTracker call: the project, the active task, refreshes after writes, status notifications. */
463async function followTaskTracker($: EngineInterface, tool: string, args: Record<string, unknown>, ran: ToolCallResult): Promise<ToolCallResult> {
464 if (ran.deny !== undefined || ran.isError === true) {
465 return ran
466 }
467
468 const name = tool.slice(PREFIX.length)
469 hud.isQuiet = name === 'pauseActiveTask'
470 if (typeof args.projectId === 'string' && /^[0-9a-f-]{36}$/.test(args.projectId) && args.projectId !== hud.projectId) {
471 hud.projectId = args.projectId
472 await $.store.set('projectId', hud.projectId)
473 }
474
475 if (name === 'setActiveTask' && typeof args.taskId === 'string') {
476 const taskId = args.taskId
477 await update($, activeTaskAtom, () => taskId)
478 requestRefresh(false, false)
479 } else if (name === 'clearActiveTask') {
480 await update($, activeTaskAtom, () => null)
481 await update($, phaseAtom, p => (p === null ? p : { ...p, isIdle: true }))
482 } else if (WRITES.has(name) || LINK_WRITES.has(name)) {
483 requestRefresh(LINK_WRITES.has(name), false)
484 }
485
486 if (name === 'updateTaskStatus' && typeof args.taskId === 'string') {
487 await notifyStatus($, args.taskId, args.status)
488 }
489
490 return ran
491}
492
493/** Orchestrator mode follows the main thread's skills; a fork that returns with agents still running is said so. */
494async function followSkill($: EngineInterface, skill: string, ran: ToolCallResult, before: Set<string>): Promise<ToolCallResult> {
495 if (ran.deny !== undefined) {
496 return ran
497 }
498
499 const started = (await $.agent.list()).filter(agent => isRunning(agent) && !before.has(agent.id))
500 if (started.length === 0) {
501 return ran
502 }
503
504 const names = started.map(agent => `${agent.type} "${agent.description}"`).join(', ')
505 return {
506 ...ran,
507 context: [
508 ...(ran.context ?? []),
509 `devflow-hud: the skill ${skill} returned while ${started.length} agent(s) it dispatched still run: ${names}. ` +
510 'Their reports arrive in this session as messages when they finish. Do not report on their work, or start work that depends on it, before then.',
511 ],
512 }
513}
514
515async function noteSkill($: EngineInterface, skill: string): Promise<Set<string>> {
516 if (hud.orchestrator.test(skill)) {
517 await update($, guardAtom, guard => ({ ...guard, isActive: true, skill }))
518 } else if (!/(^|:)(tt-|adr$|continuous-learning$)/.test(skill)) {
519 await update($, guardAtom, guard => ({ ...guard, isActive: false, skill: '' }))
520 }
521
522 return new Set((await $.agent.list()).filter(isRunning).map(agent => agent.id))
523}
524
525async function runCommand($: EngineInterface, args: string): Promise<{ text: string }> {
526 const [verb = '', arg = ''] = args.trim().split(/\s+/)
527 switch (verb) {
528 case '':
529 await openPane($)
530 requestRefresh(false, true)
531 return { text: 'devflow HUD pane opened.' }
532 case 'refresh':
533 requestRefresh(true, true)
534 return { text: 'devflow HUD: reading TaskTracker again.' }
535 case 'guard':
536 if (arg !== 'on' && arg !== 'off') return { text: USAGE }
537 await setGuardEnabled($, arg === 'on')
538 return { text: `devflow HUD: orchestrator guard ${arg}.` }
539 case 'band':
540 if (arg !== 'on' && arg !== 'off') return { text: USAGE }
541 await update($, bandHiddenAtom, () => arg === 'off')
542 return { text: `devflow HUD: band ${arg === 'on' ? 'shown' : 'hidden'}.` }
543 case 'project':
544 if (!/^[0-9a-f-]{36}$/.test(arg)) return { text: USAGE }
545 hud.projectId = arg
546 await $.store.set('projectId', hud.projectId)
547 requestRefresh(true, true)
548 return { text: `devflow HUD: project ${arg}.` }
549 default:
550 return { text: USAGE }
551 }
552}
553
554export const register: Register = (on, options) => {
555 hud.orchestrator = new RegExp(String(options.orchestratorSkills || 'implement-plan|implement-phase|tt-workflow-run'))
556 hud.projectId = String(options.projectId ?? '')
557
558 on('session.start', async ($, e, next) => {
559 await start($)
560 return next(e)
561 })
562
563 on('tool.call', async ($, e, next) => {
564 const tool = String(e.tool)
565 if (!tool.startsWith(PREFIX)) {
566 return next(e)
567 }
568
569 return followTaskTracker($, tool, e as unknown as Record<string, unknown>, await next(e))
570 }).catch(($, e, next) => next(e))
571
572 on('tool.call', { tool: 'Skill' }, async ($, e, next) => {
573 if (e.agentId !== undefined) {
574 return next(e)
575 }
576
577 const before = await noteSkill($, e.skill)
578 return followSkill($, e.skill, await next(e), before)
579 }).catch(($, e, next) => next(e))
580
581 // The orchestrator guard: main-thread file writes are refused while an orchestrator skill leads.
582 on('tool.call', { tool: 'Write' }, async ($, e, next) => {
583 const deny = await guardRefusal($, e.file_path, e.agentId)
584 return deny === undefined ? next(e) : { deny }
585 }).catch(($, e, next) => next(e))
586
587 on('tool.call', { tool: 'Edit' }, async ($, e, next) => {
588 const deny = await guardRefusal($, e.file_path, e.agentId)
589 return deny === undefined ? next(e) : { deny }
590 }).catch(($, e, next) => next(e))
591
592 on('tool.call', { tool: 'NotebookEdit' }, async ($, e, next) => {
593 const deny = await guardRefusal($, e.notebook_path, e.agentId)
594 return deny === undefined ? next(e) : { deny }
595 }).catch(($, e, next) => next(e))
596
597 // A question to the person is a wait: pause the active task's timer first.
598 on('tool.call', { tool: 'AskUserQuestion' }, async ($, e, next) => {
599 if (e.agentId === undefined) {
600 await pauseIfIdle($)
601 }
602
603 return next(e)
604 }).catch(($, e, next) => next(e))
605
606 // Check alarms: build, test and verify results in any loop's shell output.
607 on('tool.call', { tool: 'Bash' }, async ($, e, next) => readChecks($, e.command, e.agentId, await next(e))).catch(($, e, next) => next(e))
608 on('tool.call', { tool: 'PowerShell' }, async ($, e, next) => readChecks($, e.command, e.agentId, await next(e))).catch(($, e, next) => next(e))
609
610 // The end of a main-thread turn is a wait for the person, unless agents still work.
611 on('turn.complete', async ($, e, next) => {
612 if (e.agentId === undefined) {
613 await pauseIfIdle($)
614 }
615
616 return next(e)
617 })
618
619 // The person's prompt ends the wait; it reminds the model of agents still running.
620 on('prompt.submit', async ($, e, next) => {
621 hud.isQuiet = false
622 const note = await runningNote($, 'as this prompt arrives,')
623 return next(note === undefined ? e : { ...e, context: [...(e.context ?? []), note] })
624 }).catch(($, e, next) => next(e))
625
626 on('command.run', { command: 'hud' }, async ($, e) => runCommand($, e.args))
627
628 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
629 if (e.props.hasSurvey || (await read($, bandHiddenAtom))) {
630 return next(e)
631 }
632
633 await read($, tickAtom)
634 const view = await viewOf($)
635 return hasContent(view) ? band($.ui.resolve(e), view, e.props.bodyColumns) : next(e)
636 })
637
638 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
639 await read($, tickAtom)
640 const actions = {
641 refresh: () => requestRefresh(true, true),
642 toggleGuard: () => void toggleGuard($),
643 toggleBand: () => void toggleBand($),
644 }
645
646 return pane($.ui.resolve(e), await viewOf($), actions, await read($, bandHiddenAtom))
647 })
648}
649hooks/parse.ts 144 lines1// Pure parsers: TaskTracker MCP texts, dotnet test/build output, .sln files and shell commands.
2// No `$` here, so every function is unit-tested in parse.test.ts.
3
4import type { HudCriterion, HudTask } from '../types'
5
6const UUID = '[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}'
7
8/** Capture group `i` of a match; '' when it did not take part. */
9function group(m: RegExpMatchArray, i: number): string {
10 return m[i] ?? ''
11}
12
13/** `Task "<title>" (id <uuid>) v<n> [<type>/<status>/<priority>]`, the head of getTask's answer. */
14export function parseTask(text: string): HudTask | null {
15 const m = new RegExp(`^Task "(.*)" \\(id (${UUID})\\) v\\d+ \\[([a-z_]+)/([a-z_]+)`, 'm').exec(text)
16 return m ? { title: group(m, 1), id: group(m, 2), type: group(m, 3), status: group(m, 4) } : null
17}
18
19/** The `- [<type>/<status>] <title> (id <uuid>)` lines of getChildTasks and getTaskAncestors. */
20export function parseTaskList(text: string): HudTask[] {
21 const line = new RegExp(`^\\s*-\\s+\\[([a-z_]+)/([a-z_]+)\\]\\s+(.*?)\\s+\\(id (${UUID})\\)\\s*$`, 'gm')
22 return [...text.matchAll(line)].map(m => ({ type: group(m, 1), status: group(m, 2), title: group(m, 3), id: group(m, 4) }))
23}
24
25/** The `- [<status>/<priority>] <slug>: <title> (id <uuid>)` lines of listRequirements. */
26export function parseRequirements(text: string): { id: string; slug: string }[] {
27 const line = new RegExp(`^\\s*-\\s+\\[[a-z_]+/[a-z_]+\\]\\s+([^:\\s]+):.*\\(id (${UUID})\\)\\s*$`, 'gm')
28 return [...text.matchAll(line)].map(m => ({ slug: group(m, 1), id: group(m, 2) }))
29}
30
31/** The `- [x] (<n>) <text> (id <uuid>)` lines of listAcceptanceCriteria; any mark but a space is satisfied. */
32export function parseCriteria(text: string): HudCriterion[] {
33 const line = new RegExp(`^\\s*-\\s+\\[(.)\\]\\s+\\(\\d+\\)\\s+(.*?)\\s+\\(id ${UUID}\\)\\s*$`, 'gm')
34 return [...text.matchAll(line)].map(m => ({ isSatisfied: group(m, 1) !== ' ', text: group(m, 2) }))
35}
36
37/** Whether listRequirementTaskLinks' answer links the requirement to task `taskId`. */
38export function linksTask(text: string, taskId: string): boolean {
39 return new RegExp(`\\btask ${taskId}\\b`).test(text)
40}
41
42export type TestSummary = { project: string; failed: number; passed: number; skipped: number }
43
44/** dotnet test's `Passed!/Failed! - Failed: n, Passed: n, Skipped: n, Total: n, Duration: … - X.dll` lines, last per project. */
45export function parseTestSummaries(text: string): TestSummary[] {
46 const line = /^[^\S\n]*(?:\d+[:-])?\s*(?:Passed!|Failed!)\s+-\s+Failed:\s+(\d+),\s+Passed:\s+(\d+),\s+Skipped:\s+(\d+),\s+Total:\s+\d+,.*?-\s+([^\s\\/]+?)\.dll\b/gm
47 const byProject = new Map<string, TestSummary>()
48 for (const m of text.matchAll(line)) {
49 byProject.set(group(m, 4), { project: group(m, 4), failed: Number(group(m, 1)), passed: Number(group(m, 2)), skipped: Number(group(m, 3)) })
50 }
51
52 return [...byProject.values()].sort((a, b) => (a.project < b.project ? -1 : a.project > b.project ? 1 : 0))
53}
54
55/** The projects dotnet test started: `Test run for <path>/<X>.dll (…)`. */
56export function parseTestRuns(text: string): string[] {
57 return [...new Set([...text.matchAll(/Test run for .*?[\\/]([^\\/\s]+?)\.dll\b/g)].map(m => group(m, 1)))].sort()
58}
59
60/** MSBuild's closing `n Warning(s)` / `n Error(s)` lines (the last of each). */
61export function parseBuildSummary(text: string): { warnings: number; errors: number } | null {
62 const warnings = [...text.matchAll(/^\s*(\d+) Warning\(s\)\s*$/gm)].at(-1)
63 const errors = [...text.matchAll(/^\s*(\d+) Error\(s\)\s*$/gm)].at(-1)
64 return warnings && errors ? { warnings: Number(group(warnings, 1)), errors: Number(group(errors, 1)) } : null
65}
66
67/** The test projects a .sln lists: projects whose name contains "Test". */
68export function parseSlnTestProjects(sln: string): string[] {
69 const line = /^Project\("\{[^}]+\}"\)\s*=\s*"([^"]+)",\s*"([^"]+\.(?:cs|fs|vb)proj)"/gm
70 return [...sln.matchAll(line)].map(m => group(m, 1)).filter(name => /test/i.test(name)).sort()
71}
72
73/** What a shell command runs, for the check alarms. */
74export type CommandShape = {
75 isDotnetTest: boolean
76 isBuild: boolean
77 isVerify: boolean
78 /** The solution `dotnet test` names, resolved against a leading `cd <dir> &&`; null when none. */
79 sln: string | null
80 /** A filter, a pipe or a redirect: the output may not show every project. */
81 isPartial: boolean
82}
83
84/** Reads a Bash or PowerShell command line. */
85export function shapeOf(command: string): CommandShape {
86 const test = /\bdotnet\s+test\b([^|&;>]*)/.exec(command)
87 const testArgs = test === null ? '' : group(test, 1)
88 let sln: string | null = null
89 const named = /(\S+\.slnx?)\b/.exec(testArgs)
90 if (named) {
91 sln = group(named, 1).replace(/^["']|["']$/g, '')
92 const cd = /^\s*cd\s+(["']?)([^"'&;]+?)\1\s*&&/.exec(command)
93 if (cd && !isAbsolute(sln)) {
94 sln = `${group(cd, 2).replace(/[\\/]+$/, '')}/${sln}`
95 }
96
97 sln = fromGitBash(sln)
98 }
99
100 return {
101 isDotnetTest: test !== null,
102 isBuild: /\bdotnet\s+build\b/.test(command),
103 isVerify: /verify\.ps1\b/.test(command),
104 sln,
105 isPartial: test !== null && (/--filter\b/.test(testArgs) || /[|>]/.test(command.slice(test.index))),
106 }
107}
108
109function isAbsolute(path: string): boolean {
110 return /^([a-zA-Z]:)?[\\/]/.test(path)
111}
112
113/** `/c/projects/x` (Git Bash) as `C:/projects/x`; other paths unchanged. */
114export function fromGitBash(path: string): string {
115 return path.replace(/^\/([a-zA-Z])\//, (_, drive: string) => `${drive.toUpperCase()}:/`)
116}
117
118/** Whether a Write/Edit path is outside the orchestrator guard: Claude's own folders (memory, plans, scratchpad). */
119export function isGuardExempt(path: string): boolean {
120 const p = path.replace(/\\/g, '/').toLowerCase()
121 return p.includes('/.claude/') || p.includes('/appdata/local/temp/claude/') || p.startsWith('/tmp/claude')
122}
123
124/** `12m`, `1h05m`, `40s`. */
125export function formatDuration(ms: number): string {
126 const s = Math.max(0, Math.round(ms / 1000))
127 if (s < 60) return `${s}s`
128 const m = Math.floor(s / 60)
129 if (m < 60) return `${m}m`
130 return `${Math.floor(m / 60)}h${String(m % 60).padStart(2, '0')}m`
131}
132
133/** `▮▮▯▯` for done of total, `width` cells. */
134export function progressBar(done: number, total: number, width = 5): string {
135 if (total <= 0) return ''
136 const full = Math.round((done / total) * width)
137 return '▮'.repeat(full) + '▯'.repeat(width - full)
138}
139
140/** A title cut to `max` characters with an ellipsis. */
141export function short(text: string, max: number): string {
142 return text.length <= max ? text : `${text.slice(0, Math.max(1, max - 1))}…`
143}
144hooks/state.ts 10 lines1// Shared constants and pure helpers. The state atoms themselves are declared in register.tsx: the engine
2// reads `$.state` references only where they are written in the file that uses them.
3
4export const PANE = 'devflow-hud'
5
6/** Statuses of an agent whose loop still works (a HudAgent or a $.agent.list() row). */
7export function isRunning(agent: { status: string }): boolean {
8 return agent.status === 'pending' || agent.status === 'running' || agent.status === 'waiting'
9}
10hooks/view.tsx 187 lines1// The band above the prompt (one line) and the /hud pane: drawn from the session state alone.
2
3import type { ElementTable } from 'claude-code'
4
5import type { HudAgent, HudCheck, HudGuard, HudPhase } from '../types'
6import { formatDuration, progressBar, short } from './parse'
7import { isRunning } from './state'
8
9export type HudView = {
10 phase: HudPhase | null
11 agents: HudAgent[]
12 checks: HudCheck[]
13 guard: HudGuard
14 now: number
15}
16
17/** Done and total of the phase's direct sub-tasks (archived ones are not listed by TaskTracker). */
18function subtaskProgress(phase: HudPhase): { done: number; total: number } {
19 return { done: phase.subtasks.filter(t => t.status === 'completed').length, total: phase.subtasks.length }
20}
21
22function acProgress(phase: HudPhase): { done: number; total: number } {
23 const all = phase.requirements.flatMap(r => r.criteria)
24 return { done: all.filter(c => c.isSatisfied).length, total: all.length }
25}
26
27function checkLabel(check: HudCheck): string {
28 if (check.label === 'build') return `build ${check.warnings ?? 0}w/${check.errors ?? 0}e`
29 if (check.label === 'verify') return 'verify.ps1'
30 return `tests ${check.failed}✗/${check.passed}✓`
31}
32
33/** Whether there is anything to show. */
34export function hasContent(view: HudView): boolean {
35 return view.phase !== null || view.agents.some(isRunning) || view.checks.length > 0 || (view.guard.isActive && view.guard.isEnabled)
36}
37
38/** The one-line band, drawn with the surface's element table. */
39export function band(elements: ElementTable, view: HudView, columns: number) {
40 const { Box, Text } = elements
41 const { phase, agents, checks, guard, now } = view
42 const running = agents.filter(isRunning)
43 const first = running[0]
44 const last = checks.at(-1)
45 const titleRoom = Math.max(12, Math.floor(columns / 4))
46
47 return (
48 <Box flexDirection="row">
49 <Text wrap="truncate-end">
50 {phase !== null && (
51 <Text color={phase.isIdle ? 'subtle' : 'claude'} bold={!phase.isIdle}>
52 ◆ {short(phase.phase.title, titleRoom)}
53 </Text>
54 )}
55 {phase?.active && <Text dimColor> › {short(phase.active.title, titleRoom)}</Text>}
56 {phase !== null && phase.subtasks.length > 0 && (
57 <Text>
58 {' '}
59 {progressBar(subtaskProgress(phase).done, subtaskProgress(phase).total)} {subtaskProgress(phase).done}/
60 {subtaskProgress(phase).total}
61 </Text>
62 )}
63 {phase !== null && acProgress(phase).total > 0 && (
64 <Text color={acProgress(phase).done === acProgress(phase).total ? 'success' : undefined}>
65 {' '}AC {acProgress(phase).done}/{acProgress(phase).total}
66 </Text>
67 )}
68 {first !== undefined && (
69 <Text color="suggestion">
70 {' '}⚙ {running.length > 1 ? `${running.length}× ` : ''}
71 {short(first.type, 14)} "{short(first.description, 28)}" {formatDuration(now - first.startedAt)}
72 </Text>
73 )}
74 {last !== undefined && (
75 <Text color={last.isOk ? 'success' : 'error'}>
76 {' '}
77 {last.isOk ? '✔' : '✗'} {checkLabel(last)}
78 <Text dimColor> {formatDuration(now - last.at)}</Text>
79 </Text>
80 )}
81 {last !== undefined && last.missing.length > 0 && <Text color="warning"> ⚠ {last.missing.length} not run</Text>}
82 {guard.isActive && guard.isEnabled && <Text color="permission">{' '}◈ orchestrator</Text>}
83 </Text>
84 </Box>
85 )
86}
87
88const STATUS_GLYPH: Record<string, string> = {
89 completed: '✔',
90 in_progress: '▶',
91 pending: '·',
92 blocked: '⛔',
93}
94
95/** The /hud pane: the phase tree, every AC, the agents, the checks and the guard. */
96export function pane(
97 elements: ElementTable,
98 view: HudView,
99 actions: { refresh: () => void; toggleGuard: () => void; toggleBand: () => void },
100 isBandHidden: boolean,
101) {
102 const { Box, Text, Button } = elements
103 const { phase, agents, checks, guard, now } = view
104
105 return (
106 <Box flexDirection="column">
107 <Box flexDirection="row" gap={1}>
108 <Button key="refresh" hotkey="r" label="Refresh" onPress={actions.refresh} />
109 <Button key="guard" hotkey="g" label={guard.isEnabled ? 'Guard: on' : 'Guard: off'} onPress={actions.toggleGuard} />
110 <Button key="band" hotkey="b" label={isBandHidden ? 'Band: hidden' : 'Band: shown'} onPress={actions.toggleBand} />
111 </Box>
112
113 <Text bold> </Text>
114 {phase === null ? (
115 <Text dimColor>No TaskTracker phase seen yet. It appears when a task is activated (setActiveTask), or press Refresh.</Text>
116 ) : (
117 <Box flexDirection="column">
118 <Text bold color="claude">
119 {STATUS_GLYPH[phase.phase.status] ?? '?'} {phase.phase.title} <Text dimColor>({phase.phase.status}{phase.isIdle ? ', no active task' : ''})</Text>
120 </Text>
121 {phase.error && <Text color="error"> last refresh failed: {phase.error}</Text>}
122 {phase.subtasks.map(task => (
123 <Text color={phase.active?.id === task.id ? 'suggestion' : undefined} dimColor={task.status === 'completed'}>
124 {' '}
125 {STATUS_GLYPH[task.status] ?? '?'} {task.title}
126 {phase.active?.id === task.id ? ' ← active' : ''}
127 </Text>
128 ))}
129 {phase.requirements.length === 0 && (
130 <Text dimColor> No linked requirements read (set the projectId option, or let a TaskTracker call with a projectId pass).</Text>
131 )}
132 {phase.requirements.map(requirement => (
133 <Box flexDirection="column">
134 <Text bold>
135 {' '}
136 {requirement.slug} {requirement.criteria.filter(c => c.isSatisfied).length}/{requirement.criteria.length}
137 </Text>
138 {requirement.criteria.map(criterion => (
139 <Text wrap="truncate-end" dimColor={criterion.isSatisfied} color={criterion.isSatisfied ? 'success' : undefined}>
140 {' '}
141 {criterion.isSatisfied ? '[x]' : '[ ]'} {criterion.text}
142 </Text>
143 ))}
144 </Box>
145 ))}
146 <Text dimColor> read {formatDuration(now - phase.refreshedAt)} ago</Text>
147 </Box>
148 )}
149
150 <Text bold> </Text>
151 <Text bold>Agents</Text>
152 {agents.length === 0 && <Text dimColor> none this session</Text>}
153 {agents.slice(-12).map(agent => (
154 <Text color={isRunning(agent) ? 'suggestion' : agent.status === 'completed' ? undefined : 'error'} dimColor={!isRunning(agent) && agent.status === 'completed'}>
155 {' '}
156 {isRunning(agent) ? '⚙' : agent.status === 'completed' ? '✔' : '✗'} {agent.type} "{agent.description}" {agent.status}{' '}
157 {formatDuration((agent.endedAt ?? now) - agent.startedAt)}
158 </Text>
159 ))}
160
161 <Text bold> </Text>
162 <Text bold>Checks</Text>
163 {checks.length === 0 && <Text dimColor> none read yet (dotnet build/test and verify.ps1 output is read from shell commands)</Text>}
164 {checks.slice(-10).map(check => (
165 <Text color={check.isOk ? undefined : 'error'}>
166 {' '}
167 {check.isOk ? '✔' : '✗'} {checkLabel(check)} <Text dimColor>({check.loop}, {formatDuration(now - check.at)} ago{check.projects > 0 ? `, ${check.projects} projects` : ''})</Text>
168 {check.missing.length > 0 && <Text color="warning"> ⚠ not run: {check.missing.join(', ')}</Text>}
169 {check.summary && <Text dimColor> {check.summary}</Text>}
170 </Text>
171 ))}
172
173 <Text bold> </Text>
174 <Text>
175 <Text bold>Guard </Text>
176 {!guard.isEnabled ? (
177 <Text dimColor>off (/hud guard on)</Text>
178 ) : guard.isActive ? (
179 <Text color="permission">active: {guard.skill} leads the main thread; its Write/Edit/NotebookEdit are refused</Text>
180 ) : (
181 <Text dimColor>armed; turns on when an orchestrator skill is invoked</Text>
182 )}
183 </Text>
184 </Box>
185 )
186}
187types/index.d.ts 67 lines1/** A TaskTracker task as the HUD shows it. */
2export type HudTask = { id: string; title: string; type: string; status: string }
3
4/** One acceptance criterion of a requirement linked to the phase. */
5export type HudCriterion = { text: string; isSatisfied: boolean }
6
7/** A requirement linked to the phase, with its acceptance criteria. */
8export type HudRequirement = { id: string; slug: string; criteria: HudCriterion[] }
9
10/** The phase the active task belongs to, as last read from TaskTracker. */
11export type HudPhase = {
12 phase: HudTask
13 /** The active task when it is a task under the phase; null when the phase itself is active. */
14 active: HudTask | null
15 subtasks: HudTask[]
16 requirements: HudRequirement[]
17 refreshedAt: number
18 /** True when no task is active any more (the view is the last one seen). */
19 isIdle: boolean
20 error?: string
21}
22
23/** A subagent of this session. */
24export type HudAgent = {
25 id: string
26 description: string
27 type: string
28 status: string
29 startedAt: number
30 endedAt?: number
31}
32
33/** A build, test or verify result read from a shell command's output. */
34export type HudCheck = {
35 label: string
36 isOk: boolean
37 failed: number
38 passed: number
39 warnings?: number
40 errors?: number
41 /** Test projects of the solution that reported no result. */
42 missing: string[]
43 projects: number
44 at: number
45 /** "main" or the agent's description. */
46 loop: string
47 summary?: string
48}
49
50/** The orchestrator guard: on while an orchestrator skill leads the main thread. */
51export type HudGuard = { isEnabled: boolean; isActive: boolean; skill: string }
52
53declare module 'claude-code' {
54 interface PluginState {
55 devflow: {
56 phase: HudPhase | null
57 activeTaskId: string | null
58 agents: HudAgent[]
59 checks: HudCheck[]
60 guard: HudGuard
61 isBandHidden: boolean
62 /** Bumped every 30 s while agents run, so elapsed times redraw. */
63 tick: number
64 }
65 }
66}
67