SLOPSHOPPER

DevFlow Skills Collection

A verification-first development pipeline for Claude Code: brainstorm to plan to phased implementation, with quality gates (verification-loop, code-review…

newpanebandguardcommandtoast
★ 20v2.1.0MITupdated 2026-10-09mhylle/claude-skills-collection
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · devflow
│ ┃ devflow HUD ✕ › fix the failing auth test and add an audit log call │ ┃ [ Refresh ] [ Guard: on ] [ Band: shown ] │ ┃ ⏺ Read(src/auth.ts) │ ┃ No TaskTracker phase seen yet. It appears ⎿ Read 6 lines │ ┃ when a task is activated (setActiveTask), or ⏺ Update(src/auth.ts) │ ┃ press Refresh. ⎿ Added 2 lines, removed 1 line │ ┃ ⏺ Bash(bun test) │ ┃ Agents ⎿ 3 pass, 1 fail │ ┃ none this session │ ┃ ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ Checks │ ┃ none read yet (dotnet build/test and ✻ Worked for 42s · done 4:20 PM │ ┃ verify.ps1 output is read from shell │ ┃ commands) › /hud │ ┃ ⎿ devflow: devflow HUD pane opened. │ ┃ Guard armed; turns on when an orchestrator │ ┃ skill is invoked │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · devflow HUD
[ Refresh ] [ Guard: on ] [ Band: shown ] No TaskTracker phase seen yet. It appears when a task is activated (setActiveTask), or press Refresh. Agents none this session Checks none read yet (dotnet build/test and verify.ps1 output is read from shell commands) Guard armed; turns on when an orchestrator skill is invoked
README

Claude Code Skills Collection

Custom skills and agents for Claude Code that enhance codebase research, context management, and implementation planning workflows.

New to this workflow? See the Workflow Overview for a step-by-step guide.

Installing? This repo ships as the devflow Claude Code plugin — see Installation and Plugin Distribution.

Built on Claude Code Task Tools

This skill collection uses Claude Code's native Task tools for progress tracking.

Progress tracking is handled entirely through Claude Code's built-in Task system:

ToolPurpose
TaskCreateCreate tasks for each phase with dependencies
TaskUpdateMark tasks as in_progress or completed
TaskListView all tasks with status and blockers
TaskGetGet full task details including description

Benefits:

  • Persistent - Tasks survive session restarts
  • Cross-session - Share progress across multiple terminals
  • Dependency tracking - Blocked tasks visible, execute in order
  • No file pollution - Progress tracked in memory, not plan files

Plans remain pure specification documents. See Progress Tracking for details.

Claude Code 2.1.x Feature Alignment

This collection uses modern Claude Code skill features (v2.1.16+):

FeatureSkills Using ItPurpose
context: forkcodebase-researchRun in isolated subagent context. Only for skills that run to completion without the user and don't wait on agents of their own: agent reports arrive in the top-level session, so the phase leads (implement-phase, tt-implement-phase) run inline.
agent: Explore/Plan(none)Subagent type for a forked skill. Both built-ins are read-only and lack the Agent tool, so any skill that writes files or spawns subagents must leave this unset.
allowed-toolscode-review, verification-loop, security-review, adversarial-reviewer, codebase-research, strategic-compactRestrict available tools (read-only enforcement)
argument-hintimplement-plan, implement-phase, adr, e2e-testing, code-review, adversarial-reviewer, context-saver, prompt-generatorShow usage hints in autocomplete
disable-model-invocationcontext-saver, prompt-generatorUser-only invocation (no auto-trigger)
user-invocable: falseimplement-phaseHide from user menu (internal skill)

Never add context: fork to an interactive skill. A forked skill runs in a subagent with no access to the conversation history, and since Claude Code v2.1.218 it is backgrounded by default — so a skill that asks the user questions and waits for answers (brainstorm) simply never reaches the user. agent: Explore compounds it: that agent type is read-only, so the skill cannot write its output document either. tests/test-interactive-skills.sh guards this.

Argument Substitution

Skills support the new argument syntax:

  • $0, $1, $2 - Positional arguments
  • $ARGUMENTS - All arguments
  • ${CLAUDE_SESSION_ID} - Session tracking

Example: /implement-plan docs/plans/my-feature.md passes the path as $0.

Implementation Workflow

The core workflow for implementing features follows this hierarchy:

                              ┌─────────────────┐
                              │   brainstorm    │
                              │   (ideation)    │
                              └────────┬────────┘
                                       │
                           ┌───────────┴───────────┐
                           ▼                       ▼
                  ┌─────────────────┐   ┌──────────────────┐
                  │      adr        │   │   user-story     │
                  │  (decisions)    │   │ (requirements)   │
                  └────────┬────────┘   └──────────┬───────┘
                           └────────────┬──────────┘
                                       │
                                       ▼
                              ┌─────────────────┐
                              │  create-plan    │
                              │ (spec → phases) │
                              └────────┬────────┘
                                       │
                                       ▼
┌──────────────────────────────────────────────────────────────────────────────┐
│                              implement-plan                                   │
│                           (orchestrates all phases)                           │
│                                                                               │
│   ┌─────────────┐    ┌─────────────┐    ┌─────────────┐                      │
│   │   Phase 1   │───▶│   Phase 2   │───▶│   Phase N   │───▶ Complete        │
│   └──────┬──────┘    └──────┬──────┘    └──────┬──────┘                      │
│          │                  │                  │                              │
│          ▼                  ▼                  ▼                              │
│   ┌─────────────────────────────────────────────────────────────────────┐    │
│   │                    implement-phase (per phase)                       │    │
│   │  ┌────────────────────────────────────────────────────────────┐     │    │
│   │  │ Step 1: Implementation (subagents) [TDD mode: tests first] │     │    │
│   │  │ Step 2: verification-loop (6-phase exit conditions) ─────┐│     │    │
│   │  │ Step 3: Integration Testing (API/UI via Playwright)      ││     │    │
│   │  │ Step 4: code-review ─────────────────────────────────────┼┼──┐  │    │
│   │  │         ├─► security-review (optional OWASP audit) ──────┼┼──┼─┐│    │
│   │  │ Step 5: ADR Compliance ──────────────────────────────────┼┼──┼─┼┤    │
│   │  │ Step 6: Plan Sync                                        ││  │ ││    │
│   │  │ Step 7: Prompt Archival                                  ││  │ ││    │
│   │  │ Step 8: Completion Report                                ││  │ ││    │
│   │  └──────────────────────────────────────────────────────────┼┼──┼─┼┘    │
│   └─────────────────────────────────────────────────────────────┼┼──┼─┼─────┘
│                                                                 ││  │ │      │
│         ┌───────────────────────────────────────────────────────┘│  │ │      │
│         │              ┌─────────────────────────────────────────┘  │ │      │
│         │              │              ┌──────────────────────────────┘ │      │
│         │              │              │              ┌─────────────────┘      │
│         ▼              ▼              ▼              ▼                        │
│   ┌───────────┐  ┌───────────┐  ┌───────────┐  ┌─────────────────┐           │
│   │verification│  │code-review│  │security-  │  │      adr        │           │
│   │   -loop   │  │           │  │  review   │  │                 │           │
│   └───────────┘  └───────────┘  └───────────┘  └─────────────────┘           │
│    (default)                      (optional)                                  │
└───────────────────────────────────────────────────────────────────────────────┘
                                       │
                                       ▼
                              ┌─────────────────┐
                              │   e2e-testing   │
                              │  (validation)   │
                              └────────┬────────┘
                                       │
                                       ▼
                        ┌──────────────────────────┐
                        │   continuous-learning    │
                        │ (extract session patterns)│
                        └──────────────────────────┘

Workflow Stages

StageSkillPurpose
IdeationbrainstormRefine rough ideas through Socratic questioning
Requirementsuser-storyGenerate hierarchical user stories with acceptance criteria
Planningcreate-planCreate detailed, phased implementation plans
Iterationiterate-planUpdate plans based on feedback
Executionimplement-planOrchestrate full plan execution
Phase Workimplement-phaseExecute single phase with quality gates
Qualitycode-reviewVerify code quality, patterns, ADR compliance
Adversarial Qualityadversarial-reviewerSubagent-based hostile review (Saboteur, New Hire, Security Auditor) to break self-review blind spots
Comprehensive Auditcodebase-auditLong-running full-codebase audit — partitions the repo, delegates to adversarial-reviewer per partition, synthesizes a written remediation report
Securitysecurity-reviewOWASP-aligned security audit (optional step)
Verificationverification-loop6-phase verification: build, type, lint, test, security, diff
Metrics Gatecode-quality-auditCoverage, complexity, module size, deps, mutation — gate or on-demand
DecisionsadrDocument architectural decisions
Testinge2e-testingEnd-to-end validation with Playwright
Evaluationeval-harnessFormal capability/regression testing with metrics
Learningcontinuous-learningExtract patterns from sessions for reuse

Skills

Skills are invoked via the Skill tool or /skill-name shorthand.

Planning & Implementation

SkillTriggerDescription
brainstorm/brainstorm, "explore this idea"Interactive idea refinement using Socratic questioning, written to docs/brainstorms/ for create-plan or tt-create-plan
user-story/user-story, "create user stories"Generate hierarchical user stories (epics/features/tasks) with Given/When/Then acceptance criteria
create-plan/create-plan, "plan the implementation"Creates detailed implementation plans through research
iterate-plan"update the plan", "iterate on this plan"Updates existing plans based on feedback
implement-plan/implement-plan, "implement the plan"Orchestrates execution of complete plans (subagent mode)
implement-phaseCalled by implement-planExecutes single phase with all quality gates
tt-create-plan/tt-create-plan, "plan this in tasktracker"TaskTracker-native planning: requirements with acceptance criteria and phase tasks from templates instead of a docs/plans/*.md file
tt-implement-plan/tt-implement-plan, "execute the tasktracker plan"Runs a TaskTracker plan phase by phase, delegating each phase to tt-implement-phase, with active-task time tracking and insight logging
tt-workflow-run/tt-workflow-run, "drive the backlog to done"Autonomous gated run over a TaskTracker backlog: next ready slice → tt-implement-phase → drift and defect gates → measured-time projection, until the backlog drains
tt-workflow-build/tt-workflow-build, "parallel build (tasktracker)"Parallel build of a TaskTracker project, on the Workflow tool when it is enabled and on parallel subagents otherwise, driven to done
tt-create-build-loop/tt-create-build-loop, "set up an autonomous build loop"Writes prompts/autonomous-build-loop.md, the prompt a /loop re-reads every iteration to plan, implement and verify TaskTracker phases unattended under a zero-error gate, with a human-only queue and a per-phase cost ledger. Detects the repo, stack, CI and TaskTracker phases, asks only what it can't, creates missing queue/ledger phases, optionally installs the bundled token-usage tooling, and installs a clean-worktree gate runner
workflow-guide/workflow-guide, "which workflow should I use"Routes work to a lane: file-based or TaskTracker pipeline, autonomous run, unattended loop, parallel build or audit, or issue-to-merge

Quality & Documentation

SkillTriggerDescription
code-review/devflow:code-review (bare /code-review is Claude Code's built-in review), Step 4 of implement-phase and tt-implement-phaseSystematic review: SRP, patterns, ADR compliance
adversarial-reviewer/adversarial-reviewer, "adversarial review", "critical review", "audit this repo"Spawns three hostile-persona subagents (Saboteur, New Hire, Security Auditor) in parallel; each must find ≥1 issue; cross-persona findings get severity-promoted. Default mode reviews a diff; --codebase [path] reviews a whole repo/subtree with strategic per-persona deep-dives
grumpy-reviewer/grumpy-reviewer, "structural review", "is this well factored"One isolated-subagent reviewer that judges only maintainability (separation of concerns, helper extraction, small files, the rule of 7) and never learns how the code was produced
codebase-audit/codebase-audit, "comprehensive codebase review", "thorough audit", "code due diligence"Long-running full-coverage audit. Partitions the repo, delegates to /adversarial-reviewer --codebase per partition, synthesizes systemic findings, produces written remediation report. Resumable. Pairs with code-quality-audit for qualitative + quantitative picture
tt-workflow-audit/tt-workflow-audit, "parallel audit (tasktracker)"Read-only parallel audit of a TaskTracker project's repo, backlog or architecture: a ranked risk register, with fix tasks written back by the parent on approval. Resumable
adr/adr, "document decision"Creates Architecture Decision Records
e2e-testing/e2e-testing, "test my webapp"E2E testing with Playwright MCP
security-review/devflow:security-review (bare /security-review is the built-in), auth/input code10-category OWASP-aligned security audit
verification-loop/verification-loop, "verify implementation"6-phase verification: build, type, lint, test, security, diff
code-quality-audit/code-quality-audit, "audit code quality", "run mutation testing"Coverage + complexity + module size + dependency cycles + mutation score. Gate or on-demand modes
eval-harness/eval-harness, "run evals"Formal evaluation framework with pass@k metrics

Research & Context

SkillTriggerDescription
codebase-research"how does X work"Parallel codebase research with sub-agents
context-saver/context-saver, "save context"Preserves session state for continuation
prompt-generator/prompt, "generate prompt"Creates implementation prompts for phases

Learning & Optimization

SkillTriggerDescription
continuous-learning"save what we learned", /continuous-learningPromotes reusable procedures from a session into learned skills at ~/.claude/skills/learned-<slug>/ (on demand; one-liners go to auto-memory)
strategic-compactPreToolUse hookSuggests /compact at logical boundaries, not arbitrary thresholds
skill-visualizer/skill-visualizer, "visualize skills"Generate interactive HTML visualizations of skills and codebase

Development

SkillTriggerDescription
agent-creator"create agent", "build agent"Creates composable AI agent systems in NestJS
design-language"design my own design language", "build my design system", "I want my own look not the Claude default"Guided design-brief → living-styleguide compiler. Reaction-first elicitation (web research + real references + live archetypes rendered in your own tokens → you react → targeted refinement) with a per-dimension 1–5 "safeness" gradient (conventional→experimental) across the entire system (color, type, spacing, shape, elevation, motion, imagery, components, patterns, voice). Resumable multi-phase flow; emits a self-contained, portable ./design-system/ (interactive dashboard + css/tokens.css single-source-of-truth + W3C design-tokens.json + DESIGN_LANGUAGE.md contract with an explicit anti-Claude reference + pre-ship checklist + a default-vs-yours compare.html) that future sessions read and build from. Different from frontend-design (which styles one UI in the moment).

Media & Content

SkillTriggerDescription
video-explainer/video-explainer <url-or-path>, "make an HTML explainer for this video", "turn this talk into visual notes"Turns a video (YouTube, any yt-dlp URL, or a local file) into one self-contained HTML explainer page: headline + TL;DR, key takeaways, a section per part of the video, and diagrams built in pure HTML + CSS (17 components: flows, timelines, bar/line/dumbbell charts, threshold scales, comparisons, cycles, layers…). Watches the video via the watch plugin (frames + transcript), adds screenshots only when the video's own image is the information, renders the page headless at desktop and phone width to check its own layout, and links every claim to its timestamp. Requires the watch plugin: see the Video Explainer Guide for install and usage.

Pipeline (Issue → Merge)

SkillTriggerDescription
ship-issue/ship-issue <issue-number-or-url>One-command GitHub issue → merged PR pipeline across nine stages (preflight, plan, implement, review, ci, cloud_review, deploy, e2e, logs) with exactly two human gates — plan approval and merge confirmation. Model-tiered: Fable 5 plans, orchestrates, and runs the merge-gate review; Opus 4.8 implements (TDD); Sonnet 4.6 runs staging E2E and log checks. File-based run state gives lossless crash-resume; per-stage time tracking (work / gate-wait / crash-gap, with a per-model-tier rollup) is embedded in the Gate 2 merge brief. Pair with the single-file dashboard.py for a live view.

Overview page: open skills/ship-issue/pipeline.html in a browser for a one-page tour — the skill, its five agents, and the full stage-flow diagram with model-tier colour-coding (self-contained, no dependencies).

Agents

Agents are specialized subagents launched with the Agent tool. When the collection is installed as a plugin they are named devflow:<agent>.

AgentPurpose
implementersonnet, medium effort: writes code and tests for one briefed unit of work; used by tt-implement-phase for implementation and fix rounds
mechanichaiku, low effort: runs build/type/lint/test checks and verification-loop, report-only, never edits
revieweropus, high effort (callers pass fable for security reviews): runs code-review / security-review on code it did not write
codebase-analyzerTraces implementation with file:line references
codebase-locatorFinds files by topic/feature ("Super Grep/Glob")
codebase-pattern-finderFinds concrete code examples and patterns
docs-analyzerExtracts insights from docs, ADRs, design docs
docs-locatorFinds documentation and research notes
web-search-researcherWeb research for APIs, libraries, troubleshooting
browser-verification-agentUI testing via Playwright MCP with screenshot evidence
design-researcherclaude-sonnet-4-6 — per-dimension inspiration research for the design-language skill (WebSearch → Playwright screenshots → cached manifest), steering away from the default Claude look

ship-issue pipeline agents

Each agent pins its model in frontmatter — the tier is part of the contract, never switched mid-task (prompt caches are model-scoped; ADR-0007). A fix cycle is always a fresh task on the same tier, never a resumed task on a different model.

AgentModelPurpose
issue-plannerclaude-fable-5 (Fable 5)Turns a GitHub issue + codebase into the plan presented at Gate 1. Outcome-prompted (ADR-0008).
merge-gate-reviewerclaude-fable-5 (Fable 5)Last-line diff review at the merge gate; verdict contract APPROVE / FIX (itemized blockers). Outcome-prompted.
tdd-implementerclaude-opus-4-8 (Opus 4.8)Tests-first implementation, UI components, API routes; receives reviewer/CI/E2E blockers verbatim on fix cycles.
staging-e2e-verifierclaude-sonnet-4-6 (Sonnet 4.6)Runs the plan's E2E scenarios against the live staging URL via Playwright MCP; PASS/FAIL/FLAKY/BLOCKED with screenshot evidence.
staging-log-verifierclaude-sonnet-4-6 (Sonnet 4.6)Scans staging service logs over the deploy window; CLEAN vs ERRORS_FOUND with cited lines.

Installation

This repository is distributed exclusively as a Claude Code plugin named devflow, published through a marketplace named mhylle. Inside Claude Code:

/plugin marketplace add mhylle/claude-skills-collection
/plugin install devflow@mhylle

Then /reload-plugins (or restart) if the install summary asks for it.

Plugin skills are namespaced, so every skill is invoked as /devflow:<skill>:

/devflow:brainstorm
/devflow:create-plan
/devflow:implement-plan

Nothing is copied into ~/.claude/ — the plugin lives in Claude Code's plugin cache, carries a version, and updates in place:

/plugin marketplace update mhylle

To remove it: /plugin uninstall devflow.

See Plugin Distribution for the manifest layout, versioning, publishing, and local development.

Migrating from the old install script

Earlier versions shipped an install.sh that copied skills, agents, and hooks into ~/.claude/. It is gone (ADR-0011).

Those copies do not disappear on their own, and plugin skills do not override same-named personal skills — so until you remove them you will have both /brainstorm and /devflow:brainstorm pointing at two copies that drift apart, with Claude free to auto-invoke either. Remove the copies once, after installing the plugin:

# from a checkout of this repo
for s in $(ls skills); do rm -rf "$HOME/.claude/skills/$s"; done
for a in agents/*.md; do rm -f "$HOME/.claude/agents/$(basename "$a")"; done

~/.claude/hooks.json is your own user configuration, not a copy of this repo's file. The old script overwrote it (leaving hooks.json.backup); the plugin no longer touches it. Review it by hand and delete only the entries you recognise from this collection.

What the Plugin Ships

Skills, agents, and these hooks — all discovered automatically from the plugin's default component locations (skills/, agents/, hooks/hooks.json).

Hooks (automatic behaviors):

/devflow:strategic-compact is invoked on demand rather than by a hook — see ADR-0011.

Hook TypeNameTriggerPurpose
PreToolUsetmux-dev-blocknpm run dev etc.Block dev servers outside tmux
PreToolUsetmux-reminderLong-running commandsSuggest tmux for session persistence
PreToolUsegit-push-reviewgit pushReminder to review before push
PreToolUsedoc-file-warn.md/.txt creationWarn about docs outside docs/ structure
PostToolUsepr-url-loggergh pr createLog PR URL and review command
PostToolUseprettier-formatJS/TS file editsAuto-format with Prettier
PostToolUsetypescript-check.ts/.tsx editsRun tsc --noEmit and show errors

| PostToolUse | console-log-warn | JS/TS file edits | Wa

Source 5 files
hooks/register.tsx 649 lines
1// devflow HUD (part of the devflow plugin): a progress band and a /hud pane for the devflow TaskTracker skills, and the functions
2// the skills lean on: the orchestrator guard, the in-flight watch, the time-tracking heartbeat and
3// pause, and the test-run alarms. See README.md.
4
5import { atom, read, update } from 'claude-code'
6import type { EngineInterface, Register, ToolCallResult } from 'claude-code'
7
8import type { HudAgent, HudCheck, HudGuard, HudPhase, HudRequirement, HudTask } from '../types'
9import {
10  formatDuration,
11  isGuardExempt,
12  linksTask,
13  parseBuildSummary,
14  parseCriteria,
15  parseRequirements,
16  parseSlnTestProjects,
17  parseTask,
18  parseTaskList,
19  parseTestRuns,
20  parseTestSummaries,
21  shapeOf,
22} from './parse'
23import { isRunning, PANE } from './state'
24import { band, hasContent, pane, type HudView } from './view'
25
26// Every function that takes `$`, and every state atom, lives in this file: the engine follows `$` and reads
27// `$.state` references only where they are declared in the file that uses them.
28
29// The session state the band and the pane draw from (declared in ../types, PluginState).
30const phaseAtom = atom({ plugin: 'devflow', key: 'phase' } as const, null as HudPhase | null)
31const activeTaskAtom = atom({ plugin: 'devflow', key: 'activeTaskId' } as const, null as string | null)
32const agentsAtom = atom({ plugin: 'devflow', key: 'agents' } as const, [] as HudAgent[])
33const checksAtom = atom({ plugin: 'devflow', key: 'checks' } as const, [] as HudCheck[])
34const guardAtom = atom({ plugin: 'devflow', key: 'guard' } as const, { isEnabled: true, isActive: false, skill: '' } as HudGuard)
35const bandHiddenAtom = atom({ plugin: 'devflow', key: 'isBandHidden' } as const, false)
36const tickAtom = atom({ plugin: 'devflow', key: 'tick' } as const, 0)
37
38/** How often the background tick polls the agents and runs a wanted refresh. */
39const TICK_MS = 2_000
40
41/** TaskTracker closes a time segment after ~5 minutes without a call; heartbeat inside that. */
42const HEARTBEAT_MS = 240_000
43
44/** TaskTracker tools after which the view is read again. */
45const WRITES = new Set([
46  'updateTaskStatus',
47  'batchUpdateStatus',
48  'createTask',
49  'batchCreateTasks',
50  'updateTask',
51  'batchUpdateTasks',
52  'archiveTask',
53  'deleteTask',
54  'reparentTask',
55  'completeWithCaveat',
56  'addAcceptanceCriterion',
57  'updateAcceptanceCriterion',
58  'deleteAcceptanceCriterion',
59])
60
61/** Writes that change which requirements a phase links (the cached links are read again). */
62const LINK_WRITES = new Set(['linkRequirementToTask', 'unlinkRequirementFromTask'])
63
64const PREFIX = 'mcp__tasktracker__tasktracker_'
65
66const USAGE = '/hud opens the pane. /hud refresh | guard on|off | band on|off | project <uuid>'
67
68/** Module state; a reload starts it over (what is drawn lives in $.state). */
69const hud = {
70  orchestrator: /implement-plan|implement-phase|tt-workflow-run/,
71  projectId: '',
72  /** After an automatic pause: no TaskTracker call of ours until the next turn, or it would book the wait. */
73  isQuiet: false,
74  isRefreshWanted: false,
75  isForcedRefresh: false,
76  isUserRefresh: false,
77  isTicking: false,
78  lastHeartbeat: 0,
79  lastRedraw: 0,
80  slnTests: new Map<string, string[]>(),
81}
82
83// TaskTracker, through the session's own MCP connection. Every call heartbeats the active task (TaskTracker
84// books time from calls), so none is made during a wait for the person (hud.isQuiet).
85
86/** One TaskTracker tool call; its text blocks joined. Rejects when the server reports an error. */
87async function call($: EngineInterface, tool: string, args: Record<string, unknown>): Promise<string> {
88  const result = await $.mcp.call('tasktracker', `tasktracker_${tool}`, args)
89  const text = result.content.map(block => block.text ?? '').join('\n')
90  if (result.isError) {
91    throw new Error(`${tool}: ${text.slice(0, 200)}`)
92  }
93
94  return text
95}
96
97/** The requirements linked to a phase, cached in $.store for an hour (links rarely change). */
98async function linkedRequirements($: EngineInterface, phaseId: string, isForced: boolean): Promise<{ id: string; slug: string }[]> {
99  const key = `links:${phaseId}`
100  const cached = (await $.store.get(key)) as { at: number; requirements: { id: string; slug: string }[] } | undefined
101  const now = await $.clock.now()
102  if (cached && !isForced && now - cached.at < 3_600_000) {
103    return cached.requirements
104  }
105
106  const all = parseRequirements(await call($, 'listRequirements', { projectId: hud.projectId, status: 'approved' }))
107  const linked: { id: string; slug: string }[] = []
108  for (const requirement of all) {
109    if (linksTask(await call($, 'listRequirementTaskLinks', { requirementId: requirement.id }), phaseId)) {
110      linked.push(requirement)
111    }
112  }
113
114  await $.store.set(key, { at: now, requirements: linked })
115  return linked
116}
117
118/** The phase of task `taskId` with its sub-tasks and (when the project is known) its requirements' ACs. */
119async function loadPhase($: EngineInterface, taskId: string, isForced: boolean): Promise<HudPhase> {
120  const task = parseTask(await call($, 'getTask', { taskId }))
121  if (task === null) {
122    throw new Error(`getTask ${taskId}: unexpected answer`)
123  }
124
125  let phase: HudTask = task
126  if (task.type !== 'phase') {
127    phase = parseTaskList(await call($, 'getTaskAncestors', { taskId })).find(t => t.type === 'phase') ?? task
128  }
129
130  const subtasks = parseTaskList(await call($, 'getChildTasks', { taskId: phase.id }))
131  const requirements: HudRequirement[] = []
132  if (hud.projectId !== '' && phase.type === 'phase') {
133    for (const requirement of await linkedRequirements($, phase.id, isForced)) {
134      const criteria = parseCriteria(await call($, 'listAcceptanceCriteria', { requirementId: requirement.id }))
135      requirements.push({ ...requirement, criteria })
136    }
137  }
138
139  return { phase, active: task.id === phase.id ? null : task, subtasks, requirements, refreshedAt: await $.clock.now(), isIdle: false }
140}
141
142function requestRefresh(isForced: boolean, isUser: boolean) {
143  hud.isRefreshWanted = true
144  hud.isForcedRefresh ||= isForced
145  hud.isUserRefresh ||= isUser
146}
147
148async function viewOf($: EngineInterface): Promise<HudView> {
149  return {
150    phase: await read($, phaseAtom),
151    agents: await read($, agentsAtom),
152    checks: await read($, checksAtom),
153    guard: await read($, guardAtom),
154    now: await $.clock.now(),
155  }
156}
157
158/** Reads the phase of the active task (or, on a user's refresh with none active, of the last phase seen). */
159async function refresh($: EngineInterface) {
160  const active = await read($, activeTaskAtom)
161  const shown = await read($, phaseAtom)
162  const target = active ?? (hud.isUserRefresh ? (shown?.phase.id ?? null) : null)
163  const wasQuiet = hud.isQuiet
164  const isForced = hud.isForcedRefresh
165  hud.isRefreshWanted = false
166  hud.isForcedRefresh = false
167  hud.isUserRefresh = false
168  if (target === null) {
169    return
170  }
171
172  try {
173    const phase: HudPhase = { ...(await loadPhase($, target, isForced)), isIdle: active === null }
174    await update($, phaseAtom, () => phase)
175    await $.store.set('lastPhase', phase)
176  } catch (error) {
177    const message = error instanceof Error ? error.message : String(error)
178    await update($, phaseAtom, p => (p === null ? p : { ...p, error: message }))
179  }
180
181  if (wasQuiet) {
182    // A refresh the person asked for during a wait: close the segment its calls opened.
183    await call($, 'pauseActiveTask', {}).catch(() => undefined)
184  }
185}
186
187/** Mirrors $.agent.list() into the state; a toast for each agent that finished. */
188async function pollAgents($: EngineInterface) {
189  const listed = await $.agent.list()
190  const known = await read($, agentsAtom)
191  const now = await $.clock.now()
192  const byId = new Map(known.map(agent => [agent.id, agent]))
193  const finished: HudAgent[] = []
194  let isChanged = false
195  for (const info of listed) {
196    const previous = byId.get(info.id)
197    if (previous === undefined) {
198      byId.set(info.id, {
199        id: info.id,
200        description: info.description,
201        type: info.type,
202        status: info.status,
203        startedAt: now,
204        ...(isRunning(info) ? {} : { endedAt: now }),
205      })
206      isChanged = true
207    } else if (previous.status !== info.status) {
208      const changed: HudAgent = { ...previous, status: info.status }
209      if (isRunning(previous) && !isRunning(info)) {
210        changed.endedAt = now
211        finished.push(changed)
212      }
213
214      byId.set(info.id, changed)
215      isChanged = true
216    }
217  }
218
219  for (const previous of known) {
220    if (isRunning(previous) && !listed.some(info => info.id === previous.id)) {
221      const ended: HudAgent = { ...previous, status: 'completed', endedAt: now }
222      byId.set(previous.id, ended)
223      finished.push(ended)
224      isChanged = true
225    }
226  }
227
228  if (isChanged) {
229    await update($, agentsAtom, () => [...byId.values()].sort((a, b) => a.startedAt - b.startedAt).slice(-30))
230  }
231
232  for (const agent of finished) {
233    const mark = agent.status === 'completed' ? '✔' : '✗'
234    $.ui.toast(`${mark} ${agent.type} "${agent.description}" ${agent.status} after ${formatDuration((agent.endedAt ?? now) - agent.startedAt)}`)
235  }
236}
237
238/** The background tick: agents, elapsed-time redraws, wanted refreshes and the heartbeat. */
239async function tick($: EngineInterface) {
240  if (hud.isTicking) {
241    return
242  }
243
244  hud.isTicking = true
245  try {
246    await pollAgents($)
247    const now = await $.clock.now()
248    const isAnyRunning = (await read($, agentsAtom)).some(isRunning)
249    if (isAnyRunning && now - hud.lastRedraw >= 30_000) {
250      hud.lastRedraw = now
251      await update($, tickAtom, n => n + 1)
252    }
253
254    if (hud.isRefreshWanted && (!hud.isQuiet || hud.isUserRefresh)) {
255      await refresh($)
256    }
257
258    const active = await read($, activeTaskAtom)
259    if (!hud.isQuiet && isAnyRunning && active !== null && now - hud.lastHeartbeat >= HEARTBEAT_MS) {
260      // Agents work while the main thread waits for them: keep the active task's segment open.
261      hud.lastHeartbeat = now
262      await call($, 'getCurrentTimer', { taskId: active }).catch(() => undefined)
263    }
264  } finally {
265    hud.isTicking = false
266  }
267}
268
269/** Pauses the active task's timer before a wait for the person, unless agents still work. */
270async function pauseIfIdle($: EngineInterface) {
271  if (hud.isQuiet || (await read($, activeTaskAtom)) === null) {
272    return
273  }
274
275  if ((await $.agent.list()).some(isRunning)) {
276    return
277  }
278
279  await call($, 'pauseActiveTask', {}).catch(() => undefined)
280  hud.isQuiet = true
281}
282
283async function agentLabel($: EngineInterface, agentId: string | undefined): Promise<string> {
284  if (agentId === undefined) return 'main'
285  const agent = (await read($, agentsAtom)).find(a => a.id === agentId)
286  return agent === undefined ? 'agent' : `${agent.type} "${agent.description}"`
287}
288
289async function testProjectsOf($: EngineInterface, sln: string): Promise<string[]> {
290  let projects = hud.slnTests.get(sln)
291  if (projects === undefined) {
292    projects = parseSlnTestProjects(await $.fs.read(sln))
293    hud.slnTests.set(sln, projects)
294  }
295
296  return projects
297}
298
299/** Reads build, test and verify results from a shell command's output; notes for the model when a project did not run. */
300async function readChecks($: EngineInterface, command: string, agentId: string | undefined, ran: ToolCallResult): Promise<ToolCallResult> {
301  if (ran.deny !== undefined) {
302    return ran
303  }
304
305  const text = typeof ran.text === 'string' ? ran.text : ''
306  const shape = shapeOf(command)
307  const summaries = parseTestSummaries(text)
308  const build = parseBuildSummary(text)
309  if (summaries.length === 0 && build === null && !shape.isVerify) {
310    return ran
311  }
312
313  const now = await $.clock.now()
314  const loop = await agentLabel($, agentId)
315  const checks: HudCheck[] = []
316  const notes: string[] = []
317  if (build !== null && summaries.length === 0 && !shape.isVerify) {
318    checks.push({
319      label: 'build',
320      isOk: build.errors === 0 && build.warnings === 0,
321      failed: 0,
322      passed: 0,
323      warnings: build.warnings,
324      errors: build.errors,
325      missing: [],
326      projects: 0,
327      at: now,
328      loop,
329    })
330  }
331
332  if (summaries.length > 0) {
333    let missing: string[] = []
334    const started = parseTestRuns(text)
335    if (shape.isDotnetTest && shape.sln !== null && (!shape.isPartial || started.length > 0)) {
336      const expected = await testProjectsOf($, shape.sln).catch(() => [] as string[])
337      const seen = new Set([...summaries.map(s => s.project), ...started])
338      missing = expected.filter(project => !seen.has(project))
339    }
340
341    const failed = summaries.reduce((n, s) => n + s.failed, 0)
342    checks.push({
343      label: 'tests',
344      isOk: failed === 0 && missing.length === 0,
345      failed,
346      passed: summaries.reduce((n, s) => n + s.passed, 0),
347      missing,
348      projects: summaries.length,
349      at: now,
350      loop,
351    })
352    if (missing.length > 0) {
353      notes.push(
354        `devflow-hud: ${missing.length} test project(s) of ${shape.sln} reported no result: ${missing.join(', ')}. ` +
355          'dotnet test can skip a project (for example while a test project it references fails): run each one separately before counting the totals.',
356      )
357    }
358  }
359
360  if (shape.isVerify) {
361    const summary = text.split('\n').reverse().find(line => /\b(passed|failed|FAIL)\b/i.test(line))?.trim()
362    checks.push({ label: 'verify', isOk: ran.isError !== true, failed: 0, passed: 0, missing: [], projects: 0, at: now, loop, summary })
363  }
364
365  await update($, checksAtom, list => [...list, ...checks].slice(-20))
366  const bad = checks.find(check => !check.isOk)
367  if (bad !== undefined) {
368    const detail = bad.missing.length > 0 ? `not run: ${bad.missing.join(', ')}` : bad.label === 'tests' ? `${bad.failed} failed` : bad.label
369    $.ui.toast(`✗ ${bad.label} (${loop}): ${detail}`)
370  }
371
372  return notes.length === 0 ? ran : { ...ran, context: [...(ran.context ?? []), ...notes] }
373}
374
375/** The guard's refusal of a main-thread write while an orchestrator skill leads, or undefined. */
376async function guardRefusal($: EngineInterface, path: string, agentId: string | undefined): Promise<string | undefined> {
377  if (agentId !== undefined || isGuardExempt(path)) {
378    return undefined
379  }
380
381  const guard = await read($, guardAtom)
382  if (!guard.isEnabled || !guard.isActive) {
383    return undefined
384  }
385
386  return (
387    `devflow-hud orchestrator guard: ${guard.skill} leads this session, and a plan orchestrator never writes code itself. ` +
388    'Dispatch this change to the implementer role agent (devflow:implementer), as /tt-implement-phase describes; do not work around this with shell commands. ' +
389    'Only the user can switch the guard off, with /hud guard off.'
390  )
391}
392
393/** Notifies when a task goes blocked or the shown phase completes. */
394async function notifyStatus($: EngineInterface, taskId: string, status: unknown) {
395  const phase = await read($, phaseAtom)
396  const task = [phase?.phase, phase?.active, ...(phase?.subtasks ?? [])].find(t => t?.id === taskId)
397  const title = task?.title ?? taskId
398  if (status === 'blocked') {
399    await $.ui.notify(`Blocked: ${title}`, { title: 'devflow' })
400  } else if (status === 'completed' && phase?.phase.id === taskId) {
401    await $.ui.notify(`Phase completed: ${title}`, { title: 'devflow' })
402  }
403}
404
405/** Reminder for the model of the agents still running, or undefined when none. */
406async function runningNote($: EngineInterface, lead: string): Promise<string | undefined> {
407  const running = (await $.agent.list()).filter(isRunning)
408  if (running.length === 0) {
409    return undefined
410  }
411
412  const names = running.map(agent => `${agent.type} "${agent.description}"`).join(', ')
413  return `devflow-hud: ${lead} ${running.length} agent(s) still run: ${names}. Their reports arrive in this session as messages when they finish; do not report on their work, predict it, or start work that depends on it before then.`
414}
415
416async function openPane($: EngineInterface) {
417  await $.ui.open({ id: PANE, title: 'devflow HUD' })
418}
419
420async function setGuardEnabled($: EngineInterface, isEnabled: boolean) {
421  await update($, guardAtom, guard => ({ ...guard, isEnabled }))
422  await $.store.set('guardEnabled', isEnabled)
423}
424
425async function toggleGuard($: EngineInterface) {
426  await setGuardEnabled($, !(await read($, guardAtom)).isEnabled)
427}
428
429async function toggleBand($: EngineInterface) {
430  await update($, bandHiddenAtom, isHidden => !isHidden)
431}
432
433/** Loads what the HUD keeps across sessions and starts the background tick. */
434async function start($: EngineInterface) {
435  await $.command.register({
436    name: 'hud',
437    description: 'devflow HUD: the TaskTracker phase, agents and checks',
438    argumentHint: '[refresh|guard on|off|band on|off|project <id>]',
439  })
440  const storedProject = await $.store.get('projectId')
441  if (hud.projectId === '' && typeof storedProject === 'string') {
442    hud.projectId = storedProject
443  }
444
445  const isGuardEnabled = await $.store.get('guardEnabled')
446  if (typeof isGuardEnabled === 'boolean') {
447    await update($, guardAtom, guard => ({ ...guard, isEnabled: isGuardEnabled }))
448  }
449
450  if ((await read($, phaseAtom)) === null) {
451    const last = (await $.store.get('lastPhase')) as HudPhase | undefined | null
452    if (last !== undefined && last !== null) {
453      await update($, phaseAtom, () => ({ ...last, isIdle: true }))
454    }
455  }
456
457  $.clock.every(TICK_MS, () => {
458    void tick($)
459  })
460}
461
462/** Follows a TaskTracker call: the project, the active task, refreshes after writes, status notifications. */
463async function followTaskTracker($: EngineInterface, tool: string, args: Record<string, unknown>, ran: ToolCallResult): Promise<ToolCallResult> {
464  if (ran.deny !== undefined || ran.isError === true) {
465    return ran
466  }
467
468  const name = tool.slice(PREFIX.length)
469  hud.isQuiet = name === 'pauseActiveTask'
470  if (typeof args.projectId === 'string' && /^[0-9a-f-]{36}$/.test(args.projectId) && args.projectId !== hud.projectId) {
471    hud.projectId = args.projectId
472    await $.store.set('projectId', hud.projectId)
473  }
474
475  if (name === 'setActiveTask' && typeof args.taskId === 'string') {
476    const taskId = args.taskId
477    await update($, activeTaskAtom, () => taskId)
478    requestRefresh(false, false)
479  } else if (name === 'clearActiveTask') {
480    await update($, activeTaskAtom, () => null)
481    await update($, phaseAtom, p => (p === null ? p : { ...p, isIdle: true }))
482  } else if (WRITES.has(name) || LINK_WRITES.has(name)) {
483    requestRefresh(LINK_WRITES.has(name), false)
484  }
485
486  if (name === 'updateTaskStatus' && typeof args.taskId === 'string') {
487    await notifyStatus($, args.taskId, args.status)
488  }
489
490  return ran
491}
492
493/** Orchestrator mode follows the main thread's skills; a fork that returns with agents still running is said so. */
494async function followSkill($: EngineInterface, skill: string, ran: ToolCallResult, before: Set<string>): Promise<ToolCallResult> {
495  if (ran.deny !== undefined) {
496    return ran
497  }
498
499  const started = (await $.agent.list()).filter(agent => isRunning(agent) && !before.has(agent.id))
500  if (started.length === 0) {
501    return ran
502  }
503
504  const names = started.map(agent => `${agent.type} "${agent.description}"`).join(', ')
505  return {
506    ...ran,
507    context: [
508      ...(ran.context ?? []),
509      `devflow-hud: the skill ${skill} returned while ${started.length} agent(s) it dispatched still run: ${names}. ` +
510        'Their reports arrive in this session as messages when they finish. Do not report on their work, or start work that depends on it, before then.',
511    ],
512  }
513}
514
515async function noteSkill($: EngineInterface, skill: string): Promise<Set<string>> {
516  if (hud.orchestrator.test(skill)) {
517    await update($, guardAtom, guard => ({ ...guard, isActive: true, skill }))
518  } else if (!/(^|:)(tt-|adr$|continuous-learning$)/.test(skill)) {
519    await update($, guardAtom, guard => ({ ...guard, isActive: false, skill: '' }))
520  }
521
522  return new Set((await $.agent.list()).filter(isRunning).map(agent => agent.id))
523}
524
525async function runCommand($: EngineInterface, args: string): Promise<{ text: string }> {
526  const [verb = '', arg = ''] = args.trim().split(/\s+/)
527  switch (verb) {
528    case '':
529      await openPane($)
530      requestRefresh(false, true)
531      return { text: 'devflow HUD pane opened.' }
532    case 'refresh':
533      requestRefresh(true, true)
534      return { text: 'devflow HUD: reading TaskTracker again.' }
535    case 'guard':
536      if (arg !== 'on' && arg !== 'off') return { text: USAGE }
537      await setGuardEnabled($, arg === 'on')
538      return { text: `devflow HUD: orchestrator guard ${arg}.` }
539    case 'band':
540      if (arg !== 'on' && arg !== 'off') return { text: USAGE }
541      await update($, bandHiddenAtom, () => arg === 'off')
542      return { text: `devflow HUD: band ${arg === 'on' ? 'shown' : 'hidden'}.` }
543    case 'project':
544      if (!/^[0-9a-f-]{36}$/.test(arg)) return { text: USAGE }
545      hud.projectId = arg
546      await $.store.set('projectId', hud.projectId)
547      requestRefresh(true, true)
548      return { text: `devflow HUD: project ${arg}.` }
549    default:
550      return { text: USAGE }
551  }
552}
553
554export const register: Register = (on, options) => {
555  hud.orchestrator = new RegExp(String(options.orchestratorSkills || 'implement-plan|implement-phase|tt-workflow-run'))
556  hud.projectId = String(options.projectId ?? '')
557
558  on('session.start', async ($, e, next) => {
559    await start($)
560    return next(e)
561  })
562
563  on('tool.call', async ($, e, next) => {
564    const tool = String(e.tool)
565    if (!tool.startsWith(PREFIX)) {
566      return next(e)
567    }
568
569    return followTaskTracker($, tool, e as unknown as Record<string, unknown>, await next(e))
570  }).catch(($, e, next) => next(e))
571
572  on('tool.call', { tool: 'Skill' }, async ($, e, next) => {
573    if (e.agentId !== undefined) {
574      return next(e)
575    }
576
577    const before = await noteSkill($, e.skill)
578    return followSkill($, e.skill, await next(e), before)
579  }).catch(($, e, next) => next(e))
580
581  // The orchestrator guard: main-thread file writes are refused while an orchestrator skill leads.
582  on('tool.call', { tool: 'Write' }, async ($, e, next) => {
583    const deny = await guardRefusal($, e.file_path, e.agentId)
584    return deny === undefined ? next(e) : { deny }
585  }).catch(($, e, next) => next(e))
586
587  on('tool.call', { tool: 'Edit' }, async ($, e, next) => {
588    const deny = await guardRefusal($, e.file_path, e.agentId)
589    return deny === undefined ? next(e) : { deny }
590  }).catch(($, e, next) => next(e))
591
592  on('tool.call', { tool: 'NotebookEdit' }, async ($, e, next) => {
593    const deny = await guardRefusal($, e.notebook_path, e.agentId)
594    return deny === undefined ? next(e) : { deny }
595  }).catch(($, e, next) => next(e))
596
597  // A question to the person is a wait: pause the active task's timer first.
598  on('tool.call', { tool: 'AskUserQuestion' }, async ($, e, next) => {
599    if (e.agentId === undefined) {
600      await pauseIfIdle($)
601    }
602
603    return next(e)
604  }).catch(($, e, next) => next(e))
605
606  // Check alarms: build, test and verify results in any loop's shell output.
607  on('tool.call', { tool: 'Bash' }, async ($, e, next) => readChecks($, e.command, e.agentId, await next(e))).catch(($, e, next) => next(e))
608  on('tool.call', { tool: 'PowerShell' }, async ($, e, next) => readChecks($, e.command, e.agentId, await next(e))).catch(($, e, next) => next(e))
609
610  // The end of a main-thread turn is a wait for the person, unless agents still work.
611  on('turn.complete', async ($, e, next) => {
612    if (e.agentId === undefined) {
613      await pauseIfIdle($)
614    }
615
616    return next(e)
617  })
618
619  // The person's prompt ends the wait; it reminds the model of agents still running.
620  on('prompt.submit', async ($, e, next) => {
621    hud.isQuiet = false
622    const note = await runningNote($, 'as this prompt arrives,')
623    return next(note === undefined ? e : { ...e, context: [...(e.context ?? []), note] })
624  }).catch(($, e, next) => next(e))
625
626  on('command.run', { command: 'hud' }, async ($, e) => runCommand($, e.args))
627
628  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
629    if (e.props.hasSurvey || (await read($, bandHiddenAtom))) {
630      return next(e)
631    }
632
633    await read($, tickAtom)
634    const view = await viewOf($)
635    return hasContent(view) ? band($.ui.resolve(e), view, e.props.bodyColumns) : next(e)
636  })
637
638  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
639    await read($, tickAtom)
640    const actions = {
641      refresh: () => requestRefresh(true, true),
642      toggleGuard: () => void toggleGuard($),
643      toggleBand: () => void toggleBand($),
644    }
645
646    return pane($.ui.resolve(e), await viewOf($), actions, await read($, bandHiddenAtom))
647  })
648}
649
hooks/parse.ts 144 lines
1// Pure parsers: TaskTracker MCP texts, dotnet test/build output, .sln files and shell commands.
2// No `$` here, so every function is unit-tested in parse.test.ts.
3
4import type { HudCriterion, HudTask } from '../types'
5
6const UUID = '[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}'
7
8/** Capture group `i` of a match; '' when it did not take part. */
9function group(m: RegExpMatchArray, i: number): string {
10  return m[i] ?? ''
11}
12
13/** `Task "<title>" (id <uuid>) v<n> [<type>/<status>/<priority>]`, the head of getTask's answer. */
14export function parseTask(text: string): HudTask | null {
15  const m = new RegExp(`^Task "(.*)" \\(id (${UUID})\\) v\\d+ \\[([a-z_]+)/([a-z_]+)`, 'm').exec(text)
16  return m ? { title: group(m, 1), id: group(m, 2), type: group(m, 3), status: group(m, 4) } : null
17}
18
19/** The `- [<type>/<status>] <title> (id <uuid>)` lines of getChildTasks and getTaskAncestors. */
20export function parseTaskList(text: string): HudTask[] {
21  const line = new RegExp(`^\\s*-\\s+\\[([a-z_]+)/([a-z_]+)\\]\\s+(.*?)\\s+\\(id (${UUID})\\)\\s*$`, 'gm')
22  return [...text.matchAll(line)].map(m => ({ type: group(m, 1), status: group(m, 2), title: group(m, 3), id: group(m, 4) }))
23}
24
25/** The `- [<status>/<priority>] <slug>: <title> (id <uuid>)` lines of listRequirements. */
26export function parseRequirements(text: string): { id: string; slug: string }[] {
27  const line = new RegExp(`^\\s*-\\s+\\[[a-z_]+/[a-z_]+\\]\\s+([^:\\s]+):.*\\(id (${UUID})\\)\\s*$`, 'gm')
28  return [...text.matchAll(line)].map(m => ({ slug: group(m, 1), id: group(m, 2) }))
29}
30
31/** The `- [x] (<n>) <text> (id <uuid>)` lines of listAcceptanceCriteria; any mark but a space is satisfied. */
32export function parseCriteria(text: string): HudCriterion[] {
33  const line = new RegExp(`^\\s*-\\s+\\[(.)\\]\\s+\\(\\d+\\)\\s+(.*?)\\s+\\(id ${UUID}\\)\\s*$`, 'gm')
34  return [...text.matchAll(line)].map(m => ({ isSatisfied: group(m, 1) !== ' ', text: group(m, 2) }))
35}
36
37/** Whether listRequirementTaskLinks' answer links the requirement to task `taskId`. */
38export function linksTask(text: string, taskId: string): boolean {
39  return new RegExp(`\\btask ${taskId}\\b`).test(text)
40}
41
42export type TestSummary = { project: string; failed: number; passed: number; skipped: number }
43
44/** dotnet test's `Passed!/Failed!  - Failed: n, Passed: n, Skipped: n, Total: n, Duration: … - X.dll` lines, last per project. */
45export function parseTestSummaries(text: string): TestSummary[] {
46  const line = /^[^\S\n]*(?:\d+[:-])?\s*(?:Passed!|Failed!)\s+-\s+Failed:\s+(\d+),\s+Passed:\s+(\d+),\s+Skipped:\s+(\d+),\s+Total:\s+\d+,.*?-\s+([^\s\\/]+?)\.dll\b/gm
47  const byProject = new Map<string, TestSummary>()
48  for (const m of text.matchAll(line)) {
49    byProject.set(group(m, 4), { project: group(m, 4), failed: Number(group(m, 1)), passed: Number(group(m, 2)), skipped: Number(group(m, 3)) })
50  }
51
52  return [...byProject.values()].sort((a, b) => (a.project < b.project ? -1 : a.project > b.project ? 1 : 0))
53}
54
55/** The projects dotnet test started: `Test run for <path>/<X>.dll (…)`. */
56export function parseTestRuns(text: string): string[] {
57  return [...new Set([...text.matchAll(/Test run for .*?[\\/]([^\\/\s]+?)\.dll\b/g)].map(m => group(m, 1)))].sort()
58}
59
60/** MSBuild's closing `n Warning(s)` / `n Error(s)` lines (the last of each). */
61export function parseBuildSummary(text: string): { warnings: number; errors: number } | null {
62  const warnings = [...text.matchAll(/^\s*(\d+) Warning\(s\)\s*$/gm)].at(-1)
63  const errors = [...text.matchAll(/^\s*(\d+) Error\(s\)\s*$/gm)].at(-1)
64  return warnings && errors ? { warnings: Number(group(warnings, 1)), errors: Number(group(errors, 1)) } : null
65}
66
67/** The test projects a .sln lists: projects whose name contains "Test". */
68export function parseSlnTestProjects(sln: string): string[] {
69  const line = /^Project\("\{[^}]+\}"\)\s*=\s*"([^"]+)",\s*"([^"]+\.(?:cs|fs|vb)proj)"/gm
70  return [...sln.matchAll(line)].map(m => group(m, 1)).filter(name => /test/i.test(name)).sort()
71}
72
73/** What a shell command runs, for the check alarms. */
74export type CommandShape = {
75  isDotnetTest: boolean
76  isBuild: boolean
77  isVerify: boolean
78  /** The solution `dotnet test` names, resolved against a leading `cd <dir> &&`; null when none. */
79  sln: string | null
80  /** A filter, a pipe or a redirect: the output may not show every project. */
81  isPartial: boolean
82}
83
84/** Reads a Bash or PowerShell command line. */
85export function shapeOf(command: string): CommandShape {
86  const test = /\bdotnet\s+test\b([^|&;>]*)/.exec(command)
87  const testArgs = test === null ? '' : group(test, 1)
88  let sln: string | null = null
89  const named = /(\S+\.slnx?)\b/.exec(testArgs)
90  if (named) {
91    sln = group(named, 1).replace(/^["']|["']$/g, '')
92    const cd = /^\s*cd\s+(["']?)([^"'&;]+?)\1\s*&&/.exec(command)
93    if (cd && !isAbsolute(sln)) {
94      sln = `${group(cd, 2).replace(/[\\/]+$/, '')}/${sln}`
95    }
96
97    sln = fromGitBash(sln)
98  }
99
100  return {
101    isDotnetTest: test !== null,
102    isBuild: /\bdotnet\s+build\b/.test(command),
103    isVerify: /verify\.ps1\b/.test(command),
104    sln,
105    isPartial: test !== null && (/--filter\b/.test(testArgs) || /[|>]/.test(command.slice(test.index))),
106  }
107}
108
109function isAbsolute(path: string): boolean {
110  return /^([a-zA-Z]:)?[\\/]/.test(path)
111}
112
113/** `/c/projects/x` (Git Bash) as `C:/projects/x`; other paths unchanged. */
114export function fromGitBash(path: string): string {
115  return path.replace(/^\/([a-zA-Z])\//, (_, drive: string) => `${drive.toUpperCase()}:/`)
116}
117
118/** Whether a Write/Edit path is outside the orchestrator guard: Claude's own folders (memory, plans, scratchpad). */
119export function isGuardExempt(path: string): boolean {
120  const p = path.replace(/\\/g, '/').toLowerCase()
121  return p.includes('/.claude/') || p.includes('/appdata/local/temp/claude/') || p.startsWith('/tmp/claude')
122}
123
124/** `12m`, `1h05m`, `40s`. */
125export function formatDuration(ms: number): string {
126  const s = Math.max(0, Math.round(ms / 1000))
127  if (s < 60) return `${s}s`
128  const m = Math.floor(s / 60)
129  if (m < 60) return `${m}m`
130  return `${Math.floor(m / 60)}h${String(m % 60).padStart(2, '0')}m`
131}
132
133/** `▮▮▯▯` for done of total, `width` cells. */
134export function progressBar(done: number, total: number, width = 5): string {
135  if (total <= 0) return ''
136  const full = Math.round((done / total) * width)
137  return '▮'.repeat(full) + '▯'.repeat(width - full)
138}
139
140/** A title cut to `max` characters with an ellipsis. */
141export function short(text: string, max: number): string {
142  return text.length <= max ? text : `${text.slice(0, Math.max(1, max - 1))}…`
143}
144
hooks/state.ts 10 lines
1// Shared constants and pure helpers. The state atoms themselves are declared in register.tsx: the engine
2// reads `$.state` references only where they are written in the file that uses them.
3
4export const PANE = 'devflow-hud'
5
6/** Statuses of an agent whose loop still works (a HudAgent or a $.agent.list() row). */
7export function isRunning(agent: { status: string }): boolean {
8  return agent.status === 'pending' || agent.status === 'running' || agent.status === 'waiting'
9}
10
hooks/view.tsx 187 lines
1// The band above the prompt (one line) and the /hud pane: drawn from the session state alone.
2
3import type { ElementTable } from 'claude-code'
4
5import type { HudAgent, HudCheck, HudGuard, HudPhase } from '../types'
6import { formatDuration, progressBar, short } from './parse'
7import { isRunning } from './state'
8
9export type HudView = {
10  phase: HudPhase | null
11  agents: HudAgent[]
12  checks: HudCheck[]
13  guard: HudGuard
14  now: number
15}
16
17/** Done and total of the phase's direct sub-tasks (archived ones are not listed by TaskTracker). */
18function subtaskProgress(phase: HudPhase): { done: number; total: number } {
19  return { done: phase.subtasks.filter(t => t.status === 'completed').length, total: phase.subtasks.length }
20}
21
22function acProgress(phase: HudPhase): { done: number; total: number } {
23  const all = phase.requirements.flatMap(r => r.criteria)
24  return { done: all.filter(c => c.isSatisfied).length, total: all.length }
25}
26
27function checkLabel(check: HudCheck): string {
28  if (check.label === 'build') return `build ${check.warnings ?? 0}w/${check.errors ?? 0}e`
29  if (check.label === 'verify') return 'verify.ps1'
30  return `tests ${check.failed}✗/${check.passed}✓`
31}
32
33/** Whether there is anything to show. */
34export function hasContent(view: HudView): boolean {
35  return view.phase !== null || view.agents.some(isRunning) || view.checks.length > 0 || (view.guard.isActive && view.guard.isEnabled)
36}
37
38/** The one-line band, drawn with the surface's element table. */
39export function band(elements: ElementTable, view: HudView, columns: number) {
40  const { Box, Text } = elements
41  const { phase, agents, checks, guard, now } = view
42  const running = agents.filter(isRunning)
43  const first = running[0]
44  const last = checks.at(-1)
45  const titleRoom = Math.max(12, Math.floor(columns / 4))
46
47  return (
48    <Box flexDirection="row">
49      <Text wrap="truncate-end">
50        {phase !== null && (
51          <Text color={phase.isIdle ? 'subtle' : 'claude'} bold={!phase.isIdle}>
52            ◆ {short(phase.phase.title, titleRoom)}
53          </Text>
54        )}
55        {phase?.active && <Text dimColor> › {short(phase.active.title, titleRoom)}</Text>}
56        {phase !== null && phase.subtasks.length > 0 && (
57          <Text>
58            {'  '}
59            {progressBar(subtaskProgress(phase).done, subtaskProgress(phase).total)} {subtaskProgress(phase).done}/
60            {subtaskProgress(phase).total}
61          </Text>
62        )}
63        {phase !== null && acProgress(phase).total > 0 && (
64          <Text color={acProgress(phase).done === acProgress(phase).total ? 'success' : undefined}>
65            {'  '}AC {acProgress(phase).done}/{acProgress(phase).total}
66          </Text>
67        )}
68        {first !== undefined && (
69          <Text color="suggestion">
70            {'  '}⚙ {running.length > 1 ? `${running.length}× ` : ''}
71            {short(first.type, 14)} "{short(first.description, 28)}" {formatDuration(now - first.startedAt)}
72          </Text>
73        )}
74        {last !== undefined && (
75          <Text color={last.isOk ? 'success' : 'error'}>
76            {'  '}
77            {last.isOk ? '✔' : '✗'} {checkLabel(last)}
78            <Text dimColor> {formatDuration(now - last.at)}</Text>
79          </Text>
80        )}
81        {last !== undefined && last.missing.length > 0 && <Text color="warning"> ⚠ {last.missing.length} not run</Text>}
82        {guard.isActive && guard.isEnabled && <Text color="permission">{'  '}◈ orchestrator</Text>}
83      </Text>
84    </Box>
85  )
86}
87
88const STATUS_GLYPH: Record<string, string> = {
89  completed: '✔',
90  in_progress: '▶',
91  pending: '·',
92  blocked: '⛔',
93}
94
95/** The /hud pane: the phase tree, every AC, the agents, the checks and the guard. */
96export function pane(
97  elements: ElementTable,
98  view: HudView,
99  actions: { refresh: () => void; toggleGuard: () => void; toggleBand: () => void },
100  isBandHidden: boolean,
101) {
102  const { Box, Text, Button } = elements
103  const { phase, agents, checks, guard, now } = view
104
105  return (
106    <Box flexDirection="column">
107      <Box flexDirection="row" gap={1}>
108        <Button key="refresh" hotkey="r" label="Refresh" onPress={actions.refresh} />
109        <Button key="guard" hotkey="g" label={guard.isEnabled ? 'Guard: on' : 'Guard: off'} onPress={actions.toggleGuard} />
110        <Button key="band" hotkey="b" label={isBandHidden ? 'Band: hidden' : 'Band: shown'} onPress={actions.toggleBand} />
111      </Box>
112
113      <Text bold> </Text>
114      {phase === null ? (
115        <Text dimColor>No TaskTracker phase seen yet. It appears when a task is activated (setActiveTask), or press Refresh.</Text>
116      ) : (
117        <Box flexDirection="column">
118          <Text bold color="claude">
119            {STATUS_GLYPH[phase.phase.status] ?? '?'} {phase.phase.title} <Text dimColor>({phase.phase.status}{phase.isIdle ? ', no active task' : ''})</Text>
120          </Text>
121          {phase.error && <Text color="error">  last refresh failed: {phase.error}</Text>}
122          {phase.subtasks.map(task => (
123            <Text color={phase.active?.id === task.id ? 'suggestion' : undefined} dimColor={task.status === 'completed'}>
124              {'  '}
125              {STATUS_GLYPH[task.status] ?? '?'} {task.title}
126              {phase.active?.id === task.id ? '  ← active' : ''}
127            </Text>
128          ))}
129          {phase.requirements.length === 0 && (
130            <Text dimColor>  No linked requirements read (set the projectId option, or let a TaskTracker call with a projectId pass).</Text>
131          )}
132          {phase.requirements.map(requirement => (
133            <Box flexDirection="column">
134              <Text bold>
135                {'  '}
136                {requirement.slug} {requirement.criteria.filter(c => c.isSatisfied).length}/{requirement.criteria.length}
137              </Text>
138              {requirement.criteria.map(criterion => (
139                <Text wrap="truncate-end" dimColor={criterion.isSatisfied} color={criterion.isSatisfied ? 'success' : undefined}>
140                  {'    '}
141                  {criterion.isSatisfied ? '[x]' : '[ ]'} {criterion.text}
142                </Text>
143              ))}
144            </Box>
145          ))}
146          <Text dimColor>  read {formatDuration(now - phase.refreshedAt)} ago</Text>
147        </Box>
148      )}
149
150      <Text bold> </Text>
151      <Text bold>Agents</Text>
152      {agents.length === 0 && <Text dimColor>  none this session</Text>}
153      {agents.slice(-12).map(agent => (
154        <Text color={isRunning(agent) ? 'suggestion' : agent.status === 'completed' ? undefined : 'error'} dimColor={!isRunning(agent) && agent.status === 'completed'}>
155          {'  '}
156          {isRunning(agent) ? '⚙' : agent.status === 'completed' ? '✔' : '✗'} {agent.type} "{agent.description}" {agent.status}{' '}
157          {formatDuration((agent.endedAt ?? now) - agent.startedAt)}
158        </Text>
159      ))}
160
161      <Text bold> </Text>
162      <Text bold>Checks</Text>
163      {checks.length === 0 && <Text dimColor>  none read yet (dotnet build/test and verify.ps1 output is read from shell commands)</Text>}
164      {checks.slice(-10).map(check => (
165        <Text color={check.isOk ? undefined : 'error'}>
166          {'  '}
167          {check.isOk ? '✔' : '✗'} {checkLabel(check)} <Text dimColor>({check.loop}, {formatDuration(now - check.at)} ago{check.projects > 0 ? `, ${check.projects} projects` : ''})</Text>
168          {check.missing.length > 0 && <Text color="warning"> ⚠ not run: {check.missing.join(', ')}</Text>}
169          {check.summary && <Text dimColor> {check.summary}</Text>}
170        </Text>
171      ))}
172
173      <Text bold> </Text>
174      <Text>
175        <Text bold>Guard </Text>
176        {!guard.isEnabled ? (
177          <Text dimColor>off (/hud guard on)</Text>
178        ) : guard.isActive ? (
179          <Text color="permission">active: {guard.skill} leads the main thread; its Write/Edit/NotebookEdit are refused</Text>
180        ) : (
181          <Text dimColor>armed; turns on when an orchestrator skill is invoked</Text>
182        )}
183      </Text>
184    </Box>
185  )
186}
187
types/index.d.ts 67 lines
1/** A TaskTracker task as the HUD shows it. */
2export type HudTask = { id: string; title: string; type: string; status: string }
3
4/** One acceptance criterion of a requirement linked to the phase. */
5export type HudCriterion = { text: string; isSatisfied: boolean }
6
7/** A requirement linked to the phase, with its acceptance criteria. */
8export type HudRequirement = { id: string; slug: string; criteria: HudCriterion[] }
9
10/** The phase the active task belongs to, as last read from TaskTracker. */
11export type HudPhase = {
12  phase: HudTask
13  /** The active task when it is a task under the phase; null when the phase itself is active. */
14  active: HudTask | null
15  subtasks: HudTask[]
16  requirements: HudRequirement[]
17  refreshedAt: number
18  /** True when no task is active any more (the view is the last one seen). */
19  isIdle: boolean
20  error?: string
21}
22
23/** A subagent of this session. */
24export type HudAgent = {
25  id: string
26  description: string
27  type: string
28  status: string
29  startedAt: number
30  endedAt?: number
31}
32
33/** A build, test or verify result read from a shell command's output. */
34export type HudCheck = {
35  label: string
36  isOk: boolean
37  failed: number
38  passed: number
39  warnings?: number
40  errors?: number
41  /** Test projects of the solution that reported no result. */
42  missing: string[]
43  projects: number
44  at: number
45  /** "main" or the agent's description. */
46  loop: string
47  summary?: string
48}
49
50/** The orchestrator guard: on while an orchestrator skill leads the main thread. */
51export type HudGuard = { isEnabled: boolean; isActive: boolean; skill: string }
52
53declare module 'claude-code' {
54  interface PluginState {
55    devflow: {
56      phase: HudPhase | null
57      activeTaskId: string | null
58      agents: HudAgent[]
59      checks: HudCheck[]
60      guard: HudGuard
61      isBandHidden: boolean
62      /** Bumped every 30 s while agents run, so elapsed times redraw. */
63      tick: number
64    }
65  }
66}
67