DevForgeAI spec-driven planning skills for Claude Code: brainstorm ideas, turn the promoted ones into a PRD with traceable IDs, define the architecture…

DevForgeAI is a Claude Code plugin, devforgeai, of spec-driven planning skills. It takes an idea from a brainstorm through a product requirements document (PRD) and its architecture to epics. Each step is written as a Markdown document with stable item IDs and traceable links, and every judgment call it records (which ideas to pursue, priorities, what ships now) is left to you. A context skill writes the project's stack, layout and convention documents from the architecture's decisions. Three skills work outside that chain: a documents updater for a repository's README, CHANGELOG and guides, a git workflow skill that takes work through commits and pull requests to a merge, and a spec lookup that cites what a project's specs and ADRs already decided before Claude proposes anything new. A progress tracker shows each brainstorm and architecture run's checklist steps as they happen.
It is for people who plan software with Claude Code. It is unreleased: load the plugin from this repository's source.
git skill: git, and for its pr and merge phases the GitHub CLI (gh), signed in. GitHub is the only supported host; the other phases work with any git remote.<devforgeai> with the path to this repository: claude --plugin-dir <devforgeai>/src/claude/DevForgeAI
/devforgeai:brainstorm ways to cut appointment no-shows, or ask in plain words to brainstorm. Confirm which ideas to promote and whether the brainstorm is done. The skill writes docs/specs/brainstorm/BRN-001.md in your project. /devforgeai:prd BRN-001
The skill asks only about what the brainstorm leaves open, writes docs/specs/prd/PRD-001.md (the next free number), and names the next step.
/devforgeai:architecture PRD-001
The skill settles each shared architectural question only by your decision, an accepted ADR or approved policy, writes docs/specs/arch/ARCH-001.md plus an ADR for each decision you make, and reports which requirements are ready for epics.
docs/specs/context/ from the architecture's decisions and the conventions you confirm: /devforgeai:context
The epic skill doesn't read them; they are written for the story and spec steps (ADR-004).
/devforgeai:epic PRD-001
The skill proposes a grouping of the ready, current-release requirements, writes docs/specs/epic/EPIC-NNN.md for each epic once you confirm it, and lists every requirement it left out with the reason.
| Skill | Invoke | What it does | Status | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
brainstorm | /devforgeai:brainstorm [topic] | Runs a structured brainstorm and writes a BRN document with problems, ideas, assumptions and the dispositions you confirmed | Implemented (SPEC-001) | ||||||||
prd | /devforgeai:prd [BRN-NNN] | Drafts a PRD from a converged brainstorm's promoted ideas, interviews you only for the gaps, and applies any approved policy | Implemented (SPEC-002) | ||||||||
documents-updater | /devforgeai:documents-updater [base-revision-or-range] [propose] | Updates a repository's README, CHANGELOG and guides from its git changes, or proposes the edits | Implemented (SPEC-006) | ||||||||
architecture | /devforgeai:architecture [PRD-NNN] | Identifies the architectural questions separate epics must share, settles each only by your decision, an accepted ADR or approved policy, and writes an ARCH document with ADRs and a report of which requirements are ready for epics | Implemented (SPEC-003) | ||||||||
epic | /devforgeai:epic [PRD-NNN] | Groups a PRD's ready, current-release requirements into epics you confirm, and reports every requirement left out and why | Implemented (SPEC-004) | ||||||||
context | /devforgeai:context [document] | Writes and maintains the context documents in docs/specs/context/ from accepted ADRs, approved policy, the ARCH and conventions you confirm; cites every decision and labels observed practice | Implemented (SPEC-011) | ||||||||
spec-lookup | /devforgeai:spec-lookup [ID or term] | Searches the project's docs/specs/ and cites each match by file, line, version and status; anything with no match is a question for you, never something to build. Other skills run it through the devforgeai:spec-lookup agent | Implemented (SPEC-014) | ||||||||
git | `/devforgeai:git [status\ | connect\ | start\ | commit\ | push\ | pr\ | merge\ | sync\ | prune] [details]` | Commits and pushes work from its own branch and worktree, opens or updates a GitHub pull request, merges only a PR that an independent QA session approved for its head commit and you confirm, fast-forwards the default branch without discarding local edits, and prunes merged worktrees | Implemented; the source is v3, which is in review and not yet evaluated (v2 is the last approved version) (SPEC-007, qualification) |
The planning chain is Brainstorm → PRD → Architecture Definition → Epic → Story → Spec; the first four steps are implemented. The story skill is specified as a draft (SPEC-009); until it is built, the epic skill's handoff says to write stories by hand from the story template.
The plugin includes a progress tracker for the brainstorm and architecture skills (SPEC-012, SPEC-013). It is a Claude Code mod, so it runs only where Claude Code loads plugin hook modules.
devforgeai/progress/ in the project. Files older than 30 days are removed when a later session starts its first tracked run (the plugin's retentionDays setting)..claude/devforgeai.local.md.tracking setting to off to turn the tracker off; the skills work the same either way.Skills write to docs/specs/<type>/<ID>.md in your project, such as docs/specs/prd/PRD-002.md, and allocate each ID themselves. Values you haven't decided stay null or carry a [NEEDS CLARIFICATION] marker instead of a guess. The document format, IDs, links and provenance rules are in the templates README, and the JSON Schemas are in src/schemas/.
The prd, architecture and context skills also read optional policy: approved organization or project policy documents in docs/specs/policy/ (template: src/templates/policy.md), and a user-local .claude/devforgeai.local.md for interaction defaults such as interview.max_calls. ADR-003 defines the rules. The epic skill reads policy only to check that each setting the architecture relied on is still approved, active and at the version it linked.
Each implemented skill has an eval suite in src/claude/DevForgeAI/evals/<skill>/. Run it from a plain terminal, not inside a Claude Code session, at the repository root:
A="--allow-tools Write Edit Bash --scaffold --judge-model sonnet --threshold 0.8"
claude plugin eval src/claude/DevForgeAI --tag prd $A --output-dir tmp/eval-results/prd
The run prints a score per case; the bar is 0.8 per case over three runs. CLAUDE.md covers the options, the manual checks and how the plugin is deployed.
hooks/progress.tsx 3076 lines1// DevForgeAI's progress tracker adapter for Claude Code (SPEC-013 v27).
2//
3// It records each run of a tracked skill as SPEC-012's event log, runs SPEC-012's evaluator on a timer, and shows
4// the run in the status line, a two-row band above the prompt and toasts. In enforce mode it refuses the write
5// gate's Write or Edit when that gate raised flags, refuses a question at SPEC-012's question gate (no step marked,
6// no step tag, or another step's tag), and gives the model the report gate's flags. It records Claude Code's task list as step events.
7// A tracked skill Claude loads mid-run pauses the open run on a trail, which unwinds when Claude goes back (BEH-29,
8// BEH-30). A plugin skill whose SKILL.md metadata says devforgeai-tracked "false" is untracked: its load changes nothing
9// and the turn it loads in records no tool events and no replies (BEH-02, versions 18 and 19). It fails open:
10// when it can't run, the work goes on and the user is told (ADR-006 D1). Version 22: an answer event records the form Claude
11// asked (BEH-37); a Bash command naming a document the run writes at its write gate is refused in enforce mode, and the
12// documents a Bash call wrote are recorded after it (BEH-38); the line on continuing states the draft, the ask and the forms
13// (BEH-39). Old run and session folders are pruned
14// by progress/prune.py, which the adapter starts once per session and root (BEH-19). Versions 21 to 25: another mod's tool
15// calls change nothing of the tracker (BEH-33); /progress prints the run (BEH-34); a row above the prompt warns that
16// /devforgeai:precompact will run, and the adapter runs it once as if typed (BEH-35, BEH-36, BEH-41); a turn's tokens are a
17// usage event (BEH-40) and a line of the session's odometer ledger (SPEC-016 BEH-10, ERR-06). Held with the dashboard, and not
18// built here: the pane /progress will open (BEH-34), the tile that starts a skill (BEH-41's other setter, ERR-24) and
19// BEH-01's /devforgeai:dashboard answer with tracking off. Version 27: a write of the tracker's that fails (the folder and its
20// .gitignore, a run's events.jsonl, the odometer's file) is a hold, a retry and a remedy, not a stop: the run stays open, its lines
21// stay in memory, the adapter tries again at the end of each main-loop turn and on `/progress retry`, and only after 10 failed turn
22// tries, or at 4 MiB, does it stop as earlier versions did at once (BEH-42, ERR-03; SPEC-016 version 3 for the ledger).
23//
24// Every use of `$` stays in top-level functions of this file (claude plugin validate's rule); progress-core.ts
25// holds the pure helpers.
26import { atom, read, update } from 'claude-code'
27import type { EngineInterface, Register } from 'claude-code'
28import type { ProgressMode, ProgressModeSource, ProgressPaused, ProgressPrecompact, ProgressReturned, ProgressRun, ProgressSummary, ProgressWroteSeen } from '../types'
29import {
30 adherenceText, bandRows, byteSize, compactTexts, editResult, eventLine, exitOf, finalTimeout, fit, followsTaskList, hasTaskList,
31 hintText, isAnswered, isEngine, isFailed, isPersonPrompt, isTracked, isUntrackedSkill, isWaiverQuestion, keptContent, markedStep, newFlagToasts,
32 questionRefusal, questionTag, refusalCause, refusalText, replyText, reportContext, retentionOf, reviewItems, reviewQuestion,
33 runId, runName, skillName, statusText, stepLabel, waiverAnswer, returnLine, pausedWith, keptTrail, endReason, hasRoom, trailNote, TRAIL_NOTE_START, stopsRun,
34 exitQuestion, keptText, nestedExitQuestion, nestedKeptText, isDismissal, CONFIRMED, resumePlan, resumeQuestion, resumeLine,
35 removeArgv, resumesOf, workFilePaths, workFilesDue, workFilesProblem, formOf, gatedOf, outsideWord, outsideRefusal, outsideMessage,
36 folderOf, folderOfFile, changedFiles, signature, logNames, windowClosed, wroteText, withContext, ruleMatches, pathMatches,
37 OUTSIDE_WRITE_TYPE, OUTSIDE_ADVICE, WROTE_CONTEXT_ROUTE, stepOfTask, stepStateOf, stuckAdvice, stuckText, summaryOf, taskIdOf, todoSteps, toolPath, FORMAT, IDLE_MS, LOG_LIMIT, NOTE_START, TASK_TAG,
38 isFromMod, isOwnRun, fuelSetting, measuredShare, precompactRow, precompactDue, NO_PRECOMPACT, usageFields, ledgerLine, progressReport,
39 addStart, takeStart, isCompaction,
40 HOLD_TRIES, STOPPED_LINE, NOTHING_TO_RETRY, STOP_TOAST, errorLine, remedyRow, heldToast, recoveredToast, odometerGaveUpToast, heldLog,
41 recoveredLog, gaveUpLog, savedAnswer, stillAnswer, liftedAnswer, stillStoppedAnswer, odometerOnAnswer, odometerStillAnswer,
42} from './progress-core'
43import type { Fields, Gated, HoldKind, HoldView, Listed, ProgressState, Refused, ReviewItem, ToolOutcome } from './progress-core'
44
45type E = EngineInterface
46
47const RUN = atom({ plugin: 'devforgeai', key: 'run' } as const, null as ProgressRun | null)
48const MODE = atom({ plugin: 'devforgeai', key: 'mode' } as const, 'observe' as ProgressMode)
49const SOURCE = atom({ plugin: 'devforgeai', key: 'modeSource' } as const, 'framework-default' as ProgressModeSource)
50const SUMMARY = atom({ plugin: 'devforgeai', key: 'summary' } as const, null as ProgressSummary | null)
51const LAST = atom({ plugin: 'devforgeai', key: 'lastEventAt' } as const, 0)
52const MARKED = atom({ plugin: 'devforgeai', key: 'marked' } as const, false)
53const SHOWN = atom({ plugin: 'devforgeai', key: 'shown' } as const, [] as string[])
54const SENT = atom({ plugin: 'devforgeai', key: 'contextSent' } as const, [] as number[])
55const OFF = atom({ plugin: 'devforgeai', key: 'off' } as const, null as string | null)
56const TASKS = atom({ plugin: 'devforgeai', key: 'tasks' } as const, {} as Record<string, number>)
57const TODOS = atom({ plugin: 'devforgeai', key: 'todos' } as const, {} as Record<string, string>)
58const HINTED = atom({ plugin: 'devforgeai', key: 'hinted' } as const, false)
59const ADHERED = atom({ plugin: 'devforgeai', key: 'adhered' } as const, null as string | null)
60const REFUSALS = atom({ plugin: 'devforgeai', key: 'refusals' } as const, {} as Record<string, number>)
61const REFUSED = atom({ plugin: 'devforgeai', key: 'refused' } as const, [] as Refused[])
62const REVIEWED = atom({ plugin: 'devforgeai', key: 'reviewed' } as const, null as string | null)
63const TRAIL = atom({ plugin: 'devforgeai', key: 'trail' } as const, [] as ProgressPaused[])
64const RETURNED = atom({ plugin: 'devforgeai', key: 'returned' } as const, [] as ProgressReturned[])
65const CLEANED = atom({ plugin: 'devforgeai', key: 'cleaned' } as const, [] as string[])
66const WROTE_SEEN = atom({ plugin: 'devforgeai', key: 'wroteSeen' } as const, {} as ProgressWroteSeen)
67const CLOSED = atom({ plugin: 'devforgeai', key: 'closedFor' } as const, null as string | null)
68// The precompact row's values and BEH-36's run mark (versions 21, 23 and 25). They belong to the session, not to a run, so they
69// are not part of the open run's Live values: the host empties them at /clear, /resume and /branch.
70const PRECOMPACT = atom({ plugin: 'devforgeai', key: 'precompact' } as const, NO_PRECOMPACT as ProgressPrecompact)
71
72// The open run's values, as the module holds them (version 14): every hook reads and writes these, never a $.state
73// snapshot, since each dispatch's $.state reads one moment of its own (a hook that awaits across another's write would
74// act on the old run). $.state keeps a mirror, written in order with the latest values, for a reload and the band.
75type Live = {
76 run: ProgressRun | null; summary: ProgressSummary | null; lastEventAt: number; marked: boolean; shown: string[]
77 contextSent: number[]; tasks: Record<string, number>; todos: Record<string, string>; adhered: string | null
78 refusals: Record<string, number>; refused: Refused[]; reviewed: string | null; trail: ProgressPaused[]
79 returned: ProgressReturned[]
80 /** The runs whose work files' cleanup was started (BEH-32), newest last, the last 50. */
81 cleaned: string[]
82 /** For each run, the documents a Bash call wrote that were recorded, by path: their size and mtimeMs as listed (BEH-38 (b);
83 * version 22), so an unchanged recorded file isn't recorded twice. The last 20 runs. */
84 wroteSeen: ProgressWroteSeen
85 /** The run whose last evaluation closed BEH-38's window (every step reached, ended, or its validator's step done), or null. */
86 closedFor: string | null
87}
88let live: Live | null = null
89let mirrorChain: Promise<unknown> = Promise.resolve()
90
91function emptyLive(): Live {
92 return { run: null, summary: null, lastEventAt: 0, marked: false, shown: [], contextSent: [], tasks: {}, todos: {},
93 adhered: null, refusals: {}, refused: [], reviewed: null, trail: [], returned: [], cleaned: [], wroteSeen: {}, closedFor: null }
94}
95
96/** The module's values, read once from $.state after a load or a reload (BEH-17). */
97async function hydrate($: E): Promise<Live> {
98 if (live !== null) return live
99 const run = await read($, RUN)
100 const got: Live = {
101 run, summary: await read($, SUMMARY), lastEventAt: await read($, LAST),
102 marked: await read($, MARKED), shown: await read($, SHOWN), contextSent: await read($, SENT),
103 tasks: await read($, TASKS), todos: await read($, TODOS), adhered: await read($, ADHERED),
104 refusals: await read($, REFUSALS), refused: await read($, REFUSED), reviewed: await read($, REVIEWED),
105 // An entry without its run, version 13's shape, is dropped (BEH-30); the mirror catches up at the next change.
106 // A reload between the mirror's writes of a push can leave the open run on the trail too: it is dropped.
107 trail: keptTrail(await read($, TRAIL)).filter(t => t.run!.id !== run?.id), returned: await read($, RETURNED),
108 cleaned: await read($, CLEANED), wroteSeen: await read($, WROTE_SEEN), closedFor: await read($, CLOSED),
109 }
110 if (live === null) live = got
111 return live
112}
113
114async function get<K extends keyof Live>($: E, key: K): Promise<Live[K]> {
115 return (await hydrate($))[key]
116}
117
118/** Change one value: the module's at once (no await between reading and writing it), then $.state's mirror. */
119async function put<K extends keyof Live>($: E, key: K, change: (value: Live[K]) => Live[K]): Promise<Live[K]> {
120 const l = await hydrate($)
121 const value = change(l[key])
122 l[key] = value
123 await mirrorAll($, [key])
124 return value
125}
126
127/** Change several values at once (version 14): `change` reads the module's values and returns the new ones with no await
128 * between, or null for no change; then each key is mirrored in the order given. True when it changed them. */
129async function putMany($: E, change: (l: Live) => Partial<Live> | null): Promise<boolean> {
130 const l = await hydrate($)
131 const next = change(l)
132 if (next === null) return false
133 Object.assign(l, next)
134 await mirrorAll($, Object.keys(next) as (keyof Live)[])
135 return true
136}
137
138/** putMany, only while `runId` is the open run: an evaluation, a refusal or an event of a run that a switch has paused,
139 * returned or ended changes nothing of the run open now (BEH-30). */
140async function putFor($: E, runId: string, change: (l: Live) => Partial<Live>): Promise<boolean> {
141 return putMany($, l => (l.run !== null && l.run.id === runId ? change(l) : null))
142}
143
144async function mirrorAll($: E, keys: (keyof Live)[]): Promise<void> {
145 const step = mirrorChain.then(async () => {
146 for (const key of keys) await mirror($, key)
147 })
148 mirrorChain = step.catch(() => undefined)
149 await step.catch(() => undefined)
150}
151
152/** Write the module's latest value of one key to $.state (literal refs, as claude plugin validate reads them). */
153async function mirror($: E, key: keyof Live): Promise<void> {
154 const l = live
155 if (l === null) return
156 switch (key) {
157 case 'run': await $.state.set({ plugin: 'devforgeai', key: 'run' } as const, l.run); break
158 case 'summary': await $.state.set({ plugin: 'devforgeai', key: 'summary' } as const, l.summary); break
159 case 'lastEventAt': await $.state.set({ plugin: 'devforgeai', key: 'lastEventAt' } as const, l.lastEventAt); break
160 case 'marked': await $.state.set({ plugin: 'devforgeai', key: 'marked' } as const, l.marked); break
161 case 'shown': await $.state.set({ plugin: 'devforgeai', key: 'shown' } as const, l.shown); break
162 case 'contextSent': await $.state.set({ plugin: 'devforgeai', key: 'contextSent' } as const, l.contextSent); break
163 case 'tasks': await $.state.set({ plugin: 'devforgeai', key: 'tasks' } as const, l.tasks); break
164 case 'todos': await $.state.set({ plugin: 'devforgeai', key: 'todos' } as const, l.todos); break
165 case 'adhered': await $.state.set({ plugin: 'devforgeai', key: 'adhered' } as const, l.adhered); break
166 case 'refusals': await $.state.set({ plugin: 'devforgeai', key: 'refusals' } as const, l.refusals); break
167 case 'refused': await $.state.set({ plugin: 'devforgeai', key: 'refused' } as const, l.refused); break
168 case 'reviewed': await $.state.set({ plugin: 'devforgeai', key: 'reviewed' } as const, l.reviewed); break
169 case 'trail': await $.state.set({ plugin: 'devforgeai', key: 'trail' } as const, l.trail); break
170 case 'returned': await $.state.set({ plugin: 'devforgeai', key: 'returned' } as const, l.returned); break
171 case 'cleaned': await $.state.set({ plugin: 'devforgeai', key: 'cleaned' } as const, l.cleaned); break
172 case 'wroteSeen': await $.state.set({ plugin: 'devforgeai', key: 'wroteSeen' } as const, l.wroteSeen); break
173 case 'closedFor': await $.state.set({ plugin: 'devforgeai', key: 'closedFor' } as const, l.closedFor); break
174 }
175}
176
177const EVALUATOR_TIMEOUT = 5000
178const START_TIMEOUT = 3000
179const PRUNE_TIMEOUT = 10000
180const LOG_FULL = 'event log full'
181const NO_PYTHON = 'python not found'
182const NO_WRITE = 'cannot write devforgeai/progress'
183
184// Values of the process, not the session: they outlive /clear, which empties $.state (DM-03).
185let interactive: boolean | null = null
186let python: string | null | undefined
187let timerOn = false
188let evaluating = false
189// The timer's evaluation in flight, which the review waits for (BEH-26).
190let inFlight: Promise<void> | null = null
191let turnOpen = false
192/** The skills the main loop's Skill tool calls in flight are loading, by name (BEH-29); emptied at each main-loop turn. */
193const skillsLoading = new Map<string, number>()
194/** The skills another plugin's mod is loading with a Skill tool call in flight, by name (BEH-33, Bryan 2026-10-08, "Leave the run
195 * open"): a skill.prompt for one of these, with no Claude's call and no typed command running the same name, is the mod's load
196 * and changes nothing of the tracker. */
197const modLoading = new Map<string, number>()
198/** Skills whose load inside a Skill call changed the open run (a push, an unwind, a switch): that call records nothing
199 * when it returns, in either run (BEH-29). */
200const switchedIn = new Set<string>()
201/** Bumped at each change of the open run (open, push, unwind): a hook that began under an older value began before
202 * the open run opened (BEH-30 (a), version 14). */
203let tenure = 0
204/** The skill name the person's command.run is running, until its next(e) settles: a typed skill's skill.prompt fires
205 * inside it (BEH-31, version 16; probe 2026-10-05). */
206let typedName: string | null = null
207/** The open run has recorded the event of a tool call or an answer whose hook began after it opened (BEH-30 (a)). */
208let worked = false
209/** The turn is marked (BEH-02, version 18): an untracked skill was loaded in it, by the person (the typed name), by Claude's
210 * Skill call (the in-flight set) or, from version 26, as the plugin's own precompact skill while BEH-36's pending mark is set.
211 * Set at that skill.prompt; turn.start doesn't clear it, since a typed load's skill.prompt comes before its turn starts. From
212 * version 26 it is bound to a turn (markBind) and ends by the turn-ID rule (markEndsAt). A tool call whose hook began while
213 * it was set records no tool event, and a reply arriving while it is set isn't recorded (version 19). */
214let untrackedTurn = false
215/** What the mark is bound to (BEH-02, version 26): a turn (its ID, and the IDs of the main-loop turns started since); 'next',
216 * a load bound to the next main-loop turn.start; or 'any', when no turn ID can be bound, so the next turn.complete ends it. */
217type MarkBind = { kind: 'turn'; id: string; since: string[] } | { kind: 'next' } | { kind: 'any' }
218let markBind: MarkBind | null = null
219/** Whether the mark was set at a load of the plugin's own precompact skill, which is when its end is written to adapter.log. */
220let markIsPrecompact = false
221/** The main loop's current turn: the turnId of its latest turn.start, or null when none is known (version 26). */
222let currentTurn: string | null = null
223/** Where the kept name came from: the person's own command, or BEH-41's set (the adapter's own $.command.run). The load line says. */
224let typedSource: 'person' | 'set' | null = null
225let disabled = false
226let modeSession: string | null = null
227// The root the mode was resolved for: the local preference file is per checkout (BEH-16).
228let modeRoot: string | null = null
229let pluginSkills: string[] | null = null
230let lastStatus: string | undefined
231let pendingReport: { seq: number; text: string; run: string } | null = null
232/** BEH-38 (b)'s enforce-mode texts waiting for the user's next prompt: the fallback of the P14 switch (WROTE_CONTEXT_ROUTE),
233 * and the route for a call whose answer is a deny. A separate slot from the report's, so neither clobbers the other. */
234let pendingWrote: { run: string; text: string }[] = []
235/** A run's manifests' gated rules, read once at its first Bash call (BEH-38), by run ID; the last 20. */
236const gatedCache = new Map<string, Gated>()
237// The open run's event lines: in the module and events.jsonl, since a $.state value holds at most 4,194,304
238// characters (see types/index.d.ts); read back from the file after a reload.
239// `written` is the number of those lines the file is known to hold: the rest run ahead of it while a hold lasts (version 27).
240let held: { id: string; lines: string[]; written: number } | null = null
241/** A hold (BEH-42, version 27): a write of the tracker's that failed, kept in the module's memory and not in $.state, since a reload
242 * loses the lines it keeps and a count that outlived them would describe nothing. `path` and `error` are the latest failed write's. */
243type Hold = { path: string; error: string; firstAt: number; tries: number }
244/** Paused runs whose first write never worked (Bryan, 2026-10-09, the review of 0.29.0, C1): a nested load pushed the run while its log
245 * was held and the push's write failed, so the file does not exist. Their lines stay in memory, in the hold, until a try writes them;
246 * the resume takes them back. */
247const parked = new Map<string, { run: ProgressRun; lines: string[]; written: number }>()
248/** The open run's log: `held`'s lines run ahead of its events.jsonl (the folder and its .gitignore included). */
249let runHold: Hold | null = null
250/** The session's odometer file: `ledgerHeld` runs ahead of it. */
251let odoHold: Hold | null = null
252/** The ledger stopped at the tenth failed turn try (SPEC-016 ERR-06): `/progress retry` lifts it. */
253let odoGaveUp = false
254/** A new run's pruning waits for the write that creates its folder (BEH-19, version 27). */
255let pendingPrune: { r: string; session: string; keepRun: string; continued: string | null } | null = null
256let allReachedFor: string | null = null
257let bandChain: Promise<unknown> = Promise.resolve()
258let logChain: Promise<unknown> = Promise.resolve()
259let recordChain: Promise<unknown> = Promise.resolve()
260let absorbChain: Promise<unknown> = Promise.resolve()
261// adapter.log lines wait here until a run has created devforgeai/progress/ with its .gitignore (BEH-15), so a
262// session that runs no tracked skill writes nothing in the project. Then they go to the session's folder in the
263// root of the latest run (DM-02).
264let logRoot: string | null = null
265let early: string[] = []
266const ADAPTER_LOG_LIMIT = 512 * 1024
267const noticed = new Set<string>()
268// The session IDs and roots already pruned (BEH-19), and the retentionDays setting (DM-06).
269const pruned = new Set<string>()
270let retentionDays = 30
271// Whether the open run follows the task list, read once per run from its skill-loaded line.
272let followsFor: { id: string; value: boolean } | null = null
273// Task-tool calls whose step events aren't recorded yet: a question check waits for them (BEH-21), so a question
274// sent in the same batch as the TaskUpdate that marks its step isn't refused. The wait is bounded.
275const taskWork = new Set<Promise<void>>()
276const TASK_WAIT_MS = 2000
277const TASK_TOOLS = ['TaskCreate', 'TaskUpdate', 'TodoWrite']
278
279// Versions 21 to 25. The fuel settings (DM-07, DM-08) are read when the module loads, as retentionDays is, and the notes of a
280// setting that was out of range wait here for session.start, which has the `$` that register lacks.
281let warnFuel = 30
282let runFuel = 20
283let settingNotes: string[] = []
284/** Whether this session's latest registration of /progress succeeded (BEH-34, ERR-21). */
285let progressOk = false
286/** BEH-36's run mark in memory (BEH-36, BEH-35): $.state's precompact.ran is the copy a reload reads. The check and the set
287 * happen with no await between them, so two overlapping measurements start one run. */
288let autoRun = false
289/** BEH-41's set of pending command names (version 25), in the adapter's own memory and not in $.state: BEH-36's run adds
290 * precompact before its $.command.run, and a tile's Start (held with the dashboard) will add its skill's name. The plugin's own
291 * command.run hook takes a name when it sees that run; a failure takes its own name only; a reload, /clear, /resume, /branch
292 * and the session's end empty the set (not a turn's end: the run starts once the session is idle, after the turn). */
293const starts = new Set<string>()
294
295/** Register a task-tool call as under way; the function returned ends it. */
296function startTaskWork(): () => void {
297 let finish = () => {}
298 const work = new Promise<void>(resolve => {
299 finish = resolve
300 })
301 taskWork.add(work)
302 void work.then(() => taskWork.delete(work))
303 return finish
304}
305
306function message(err: unknown): string {
307 return err instanceof Error ? err.message : String(err)
308}
309
310function firstLine(text: string): string {
311 return text.split('\n').map(l => l.trim()).filter(Boolean)[0] ?? ''
312}
313
314function progressDir(r: string): string {
315 return `${r}/devforgeai/progress`
316}
317
318const SESSION_ID = /^[A-Za-z0-9_-]+$/
319
320/** The session's own folder (DM-02): its ID changes at /clear, /resume and /branch, which start a new one. A session
321 * ID is a UUID; one of another shape (empty, or holding '/' or '..') never makes a path (BEH-15). */
322async function sessionDir($: E, r: string): Promise<string> {
323 const id = await $.session.id()
324 if (typeof id !== 'string' || !SESSION_ID.test(id)) throw new Error('unusable session ID')
325 return `${progressDir(r)}/sessions/${id}`
326}
327
328/** The root a run opened in; a run kept in $.state from before version 3 has none, so it comes from the folder. */
329function rootOf(run: ProgressRun): string {
330 return run.root ?? run.dir.replace(/\/devforgeai\/progress\/runs\/[^/]+$/, '')
331}
332
333async function hasSurface($: E): Promise<boolean> {
334 try {
335 return (await $.session.surfaces()).length > 0
336 } catch {
337 return true
338 }
339}
340
341/** A toast, and where nothing draws a dim transcript line too (BEH-01); `key` shows it once per session. */
342async function notify($: E, text: string, key?: string): Promise<void> {
343 if (key !== undefined) {
344 if (noticed.has(key)) return
345 noticed.add(key)
346 }
347 try {
348 await $.ui.toast(text)
349 if (!(await hasSurface($))) await $.ui.log(text)
350 } catch {
351 // a notice that can't be shown changes nothing
352 }
353}
354
355/** One line in the session's adapter.log (DM-02); held in memory until a run has created the folder, kept
356 * to its last half when it passes 512 KiB, and a failed write is ignored. From version 27 no write is tried while a hold lasts or
357 * after the tracker has stopped (BEH-42): the lines wait in memory, the first 200, and the first write after the hold clears or the
358 * stop is lifted writes them ahead of its own line. */
359async function adapterLog($: E, kind: string, text: string, runId?: string): Promise<void> {
360 const step = logChain.then(async () => {
361 const run = await get($, 'run')
362 const now = new Date(await $.clock.now()).toISOString().replace(/\.\d{3}Z$/, 'Z')
363 // One line per entry, whatever the text: model text can't add lines of its own (DM-02).
364 const line = `${now} ${runId ?? run?.id ?? '-'} ${kind}: ${text.replace(/\s*[\r\n]+\s*/g, ' ')}\n`
365 const root = logRoot
366 if (root === null || runHold !== null || odoHold !== null || disabled) {
367 // The first lines are kept (the early notices are the ones that matter); past 200, newer ones are dropped.
368 if (early.length < 200) early = [...early, line]
369 return
370 }
371 const lead = early.join('')
372 early = []
373 await appendLog($, root, lead + line)
374 }).catch(() => undefined)
375 logChain = step
376 await step
377}
378
379async function appendLog($: E, r: string, lines: string): Promise<void> {
380 const path = `${await sessionDir($, r)}/adapter.log`
381 let before = (await $.fs.exists(path)) ? await $.fs.read(path) : ''
382 if (byteSize(before) > ADAPTER_LOG_LIMIT) before = before.slice(Math.floor(before.length / 2)).replace(/^[^\n]*\n/, '')
383 await $.fs.write(path, before + lines)
384}
385
386/** The status line, sent only when its text changes (BEH-10). */
387async function refreshStatus($: E): Promise<void> {
388 const now = await $.clock.now()
389 const mode = await read($, MODE)
390 const off = await read($, OFF)
391 // The open run's values and the trail as one moment of the module (version 14).
392 const l = await hydrate($)
393 const idle = !turnOpen && l.lastEventAt > 0 && now - l.lastEventAt > IDLE_MS
394 const text = statusText(l.summary, mode, idle, off, l.trail, holdKind())
395 if (text === lastStatus) return
396 lastStatus = text
397 await $.ui.status(text)
398 if (text !== undefined && !(await hasSurface($))) await $.ui.log(text)
399}
400
401/** Tracking can't go on for the session (ERR-03): a hold gave up, after 10 failed turn tries or at 4 MiB (BEH-42 (f)). The run and
402 * its held lines are dropped, the trail empties and its runs get no run-end (BEH-05, version 14), and the odometer's writes stop
403 * with it (SPEC-016 BEH-10); `/progress retry` is the way back (BEH-34). */
404async function stopTracking($: E, reason: string): Promise<void> {
405 disabled = true
406 const had = (await hydrate($)).trail.length
407 held = null
408 runHold = null
409 odoHold = null
410 parked.clear()
411 pendingPrune = null
412 const dropped = ledgerHeld.length
413 ledgerHeld = []
414 await putMany($, () => ({ run: null, trail: [], returned: [], summary: null }))
415 if (had) await adapterLog($, 'trail', `empty (tracking stopped: ${reason})`)
416 if (dropped > 0) await adapterLog($, 'dashboard', `odometer: dropped ${dropped} lines: the tracker stopped`)
417 await update($, OFF, () => reason)
418 await notify($, STOP_TOAST)
419 await refreshStatus($)
420 redraw($)
421}
422
423/** The holds live in the module's memory and no $.state value changes with them, so nothing else would draw a change of them: a
424 * hold's start and clear, and the stop and its lift, redraw the band and its rows (BEH-42 (c)). */
425function redraw($: E): void {
426 try {
427 $.ui.invalidate('ui.render')
428 } catch {
429 // nothing to redraw
430 }
431}
432
433/** Which hold the status line, the row and /progress show: the run's wins when both exist (BEH-10, BEH-11). */
434function holdKind(): HoldKind | null {
435 return runHold !== null ? 'run' : odoHold !== null ? 'odometer' : null
436}
437
438/** The lines of the open run's log that the file lacks. */
439function heldLacks(): number {
440 let n = held === null ? 0 : Math.max(0, held.lines.length - held.written)
441 for (const p of parked.values()) n += Math.max(0, p.lines.length - p.written)
442 return n
443}
444
445/** What the row, the toast and /progress say of the hold shown, or null. */
446function holdView(): HoldView | null {
447 if (runHold !== null) return { kind: 'run', lines: heldLacks(), path: runHold.path, error: runHold.error }
448 if (odoHold !== null) return { kind: 'odometer', lines: ledgerHeld.length, path: odoHold.path, error: odoHold.error }
449 return null
450}
451
452/** A write that failed: where, and the host's error. */
453type Failed = { path: string; err: unknown }
454
455/** A write of the tracker's has failed (BEH-42 (a), (c)): the hold starts, with one adapter.log line, one toast and a redraw; a hold
456 * that already lasts only takes the latest failed write's path and error. `lacks` is the number of lines the file lacks. */
457async function holdStarts($: E, which: 'run' | 'odometer', path: string, err: unknown, lacks: number): Promise<void> {
458 const error = errorLine(message(err))
459 const at = await $.clock.now()
460 const known = which === 'run' ? runHold : odoHold
461 if (known !== null) {
462 known.path = path
463 known.error = error
464 return
465 }
466 const hold: Hold = { path, error, firstAt: at, tries: 0 }
467 if (which === 'run') runHold = hold
468 else odoHold = hold
469 // The kind is a literal at each call (VER-66's structure test reads it)
470 if (which === 'run') await adapterLog($, 'write', heldLog(lacks, path, error))
471 else await adapterLog($, 'dashboard', `odometer: ${heldLog(lacks, path, error)}`)
472 await notify($, heldToast({ kind: which, lines: lacks, path, error }, progressOk))
473 await refreshStatus($)
474 redraw($)
475}
476
477/** The evaluator can't run, or failed: the work goes on, the user is told (BEH-14, ERR-01, ERR-02, ERR-07). */
478async function failOpen($: E, reason: string): Promise<void> {
479 await update($, OFF, () => reason)
480 await notify($, `DevForgeAI progress: off (${reason})`, `fail-open:${reason}`)
481 await adapterLog($, 'fail-open', reason)
482 await refreshStatus($)
483}
484
485/** A hook's own code failed: log it, tell the user once, and let the event go on (BEH-14). */
486async function recover($: E, hook: string, err: unknown): Promise<void> {
487 const text = `${hook}: ${message(err)}`
488 await adapterLog($, 'error', text)
489 await notify($, `DevForgeAI progress: ${text}`, `error:${text}`)
490}
491
492/** Send the session's log to a root's folder, writing the lines held until then the first time (BEH-15). */
493async function useLogRoot($: E, r: string): Promise<void> {
494 const first = logRoot === null
495 logRoot = r
496 if (!first) return
497 // The lines keep waiting while a hold lasts or the tracker has stopped (BEH-42).
498 if (runHold !== null || odoHold !== null || disabled) return
499 const waiting = early.join('')
500 early = []
501 if (waiting) {
502 const step = logChain.then(() => appendLog($, r, waiting)).catch(() => undefined)
503 logChain = step
504 await step
505 }
506}
507
508/** The .gitignore that keeps devforgeai/progress/ out of git, written again if it was deleted (BEH-15). `force` writes it whether or
509 * not it exists: the probe of BEH-34 after a stop, since a probe that skipped an existing file would claim a recovery it never
510 * tested. The failed write, or null. */
511async function ensureIgnore($: E, r: string, force = false): Promise<Failed | null> {
512 const file = `${progressDir(r)}/.gitignore`
513 try {
514 if (force || !(await $.fs.exists(file))) await $.fs.write(file, '*\n')
515 return null
516 } catch (err) {
517 return { path: file, err }
518 }
519}
520
521/** devforgeai/progress/ and its .gitignore in a run's root, before anything else is written there (BEH-15); the failed write, or
522 * null. A failure holds (BEH-42), it no longer stops tracking; no root is known to adapter.log or the ledger until it works. */
523async function ensureDir($: E, r: string): Promise<Failed | null> {
524 const failed = await ensureIgnore($, r)
525 if (failed === null) await useLogRoot($, r)
526 return failed
527}
528
529/** The open run's lines, read back from events.jsonl when the module was reloaded (BEH-17). A failed read
530 * throws, and nothing is cached: writing a fresh log over the real one would lose the run's events. */
531async function linesOf($: E, run: ProgressRun): Promise<string[]> {
532 if (held !== null && held.id === run.id) return held.lines
533 const kept = parked.get(run.id)
534 if (kept !== undefined) {
535 // The run resumes after a push whose write failed: its lines come back from memory, not from a file that was never written.
536 parked.delete(run.id)
537 held = { id: run.id, lines: kept.lines, written: kept.written }
538 return kept.lines
539 }
540 const lines = (await $.fs.read(`${run.dir}/events.jsonl`)).split('\n').filter(Boolean)
541 held = { id: run.id, lines, written: lines.length }
542 return lines
543}
544
545/** One write of a run's whole log (no append in $.fs): the failed write, or null. */
546async function putLog($: E, run: ProgressRun, text: string): Promise<Failed | null> {
547 const path = `${run.dir}/events.jsonl`
548 try {
549 await $.fs.write(path, text)
550 return null
551 } catch (err) {
552 return { path, err }
553 }
554}
555
556/** Write the run's whole log (no append in $.fs); false when it can't be recorded: it would pass 4 MiB (ERR-11, or the end of a hold,
557 * BEH-42 (f)), or `open` is false and the write failed. `open` is false for a run that isn't the open one (a paused run's run-end):
558 * its lines aren't cached, and its full log stops nothing.
559 * Version 27 (BEH-42): the lines of the open run are the module's before the file has them. While a hold lasts, an event only
560 * advances them and writes nothing; a write that fails starts the hold and still counts as recorded; a paused run's failed write is
561 * dropped with one line and starts no hold. */
562async function writeLog($: E, run: ProgressRun, lines: string[], open = true): Promise<boolean> {
563 const text = lines.join('\n') + '\n'
564 const big = byteSize(text) > LOG_LIMIT
565 const known = held !== null && held.id === run.id ? held.written : null
566 if (open && runHold !== null) {
567 // No write for each event; the file the next try writes is whole (DM-02). The bound ends the hold, and 'event log full' isn't shown.
568 if (big) {
569 await giveUp($, 'bytes')
570 return false
571 }
572 held = { id: run.id, lines, written: known ?? 0 }
573 return true
574 }
575 if (big) {
576 if (!open) return false
577 // Tracking stops for the session: the trail empties and its runs get no run-end (BEH-05, version 14).
578 const had = (await hydrate($)).trail.length
579 await putMany($, () => ({ run: null, trail: [] }))
580 if (had) await adapterLog($, 'trail', `empty (${LOG_FULL})`)
581 await update($, OFF, () => LOG_FULL)
582 await notify($, `DevForgeAI progress: off (${LOG_FULL})`, `full:${run.id}`)
583 await refreshStatus($)
584 return false
585 }
586 const failed = await putLog($, run, text)
587 if (failed === null) {
588 if (open) held = { id: run.id, lines, written: lines.length }
589 else {
590 parked.delete(run.id)
591 if (held !== null && held.id === run.id) held = null
592 }
593 return true
594 }
595 if (!open) {
596 const kept = parked.get(run.id)
597 if (kept !== undefined) {
598 kept.lines = lines // a parked run keeps its lines, the run-end included, for the next try
599 return true
600 }
601 await adapterLog($, 'write', `dropped 1 lines of ${run.id}: ${failed.path}: ${errorLine(message(failed.err))}`)
602 return false
603 }
604 held = { id: run.id, lines, written: known ?? lines.length - 1 }
605 await holdStarts($, 'run', failed.path, failed.err, lines.length - held.written)
606 return true
607}
608
609/** A hold gave up (BEH-42 (f)): 10 failed turn tries, or the held log would pass 4 MiB. Tracking stops as ERR-03 did at the first
610 * failure before version 27. */
611async function giveUp($: E, why: 'tries' | 'bytes'): Promise<void> {
612 const hold = runHold
613 if (hold === null) return
614 const lost = heldLacks()
615 await adapterLog($, 'write', why === 'tries' ? gaveUpLog(hold.tries, hold.path, hold.error, lost) : `gave up at 4 MiB: ${hold.path}; dropped ${lost} lines`)
616 await stopTracking($, NO_WRITE)
617}
618
619/** One item of the record chain: events, and every change of which run is open (a push, an unwind, a switch, a stop,
620 * session.end's run-ends), happen one at a time, so an event never lands in a run's log after its switch (version 14).
621 * Inside an item, call only what never waits on the chain: recordNow, appendEvent, endOpenNow, linesOf, writeLog,
622 * putMany, putFor, adapterLog, prepareRun, afterOpen, returnStepNow, unwindNow, claudeLoadNow, taskStepsNow,
623 * stopTracking. Never record, settle, pendingState, review, switchRun, finishSwitch, finalEvaluation, stopIfAsked or
624 * turnEndUnwind, which wait on the chain and would wait on the item itself. */
625function chained<T>(work: () => Promise<T>): Promise<T> {
626 const step = recordChain.then(work)
627 recordChain = step.catch(() => undefined)
628 return step
629}
630
631/** The open run changes (open, push, unwind): inside the change of putMany, so a hook starting during the mirror
632 * already sees the new tenure (BEH-30 (a)). */
633function onSwitch(): void {
634 tenure += 1
635 worked = false
636 pendingReport = null
637 pendingWrote = []
638}
639
640/** Append one event to the open run (BEH-04); `mark` asks the timer to evaluate (BEH-06). Events are recorded one
641 * at a time, so two hooks that overlap can't take the same seq or lose each other's line. `began` is the tenure
642 * the event's hook began under (tool and answer events). */
643async function record($: E, kind: string, fields: Fields, mark = true, began = -1): Promise<void> {
644 await chained(() => recordNow($, kind, fields, mark, began))
645}
646
647/** One event in a run's log: the open run's (cached lines) or another's (read from its file). Null when nothing was
648 * written: a second run-end, or a log that can't be written. */
649async function appendEvent($: E, run: ProgressRun, kind: string, fields: Fields, open: boolean): Promise<{ run: ProgressRun; now: number } | null> {
650 const now = await $.clock.now()
651 // The seq comes from the run's own lines: a $.state read inside one dispatch sees that dispatch's moment, so
652 // overlapping hooks would read the same seq from it.
653 const kept = open ? undefined : parked.get(run.id)
654 const lines = open ? await linesOf($, run) : kept !== undefined ? kept.lines : (await $.fs.read(`${run.dir}/events.jsonl`)).split('\n').filter(Boolean)
655 // A run already ended (a deliberate stop, BEH-27) gets no second run-end, checked inside the chain (BEH-05, v12);
656 // events after the stop (the turn's end) may follow it, so any run-end in the log counts.
657 if (kind === 'run-end' && lines.some(l => l.includes('"kind":"run-end"'))) return null
658 const seq = lines.length + 1
659 const next: ProgressRun = { ...run, seq }
660 if (!(await writeLog($, next, [...lines, eventLine(run.id, seq, now, kind, fields)], open))) return null
661 return { run: next, now }
662}
663
664/** A paused run's run-end, written while it isn't the open run; a log that can't be read gets none, and the switch,
665 * unwind or session end goes on (BEH-05, ERR-03). The run as it ended, or null. */
666async function endPausedNow($: E, run: ProgressRun, reason: string): Promise<ProgressRun | null> {
667 try {
668 return (await appendEvent($, run, 'run-end', { reason }, false))?.run ?? null
669 } catch (err) {
670 await adapterLog($, 'error', `run-end of ${run.id}: ${firstLine(message(err))}`)
671 return null
672 }
673}
674
675/** Record into the open run, inside a chain item; the ID of the run it recorded into, or null. */
676async function recordNow($: E, kind: string, fields: Fields, mark: boolean, began = -1): Promise<string | null> {
677 const run = await get($, 'run')
678 if (run === null || disabled) return null
679 const got = await appendEvent($, run, kind, fields, true)
680 if (got === null) return null
681 await putFor($, run.id, () => ({ run: got.run, lastEventAt: got.now, ...(mark ? { marked: true } : {}) }))
682 if ((kind === 'tool' || kind === 'answer') && began === tenure) worked = true
683 return run.id
684}
685
686// ---- the retry of a hold (BEH-42 (b), (e), (h)) ----
687
688/** What a try came to, for /progress retry's answer. */
689type Tried = { kind: HoldKind; ok: boolean; saved: number; path: string; error: string; lines: number; tries: number }
690
691/** The hold of the run's log has cleared: the lines it kept are in the file (BEH-42 (e)). The run is marked, so the timer's next
692 * evaluation, the first since the hold and over the whole file, runs once (BEH-06); a new run's pruning, which waited for its
693 * folder, starts (BEH-19). */
694async function runRecovered($: E, run: ProgressRun, wrote: number): Promise<void> {
695 const hold = runHold
696 if (hold === null) return
697 runHold = null
698 await useLogRoot($, rootOf(run))
699 await adapterLog($, 'write', recoveredLog(hold.tries, `${run.dir}/events.jsonl`, wrote))
700 await putFor($, run.id, () => ({ marked: true }))
701 await notify($, recoveredToast(wrote))
702 const prune = pendingPrune
703 pendingPrune = null
704 if (prune !== null && prune.keepRun === run.id) await startPrune($, prune.r, prune.session, prune.keepRun, prune.continued)
705 await refreshStatus($)
706 redraw($)
707}
708
709/** The lines of the parked runs, written whole each into its own file (C1): the failed write, or null; `wrote` counts the lines saved. */
710async function flushParked($: E, count: { wrote: number }): Promise<Failed | null> {
711 for (const [id, p] of [...parked]) {
712 const failed = (await ensureIgnore($, rootOf(p.run))) ?? (await putLog($, p.run, p.lines.join('\n') + '\n'))
713 if (failed !== null) return failed
714 parked.delete(id)
715 count.wrote += Math.max(0, p.lines.length - p.written)
716 }
717 return null
718}
719
720/** One try at the run's hold: the folder and its .gitignore if missing (BEH-15), then the whole log. `how` is 'turn' at a main-loop
721 * turn's end (a failure adds 1 to the count, and the tenth gives up), 'command' for /progress retry (it adds nothing), and 'leave'
722 * when the open run leaves memory (BEH-42 (h)): a failure then drops the run's unwritten lines with one line, and the hold goes on.
723 * 'park' is a push (BEH-29): a run that never had a file keeps its lines for its resume instead (C1), the others are dropped as 'leave'. */
724async function tryRun($: E, how: 'turn' | 'command' | 'leave' | 'park'): Promise<Tried | null> {
725 const hold = runHold
726 if (hold === null) return null
727 const run = (await hydrate($)).run
728 if (run === null) {
729 // The run is gone (a stale hold): nothing is left to write.
730 runHold = null
731 parked.clear()
732 await refreshStatus($)
733 redraw($)
734 return null
735 }
736 const lines = await linesOf($, run)
737 const lacks = held !== null && held.id === run.id ? Math.max(0, held.lines.length - held.written) : 0
738 const count = { wrote: 0 }
739 const failed = (await ensureIgnore($, rootOf(run))) ?? (await flushParked($, count)) ?? (await putLog($, run, lines.join('\n') + '\n'))
740 if (failed === null) {
741 held = { id: run.id, lines, written: lines.length }
742 await runRecovered($, run, lacks + count.wrote)
743 return { kind: 'run', ok: true, saved: lacks + count.wrote, path: `${run.dir}/events.jsonl`, error: '', lines: 0, tries: hold.tries }
744 }
745 const error = errorLine(message(failed.err))
746 hold.path = failed.path
747 hold.error = error
748 if (how === 'park' && held !== null && held.id === run.id && held.written === 0) {
749 // The run never had a file: its lines stay in memory, in the hold, for its resume or the next try that works.
750 parked.set(run.id, { run, lines, written: 0 })
751 return null
752 }
753 if (how === 'leave' || how === 'park') {
754 // The run's unwritten lines are written off: they stay in the cache, so that the run's run-end is still seen as written when a
755 // second chain item ends the same run (finishSwitch), and nothing will write them; the next run's lines replace the cache.
756 held = { id: run.id, lines, written: lines.length }
757 await adapterLog($, 'write', `dropped ${lacks} lines of ${run.id}: ${failed.path}: ${error}`)
758 return null
759 }
760 if (how === 'turn') hold.tries += 1
761 const tried: Tried = { kind: 'run', ok: false, saved: 0, path: failed.path, error, lines: lacks, tries: hold.tries }
762 if (how === 'turn' && hold.tries >= HOLD_TRIES) await giveUp($, 'tries')
763 return tried
764}
765
766/** One try at the odometer's hold: the whole file, with the lines it holds and the lines held (SPEC-016 ERR-06). It runs in the
767 * ledger's own chain, so it never races a line being added. */
768function tryOdometer($: E, how: 'turn' | 'command'): Promise<Tried | null> {
769 return ledgerRun<Tried | null>(async () => {
770 const hold = odoHold
771 const file = ledgerFile
772 if (hold === null) return null
773 if (file === null) {
774 odoHold = null
775 await refreshStatus($)
776 redraw($)
777 return null
778 }
779 const count = ledgerHeld.length
780 const lines = [...file.lines, ...ledgerHeld]
781 const path = file.path
782 try {
783 await $.fs.write(path, lines.join('\n') + '\n')
784 } catch (err) {
785 const error = errorLine(message(err))
786 hold.path = path
787 hold.error = error
788 if (how === 'turn') hold.tries += 1
789 const tried: Tried = { kind: 'odometer', ok: false, saved: 0, path, error, lines: count, tries: hold.tries }
790 if (how === 'turn' && hold.tries >= HOLD_TRIES) await odometerGivesUp($, hold, path, error)
791 return tried
792 }
793 file.lines = lines
794 ledgerHeld = ledgerHeld.slice(count)
795 odoHold = null
796 await adapterLog($, 'dashboard', `odometer: ${recoveredLog(hold.tries, path, count)}`)
797 await notify($, recoveredToast(count))
798 await refreshStatus($)
799 redraw($)
800 return { kind: 'odometer', ok: true, saved: count, path, error: '', lines: 0, tries: hold.tries }
801 })
802}
803
804/** The odometer's hold gave up (SPEC-016 ERR-06): its held lines are dropped, nothing is written for the session until
805 * `/progress retry` lifts the stop, and the tracker goes on untouched. */
806async function odometerGivesUp($: E, hold: Hold, path: string, error: string): Promise<void> {
807 const dropped = ledgerHeld.length
808 ledgerHeld = []
809 odoHold = null
810 odoGaveUp = true
811 await adapterLog($, 'dashboard', `odometer: ${gaveUpLog(hold.tries, path, error, dropped)}`)
812 await notify($, odometerGaveUpToast(path))
813 await refreshStatus($)
814 redraw($)
815}
816
817/** Try the holds once, in the record chain, so a try never races an event being recorded (BEH-42 (b)): the run's log first, then the
818 * odometer's file. A hold that starts in a turn's own records is
819 * tried at that turn's end too, so a hold stops after exactly 10 failed turn tries (BEH-42 (b), QR-05). */
820function retryHolds($: E, how: 'turn' | 'command'): Promise<Tried[]> {
821 return chained(async () => {
822 const out: Tried[] = []
823 if (runHold !== null) {
824 const tried = await tryRun($, how)
825 if (tried !== null) out.push(tried)
826 }
827 if (odoHold !== null) {
828 const tried = await tryOdometer($, how)
829 if (tried !== null) out.push(tried)
830 }
831 return out
832 })
833}
834
835/** The run as it stands: events still being written, the evaluation in flight, then one evaluation of a marked run
836 * (BEH-26, BEH-28, BEH-30 (c)). */
837async function settle($: E): Promise<void> {
838 await recordChain
839 if (inFlight !== null) await inFlight
840 // No evaluation starts while the run's log is held: the file is behind the lines (BEH-42 (d)).
841 if (!evaluating && !disabled && runHold === null && (await get($, 'marked'))) {
842 inFlight = evaluateMarked($)
843 await inFlight
844 }
845}
846
847async function logBytes($: E): Promise<number> {
848 const run = await get($, 'run')
849 return run === null ? 0 : byteSize((await linesOf($, run)).join('\n'))
850}
851
852/** A Write's content, or the file an Edit will leave (DM-01, ERR-06). */
853async function contentOf($: E, r: string, tool: string, input: Fields): Promise<string | null> {
854 if (tool === 'Write') return typeof input.content === 'string' ? input.content : null
855 if (tool !== 'Edit' || typeof input.file_path !== 'string') return null
856 try {
857 const path = input.file_path.startsWith('/') ? input.file_path : `${r}/${input.file_path}`
858 const file = await $.fs.read(path)
859 return editResult(file, String(input.old_string ?? ''), String(input.new_string ?? ''), input.replace_all === true)
860 } catch {
861 return null
862 }
863}
864
865/** What a skill's load is (BEH-02): 'tracked' (one of the plugin's own, or a project or organization manifest's), 'untracked'
866 * (one of the plugin's own whose SKILL.md metadata has devforgeai-tracked "false", version 18) or 'other', which neither
867 * starts nor ends a run. A plugin skill's SKILL.md is read at each of its loads, never cached, so a deployed change shows
868 * at the next load; only a plugin skill can be untracked, and a manifest of its name changes nothing. A SKILL.md that can't
869 * be read leaves the skill tracked, with one adapter.log line (ERR-18). */
870async function skillKind($: E, r: string, name: string): Promise<'tracked' | 'untracked' | 'other'> {
871 if (pluginSkills === null) {
872 try {
873 pluginSkills = (await $.fs.list(`${$.plugin.root}/skills`)).filter(x => x.kind === 'dir').map(x => x.name)
874 } catch {
875 pluginSkills = []
876 }
877 }
878 if (pluginSkills.includes(name)) {
879 try {
880 return isUntrackedSkill(await $.fs.read(`${$.plugin.root}/skills/${name}/SKILL.md`)) ? 'untracked' : 'tracked'
881 } catch (err) {
882 await adapterLog($, 'skill-read', `${name}: ${firstLine(message(err))}`)
883 return 'tracked'
884 }
885 }
886 const manifests: string[] = []
887 for (const dir of [`${r}/devforgeai/manifests`, `${r}/devforgeai/manifests/organization`]) {
888 try {
889 if (await $.fs.exists(dir)) {
890 for (const x of await $.fs.list(dir)) if (x.kind === 'file' && x.name.endsWith('.json')) manifests.push(x.name.slice(0, -5))
891 }
892 } catch {
893 // an unreadable folder adds no manifest
894 }
895 }
896 return isTracked(name, pluginSkills, manifests) ? 'tracked' : 'other'
897}
898
899/** Find python3, then python (IF-03, ERR-01). */
900async function findPython($: E): Promise<void> {
901 if (python !== undefined) return
902 const r = await $.session.root()
903 for (const name of ['python3', 'python']) {
904 try {
905 const res = await $.process.run([name, '--version'], { cwd: r, timeoutMs: START_TIMEOUT })
906 if (res.exitCode === 0) {
907 python = name
908 return
909 }
910 } catch {
911 // not installed: try the next name
912 }
913 }
914 python = null
915 await failOpen($, NO_PYTHON)
916}
917
918/** Resolve progress.mode for a root with settings.py (IF-01, BEH-16). */
919async function resolveMode($: E, r: string): Promise<void> {
920 let mode: ProgressMode = 'observe'
921 let source: ProgressModeSource = 'framework-default'
922 if (python) {
923 try {
924 const res = await $.process.run([python, `${$.plugin.root}/progress/settings.py`, 'mode', '--root', r],
925 { cwd: r, timeoutMs: START_TIMEOUT })
926 const [value, from] = res.stdout.trim().split(/\s+/)
927 if (res.exitCode === 0 && (value === 'observe' || value === 'enforce')) {
928 mode = value
929 source = from === 'local' ? 'local' : 'framework-default'
930 }
931 const ignored = res.stderr.split('\n').map(l => l.trim()).filter(l => l.startsWith('ignored '))
932 for (const line of ignored) await adapterLog($, 'ignored', line)
933 if (ignored.length) await notify($, ignored.join('\n'), `ignored:${ignored.join('|')}`)
934 } catch (err) {
935 await adapterLog($, 'error', `settings.py mode: ${message(err)}`)
936 }
937 }
938 await update($, MODE, () => mode)
939 await update($, SOURCE, () => source)
940 modeSession = await $.session.id()
941 modeRoot = r
942 await adapterLog($, 'mode', `${mode} (${source})`)
943}
944
945/** Run the evaluator (IF-03) in a run's root; the state, or why it couldn't be had (ERR-01, ERR-02, ERR-07). */
946async function evaluate($: E, r: string, events: string, out: string, timeoutMs: number): Promise<{ state: ProgressState; text: string } | string> {
947 if (!python) return NO_PYTHON
948 const plugin = $.plugin.root
949 const argv = [python, `${plugin}/progress/evaluate.py`, 'evaluate', '--manifests', `${plugin}/progress/manifests`]
950 if (await $.fs.exists(`${r}/devforgeai/manifests/organization`)) argv.push('--manifests', `${r}/devforgeai/manifests/organization`)
951 if (await $.fs.exists(`${r}/devforgeai/manifests`)) argv.push('--manifests', `${r}/devforgeai/manifests`)
952 argv.push('--events', events, '--out', out, '--root', r)
953 let res
954 try {
955 res = await $.process.run(argv, { cwd: r, timeoutMs })
956 } catch (err) {
957 return /time/i.test(message(err)) ? 'evaluator timed out' : `evaluator: ${firstLine(message(err))}`
958 }
959 if (res.exitCode !== 0) return `evaluator: ${firstLine(res.stderr) || `exit ${res.exitCode}`}`
960 try {
961 const text = await $.fs.read(out)
962 return { state: JSON.parse(text) as ProgressState, text }
963 } catch {
964 return 'evaluator: its state is not readable'
965 }
966}
967
968/** Whether a run follows the task list (SPEC-012 BEH-18), read from its skill-loaded line once per run. */
969async function follows($: E, run: ProgressRun): Promise<boolean> {
970 if (followsFor === null || followsFor.id !== run.id) followsFor = { id: run.id, value: followsTaskList((await linesOf($, run))[0]) }
971 return followsFor.value
972}
973
974/** The run's lines when it follows the task list, else null: the write line and the user's toast sentence read the
975 * task list's mark from them (BEH-08, BEH-12). */
976async function followingLines($: E, run: ProgressRun): Promise<readonly string[] | null> {
977 return (await follows($, run)) ? linesOf($, run) : null
978}
979
980/** What a compaction of a run that follows the task list keeps, and its closing note (BEH-24). The marked step counts
981 * only when the last evaluation's steps have it, with its title from there; with no evaluation yet, its number alone.
982 * `reached` is true when that evaluation shows every step reached, which leaves the note out (version 8). */
983async function compactNotes($: E, run: ProgressRun): Promise<{ marked: number | null; reached: boolean; instruction: string; note: string }> {
984 const lines = await linesOf($, run)
985 let state: ProgressState | null = null
986 try {
987 state = JSON.parse(await $.fs.read(`${run.dir}/state.json`)) as ProgressState
988 } catch {
989 // no evaluation yet
990 }
991 const steps = state !== null && Array.isArray(state.steps) ? state.steps : null
992 const marked = markedStep(lines, Infinity, steps)
993 const label = marked === null ? null : steps !== null ? stepLabel(steps, marked) : `step ${marked}`
994 const reached = state !== null && state.current === null && state.ended === null
995 return { marked, reached, ...compactTexts(run.skill, label) }
996}
997
998/** Whether the session has a task list, from the tools it offers now, deferred ones included (DM-01); when the list
999 * can't be read, none, and `read` is false, so no hint blames the session's tools (ERR-14). */
1000async function taskListOf($: E): Promise<{ taskList: boolean; read: boolean }> {
1001 try {
1002 return { taskList: hasTaskList((await $.tool.list()).map(t => t.name)), read: true }
1003 } catch (err) {
1004 await adapterLog($, 'tools', firstLine(message(err)))
1005 return { taskList: false, read: false }
1006 }
1007}
1008
1009/** Step events from a task tool's call that didn't fail, recorded after its tool event, in the run it was recorded
1010 * into, inside the same chain item (BEH-20, ERR-13): the tool event, its step events and its task map land together. */
1011async function taskStepsNow($: E, into: string, tool: string, input: Fields, outcome: ToolOutcome): Promise<void> {
1012 if (isFailed(outcome)) return
1013 const l = await hydrate($)
1014 if (l.run === null || l.run.id !== into) return
1015 if (tool === 'TaskCreate') {
1016 const id = taskIdOf(outcome.text)
1017 const step = stepOfTask(input)
1018 if (id === null || step === null) {
1019 // The subject is the model's text: one line of it, so it can't add lines of its own to adapter.log.
1020 await adapterLog($, 'task', `no step for task: ${firstLine(String(input.subject ?? ''))}`)
1021 return
1022 }
1023 await putFor($, into, x => ({ tasks: { ...x.tasks, [id]: step } }))
1024 } else if (tool === 'TaskUpdate') {
1025 const step = l.tasks[String(input.taskId)]
1026 const state = stepStateOf(input.status)
1027 if (step !== undefined && state !== null) await recordNow($, 'step', { step, state }, true)
1028 } else if (tool === 'TodoWrite') {
1029 const result = (outcome as Fields).result
1030 let events: Fields[] = []
1031 await putFor($, into, x => {
1032 const got = todoSteps(input.todos, x.todos, result && typeof result === 'object' ? (result as Fields).oldTodos : undefined)
1033 events = got.events
1034 return { todos: got.statuses }
1035 })
1036 for (const e of events) await recordNow($, 'step', e, true)
1037 }
1038}
1039
1040/** Take an evaluation in: the session's current.json, the summary, toasts, the report gate's context (BEH-06, BEH-09,
1041 * BEH-12). */
1042async function absorb($: E, run: ProgressRun, got: { state: ProgressState; text: string }, cleanup = false): Promise<void> {
1043 // One at a time: the run-end evaluation and a timer's can finish together, and each reads the shown flags and
1044 // adhered before it writes them, so a notice could otherwise show twice.
1045 const step = absorbChain.then(() => absorbNow($, run, got, cleanup))
1046 absorbChain = step.catch(() => undefined)
1047 await step
1048}
1049
1050async function absorbNow($: E, run: ProgressRun, got: { state: ProgressState; text: string }, cleanup: boolean): Promise<void> {
1051 // An evaluation whose run isn't the open run changes nothing: a push, an unwind or a switch came first (BEH-30).
1052 if ((await hydrate($)).run?.id !== run.id) return
1053 try {
1054 await $.fs.write(`${await sessionDir($, rootOf(run))}/current.json`, got.text)
1055 } catch {
1056 // renderers read current.json; the run's own state.json is written
1057 }
1058 const off = await read($, OFF)
1059 if (off !== null && off !== LOG_FULL && off !== NO_PYTHON && off !== NO_WRITE) await update($, OFF, () => null)
1060 const following = await followingLines($, run)
1061 let fresh: { keys: string[]; toasts: string[] } = { keys: [], toasts: [] }
1062 // The stopping answer's step stays with a stopped run's summary (BEH-10, version 12).
1063 const open = await putFor($, run.id, l => {
1064 fresh = newFlagToasts(got.state, l.shown, following)
1065 return { summary: { ...summaryOf(got.state), stoppedAt: got.state.ended === 'stopped' ? l.summary?.stoppedAt ?? null : null },
1066 ...(fresh.keys.length ? { shown: [...l.shown, ...fresh.keys] } : {}) }
1067 })
1068 if (!open) return
1069 for (const toast of fresh.toasts) await notify($, toast)
1070 if (got.state.current === null && got.state.ended === null && allReachedFor !== run.id) {
1071 allReachedFor = run.id
1072 await notify($, `✓ ${got.state.skill}: all steps reached`)
1073 }
1074 // SPEC-012 §4's second level: a run that follows the task list and didn't keep it is told so once (BEH-22).
1075 const adherence = adherenceText(got.state)
1076 // $.state's adhered decides, so a reload of the module doesn't repeat the notice (version 6).
1077 if (adherence !== null && (await get($, 'adhered')) !== run.id && (await follows($, run))
1078 && (await putFor($, run.id, () => ({ adhered: run.id })))) {
1079 await notify($, adherence)
1080 await adapterLog($, 'adherence', adherence)
1081 }
1082 const report = reportContext(got.state)
1083 const l = await hydrate($)
1084 if (l.run?.id === run.id) {
1085 pendingReport = report !== null && !l.contextSent.includes(report.seq) && got.state.ended === null ? { ...report, run: run.id } : null
1086 }
1087 // BEH-38's window, from the last evaluated state (version 22): closed once every step is reached, the run has ended, or a
1088 // step with a script rule of target written is done; open again if a later evaluation says otherwise.
1089 const closed = windowClosed(got.state, (await gatedFor($, run)).scriptSteps)
1090 await putFor($, run.id, l => (closed === (l.closedFor === run.id) ? {} : { closedFor: closed ? run.id : null }))
1091 if (cleanup) await startCleanup($, run, got.state)
1092 await refreshStatus($)
1093}
1094
1095/** The timer's work (BEH-06): it catches every error itself, since one it let escape reaches only the debug log (ERR-10). */
1096async function tick($: E): Promise<void> {
1097 try {
1098 await refreshStatus($)
1099 // While the run's log is held the timer starts no evaluation and clears no mark (BEH-06, BEH-42 (d)).
1100 if (evaluating || disabled || runHold !== null || !(await get($, 'marked'))) return
1101 inFlight = evaluateMarked($)
1102 await inFlight
1103 } catch (err) {
1104 await failOpen($, `timer: ${message(err)}`).catch(() => undefined)
1105 }
1106}
1107
1108/** One evaluation of the open run, as the timer runs it (BEH-06); the review waits for it (BEH-26). */
1109async function evaluateMarked($: E): Promise<void> {
1110 evaluating = true // before any await, so the timer and the review never start two evaluations
1111 try {
1112 const run = await get($, 'run')
1113 if (run === null) return
1114 try {
1115 await putFor($, run.id, () => ({ marked: false }))
1116 // The folder's .gitignore may have gone with a `git clean` while the run went on (BEH-15).
1117 await ensureIgnore($, rootOf(run))
1118 const got = await evaluate($, rootOf(run), `${run.dir}/events.jsonl`, `${run.dir}/state.json`, EVALUATOR_TIMEOUT)
1119 const open = await get($, 'run')
1120 // A run that opened, paused or resumed while the evaluator ran keeps its own summary and flags (the module's values).
1121 if (open === null || open.id !== run.id) return
1122 if (typeof got === 'string') await failOpen($, got)
1123 else await absorb($, run, got, true)
1124 } finally {
1125 evaluating = false
1126 }
1127 } finally {
1128 evaluating = false
1129 inFlight = null
1130 }
1131}
1132
1133function ensureTimer($: E): void {
1134 if (timerOn) return
1135 timerOn = true
1136 $.clock.every(500, () => {
1137 void tick($)
1138 })
1139}
1140
1141/** What every interactive session does first: python, the mode, the timer (BEH-01, BEH-06, BEH-16, BEH-17). */
1142async function setup($: E): Promise<void> {
1143 await findPython($)
1144 await resolveMode($, await $.session.root())
1145 ensureTimer($)
1146 // A reload ends the holds and loses the lines they kept (BEH-42 (i)): the module's memory starts over, as it does in Claude Code
1147 // (a first session.start finds it empty already; the kit cannot reload the module, so its second session.start stands for one).
1148 const hadHold = runHold !== null || odoHold !== null
1149 if (odoHold !== null) {
1150 ledgerHeld = []
1151 ledgerFile = null
1152 }
1153 held = null
1154 runHold = null
1155 odoHold = null
1156 parked.clear()
1157 odoGaveUp = false
1158 pendingPrune = null
1159 const run = await get($, 'run')
1160 if (run !== null) {
1161 // A run opened in memory whose first write never worked has no events.jsonl: nothing is left to go on from (BEH-42 (i)).
1162 let missing = false
1163 try {
1164 missing = !(await $.fs.exists(`${run.dir}/events.jsonl`))
1165 } catch {
1166 missing = false // a read that fails for another reason is BEH-17's
1167 }
1168 if (missing) {
1169 await putMany($, () => ({ run: null }))
1170 await adapterLog($, 'write', 'dropped the run: no events.jsonl after a reload', run.id)
1171 await refreshStatus($)
1172 } else {
1173 // After a reload the module's variables start over; the open run's folder exists, so its log goes there.
1174 await useLogRoot($, rootOf(run))
1175 await put($, 'marked', () => true)
1176 }
1177 }
1178 if (hadHold) redraw($)
1179}
1180
1181/** A new run whose skill-loaded line is written but which isn't open yet (BEH-03, BEH-29). From version 27 `failure` names the write
1182 * that failed: the run then opens in memory with its line held (BEH-42 (a)). */
1183type Opening = { run: ProgressRun; line: string; now: number; session: string; r: string; listed: boolean; taskList: boolean; checklist: string
1184 failure: Failed | null }
1185
1186/** Write a new run's skill-loaded line in its own folder, in the root read as it loaded (BEH-03). A write that fails (the folder, its
1187 * .gitignore or the line) is the Opening's `failure`, not a stop: the run opens in memory, and the caller's `afterOpen` starts the
1188 * hold. Opening it is the caller's change, made with no await (version 14). */
1189async function prepareRun($: E, r: string, skill: string, checklist: string, extra: Fields = {}): Promise<Opening> {
1190 const dirFailed = await ensureDir($, r)
1191 const session = await $.session.id()
1192 // The mode is resolved here, after the folder exists, so a new root's mode line lands in that root's log (BEH-16).
1193 if (modeSession !== session || modeRoot !== r) await resolveMode($, r)
1194 const now = await $.clock.now()
1195 const id = runId(now, skill, crypto.getRandomValues(new Uint8Array(4)))
1196 const dir = `${progressDir(r)}/runs/${id}`
1197 const version = await $.session.version()
1198 const { taskList, read: listed } = await taskListOf($)
1199 const line = eventLine(id, 1, now, 'skill-loaded', {
1200 format: FORMAT, skill, checklist, host: `claude-code ${version.version}`, taskList,hooks/progress-core.ts 1512 lines1// Pure helpers of the progress tracker adapter (SPEC-013 v7). No `$` here: claude plugin validate lets `$` reach
2// only top-level functions of hooks/progress.tsx, so this file turns plain data into plain data, and its tests
3// (core.test.ts) call it directly.
4import type { ProgressMode, ProgressPaused, ProgressPrecompact, ProgressRefused, ProgressSummary } from '../types'
5
6export type Fields = Record<string, unknown>
7
8/** The 64 KiB a Write or Edit's content may take in an event, and the log sizes of ERR-11 (BEH-15). */
9export const CONTENT_LIMIT = 64 * 1024
10export const LOG_CONTENT_LIMIT = 3 * 1024 * 1024
11export const LOG_LIMIT = 4 * 1024 * 1024
12export const IDLE_MS = 30 * 60 * 1000
13export const FORMAT = 'devforgeai-events/1'
14
15/** How BEH-38 (b)'s enforce-mode text reaches the model (the P14 switch, SPEC-013 version 22): 'result' adds it to the context
16 * of the Bash call's own result, which probe P14 showed a tool.call hook can do after next(e) (Claude Code 2.1.291,
17 * 2026-10-06); 'prompt' gives it to the context of the user's next prompt instead, as BEH-09's report does. The result
18 * route falls back to the prompt when the call's answer is a deny, which can't carry context. */
19export const WROTE_CONTEXT_ROUTE: 'result' | 'prompt' = 'result'
20
21// Each kind's fields in SPEC-012 DM-02's order, after run, seq, time and kind. The order is fixed here so an
22// event line is the same bytes whichever code built it (VER-04 compares lines byte for byte).
23const ORDER: Record<string, readonly string[]> = {
24 'skill-loaded': ['format', 'skill', 'checklist', 'host', 'taskList', 'mode', 'modeSource', 'resumes', 'carried', 'answered', 'draft'],
25 tool: ['tool', 'path', 'command', 'exit', 'error', 'content', 'wrote'],
26 answer: ['answered', 'step', 'outside', 'waiver', 'questions'],
27 prompt: [],
28 reply: ['text'],
29 turn: ['phase'],
30 'run-end': ['reason'],
31 step: ['step', 'state'],
32 usage: ['turn', 'model', 'input', 'output', 'cacheRead', 'cacheWrite'],
33}
34
35/** UTC time as ISO 8601 to the second. */
36export function isoTime(ms: number): string {
37 return new Date(ms).toISOString().replace(/\.\d{3}Z$/, 'Z')
38}
39
40/** One event as a JSON line (DM-01): run, seq, time, kind, then the kind's fields; undefined fields left out. */
41export function eventLine(run: string, seq: number, ms: number, kind: string, fields: Fields = {}): string {
42 const out: Fields = { run, seq, time: isoTime(ms), kind }
43 for (const key of ORDER[kind] ?? []) if (fields[key] !== undefined) out[key] = fields[key]
44 return JSON.stringify(out)
45}
46
47/** The run ID: UTC time as yyyymmddThhmmssZ, the skill, 8 hex digits (SPEC-012 §4; BEH-03). */
48/** A skill's name as its run IDs carry it (DM-02's pattern): lower case, other characters as '-', from its first
49 * letter. The offer looks for earlier runs under the same name (BEH-31). */
50export function runName(skill: string): string {
51 return skill.toLowerCase().replace(/[^a-z0-9-]+/g, '-').replace(/^[^a-z]+/, '') || 'skill'
52}
53
54export function runId(ms: number, skill: string, random: Uint8Array): string {
55 const t = isoTime(ms).replace(/[-:]/g, '')
56 const name = runName(skill)
57 const hex = Array.from(random.slice(0, 4), b => b.toString(16).padStart(2, '0')).join('')
58 return `${t}-${name}-${hex.padEnd(8, '0')}`
59}
60
61/** A skill's name without a '<plugin>:' prefix (BEH-02; §9, P3). */
62export function skillName(raw: string): string {
63 const i = raw.lastIndexOf(':')
64 return i >= 0 ? raw.slice(i + 1) : raw
65}
66
67/** Whether a skill opens a run: one of the plugin's own, or one a project manifest names (BEH-02). */
68export function isTracked(name: string, pluginSkills: readonly string[], manifestNames: readonly string[]): boolean {
69 return pluginSkills.includes(name) || manifestNames.includes(name)
70}
71
72/** Whether a plugin skill's SKILL.md marks it untracked (BEH-02, version 18). Its frontmatter is the text between its
73 * first line, '---', and the next line that is '---'; its metadata block is the frontmatter's line 'metadata:' at
74 * column 0 and the lines after it that start with a space. The skill is untracked when that block holds a line
75 * devforgeai-tracked: "false", the value in double or single quotes, with leading and trailing spaces, a trailing
76 * '# comment' and a CR ignored (a file with CRLF line ends reads the same). Anything else is tracked: no such line, the
77 * key outside the block, false unquoted (a YAML boolean), "False" or any other value, no frontmatter. */
78export function isUntrackedSkill(skillMd: string): boolean {
79 const lines = skillMd.split('\n').map(l => (l.endsWith('\r') ? l.slice(0, -1) : l))
80 if (lines[0] !== '---') return false
81 const close = lines.indexOf('---', 1)
82 if (close < 0) return false
83 const front = lines.slice(1, close)
84 const at = front.findIndex(l => l.trimEnd() === 'metadata:')
85 if (at < 0) return false
86 for (let i = at + 1; i < front.length && front[i].startsWith(' '); i++) {
87 if (UNTRACKED_LINE.test(front[i])) return true
88 }
89 return false
90}
91
92/** The line that marks a skill untracked: the key, then the value "false" or 'false', then at most a comment. */
93const UNTRACKED_LINE = /^ *devforgeai-tracked:[ \t]+(?:"false"|'false')[ \t]*(?:[ \t]#.*)?$/
94
95/** A path relative to the project root with / separators; a path outside the root stays absolute. */
96export function relPath(root: string, path: string): string {
97 const p = path.replace(/\\/g, '/')
98 const r = root.replace(/\\/g, '/').replace(/\/+$/, '')
99 if (p === r) return '.'
100 return p.startsWith(r + '/') ? p.slice(r.length + 1) : p
101}
102
103function str(value: unknown): string | undefined {
104 return typeof value === 'string' ? value : undefined
105}
106
107/** A tool event's path field (DM-01). */
108export function toolPath(root: string, tool: string, input: Fields): string | undefined {
109 if (tool === 'Read' || tool === 'Write' || tool === 'Edit') {
110 const f = str(input.file_path)
111 return f === undefined ? undefined : relPath(root, f)
112 }
113 if (tool === 'Glob') {
114 const pattern = str(input.pattern)
115 if (pattern === undefined) return undefined
116 const base = str(input.path)
117 return base === undefined ? pattern : `${relPath(root, base).replace(/\/+$/, '')}/${pattern}`
118 }
119 if (tool === 'Grep') {
120 const base = str(input.path)
121 return base === undefined ? '.' : relPath(root, base)
122 }
123 return undefined
124}
125
126/** What next(e) resolved to for a tool call, as far as the adapter reads it. */
127export type ToolOutcome = { deny?: unknown; isError?: unknown; text?: unknown; result?: unknown }
128
129/** A call refused by anyone ({ deny }) or that failed (isError) never ran as asked: error true (DM-01). */
130export function isFailed(outcome: ToolOutcome): boolean {
131 return outcome.deny !== undefined || outcome.isError === true
132}
133
134/** The exit field: 'Exit code <n>' from a failed call's text, 0 when it succeeded, else null (§9, P6). */
135export function exitOf(outcome: ToolOutcome): number | null {
136 if (outcome.deny !== undefined) return null
137 if (outcome.isError !== true) return 0
138 const text = typeof outcome.text === 'string' ? outcome.text : typeof outcome.result === 'string' ? outcome.result : ''
139 const m = text.match(/Exit code (\d+)/)
140 return m ? Number(m[1]) : null
141}
142
143/** An AskUserQuestion was answered: it didn't fail, and its answers hold at least one entry (§9, P5). */
144export function isAnswered(outcome: ToolOutcome): boolean {
145 if (isFailed(outcome)) return false
146 const result = outcome.result as { answers?: unknown } | undefined
147 const answers = result && typeof result === 'object' ? result.answers : undefined
148 return !!answers && typeof answers === 'object' && Object.keys(answers as object).length > 0
149}
150
151/** UTF-8 size of a string. */
152export function byteSize(text: string): number {
153 return new TextEncoder().encode(text).length
154}
155
156/** The file an Edit will leave, or null when it can't be computed (ERR-06). */
157export function editResult(file: string, oldString: string, newString: string, replaceAll: boolean): string | null {
158 if (oldString === '' || byteSize(file) > CONTENT_LIMIT) return null
159 const first = file.indexOf(oldString)
160 if (first < 0) return null
161 if (replaceAll) return file.split(oldString).join(newString)
162 if (file.indexOf(oldString, first + oldString.length) >= 0) return null
163 return file.slice(0, first) + newString + file.slice(first + oldString.length)
164}
165
166/** Content kept in a tool event: at most 64 KiB, and none once the log has passed 3 MiB (BEH-15, ERR-11). */
167export function keptContent(content: string | null | undefined, logBytes: number): string | undefined {
168 if (content === null || content === undefined) return undefined
169 if (logBytes >= LOG_CONTENT_LIMIT || byteSize(content) > CONTENT_LIMIT) return undefined
170 return content
171}
172
173/** The form of an AskUserQuestion call as the record keeps it (BEH-37; version 22): for each of the input's questions, in order,
174 * its question, header, multiSelect and options (each option's label, description and preview), a field only when the input
175 * has it with that type, and nothing else: the newer shapes' fields and the result's answers are never copied. A question
176 * with no string question, and an option with no string label, is skipped (DM-02 requires them). Undefined when nothing
177 * is left, when the record would pass 64 KiB of UTF-8 as JSON, or when the log has passed 3 MiB (ERR-11): dropped whole. */
178export function formOf(input: Fields, logBytes: number): Fields[] | undefined {
179 const raw = input.questions
180 if (!Array.isArray(raw) || logBytes >= LOG_CONTENT_LIMIT) return undefined
181 const out: Fields[] = []
182 for (const q of raw) {
183 if (q === null || typeof q !== 'object' || Array.isArray(q)) continue
184 const x = q as Fields
185 if (typeof x.question !== 'string') continue
186 const item: Fields = { question: x.question }
187 if (typeof x.header === 'string') item.header = x.header
188 if (typeof x.multiSelect === 'boolean') item.multiSelect = x.multiSelect
189 if (Array.isArray(x.options)) {
190 item.options = x.options.flatMap((o: unknown) => {
191 if (o === null || typeof o !== 'object' || Array.isArray(o) || typeof (o as Fields).label !== 'string') return []
192 const y = o as Fields
193 const opt: Fields = { label: y.label }
194 if (typeof y.description === 'string') opt.description = y.description
195 if (typeof y.preview === 'string') opt.preview = y.preview
196 return [opt]
197 })
198 }
199 out.push(item)
200 }
201 if (!out.length || byteSize(JSON.stringify(out)) > CONTENT_LIMIT) return undefined
202 return out
203}
204
205/** The text of a response row: its text blocks joined by newlines; none for a row without text (§9, P7). */
206export function replyText(content: unknown): string {
207 if (typeof content === 'string') return content
208 if (!Array.isArray(content)) return ''
209 return content
210 .filter(b => !!b && typeof b === 'object' && (b as Fields).type === 'text' && typeof (b as Fields).text === 'string')
211 .map(b => (b as Fields).text as string)
212 .join('\n')
213}
214
215/** A prompt the person sent: typed at the prompt or through Remote Control (BEH-04). */
216export function isPersonPrompt(origin: unknown): boolean {
217 const kind = origin && typeof origin === 'object' ? (origin as Fields).kind : undefined
218 return kind === 'composer' || kind === 'bridge'
219}
220
221/** An event Claude Code itself fired, not a mod (next.origin; BEH-04). */
222export function isEngine(origin: unknown): boolean {
223 return !!origin && typeof origin === 'object' && (origin as Fields).plugin === 'engine'
224}
225
226/** A command.run this plugin made itself with $.command.run: origin { kind: 'plugin', name } with this plugin's name (BEH-41,
227 * version 25; the shape claude-code.d.ts documents, which the probe of 2026-10-08 confirmed for a run started in session.measure). */
228export function isOwnRun(origin: unknown, pluginName: string): boolean {
229 if (!origin || typeof origin !== 'object') return false
230 const o = origin as Fields
231 return o.kind === 'plugin' && typeof o.name === 'string' && o.name === pluginName
232}
233
234/** A tool call another plugin's mod made through $.tool.call (BEH-33, version 21): next.origin.plugin is a string other than
235 * 'engine'. A missing origin, a plugin that isn't a string and 'engine' itself count as Claude Code's (only a test double
236 * builds the first two), so nothing that was recorded stops being recorded. BEH-38 and BEH-04 still ask for 'engine'. */
237export function isFromMod(origin: unknown): boolean {
238 if (!origin || typeof origin !== 'object') return false
239 const plugin = (origin as Fields).plugin
240 return typeof plugin === 'string' && plugin !== 'engine'
241}
242
243// The parts of SPEC-012's progress state (DM-03) the adapter reads.
244export type StateStep = { n: number; title: string; state: string; userOwned?: boolean; kind?: string | null; stoppable?: boolean
245 evidence?: { type: string }[]; claim?: { state: string } | null }
246export type StateFlag = { gate: string; seq: number; step: number; type: string; message: string }
247export type ProgressState = {
248 run?: string
249 skill: string
250 current: number | null
251 ended: string | null
252 steps: StateStep[]
253 flags: StateFlag[]
254 gate: { kind: string | null; seq: number | null; refuse: boolean; reason: string | null }
255 manifest: { state: ProgressSummary['manifest'] }
256 counts?: { stepEvents?: number; unmarkedQuestions?: number }
257 /** Version 15: present when the run's manifest names workFiles (SPEC-012 BEH-21, DM-03). */
258 workFiles?: { files?: unknown; due?: unknown }
259 /** The user's answer to the waiver question, when one was answered (SPEC-012 DM-03, version 11). */
260 waiver?: string | null
261}
262
263/** The summary the status line and the band draw (DM-03). */
264export function summaryOf(state: ProgressState): ProgressSummary {
265 const cur = state.current === null ? undefined : state.steps.find(s => s.n === state.current)
266 return {
267 skill: state.skill,
268 current: state.current,
269 steps: state.steps.length,
270 flags: state.flags.length,
271 yourTurn: cur?.state === 'your-turn',
272 ended: state.ended,
273 manifest: state.manifest.state,
274 states: state.steps.map(s => s.state),
275 currentTitle: cur?.title ?? null,
276 lastFlag: state.flags.length ? state.flags[state.flags.length - 1].message : null,
277 }
278}
279
280/** The step a stopped run stopped at (BEH-10, version 12): the stopping answer's step, kept in the summary; without it,
281 * the last step past pending (the evaluator ignores everything after the run-end); null for any other run. */
282export function stopStep(summary: ProgressSummary): number | null {
283 if (summary.ended !== 'stopped') return null
284 if (typeof summary.stoppedAt === 'number') return summary.stoppedAt
285 let n = 0
286 summary.states.forEach((s, i) => { if (s !== 'pending') n = i + 1 })
287 return n > 0 ? n : null
288}
289
290/** The label that stops a run (SPEC-013 BEH-27, version 12), as SPEC-003 BEH-08 offers it at architecture's step 8. */
291export const STOP_LABEL = 'Write nothing'
292
293/** Whether an engine-fired question's answer stops the run (BEH-27): one question, tagged with a step that the latest
294 * state marks stoppable, answered with the stop label (a typed answer equal to it can't be told apart). */
295export function stopsRun(input: Fields, outcome: ToolOutcome, steps: readonly StateStep[]): boolean {
296 if (!isAnswered(outcome) || !Array.isArray(input.questions) || input.questions.length !== 1) return false
297 const step = questionTag(input).step
298 if (step === undefined || steps.find(s => s.n === step)?.stoppable !== true) return false
299 const question = (input.questions as Fields[])[0]?.question
300 const answers = (outcome.result as { answers?: Record<string, unknown> }).answers ?? {}
301 return typeof question === 'string' && Object.hasOwn(answers, question) && answers[question] === STOP_LABEL
302}
303
304/** The commands BEH-28 confirms while a run is unfinished, with the verb its dialog uses (version 12). */
305export const CONFIRMED: Record<string, string> = { clear: 'Clear', exit: 'Exit', resume: 'Resume' }
306
307/** BEH-28's question for an unfinished run. */
308export function exitQuestion(skill: string, current: number, steps: number, verb: string): string {
309 return `${skill} run is at step ${current} of ${steps} and unfinished. ${verb} anyway?`
310}
311
312/** BEH-28's text when the command is kept. */
313export function keptText(skill: string, current: number): string {
314 return `Kept working: the ${skill} run is still at step ${current}.`
315}
316
317/** BEH-28's question while runs are paused beneath the open one (version 14): a paused run is unfinished, so it asks
318 * whatever the open run's state, naming the run just beneath. `s` is the open run's summary, null with no state yet. */
319export function nestedExitQuestion(open: string, s: { current: number | null; steps: number; ended: string | null } | null,
320 b: { skill: string; step: number }, more: number, verb: string): string {
321 const where = s === null ? 'has just started'
322 : s.current === null || s.ended !== null ? 'is done' : `is at step ${s.current} of ${s.steps} and unfinished`
323 return `${open} run ${where} (${b.skill} paused at step ${b.step}${more > 0 ? `, ${more} more paused` : ''}). ${verb} anyway?`
324}
325
326/** BEH-28's text when the command is kept while runs are paused (version 14). */
327export function nestedKeptText(open: string, b: { skill: string; step: number }): string {
328 return `Kept working: the ${open} run goes on, and ${b.skill} is still paused at step ${b.step}.`
329}
330
331/** Whether a rejected $.ui.ask was the user's dismissal (Esc), as opposed to a dialog that couldn't be shown (ERR-16). */
332export function isDismissal(err: unknown): boolean {
333 const text = err instanceof Error ? err.message : String(err)
334 return /doesn't want to proceed/.test(text)
335}
336
337/** A paused run as the status line and the band name it (BEH-10, BEH-11; version 14): its skill, its return step, and
338 * the number of steps of its saved summary, if it had one. */
339export type Beneath = { skill: string; step: number; summary: { steps: number } | null }
340
341/** The status line's text (BEH-10), or undefined when there is nothing to show; `trail`, the paused runs, bottom first. From version 27
342 * `hold` says which write of the tracker's is held (BEH-42): the run's log replaces the summary and any evaluator reason; the
343 * odometer's file alone ends the summary. */
344export function statusText(summary: ProgressSummary | null, mode: ProgressMode, idle: boolean, off: string | null,
345 trail: readonly Beneath[] = [], hold: HoldKind | null = null): string | undefined {
346 if (hold === 'run') return RETRYING
347 if (off !== null) return `progress: off (${off})`
348 if (summary === null) return undefined
349 const stoppedAt = stopStep(summary)
350 let text = stoppedAt !== null
351 ? `${summary.skill} stopped at step ${stoppedAt}`
352 : summary.ended !== null
353 ? `${summary.skill} ended`
354 : summary.current === null ? `${summary.skill} done` : `${summary.skill} ${summary.current}/${summary.steps}`
355 const b = trail[trail.length - 1]
356 if (b !== undefined) {
357 text += ` · in ${b.skill} ${b.step}${b.summary !== null ? `/${b.summary.steps}` : ''}`
358 if (trail.length > 1) text += ` · ${trail.length - 1} more`
359 }
360 if (summary.yourTurn && summary.ended === null) text += ' · your turn'
361 if (summary.flags > 0) text += ` · ${summary.flags} ${summary.flags === 1 ? 'flag' : 'flags'}`
362 if (summary.manifest === 'stale' || summary.manifest === 'none') text += ' · ticks only'
363 if (idle && summary.ended === null) text += ' · idle'
364 if (mode === 'enforce') text += ' · enforce'
365 if (hold === 'odometer') text += ' · odometer retrying'
366 return text
367}
368
369const GLYPH: Record<string, string> = {
370 done: '●', current: '◆', 'your-turn': '?', pending: '○', claimed: '◐', unconfirmed: '·',
371 'skipped-with-reason': '⊘', 'not-applicable': '–', skipped: '✗', 'rule-broken': '✗', carried: '◉',
372}
373
374/** Cut a row to width characters. */
375export function fit(text: string, width: number): string {
376 const chars = Array.from(text)
377 if (chars.length <= width) return text
378 return width <= 1 ? chars.slice(0, Math.max(0, width)).join('') : chars.slice(0, width - 1).join('') + '…'
379}
380
381/** The band's two rows of text (BEH-11); the button sits on row 2 between the mode and the flag. Row 1 names the run
382 * paused just beneath, if any (version 14). */
383export function bandRows(summary: ProgressSummary, mode: ProgressMode, trail: readonly Beneath[] = []) {
384 const glyphs = summary.states.map(s => GLYPH[s] ?? '○').join('')
385 const stoppedAt = stopStep(summary)
386 const where = stoppedAt !== null
387 ? `stopped at step ${stoppedAt}`
388 : summary.ended !== null
389 ? `ended (${summary.ended})`
390 : summary.current === null
391 ? `all ${summary.steps} steps reached`
392 : `step ${summary.current} of ${summary.steps}: ${summary.currentTitle ?? ''}`
393 return {
394 row1: `${summary.skill} ${glyphs} ${where}${pausedPart(trail)}`,
395 mode: mode === 'enforce' ? 'enforce mode' : 'observe mode',
396 button: mode === 'enforce' ? 'Switch to observe' : 'Switch to enforce',
397 flag: summary.lastFlag ?? 'no flags',
398 }
399}
400
401function pausedPart(trail: readonly Beneath[]): string {
402 const b = trail[trail.length - 1]
403 return b === undefined ? '' : ` (paused: ${b.skill} at step ${b.step}${trail.length > 1 ? `, ${trail.length - 1} more` : ''})`
404}
405
406/** A flag's identity, so each is shown once (BEH-12). */
407export function flagKey(f: StateFlag): string {
408 return `${f.step}:${f.type}:${f.seq}`
409}
410
411/** Flags not shown yet, and the toast text of each (BEH-12). */
412export function newFlagToasts(state: ProgressState, shown: readonly string[], lines: readonly string[] | null = null): { keys: string[]; toasts: string[] } {
413 const fresh = state.flags.filter(f => !shown.includes(flagKey(f)))
414 return {
415 keys: fresh.map(flagKey),
416 // Written for the user (BEH-12, version 8): a decision whose typed answers went to another marked step says so,
417 // and a question gate's answer counted for nothing. `lines` is the run's, when it follows the task list.
418 toasts: fresh.map(f => {
419 const typed = lines === null ? null : typedTo(state, f, lines)
420 return `✗ Step ${f.step} ${f.type}: ${f.message}` + (typed ? ' ' + typed : '')
421 + (QUESTION_GATES.includes(f.type) ? ' Its answer, if any, counts for no step.' : '')
422 }),
423 }
424}
425
426/** The refusal text at the write gate (BEH-08), or null when the provisional state doesn't refuse at seq. It names
427 * what clears each flag: a step's own evidence or a tick in reply text, or for a decision the user's answer (the
428 * VER-15 dogfood run showed a refusal that only said "the user decides" sent Claude into the tracker's code). */
429export function refusalText(state: ProgressState, seq: number, runLines: readonly string[] | null = null): string | null {
430 const g = state.gate
431 if (g.kind !== 'write' || g.seq !== seq || !g.refuse) return null
432 const flags = state.flags.filter(f => f.seq === seq)
433 const owned = new Set(state.steps.filter(s => s.userOwned).map(s => s.n))
434 const decisions = flags.filter(f => f.type === 'rule-broken' || owned.has(f.step))
435 const steps = flags.filter(f => f.type !== 'rule-broken' && !owned.has(f.step)).map(f => f.step)
436 const said = ["DevForgeAI's progress tracker refused this write at the write gate (enforce mode):", ...flags.map(f => `- ${f.message}`)]
437 if (steps.length) {
438 const list = [...new Set(steps)].join(', ')
439 said.push(`To clear step ${list}: do the step with a tool call the run log can see (Read the files it names), `
440 + `or, if you did it, tick it as \`- [x] N.\` in your reply text. A tick only in your thinking doesn't count.`)
441 }
442 if (decisions.length) {
443 said.push('The decisions at step ' + [...new Set(decisions.map(f => f.step))].join(', ')
444 + " are the user's: ask the user, or leave those fields open.")
445 }
446 // How a decision's answer counts, whichever step the task list marks (BEH-08, version 8): `runLines` is the run's
447 // lines, when it follows the task list.
448 if (runLines !== null) said.push(...decisionLines(state, flags, runLines))
449 said.push('Then write again.' + (state.run ? ` The run's log and state are in devforgeai/progress/runs/${state.run}/.` : ''))
450 return said.join('\n')
451}
452
453/** The report gate's flags for the model (BEH-09), or null when there are none to give. */
454export function reportContext(state: ProgressState): { seq: number; text: string } | null {
455 const g = state.gate
456 if (g.kind !== 'report' || !g.refuse || g.seq === null) return null
457 const flags = state.flags.filter(f => f.seq === g.seq).map(f => `- ${f.message}`)
458 return {
459 seq: g.seq,
460 text: [`DevForgeAI's progress tracker (enforce mode) found these at the ${state.skill} run's report:`, ...flags].join('\n'),
461 }
462}
463
464/** The final evaluation's timeout at session end, from what the shared 1.5-second budget leaves, or null when
465 * there is no room for it (BEH-05): the run-end line is written first, and the log alone reproduces the state. */
466export function finalTimeout(remainingMs: number): number | null {
467 const timeout = Math.floor(remainingMs) - END_RESERVE_MS
468 return timeout >= 200 ? timeout : null
469}
470
471/** What session.end keeps for Claude Code's own work after the adapter's (BEH-05). */
472export const END_RESERVE_MS = 300
473
474/** Whether the session.end budget left has room for one more run-end (BEH-05, version 14). */
475export function hasRoom(remainingMs: number): boolean {
476 return remainingMs >= END_RESERVE_MS
477}
478
479/** The retention period in days from the retentionDays setting (DM-06): 7 to 3650, else 30. Claude Code already
480 * refuses to load the module when a stored value is out of range; this is the second guard, since the floor is
481 * what protects another, idle session's open run (BEH-19). */
482export function retentionOf(value: unknown): number {
483 const days = Number(value)
484 return Number.isInteger(days) && days >= 7 && days <= 3650 ? days : 30
485}
486
487// ---- the task list (SPEC-013 v4 and v5; SPEC-012 §4) ----
488
489/** The convention's tag: a skill whose text names it keeps its checklist in the task list (SPEC-012 §4). */
490export const TASK_TAG = 'devforgeai_step'
491
492/** Whether the session has a task list: TaskCreate and TaskUpdate, or TodoWrite (DM-01). TaskStop stops background
493 * tasks and isn't one; a model that doesn't get the task tools by default has none without the user's opt-in. */
494export function hasTaskList(names: readonly string[]): boolean {
495 return (names.includes('TaskCreate') && names.includes('TaskUpdate')) || names.includes('TodoWrite')
496}
497
498/** Whether a run follows the task list, from its skill-loaded line: taskList true and the tag in the skill's text
499 * (SPEC-012 BEH-18). */
500export function followsTaskList(line: string | undefined): boolean {
501 if (!line) return false
502 try {
503 const e = JSON.parse(line) as Fields
504 return e.taskList === true && typeof e.checklist === 'string' && e.checklist.includes(TASK_TAG)
505 } catch {
506 return false
507 }
508}
509
510/** A TaskCreate's task ID, from its result text 'Task #<id> created successfully' (BEH-20, ERR-13). */
511export function taskIdOf(text: unknown): string | null {
512 if (typeof text !== 'string') return null
513 const m = text.match(/Task #([^\s:]+) created successfully/)
514 return m ? m[1] : null
515}
516
517/** A task's step: its metadata devforgeai_step when that is a whole number, else the number its subject starts
518 * with ('<N>. <title>'), else null (BEH-20, ERR-13). */
519export function stepOfTask(input: Fields): number | null {
520 const meta = input.metadata && typeof input.metadata === 'object' ? (input.metadata as Fields)[TASK_TAG] : undefined
521 if (typeof meta === 'number' && Number.isInteger(meta) && meta >= 1) return meta
522 const m = typeof input.subject === 'string' ? input.subject.match(/^\s*(\d+)\./) : null
523 return m && Number(m[1]) >= 1 ? Number(m[1]) : null
524}
525
526/** A TaskUpdate's status as a step event's state: in_progress starts the step, completed ends it (BEH-20). */
527export function stepStateOf(status: unknown): 'started' | 'done' | null {
528 return status === 'in_progress' ? 'started' : status === 'completed' ? 'done' : null
529}
530
531/** A TodoWrite's step events and the statuses to keep (BEH-20): each '<N>.' todo is compared with its own entry in
532 * the list the call replaced (the result's oldTodos: the same content at the same place, else the first unused entry
533 * with that content), or with the kept statuses when the result has none, so an earlier run's completed todos left in
534 * the list, even beside the new run's of the same numbers, claim nothing. Becoming in_progress starts a step,
535 * becoming completed ends it; straight from pending to completed gives done only. */
536export function todoSteps(todos: unknown, last: Readonly<Record<string, string>>, oldTodos?: unknown): {
537 events: Array<{ step: number; state: 'started' | 'done' }>
538 statuses: Record<string, string>
539} {
540 const old = Array.isArray(oldTodos) ? (oldTodos as unknown[]) : null
541 const used = new Set<number>()
542 const contentOf = (x: unknown) => (x && typeof x === 'object' ? (x as Fields).content : undefined)
543 const statuses: Record<string, string> = { ...last }
544 const events: Array<{ step: number; state: 'started' | 'done' }> = []
545 if (!Array.isArray(todos)) return { events, statuses }
546 todos.forEach((todo, i) => {
547 if (!todo || typeof todo !== 'object') return
548 const { content, status } = todo as Fields
549 const m = typeof content === 'string' ? content.match(/^\s*(\d+)\./) : null
550 if (!m || typeof status !== 'string' || Number(m[1]) < 1) return
551 const step = Number(m[1])
552 let before: unknown
553 if (old !== null) {
554 const j = !used.has(i) && contentOf(old[i]) === content ? i : old.findIndex((o, k) => !used.has(k) && contentOf(o) === content)
555 if (j >= 0) used.add(j)
556 before = j >= 0 ? (old[j] as Fields).status : undefined
557 } else {
558 before = statuses[String(step)]
559 }
560 if (status === 'in_progress' && before !== 'in_progress') events.push({ step, state: 'started' })
561 if (status === 'completed' && before !== 'completed') events.push({ step, state: 'done' })
562 statuses[String(step)] = status
563 })
564 return { events, statuses }
565}
566
567/** The refusal of a question asked with no step in progress (BEH-21), with how to recover. */
568export const QUESTION_REFUSAL = "DevForgeAI's progress tracker refused this question (enforce mode): no step of this run "
569 + "is marked in progress in your task list. Tasks from an earlier run don't count: if this run's checklist isn't in "
570 + 'your task list yet, turn it into tasks first as the skill says (one task per step, subject <N>. <title>, metadata '
571 + 'devforgeai_step: N). Then mark the step this question belongs to in_progress (TaskUpdate, or TodoWrite), and ask '
572 + 'again.'
573
574/** What an unmarked question with no step tag is also told (BEH-21, version 8). */
575export const QUESTION_TAG = ' Tag the question too: add metadata: {"source": "devforgeai_step:N"} to the AskUserQuestion '
576 + "call, N being its step. If the question isn't part of this skill's checklist, give it a source of its own instead; "
577 + 'it then needs no step and counts for none.'
578
579/** SPEC-012's question-gate flag types (version 9). */
580export const QUESTION_GATES = ['unmarked-question', 'untagged-question', 'mismatched-question']
581
582/** The question refusal when the provisional state's question gate refuses at seq, else null (BEH-21). The text follows
583 * the type of the gate's flag, never its message: `marked` is the step the task list marks, `tagged` whether the
584 * question names a step (version 8). */
585export function questionRefusal(state: ProgressState, seq: number, marked: number | null = null, tagged = false): string | null {
586 const g = state.gate
587 if (g.kind !== 'question' || g.seq !== seq || !g.refuse) return null
588 const flag = state.flags.find(f => f.gate === 'question' && f.seq === seq)
589 const head = "DevForgeAI's progress tracker refused this question (enforce mode): "
590 if (flag?.type === 'mismatched-question') {
591 const n = flag.step
592 const k = marked ?? n
593 return head + `it is tagged for step ${n}, but your task list marks step ${k} in progress. If the question belongs `
594 + `to step ${n}, mark step ${n} in_progress (TaskUpdate, or TodoWrite) and ask again; if it belongs to step ${k}, `
595 + `tag it devforgeai_step:${k} and ask again.`
596 }
597 if (flag?.type === 'untagged-question') {
598 return head + "it doesn't name a step of this skill's checklist. Add metadata: {\"source\": \"devforgeai_step:N\"} to "
599 + `the AskUserQuestion call, N being the step it belongs to; your task list marks step ${marked ?? flag.step} in `
600 + 'progress, so if the question belongs to another step, mark that step in_progress first. If the question '
601 + "isn't part of this skill's checklist, give it a source of its own instead; it then counts for no step. Then ask "
602 + 'again.'
603 }
604 return QUESTION_REFUSAL + (tagged ? '' : QUESTION_TAG)
605}
606
607/** The waiver question's own tag (DM-01, version 10). */
608export const WAIVER_TAG = 'devforgeai_waiver'
609const WAIVER_LABELS: Record<string, 'proceed' | 'ask'> = { 'Proceed without questions': 'proceed', 'Ask me as usual': 'ask' }
610
611function sourceOf(input: Fields): unknown {
612 const meta = input.metadata
613 return meta !== null && typeof meta === 'object' ? (meta as Fields).source : undefined
614}
615
616/** The waiver question (DM-01, version 10): source exactly devforgeai_waiver, in a call that asks exactly one
617 * question. Never checked at the question gate (BEH-21). */
618export function isWaiverQuestion(input: Fields): boolean {
619 return sourceOf(input) === WAIVER_TAG && Array.isArray(input.questions) && input.questions.length === 1
620}
621
622/** Which fixed label a waiver question's answer picked: proceed, ask, or other for anything typed, any other text and
623 * a dismissal; a typed answer equal to a label is that label, since the result can't tell them apart (DM-01). */
624export function waiverAnswer(input: Fields, outcome: ToolOutcome): 'proceed' | 'ask' | 'other' {
625 if (!isAnswered(outcome)) return 'other'
626 const question = (input.questions as Fields[])[0]?.question
627 const answers = (outcome.result as { answers?: Record<string, unknown> }).answers ?? {}
628 const answer = typeof question === 'string' ? answers[question] : undefined
629 return typeof answer === 'string' && Object.hasOwn(WAIVER_LABELS, answer) ? WAIVER_LABELS[answer] : 'other'
630}
631
632/** A question's step tag from its input's metadata.source (DM-01, version 8): `step` for devforgeai_step:N, `outside`
633 * for a source naming something else, nothing for no source or a malformed tag. The waiver source with more than one
634 * question is `outside` too, so its questions can't take the waiver's exemption (version 10). */
635export function questionTag(input: Fields): { step?: number; outside?: true } {
636 const source = sourceOf(input)
637 if (typeof source !== 'string') return {}
638 const m = /^devforgeai_step:([1-9][0-9]*)$/.exec(source)
639 if (m) return { step: Number(m[1]) }
640 return source.startsWith(TASK_TAG) ? {} : { outside: true }
641}
642
643/** The adherence notice for a run that follows the task list, once it ends or reaches its report gate with no step
644 * event or an unmarked question (BEH-22), else null; a state without the counts says nothing. The report step's own
645 * state shows the report was reached even after a later event moved the gate on. */
646export function adherenceText(state: ProgressState): string | null {
647 const reported = state.gate.kind === 'report'
648 || state.steps.some(s => s.kind === 'report' && (s.state === 'done' || s.state === 'claimed'))
649 if (state.ended === null && !reported) return null
650 const n = state.counts?.stepEvents
651 const m = state.counts?.unmarkedQuestions
652 if (typeof n !== 'number' || typeof m !== 'number' || (n > 0 && m === 0)) return null
653 return `${state.skill} didn't keep its task list: ${n} step events, ${m} questions asked without their step marked and tagged. `
654 + 'Recommended: fix the skill so it keeps its checklist in the task list (DevForgeAI SPEC-012 §4)'
655}
656
657/** The once-per-session hint for a session with no task tools (BEH-23). */
658export function hintText(skill: string): string {
659 return `${skill}: this session has no task list, so DevForgeAI places your answers by guessing. For exact step `
660 + 'tracking, start Claude Code with CLAUDE_CODE_ENABLE_TODO_TOOLS=1 (DevForgeAI SPEC-012 §4)'
661}
662
663// ---- the task list's mark (SPEC-013 v7 and v8) ----
664
665function parsed(lines: readonly string[]): Fields[] {
666 const out: Fields[] = []
667 for (const line of lines) {
668 try {
669 out.push(JSON.parse(line) as Fields)
670 } catch {
671 // a line that isn't JSON says nothing about the mark
672 }
673 }
674 return out
675}
676
677/** The step the task list marks in progress, from a run's event lines before seq `upto`: the step whose latest step
678 * event is started, the latest started when several are (SPEC-012 BEH-18); null when none is. With `steps`, the step
679 * events naming a step the state doesn't have are left out first, as the evaluator leaves them out (SPEC-012 ERR-06),
680 * so the two agree on the mark (BEH-24, version 8). */
681export function markedStep(lines: readonly string[], upto = Infinity, steps: readonly StateStep[] | null = null): number | null {
682 const known = steps === null ? null : new Set(steps.map(s => s.n))
683 const latest = new Map<number, { state: unknown; seq: number }>()
684 for (const e of parsed(lines)) {
685 if (typeof e.seq !== 'number' || e.seq >= upto) continue
686 if (e.kind === 'step' && typeof e.step === 'number' && (known === null || known.has(e.step))) {
687 latest.set(e.step, { state: e.state, seq: e.seq })
688 }
689 }
690 let best: { n: number; seq: number } | null = null
691 for (const [n, v] of latest) if (v.state === 'started' && (best === null || v.seq > best.seq)) best = { n, seq: v.seq }
692 return best === null ? null : best.n
693}
694
695/** Whether the user typed a prompt after step n's latest started event and before seq `upto` (BEH-08, BEH-12). */
696export function typedSince(lines: readonly string[], n: number, upto = Infinity): boolean {
697 const events = parsed(lines).filter(e => typeof e.seq === 'number' && e.seq < upto)
698 const starts = events.filter(e => e.kind === 'step' && e.step === n && e.state === 'started').map(e => e.seq as number)
699 if (!starts.length) return false
700 const from = Math.max(...starts)
701 return events.some(e => e.kind === 'prompt' && (e.seq as number) > from)
702}
703
704/** 'step N (<title>)', or 'step N' when the state has no title for it. */
705export function stepLabel(steps: readonly StateStep[], n: number): string {
706 const s = steps.find(x => x.n === n)
707 return s ? `step ${n} (${s.title})` : `step ${n}`
708}
709
710/** The first skipped flag for a user-owned step among flags: the decision's step, from the flag's step and the state's
711 * userOwned, never from message text (BEH-08, BEH-12). */
712function decisionFlag(state: ProgressState, flags: readonly StateFlag[]): StateFlag | null {
713 const owned = new Set(state.steps.filter(s => s.userOwned).map(s => s.n))
714 return flags.find(f => f.type === 'skipped' && owned.has(f.step)) ?? null
715}
716
717/** The write refusal's lines for a decision (BEH-08, version 8): how step M's answer counts, whichever step is marked,
718 * and before it, when another marked step N took what the user typed, that step; none when step M itself is marked. */
719export function decisionLines(state: ProgressState, flags: readonly StateFlag[], lines: readonly string[]): string[] {
720 const flag = decisionFlag(state, flags)
721 if (flag === null) return []
722 const m = flag.step
723 const n = markedStep(lines, Infinity, state.steps)
724 if (n === m) return []
725 const label = stepLabel(state.steps, m)
726 const out: string[] = []
727 if (n !== null && typedSince(lines, n)) {
728 out.push(`Your task list marks ${stepLabel(state.steps, n)} in progress, so what the user typed since then counted for step ${n}.`)
729 }
730 out.push(`${label[0].toUpperCase()}${label.slice(1)} is the user's decision: an answer counts for it only while step ${m} `
731 + `is marked in progress, and an answer to a question only when the question is also tagged devforgeai_step:${m}. `
732 + `Mark step ${m} in_progress, ask the user with the question tagged devforgeai_step:${m}, and mark step ${m} completed.`)
733 return out
734}
735
736/** The user's sentence on a decision's flag toast (BEH-12, version 8), judged at the flag's seq: another marked step N
737 * took what the user typed; null otherwise. */
738function typedTo(state: ProgressState, flag: StateFlag, lines: readonly string[]): string | null {
739 if (decisionFlag(state, [flag]) === null) return null
740 const n = markedStep(lines, flag.seq, state.steps)
741 if (n === null || n === flag.step || !typedSince(lines, n, flag.seq)) return null
742 return `What you typed since ${stepLabel(state.steps, n)} was marked in progress counted for step ${n}: the `
743 + `${state.skill} skill didn't keep its task list in step with its work (DevForgeAI SPEC-012 §4).`
744}
745
746/** The cause of a refusal, for the stuck notice (BEH-25, versions 8 and 9): the gate's kind and the first flag raised
747 * at seq, with that flag's type and whether its step is user-owned in the state's steps (a step it doesn't list isn't). */
748export function refusalCause(state: ProgressState, seq: number):
749 { key: string; step: number; message: string; type: string; userOwned: boolean } | null {
750 const flag = state.flags.find(f => f.seq === seq)
751 if (flag === undefined || state.gate.kind === null) return null
752 const userOwned = state.steps.some(s => s.n === flag.step && s.userOwned === true)
753 return { key: `${state.gate.kind}:${flag.type}:${flag.step}`, step: flag.step, message: flag.message, type: flag.type,
754 userOwned }
755}
756
757/** One refusal kept for the run's review (BEH-25, BEH-26; version 10). */
758export type Refused = ProgressRefused
759
760/** One review item: a cause (the gate's kind, the flag's type and step) with its first message and its number of
761 * refusals, 0 for a flag no refusal has (BEH-26). */
762export type ReviewItem = { gate: string; seq: number; step: number; type: string; message: string; count: number }
763
764/** The run's review items, one per cause: the refusals' causes in the order of their first refusal, then the flags'
765 * causes no refusal has, in the order of their first flag (BEH-26). */
766export function reviewItems(refused: readonly Refused[], flags: readonly StateFlag[]): ReviewItem[] {
767 const items: ReviewItem[] = []
768 const byKey = new Map<string, ReviewItem>()
769 const key = (x: { gate: string; type: string; step: number }) => `${x.gate}:${x.type}:${x.step}`
770 for (const r of refused) {
771 const k = key(r)
772 const item = byKey.get(k)
773 if (item) item.count += 1
774 else {
775 const fresh = { gate: r.gate, seq: r.seq, step: r.step, type: r.type, message: r.message, count: 1 }
776 byKey.set(k, fresh)
777 items.push(fresh)
778 }
779 }
780 for (const f of [...flags].sort((a, b) => a.seq - b.seq)) {
781 const k = key(f)
782 if (byKey.has(k)) continue
783 const fresh = { gate: f.gate, seq: f.seq, step: f.step, type: f.type, message: f.message, count: 0 }
784 byKey.set(k, fresh)
785 items.push(fresh)
786 }
787 return items
788}
789
790/** A review item's question (BEH-26). */
791export function reviewQuestion(skill: string, i: number, n: number, item: ReviewItem): string {
792 const what = item.count > 0 ? `refused ${item.count} time(s)` : 'flagged'
793 return `${skill} run, item ${i} of ${n}: ${what} at step ${item.step} (${item.gate} gate): ${item.message}. `
794 + 'Accept it, or challenge it?'
795}
796
797/** The question gate's flag types (SPEC-012 BEH-08), whose stuck notice keeps the task-list advice. */
798const QUESTION_GATE_TYPES = new Set(['unmarked-question', 'untagged-question', 'mismatched-question'])
799
800/** The stuck notice's last sentence (BEH-25, version 9), chosen by the refused flag's type and whether its step is
801 * user-owned, never by message text: the task list for a question gate, the decision for a user-owned step's skipped
802 * flag or a rule-broken one, and otherwise the step's evidence. */
803export function stuckAdvice(type: string, userOwned: boolean): string {
804 if (type === OUTSIDE_WRITE_TYPE) return OUTSIDE_ADVICE
805 if (QUESTION_GATE_TYPES.has(type)) {
806 return "Help Claude bring its task list in step, or switch to observe mode with the band's button."
807 }
808 if (type === 'rule-broken' || (type === 'skipped' && userOwned)) {
809 return "The refused write records a decision that needs your answer: answer Claude's question about it, or ask "
810 + "Claude to leave it open, or switch to observe mode with the band's button."
811 }
812 return "Claude hasn't done that step in a way the tracker can see: ask Claude to do it as the message says, or "
813 + "switch to observe mode with the band's button."
814}
815
816/** The notice for the user when the same refusal comes twice in a run (BEH-25, versions 8 and 9). */
817export function stuckText(skill: string, step: number, message: string, advice: string): string {
818 return `${skill}: the progress tracker refused Claude twice at step ${step} for the same reason: ${message}. ${advice}`
819}
820
821/** The start of the note a compaction ends with (BEH-24), by which an earlier one is found and removed (version 8). */
822export const NOTE_START = "DevForgeAI's progress tracker: when this conversation was compacted"
823
824/** What a compaction keeps and the note it ends with (BEH-24); `label` is 'step N (<title>)', or null for none. */
825export function compactTexts(skill: string, label: string | null): { instruction: string; note: string } {
826 const marked = label ?? 'no step'
827 return {
828 instruction: `Keep, for DevForgeAI's progress tracker: in the ${skill} run, the task list marks ${marked} in progress.`,
829 note: `DevForgeAI's progress tracker: when this conversation was compacted, your task list marked ${marked} in `
830 + "progress. Before you ask anything or go on, check your task list and bring it in step with the work: mark each "
831 + "finished step done and the step you're on in_progress.",
832 }
833}
834
835// ---- the trail of paused runs (SPEC-013 BEH-29, BEH-30; version 13, paused and resumed from version 14) ----
836
837/** A trail entry as the compaction note names it. */
838export type TrailEntry = { skill: string; step: number }
839
840/** The line added to a skill Claude loads mid-run (BEH-29, version 14). */
841export function returnLine(skill: string, step: number): string {
842 return `This skill was loaded by ${skill} at step ${step}. When this skill's work is done, mark ${skill}'s step ${step} task in progress again and continue ${skill} at step ${step}.`
843}
844
845export const TRAIL_NOTE_START = 'Return points (from the progress tracker):'
846
847/** The compaction note naming the whole trail, top first (BEH-29). */
848export function trailNote(open: string, trail: readonly TrailEntry[]): string {
849 const parts = [...trail].reverse().map((t, i) => `${i === 0 ? 'continue' : 'then'} ${t.skill} at step ${t.step}`)
850 return `${TRAIL_NOTE_START} when ${open} is done, ${parts.join('; ')}.`
851}
852
853/** The index of the topmost paused run whose task IDs hold a TaskUpdate's task ID, or -1 (BEH-30 (a), version 14):
854 * whatever the update's status, Claude has gone back to that skill. */
855export function pausedWith(trail: readonly { tasks: Record<string, number> }[], taskId: unknown): number {
856 if (typeof taskId !== 'string') return -1
857 for (let i = trail.length - 1; i >= 0; i--) if (Object.hasOwn(trail[i].tasks, taskId)) return i
858 return -1
859}
860
861/** The trail as $.state gives it back after a load (BEH-30): an entry without its run (version 13's shape) is dropped. */
862export function keptTrail(raw: unknown): ProgressPaused[] {
863 if (!Array.isArray(raw)) return []
864 return raw.filter((t): t is ProgressPaused => t !== null && typeof t === 'object' && typeof t.skill === 'string'
865 && typeof t.step === 'number' && t.tasks !== null && typeof t.tasks === 'object'
866 && t.run !== null && typeof t.run === 'object' && typeof t.run.id === 'string' && typeof t.run.dir === 'string')
867}
868
869/** The reason of a run's run-end in its log, or null when it hasn't ended (BEH-05). */
870export function endReason(lines: readonly string[]): string | null {
871 for (const line of lines) {
872 if (!line.includes('"kind":"run-end"')) continue
873 try {
874 const reason = (JSON.parse(line) as Fields).reason
875 return typeof reason === 'string' ? reason : null
876 } catch {
877 return null
878 }
879 }
880 return null
881}
882
883// ---- the offer to continue an earlier run (SPEC-013 BEH-31, version 16) ----
884
885/** What a run-end's reason reads as in the offer (BEH-31). */
886const WHY: Record<string, string> = {
887 'session-end': 'session end', clear: '/clear', stopped: 'stopped', 'another-skill': 'another skill loaded',
888 returned: 'returned to the skill beneath',
889}
890
891/** An earlier run's offer: the step to continue at, the steps carried, the user's decisions that stand (answered) and
892 * those to confirm again (owned), the files it wrote, and how it ended (`when`). */
893export type ResumePlan = {
894 run: string; step: number; steps: number; carried: number[]; answered: number[]; owned: number[]; files: string[]
895 when: string
896 /** BEH-39 (version 22): the work file to continue in when it still exists (the adapter checks that), else null. */
897 draft: string | null
898 /** The draft the state names, before the adapter has checked that it exists. */
899 draftCandidate: string | null
900 /** BEH-39's <ask> and <forms>: ready-made text, '' for nothing. */
901 ask: string
902 forms: string
903}
904
905/** A step is reached when it has evidence other than waiver evidence, or a done claim (BEH-31; SPEC-012 BEH-07). */
906function reachedStep(s: StateStep): boolean {
907 return (s.evidence ?? []).some(x => x.type !== 'waiver') || s.claim?.state === 'done'
908}
909
910/** An age as the offer says it: minutes, hours or days. */
911export function ageText(ms: number): string {
912 if (!Number.isFinite(ms)) return 'an unknown time'
913 const minutes = Math.max(0, Math.floor(ms / 60000))
914 if (minutes < 60) return `${minutes} ${minutes === 1 ? 'minute' : 'minutes'}`
915 const hours = Math.floor(minutes / 60)
916 if (hours < 48) return `${hours} ${hours === 1 ? 'hour' : 'hours'}`
917 const days = Math.floor(hours / 24)
918 return `${days} ${days === 1 ? 'day' : 'days'}`
919}
920
921/** Steps as the offer names them: '1, 2, 3'. */
922export function stepList(ns: readonly number[]): string {
923 return ns.join(', ')
924}
925
926/** The offer for an earlier run, from its state evaluated once more and its log, or null when it isn't offered: its
927 * manifest isn't matched, its last step is reached (unless it ended stopped), or nothing would be carried (BEH-31).
928 * `writeGate` is the number of the step with the write gate, `now` the time in ms. */
929export function resumePlan(run: string, state: ProgressState, lines: readonly string[], writeGate: number | null,
930 now: number): ResumePlan | null {
931 if (state.manifest?.state !== 'matched') return null
932 const steps = state.steps
933 if (!steps.length) return null
934 if (reachedStep(steps[steps.length - 1]) && state.ended !== 'stopped') return null
935 const reached = steps.filter(reachedStep).map(s => s.n)
936 const highest = reached.length ? Math.max(...reached) : null
937 let step = markedStep(lines, Infinity, steps)
938 ?? (highest === null ? steps[0].n : (steps.find(s => s.n > highest)?.n ?? steps[steps.length - 1].n))
939 // Never past the first step with the write gate that has no write evidence: its document was never written. A write
940 // step carried from a run before counts as written there (review S1).
941 const gate = writeGate === null ? undefined : steps.find(s => s.n === writeGate)
942 const wrote = gate !== undefined && (gate.evidence ?? []).some(x => x.type === 'write' || x.type === 'carried')
943 if (gate !== undefined && gate.n < step && !wrote) step = gate.n
944 const carried = steps.filter(s => s.n < step).map(s => s.n)
945 if (!carried.length) return null
946 // The decisions that stand: written under the write gate (not a rule-broken write), each answered and not skipped
947 // there, in this run or, carried, in the run it continued (review S1, S2).
948 const written = gate !== undefined && carried.includes(gate.n) && gate.state !== 'rule-broken'
949 let before: number[] = []
950 try {
951 const first = JSON.parse(lines[0] ?? '{}') as Fields
952 if (Array.isArray(first.answered)) before = first.answered.filter((n): n is number => typeof n === 'number')
953 } catch {
954 // no earlier answers
955 }
956 const owned = steps.filter(s => s.n < step && s.userOwned === true)
957 const answered = written
958 ? owned.filter(s => s.state !== 'skipped' && ((s.evidence ?? []).some(x => x.type === 'answer') || before.includes(s.n))).map(s => s.n)
959 : []
960 const files: string[] = []
961 let ended: string | null = null
962 let lastTime: string | null = null
963 for (const line of lines) {
964 let e: Fields
965 try {
966 e = JSON.parse(line) as Fields
967 } catch {
968 continue
969 }
970 if (typeof e.time === 'string') lastTime = e.time
971 if (e.kind === 'run-end' && typeof e.reason === 'string' && ended === null) ended = e.reason
972 // A path is shown in the dialog and Claude's text: one line of it (review note).
973 const path = typeof e.path === 'string' ? e.path.replace(/[\u0000-\u001f\u007f]+/g, ' ') : null
974 if (e.kind === 'tool' && (e.tool === 'Write' || e.tool === 'Edit' || e.wrote === true) && e.error !== true && path !== null
975 && !files.includes(path)) files.push(path)
976 }
977 const of = `step ${step} of ${steps.length}`
978 const when = ended !== null
979 ? `ended at ${of} on ${(lastTime ?? '').slice(0, 10)} (${WHY[ended] ?? ended})`
980 : `was at ${of} with no end recorded, its last event ${ageText(now - Date.parse(lastTime ?? ''))} ago (it may still be open in another session)`
981 const forms = formsText(run, step, lines)
982 const decision = steps.find(s => s.n === step)
983 const ask = decision?.userOwned === true && state.waiver !== 'proceed'
984 ? ` Step ${step} (${decision.title}) is the user's decision: ask it, marking it in progress and tagging the question with it`
985 + `${forms.inline ? ', with the proposals in the questions below' : ''}.`
986 : ''
987 return { run, step, steps: steps.length, carried, answered, owned: owned.map(s => s.n).filter(n => !answered.includes(n)),
988 files, when, draft: null, draftCandidate: draftCandidate(state, step, writeGate), ask, forms: forms.text }
989}
990
991/** The work file BEH-39's <draft> names, before it is known to exist: the last of the state's workFiles.files while due is
992 * false, when the step to continue at is at or before the step with the write gate; null otherwise. A path that isn't a
993 * project-relative one, as the evaluator lists them, is never one. */
994export function draftCandidate(state: ProgressState, step: number, writeGate: number | null): string | null {
995 const w = state.workFiles
996 if (typeof w !== 'object' || w === null || w.due !== false || !Array.isArray(w.files) || writeGate === null || step > writeGate) return null
997 const last = w.files[w.files.length - 1]
998 if (typeof last !== 'string' || last === '' || last.startsWith('/') || last.split('/').some(p => p === '' || p === '.' || p === '..')) return null
999 return last
1000}
1001
1002const FORMS_LIMIT = 8 * 1024
1003
1004/** BEH-39's <forms>: the questions of the last answer event of the run whose step is `step` and that carries questions,
1005 * rendered inline, or one sentence pointing at them when the rendering passes 8 KiB; nothing when there is no such event.
1006 * `inline` is true when the questions themselves are in the text. */
1007export function formsText(run: string, step: number, lines: readonly string[]): { text: string; inline: boolean } {
1008 let found: unknown[] | null = null
1009 for (const line of lines) {
1010 let e: Fields
1011 try {
1012 e = JSON.parse(line) as Fields
1013 } catch {
1014 continue
1015 }
1016 if (e.kind === 'answer' && e.step === step && Array.isArray(e.questions) && e.questions.length) found = e.questions
1017 }
1018 if (found === null) return { text: '', inline: false }
1019 const out: string[] = []
1020 found.forEach((q, i) => {
1021 if (q === null || typeof q !== 'object') return
1022 const x = q as Fields
1023 out.push(`Q${i + 1}. ${String(x.question)}${typeof x.header === 'string' ? ` [${x.header}]` : ''}`)
1024 for (const o of Array.isArray(x.options) ? x.options : []) {
1025 if (o === null || typeof o !== 'object') continue
1026 const y = o as Fields
1027 out.push(`- ${String(y.label)}${typeof y.description === 'string' && y.description !== '' ? `: ${y.description}` : ''}`)
1028 if (typeof y.preview === 'string' && y.preview !== '') for (const l of y.preview.split('\n')) out.push(` ${l}`)
1029 }
1030 })
1031 const text = ` The questions last shown at step ${step} (Claude's proposals, not the user's answers):\n${out.join('\n')}`
1032 if (byteSize(text) <= FORMS_LIMIT) return { text, inline: true }
1033 return { text: ` The questions last shown at step ${step} are in devforgeai/progress/runs/${run}/events.jsonl, on the events of kind answer with step ${step} (field questions).`, inline: false }
1034}
1035
1036/** BEH-31's question. */
1037export function resumeQuestion(skill: string, plan: ResumePlan): string {
1038 return `${skill}: an earlier run ${plan.when}. It wrote ${plan.files.length ? plan.files.join(', ') : 'nothing'}. Continue it?`
1039}
1040
1041/** BEH-31's line at the end of the text Claude reads. */
1042export function resumeLine(skill: string, plan: ResumePlan): string {
1043 const owned = plan.owned.length
1044 ? ` Steps ${stepList(plan.owned)} were the user's decisions, which the record doesn't keep: before any document records them, confirm each with the user again, in order, marking its step in progress and tagging the question with it.`
1045 : ''
1046 return `This run continues the earlier ${skill} run ${plan.run}, which ${plan.when}. Steps ${stepList(plan.carried)} are `
1047 + `carried over: the tracker counts them reached. Create the task list with those steps' tasks completed, mark step `
1048 + `${plan.step} in progress, and continue at step ${plan.step}.${owned} Files it wrote: `
1049 + `${plan.files.length ? plan.files.join(', ') : 'nothing'}.${plan.draft === null ? '' : draftSentence(plan.draft)} Its replies and questions are in `
1050 + `devforgeai/progress/runs/${plan.run}/events.jsonl, the events of kind reply and answer: use them to show the user what `
1051 + `was proposed, never as a decision.${plan.ask}${plan.forms}`
1052}
1053
1054/** BEH-39's <draft>: the sentence that names the work file to continue in. */
1055export function draftSentence(path: string): string {
1056 const shown = path.replace(/[\u0000-\u001f\u007f]+/g, ' ')
1057 return ` Its draft is ${shown}: load it, keep saving to that path, and keep what it holds (its items, scores and IDs); `
1058 + `what it proposes is not yet the user's decision.`
1059}
1060
1061// ---- The work files' cleanup (BEH-32, IF-05; version 20) ----
1062
1063/** A run ID as run folders are named (SPEC-012's pattern): the only shape the adapter lets into a path. */
1064const RUN_ID = /^[0-9]{8}T[0-9]{6}Z-[a-z][a-z0-9-]*-[0-9a-f]{8}$/
1065
1066/** Whether an evaluation says the run's work files are due (SPEC-012 BEH-21: workFiles.due is true). */
1067export function workFilesDue(state: ProgressState): boolean {
1068 const w = state.workFiles
1069 return typeof w === 'object' && w !== null && !Array.isArray(w) && w.due === true
1070}
1071
1072/** The paths of a state's workFiles.files that the adapter passes on: the non-empty strings and nothing else, in order.
1073 * A state.json is a file the model's tools can write, so whatever isn't a list gives none; prune.py judges the paths
1074 * themselves (ERR-20). */
1075export function workFilePaths(state: unknown): string[] {
1076 if (typeof state !== 'object' || state === null || Array.isArray(state)) return []
1077 const w = (state as { workFiles?: unknown }).workFiles
1078 if (typeof w !== 'object' || w === null || Array.isArray(w)) return []
1079 const files = (w as { files?: unknown }).files
1080 return Array.isArray(files) ? files.filter((f): f is string => typeof f === 'string' && f !== '') : []
1081}
1082
1083/** Why a continued run's state.json gives no work files, or null when it names a workFiles object (ERR-19). */
1084export function workFilesProblem(state: unknown): string | null {
1085 if (typeof state !== 'object' || state === null || Array.isArray(state)) return 'its state is not an object'
1086 const w = (state as { workFiles?: unknown }).workFiles
1087 if (w === undefined) return 'its state has no workFiles'
1088 if (typeof w !== 'object' || w === null || Array.isArray(w)) return 'its workFiles is not an object'
1089 return null
1090}
1091
1092/** IF-05's argv (BEH-32): every path once, in the order given across the lists, each as one --file=<path> token, so a
1093 * path that begins with - is a value and never an option. */
1094export function removeArgv(python: string, plugin: string, root: string, ...lists: readonly (readonly string[])[]): string[] {
1095 const paths = [...new Set(lists.flat().filter(p => typeof p === 'string' && p !== ''))]
1096 return [python, `${plugin}/progress/prune.py`, 'remove', '--root', root, '--manifests', `${plugin}/progress/manifests`,
1097 ...paths.map(p => `--file=${p}`)]
1098}
1099
1100/** The run a run's skill-loaded line says it continues (BEH-31), when it is shaped as a run ID. */
1101export function resumesOf(line: string | undefined): string | null {
1102 if (line === undefined) return null
1103 try {
1104 const v = (JSON.parse(line) as { resumes?: unknown }).resumes
1105 return typeof v === 'string' && RUN_ID.test(v) ? v : null
1106 } catch {
1107 return null
1108 }
1109}
1110
1111// ---- Bash writes (BEH-38; version 22) ----
1112
1113/** The refused entry's type and BEH-25's advice for it (BEH-38 (a)). */
1114export const OUTSIDE_WRITE_TYPE = 'outside-write'
1115export const OUTSIDE_ADVICE = "Ask Claude to write the document with the Write tool, or switch to observe mode with the band's button."
1116
1117/** What a run's manifests make gated (BEH-38): the write rules (pattern and the step each belongs to), the patterns of the
1118 * script rules, the steps with a script rule of target written, and the work-file patterns, across the plugin, organization
1119 * and project layers (their union). */
1120export type Gated = { writes: { step: number; pattern: string }[]; scripts: string[]; scriptSteps: number[]; workFiles: string[] }
1121
1122/** The gated rules of the manifests' parsed JSON (the layers' files); anything that isn't the expected shape adds nothing. */
1123export function gatedOf(manifests: readonly unknown[]): Gated {
1124 const g: Gated = { writes: [], scripts: [], scriptSteps: [], workFiles: [] }
1125 for (const m of manifests) {
1126 if (m === null || typeof m !== 'object') continue
1127 const steps = (m as Fields).steps
1128 if (steps !== null && typeof steps === 'object' && !Array.isArray(steps)) {
1129 for (const [key, step] of Object.entries(steps as Fields)) {
1130 const n = Number(key)
1131 const evidence = step !== null && typeof step === 'object' ? (step as Fields).evidence : undefined
1132 if (!Array.isArray(evidence) || !Number.isInteger(n)) continue
1133 for (const rule of evidence) {
1134 if (rule === null || typeof rule !== 'object' || typeof (rule as Fields).pattern !== 'string') continue
1135 const r = rule as Fields
1136 const pattern = r.pattern as string
1137 if (r.type === 'write' && !g.writes.some(w => w.step === n && w.pattern === pattern)) g.writes.push({ step: n, pattern })
1138 if (r.type === 'script') {
1139 if (!g.scripts.includes(pattern)) g.scripts.push(pattern)
1140 if (r.target === 'written' && !g.scriptSteps.includes(n)) g.scriptSteps.push(n)
1141 }
1142 }
1143 }
1144 }
1145 const work = (m as Fields).workFiles
1146 if (Array.isArray(work)) for (const p of work) if (typeof p === 'string' && !g.workFiles.includes(p)) g.workFiles.push(p)
1147 }
1148 return g
1149}
1150
1151/** Python's fnmatch.fnmatchcase (the evaluator's matching): * is any run of characters, / included, ? one character,
1152 * [seq] and [!seq] a class. */
1153export function fnmatchcase(name: string, pattern: string): boolean {
1154 let re = ''
1155 for (let i = 0; i < pattern.length; i++) {
1156 const c = pattern[i]
1157 if (c === '*') re += '[\\s\\S]*'
1158 else if (c === '?') re += '[\\s\\S]'
1159 else if (c === '[') {
1160 let j = i + 1
1161 if (pattern[j] === '!') j += 1
1162 if (pattern[j] === ']') j += 1
1163 while (j < pattern.length && pattern[j] !== ']') j += 1
1164 if (j >= pattern.length) re += '\\['
1165 else {
1166 let set = pattern.slice(i + 1, j).replace(/\\/g, '\\\\')
1167 if (set[0] === '!') set = '^' + set.slice(1)
1168 else if (set[0] === '^') set = '\\' + set
1169 re += `[${set}]`
1170 i = j
1171 }
1172 } else re += c.replace(/[.*+?^${}()|[\]\\/]/g, '\\$&')
1173 }
1174 try {
1175 return new RegExp(`^${re}$`).test(name)
1176 } catch {
1177 return false
1178 }
1179}
1180
1181/** evaluate.py's path_matches: a pattern ending in / means anything inside that folder. */
1182export function pathMatches(path: string, pattern: string): boolean {
1183 if (pattern.endsWith('/')) return path.startsWith(pattern) || path === pattern.slice(0, -1)
1184 return fnmatchcase(path, pattern)
1185}
1186
1187/** A project-relative path against a rule's pattern: one starting with / or ../ matches none (SPEC-012 BEH-06). */
1188export function ruleMatches(path: string, pattern: string): boolean {
1189 return !path.startsWith('/') && !path.startsWith('../') && pathMatches(path, pattern)
1190}
1191
1192const INTERPRETERS = ['python', 'python3', 'bash', 'sh', 'node']
1193const PYTHON_N = /^python3\.[0-9]+$/
1194const ENV_WORD = /^[A-Za-z_][A-Za-z0-9_]*=/
1195
1196/** The last path segment as Python's PurePosixPath(word).name gives it. */
1197function baseName(word: string): string {
1198 const p = word.replace(/\/+$/, '')
1199 const name = p.slice(p.lastIndexOf('/') + 1)
1200 return name === '.' ? '' : nametypes/index.d.ts 117 lines1// The progress tracker adapter's $.state contract (SPEC-013 v27 DM-03; version 27 adds no key: a hold of BEH-42 lives in the module's
2// memory with the lines it keeps, since a reload loses those lines and a count that outlived them would describe nothing). $.state survives a reload of the module
3// and empties on /clear, /resume and /branch; whether the session is interactive and the evaluation timer are
4// module variables instead (BEH-01, BEH-06).
5
6/** The open run: its ID, skill, the last seq used, its folder under devforgeai/progress/runs/, and the root it
7 * opened in, which its paths, manifests and files use (BEH-03). Its event lines
8 * live in the module and events.jsonl, not here: one $.state value holds at most 4,194,304 characters, which a
9 * run's lines can pass before the log reaches its 4 MiB (found by the build's ERR-11 test). */
10export type ProgressRun = {
11 id: string
12 skill: string
13 seq: number
14 dir: string
15 root: string
16}
17
18/** What the status line and the band draw, taken from the last evaluation (SPEC-012 DM-03). */
19export type ProgressSummary = {
20 skill: string
21 current: number | null
22 steps: number
23 flags: number
24 yourTurn: boolean
25 ended: string | null
26 manifest: 'matched' | 'stale' | 'none' | 'unverified'
27 states: string[]
28 currentTitle: string | null
29 lastFlag: string | null
30 stoppedAt?: number | null // version 12: the step a deliberate stop's answer was tagged with (SPEC-013 BEH-10, BEH-27)
31}
32
33/** One enforce refusal with its gate's kind, seq and first flag, kept for the run's review (BEH-25, BEH-26). */
34export type ProgressRefused = { gate: string; seq: number; step: number; type: string; message: string }
35
36/** A run paused on the trail (BEH-29, version 14; version 13 held skill, step and tasks only): its return step, its
37 * task IDs, and the values it had while open, which it gets back when the trail unwinds to it (BEH-30). */
38export type ProgressPaused = {
39 skill: string
40 step: number
41 tasks: Record<string, number>
42 run: ProgressRun | null
43 summary: ProgressSummary | null
44 marked: boolean
45 shown: string[]
46 contextSent: number[]
47 todos: Record<string, string>
48 adhered: string | null
49 refusals: Record<string, number>
50 refused: ProgressRefused[]
51 reviewed: string | null
52}
53
54/** A run that ended returned or stopped while nested, kept for the turn's review (BEH-26, BEH-30; version 14). */
55export type ProgressReturned = {
56 run: ProgressRun | null
57 refused: ProgressRefused[]
58 reason: 'returned' | 'stopped'
59}
60
61/** The documents a Bash call wrote that were recorded, by run ID and path, with their size and mtimeMs as listed (BEH-38 (b); version 22). */
62export type ProgressWroteSeen = { [run: string]: { [path: string]: string } }
63
64/** The precompact row's values (SPEC-013 BEH-35, BEH-36, ERR-22; versions 21, 23 and 25): the measured share of the context
65 * window (a whole number from 0 to 100) or null; the row hidden by a load of the precompact skill; BEH-36's run mark (an
66 * automatic run has started since the last compaction); ERR-22's failure; and `pending` (version 26), BEH-36's one-shot mark that
67 * the plugin's own precompact skill is about to load, set in the same write as `ran`, consumed by that load, cleared by ERR-22.
68 * A compaction, /clear, /resume and /branch empty them. */
69export type ProgressPrecompact = { percent: number | null; hidden: boolean; ran: boolean; failed: boolean; pending: boolean }
70
71export type ProgressMode = 'observe' | 'enforce'
72
73export type ProgressModeSource = 'framework-default' | 'local'
74
75declare module 'claude-code' {
76 interface PluginState {
77 devforgeai: {
78 run: ProgressRun | null
79 mode: ProgressMode
80 modeSource: ProgressModeSource
81 summary: ProgressSummary | null
82 lastEventAt: number
83 marked: boolean
84 shown: string[]
85 contextSent: number[]
86 off: string | null
87 /** The open run's task IDs and their step numbers (BEH-20); a new run starts empty. */
88 tasks: Record<string, number>
89 /** TodoWrite: each step's last status, by step number (BEH-20). */
90 todos: Record<string, string>
91 /** The task-tools hint was shown this session (BEH-23). */
92 hinted: boolean
93 /** The run given the adherence notice (BEH-22), so a reload doesn't repeat it. */
94 adhered: string | null
95 /** The open run's enforce refusals by cause, '<gate kind>:<flag type>:<step>' (BEH-25); a new run starts empty. */
96 refusals: Record<string, number>
97 /** The open run's refusals, each with its gate's kind, seq and first flag, for the review (BEH-25, BEH-26). */
98 refused: ProgressRefused[]
99 /** The run whose review was asked (BEH-26), so a reload doesn't repeat it. */
100 reviewed: string | null
101 /** The paused runs, bottom first (BEH-29; version 13, the full values from version 14). */
102 trail: ProgressPaused[]
103 /** Runs that ended returned or stopped while nested, for the turn's review (BEH-30; version 14). */
104 returned: ProgressReturned[]
105 /** The runs whose work files cleanup has started, at most once per run (SPEC-013 BEH-32; version 20). */
106 cleaned: string[]
107 /** For each of the last 20 runs, the documents a Bash call wrote that were recorded, by path, with their size and mtimeMs as
108 * listed, so an unchanged recorded file isn't recorded twice (SPEC-013 BEH-38 (b); version 22). */
109 wroteSeen: ProgressWroteSeen
110 /** The run whose last evaluation closed BEH-38's window (every step reached, ended, or its validator's step done), or null (version 22). */
111 closedFor: string | null
112 /** The precompact row's values and BEH-36's run mark (SPEC-013 BEH-35, BEH-36, ERR-22; versions 21, 23 and 25). */
113 precompact: ProgressPrecompact
114 }
115 }
116}
117