Keeps a GitHub Spec Kit feature and its code tied together inside Claude Code: the spec in the system prompt, every edit booked against a task, a proof ladder…

A Claude Code mod that keeps a GitHub Spec Kit feature and its code tied together while the agent works: the spec rides every request, every edit is booked against a task, and after each turn a drift check asks whether the change still matches what the user said and what the spec says.
It complements read-only viewers such as SpecKit Companion: they show the run; this one keeps the run on the spec.
| How | |
|---|---|
| Spec in the system prompt | prompt.compose adds one section: active feature, the user's original words, user stories, constitution MUST rules, out of scope, open NEEDS CLARIFICATION, and the rules of the game. It changes only when the spec does, so the prompt cache survives. |
| Current task on every prompt | prompt.submit attaches a note: current task, its planned files, the requirements it serves, the next tasks, the drift status. |
| Intent log | Every prompt the person types is logged (intents in xref.json): their own words, the ground truth of the idea. |
| Edit booking | tool.call on Edit/Write/NotebookEdit books the file against the current task (paths come from tasks.md). A file planned for another task makes that task current. A file planned for none is drift: the model reads a note beside the tool result right away. Files Bash changed during a turn (sed, a generator, npm) are found at its end and booked the same way. Lockfiles and build output never count; manifests and configs are unclear, not drift; .xrefignore adds your own patterns. |
| Strict mode | Optional: refuses an unplanned edit and names the way forward (focus another task, link the file, or extend the spec). If the guard itself fails, it refuses too. |
| Proof ladder | Each requirement is specified, planned, implemented, tested or passing. Tested means a test file with real tests anchors it; passing means the test run proved it for its current text. A changed requirement drops back and asks to be re-verified. |
| Anchors | // @spec 001-feature/FR-003 comments count as coverage; one that names a requirement the spec no longer has is drift. |
| Intent check | After a turn that wrote files, one tool-less $.model.fork over the session compares the user's words, the spec and the diff. It returns a score, reasons and requests the spec lacks. With no transcript to fork yet, one completion fed by the intent log stands in. |
| Fix it from the pane | To spec runs /speckit-clarify with the request; As task appends a [Drift] task to tasks.md; Map FR→tasks lets a small model map requirements to tasks once and keeps the map. |
| Band and pane | A line above the prompt (▲ xref T004 · tasks 3/8 · FR 4/5 · watch (1)) that stacks with other mods' bands; while the autopilot waits, it shows the whole question with Resume, Stop and Pane. The pane: a phase strip, Goal, Vision, Now, Todo, Status, Proof, Next, a review card at the spec gate, "Away" after an autopilot run, open decisions, Auto, Tests and Drift. Each unplanned file is a card with Link to (a task picker) and Accept this; Accept all asks first. Under 70 columns the pane is compact; d shows the rest. |
| Transcript | Each booked write carries a chip (● T006 · FR-003, ▲ unplanned), and the autopilot's long prompts fold to one line (▶ auto 3/25 · implement · …); ctrl+o shows them whole. |
| Spec Kit's own commands | /speckit-implement, -plan, -tasks, -clarify, -analyze and -converge get the current task, its files, coverage and drift appended. A compaction keeps the task in focus, your open requests and the decisions. |
The model gets five tools: mcp__speckit-xref__status (CLI, setup, integration, workflow phase, autopilot and the next command), mcp__speckit-xref__ask (a question for the person, answered in place where they are there; blocks names the stories it holds up), mcp__speckit-xref__focus (make a task current), __where (which tasks and requirements a file belongs to) and __link (tie a file to a task, requirement or scenario). In a repository without Spec Kit they wait behind ToolSearch, and nothing polls.
What the mod keeps per feature lives in two files:
specs/<feature>/xref.json, committed: which tasks serve which requirement, which files are linked to what, which unplanned files were accepted, and each requirement's fingerprint. Sorted, without timestamps, so it merges and reviews cleanly..specify/xref/local/<feature>.json, git-ignored by a .gitignore of its own: your logged prompts, the intent verdicts, test results, approvals, decisions, and the autopilot's run log (<feature>.run.jsonl). Your words stay on your machine.A 0.3 xref.json is migrated on first load.
At the prompt of a terminal session:
/plugin install speckit-xref --marketplace moinsen-dev/speckit-xref
/xref auto on [steps] lets the work run: after every turn the mod hands the next Spec Kit step to Claude by itself, so nobody has to answer "next is X, shall I go on?". It follows Spec Kit's whole loop: constitution, specify, clarify, your review of the spec, plan, tasks, map, analyze, implement one phase per step, converge until the tasks stop changing, verify.
It stops only for what is yours:
/xref approve, or the dialog) lets it go on; the plugin option review sets spec+plan or none. An approval holds for that text: a changed spec asks again;[NEEDS CLARIFICATION] question, or open checklist items;git push, reset --hard, rm -rf and writes to .env files are refused and turned into a question.Claude asks through mcp__speckit-xref__ask. Where you are there, the question opens as a dialog and the run goes on with your answer in the same turn. A question that blocks only some stories waits in the pane's Decide row while the rest goes on. When an answer ends on a question anyway, a small model decides whether it is a real decision or just "shall I go on?".
Tests gate the progress. After each implement step the mod runs the test command (the testCommand option, else derived from plan.md's Testing: line). A failure becomes a repair step, at most three in a row; then it is your call. A passing suite proves every requirement with a real test (with junitPath, each one by the tests that name or anchor it). The run ends only when the suite passes once more.
It also stops on Esc or /xref-stop (both end the running turn), on an API error or a refusal, when the drift turns red, after three steps without progress, at the step budget (25 by default) and when the feature is done. After five minutes of waiting it sends a notification. The band shows auto ▶ 3/25 · $1.80 · 23m while it runs and the whole question when it waits; the pane tab reads Spec X-Ref ⏸. Back at the keyboard, the pane's Away card says what happened.
More, all off by default:
/xref auto night: a night run of up to 100 steps; it writes .specify/xref/local/briefing.md when it stops.commitPerTask: on: a step whose tests pass commits the tasks it checked off, on a feature branch only; nothing you staged is mixed in, and files outside the plan or .env files stop it.parallel: on: a phase's open [P] tasks go to speckit-xref:task-runner subagents, one per task.SPECKIT_XREF_AUTOPILOT=on|night|<steps> claude -p "…" runs the autopilot in a -p session; each next step rides the Stop hook.The pane's Auto row switches it:
p);p);r) and Stop. A spent budget restarts on Resume.Click the buttons, or give the pane the keyboard with ctrl+x tab (once more if the band takes it first). /xref pane looks at the folder first, opens the pane without taking the keyboard (so the next key never presses a button), and says where things stand.
The autopilot never sets Spec Kit up by itself: specify init and specify integration install write into the repository, so it waits there for you. To start every session with the autopilot on, set the plugin options autopilot: on and autopilotMaxSteps. It is off by default, because every step is a model turn.
/xref pane and the status tool tell the folder apart, and the pane offers the one fitting start:
| Folder | Pane offers | What Claude then does |
|---|---|---|
| empty (only dotfiles, a README, a license) | Start from an idea | asks what you want to build, sets Spec Kit up, writes the spec in your words |
existing code, no .specify/ | Set up Spec Kit here | runs specify init, drafts the constitution from the code, asks which change comes first |
| Spec Kit without its Claude Code commands | Add Claude integration | runs specify integration install claude |
| inside a Spec Kit project | the project itself | the mod looks upward for .specify/, as Spec Kit does |
Pressing one of these counts as your go. Without a press or your request nothing is written. A repository without Spec Kit stays quiet otherwise, with no band and no unasked pane.
speckit skill/speckit-xref:speckit, or any request to set up or run Spec Kit, loads a skill that:
specify CLI or runs it through uvx, says what specify init writes before it runs, then moves on to the constitution;/xref status of the active feature
/xref check run the intent check now
/xref map map requirements to tasks with a small model
/xref ack accept every edit outside the plan
/xref approve approve what the autopilot waits on (spec, plan, open checklists)
/xref focus T004 make a task current
/xref auto on [n] autopilot: work through Spec Kit by itself (night: a night run, off: stop)
/xref pane open the pane
/xref-stop stop the autopilot now, mid-turn
Options (/config or pluginConfigs in settings): mode (advisory | strict), driftCheck (fork | off), mapModel (default haiku), autopilot (off | on), autopilotMaxSteps (default 25), review (spec | spec+plan | none), testCommand, junitPath, commitPerTask (off | on), parallel (off | on).
The active feature is found as Spec Kit finds it, then by the git branch (001-…), else the spec written last.
claude --plugin-dir ./mod # load this checkout, reloading on save
claude plugin validate --strict ./mod
claude plugin test ./mod
node scripts/build-fixture.mjs # after changing examples/demo or scripts/fixtures/cases.json
examples/demo is a small Spec Kit project (magic-link login) to try it on. In the desktop app's Code tab, name the folder in CLAUDE_CODE_PLUGIN_DIRS under env in ~/.claude/settings.json.
Tested on Claude Code 2.1.295 and Spec Kit 1.1.2. The mods API is early access and can change between releases.
hooks/register.tsx 1839 lines1import { atom, read, update } from 'claude-code'
2import type { Color, EngineInterface, Register } from 'claude-code'
3
4import type { Autopilot, Ledger, RunEntry, Snapshot } from '../types'
5import { LOCAL_IGNORE, ledgerFromParts, ledgerToParts, localPath, runLogPath } from './ledger'
6import { LADDER, levelOf, parseJunit, verificationFrom, verificationFromExit } from './proof'
7import { featureFromBranch, fingerprint, hasRealTests, isTestFile, matchesAny, rulesFrom } from './rules'
8import { invokeSeparator, parseConstitution, parseFeatureJson, parseSpec, parseTasks } from './speckit'
9import type { Flow, Step } from './workflow'
10import {
11 AUTONOMY_RULES,
12 QUESTION_LABELS,
13 autopilotPrompt,
14 endsWithQuestion,
15 idleAutopilot,
16 nextStep,
17 openChecklistItems,
18 openPhase,
19 persistenceModel,
20 progressKey,
21 MAX_REPAIRS,
22 phaseStrip,
23 setupNote,
24 stepLine,
25 takesIdea,
26 tasksFingerprint,
27 testCommandFrom,
28} from './workflow'
29import {
30 acknowledgeAll,
31 anchorsIn,
32 appendRemediation,
33 applyMapping,
34 applySemantic,
35 classify,
36 composeSection,
37 currentTask,
38 editNote,
39 emptyLedger,
40 evaluate,
41 isSpecArtifact,
42 forkPrompt,
43 link,
44 logIntent,
45 mapPrompt,
46 parseSemantic,
47 recordTouch,
48 relPath,
49 reqsOf,
50 resolveIntent,
51 short,
52 speckitCommand,
53 turnContext,
54} from './xref'
55import type { Anchor } from '../types'
56import type { Classified, Level, Report } from './xref'
57
58type $ = EngineInterface
59type Options = {
60 mode: string
61 driftCheck: string
62 mapModel: string
63 autopilot: string
64 autopilotMaxSteps: number
65 testCommand: string
66 junitPath: string
67 review: string
68 commitPerTask: string
69 parallel: string
70}
71
72const PLUGIN = 'speckit-xref'
73const PANE = 'speckit-xref'
74const TITLE = 'Spec X-Ref'
75const COMMAND = 'xref'
76const STOP_COMMAND = 'xref-stop'
77const REFRESH_MS = 4000
78const WRITE_TOOLS = new Set(['Edit', 'Write', 'NotebookEdit'])
79// Prompts the person typed, wherever they typed them; a plugin's or a peer's are not their intent.
80const PERSON = new Set(['composer', 'bridge', 'sdk'])
81const LEVEL_COLOR = { none: 'subtle', green: 'success', yellow: 'warning', red: 'error' } as const
82const RANK: Record<Level, number> = { none: 0, green: 1, yellow: 2, red: 3 }
83// What the person reads: a word and a glyph, the color only on top.
84const STATE: Record<Level, { glyph: string; word: string }> = {
85 none: { glyph: '·', word: 'no feature' },
86 green: { glyph: '●', word: 'ok' },
87 yellow: { glyph: '▲', word: 'watch' },
88 red: { glyph: '✖', word: 'off-spec' },
89}
90const ANCHOR_RG = '@spec\\s+(?:[\\w.-]+/)?(?:(?:FR|SC)-\\d{3,}|T-?\\d{3,}|US\\d+-AS\\d+)'
91const ANCHOR_GIT = '@spec[[:space:]]+([[:alnum:]_.-]+/)?((FR|SC)-[0-9]{3,}|T-?[0-9]{3,}|US[0-9]+-AS[0-9]+)'
92
93const snapshotA = atom({ plugin: 'speckit-xref', key: 'snapshot' } as const, null)
94const ledgerA = atom({ plugin: 'speckit-xref', key: 'ledger' } as const, emptyLedger())
95const activeA = atom({ plugin: 'speckit-xref', key: 'active' } as const, null)
96const checkingA = atom({ plugin: 'speckit-xref', key: 'checking' } as const, false)
97const autopilotA = atom({ plugin: 'speckit-xref', key: 'autopilot' } as const, idleAutopilot())
98// Turn state lives in atoms, so a reload or a /config change in the middle of a turn keeps it.
99const askedA = atom({ plugin: 'speckit-xref', key: 'asked' } as const, null)
100const turnFilesA = atom({ plugin: 'speckit-xref', key: 'turnFiles' } as const, [] as string[])
101const turnA = atom({ plugin: 'speckit-xref', key: 'turn' } as const, null)
102const detailsA = atom({ plugin: 'speckit-xref', key: 'details' } as const, false)
103const seenAtA = atom({ plugin: 'speckit-xref', key: 'seenAt' } as const, 0)
104const chipsA = atom({ plugin: 'speckit-xref', key: 'chips' } as const, {} as Record<string, string>)
105
106// The module's own: they start over on a reload, and session.start or register fills them again.
107let root = ''
108let signature = ''
109let offered = false
110let interactive = true
111let lastLevel: Level = 'none'
112let timer: { cancel: () => void } | null = null
113// What can run Spec Kit's CLI here, looked up once a session.
114let cli = { specify: false, uvx: false }
115let mapModel = 'haiku'
116let defaultMax = 25
117let testCommandOption = ''
118let junitPath = ''
119let review: Flow['review'] = 'spec'
120// Read by the strict guard's fallback, which has to be a top-level function.
121let strict = false
122let parallel = false
123let commitPerTask = false
124// In a `-p` run the next step rides the Stop hook's re-prompt instead of a prompt of its own.
125let pendingPrompt: string | null = null
126// When the step running now was handed over: its length in the run log, where a -p run reports none.
127let stepStartedAt = 0
128
129const at = (rel: string) => (rel.startsWith('/') ? rel : `${root}/${rel}`)
130const stamp = async ($: $) => new Date(await $.clock.now()).toISOString()
131
132async function readText($: $, rel: string): Promise<string | null> {
133 try {
134 return await $.fs.read(at(rel))
135 } catch {
136 return null
137 }
138}
139
140async function mtime($: $, rel: string): Promise<number> {
141 try {
142 return (await $.fs.stat(at(rel))).mtimeMs
143 } catch {
144 return 0
145 }
146}
147
148async function exists($: $, rel: string): Promise<boolean> {
149 try {
150 return await $.fs.exists(at(rel))
151 } catch {
152 return false
153 }
154}
155
156/**
157 * The project root: the nearest folder at or above the working directory that holds `.specify/`, as Spec Kit's
158 * own scripts find it. Claude started in a subfolder of a Spec Kit project must not be offered a nested setup.
159 */
160async function projectRoot($: $, cwd: string): Promise<string> {
161 let dir = cwd.replace(/\/+$/, '') || '/'
162 for (let depth = 0; depth < 16; depth++) {
163 if (await $.fs.exists(`${dir === '/' ? '' : dir}/.specify`).catch(() => false)) return dir
164 const parent = dir.slice(0, dir.lastIndexOf('/')) || '/'
165 if (parent === dir) break
166 dir = parent
167 }
168 return cwd
169}
170
171const IGNORABLE = /^(\.|README|LICENSE|CHANGELOG)/i
172
173/** Empty: nothing at the top but dotfiles, a README, a license or a changelog. */
174async function folderKind($: $): Promise<'empty' | 'existing'> {
175 try {
176 return (await $.fs.list(root)).some(entry => !IGNORABLE.test(entry.name)) ? 'existing' : 'empty'
177 } catch {
178 return 'existing'
179 }
180}
181
182/** Whether a program is on the PATH. */
183async function onPath($: $, program: string): Promise<boolean> {
184 try {
185 return (await $.process.run(['which', program], { cwd: root, timeoutMs: 3000 })).exitCode === 0
186 } catch {
187 return false
188 }
189}
190
191/** Every feature directory with a spec.md, and when its spec or tasks last changed. */
192async function listFeatures($: $): Promise<{ dir: string; time: number }[]> {
193 const found: { dir: string; time: number }[] = []
194 for (const base of ['specs', '.specify/specs']) {
195 let entries
196 try {
197 entries = await $.fs.list(at(base))
198 } catch {
199 continue
200 }
201 for (const entry of entries) {
202 if (entry.kind !== 'dir') continue
203 const dir = `${base}/${entry.name}`
204 const spec = await mtime($, `${dir}/spec.md`)
205 if (spec) found.push({ dir, time: Math.max(spec, await mtime($, `${dir}/tasks.md`)) })
206 }
207 }
208 return found.sort((a, b) => a.dir.localeCompare(b.dir))
209}
210
211/** The git branch checked out; null outside git or on a detached HEAD. */
212async function currentBranch($: $): Promise<string | null> {
213 try {
214 const ran = await $.process.run(['git', 'rev-parse', '--abbrev-ref', 'HEAD'], { cwd: root, timeoutMs: 3000 })
215 const branch = ran.exitCode === 0 ? ran.stdout.trim() : ''
216 return branch && branch !== 'HEAD' ? branch : null
217 } catch {
218 return null
219 }
220}
221
222/**
223 * The active feature, as docs/contract-0.4.md §7 orders it: SPECIFY_FEATURE_DIRECTORY, .specify/feature.json,
224 * SPECIFY_FEATURE, the feature the git branch names (`001-…`), else the spec written last.
225 */
226async function resolveFeature($: $, features?: { dir: string; time: number }[], branch?: string | null): Promise<string | null> {
227 const named = await $.env.get('SPECIFY_FEATURE')
228 const pointers = [
229 await $.env.get('SPECIFY_FEATURE_DIRECTORY'),
230 parseFeatureJson((await readText($, '.specify/feature.json')) ?? ''),
231 named ? `specs/${named}` : null,
232 ]
233 for (const pointer of pointers) {
234 if (!pointer) continue
235 const rel = relPath(pointer, root).replace(/\/$/, '')
236 if (await exists($, `${rel}/spec.md`)) return rel
237 }
238 const all = features ?? (await listFeatures($))
239 const fromBranch = featureFromBranch(branch === undefined ? ((await currentBranch($)) ?? '') : (branch ?? ''), all.map(f => f.dir))
240 if (fromBranch) return fromBranch
241 const latest = [...all].sort((a, b) => b.time - a.time)[0]
242 return latest?.dir ?? null
243}
244
245async function currentSignature($: $, featureDir: string | null): Promise<string> {
246 const files = ['.specify/feature.json', '.specify/memory/constitution.md', '.specify/integration.json', '.specify/extensions.yml']
247 if (featureDir) files.push(`${featureDir}/spec.md`, `${featureDir}/tasks.md`, `${featureDir}/plan.md`, '.xrefignore')
248 const times = await Promise.all(files.map(f => mtime($, f)))
249 return [featureDir ?? '-', ...times].join('|')
250}
251
252/** Reads Spec Kit's files into the snapshot; a new feature brings its own ledger. */
253async function scan($: $): Promise<void> {
254 const features = await listFeatures($)
255 const branch = await currentBranch($)
256 const featureDir = await resolveFeature($, features, branch)
257 const [specMd, tasksMd, planMd, constitutionMd, integrationJson, initOptions, ignoreText] = await Promise.all([
258 featureDir ? readText($, `${featureDir}/spec.md`) : null,
259 featureDir ? readText($, `${featureDir}/tasks.md`) : null,
260 featureDir ? readText($, `${featureDir}/plan.md`) : null,
261 readText($, '.specify/memory/constitution.md'),
262 readText($, '.specify/integration.json'),
263 readText($, '.specify/init-options.json'),
264 readText($, '.xrefignore'),
265 ])
266 let extensions: string[] = []
267 try {
268 extensions = (await $.fs.list(at('.specify/extensions'))).filter(e => e.kind === 'dir' && !e.name.startsWith('.')).map(e => e.name)
269 } catch {
270 extensions = []
271 }
272 // integration.json says how commands are invoked; without it, a commands-only layout is the old one.
273 const separator = integrationJson ? invokeSeparator(integrationJson) : null
274 const commandsOnly = separator ? separator === '.' : (await exists($, '.claude/commands/speckit.clarify.md')) && !(await exists($, '.claude/skills/speckit-clarify/SKILL.md'))
275 const tasks = tasksMd ? parseTasks(tasksMd) : []
276 const previous = await read($, snapshotA)
277 const snapshot: Snapshot = {
278 initialized: await exists($, '.specify'),
279 featureDir,
280 features: features.map(f => f.dir),
281 hasPlan: planMd !== null,
282 extensions,
283 claudeIntegration: await hasClaudeIntegration($),
284 speckitVersion: speckitVersionOf(initOptions),
285 tools: cli,
286 folder: await folderKind($),
287 spec: specMd ? parseSpec(specMd) : null,
288 tasks,
289 constitution: constitutionMd ? parseConstitution(constitutionMd) : null,
290 commandStyle: commandsOnly ? 'commands' : 'skills',
291 rules: rulesFrom(ignoreText),
292 realTests: previous?.featureDir === featureDir ? previous.realTests : [],
293 planFingerprint: planMd ? fingerprint(planMd) : null,
294 testCommand: testCommandOption || testCommandFrom(planMd),
295 branch,
296 phase: openPhase(tasks),
297 tasksFingerprint: tasksFingerprint(tasks),
298 commands: await installedCommands($),
299 checklists: featureDir ? await checklists($, featureDir) : [],
300 persistence: persistenceModel(constitutionMd),
301 }
302 if (!previous || previous.featureDir !== featureDir) {
303 const ledger = featureDir ? ledgerFromParts(await readText($, `${featureDir}/xref.json`), await readText($, localPath(featureDir))) : emptyLedger()
304 const anchors = featureDir ? await scanAnchors($) : null
305 await update($, ledgerA, () => (anchors ? { ...ledger, anchors } : ledger))
306 await update($, activeA, () => null)
307 snapshot.realTests = await findRealTests($, anchors ? { ...ledger, anchors } : ledger)
308 }
309 await update($, snapshotA, () => snapshot)
310 signature = await currentSignature($, featureDir)
311}
312
313/** Spec Kit's commands installed for Claude Code, by name without prefix: `plan`, `analyze`, `xref-check`. */
314async function installedCommands($: $): Promise<string[]> {
315 const names = new Set<string>()
316 try {
317 for (const e of await $.fs.list(at('.claude/skills'))) if (e.kind === 'dir' && e.name.startsWith('speckit-')) names.add(e.name.slice('speckit-'.length))
318 } catch {
319 // No skills folder: the commands layout, or no integration.
320 }
321 try {
322 for (const e of await $.fs.list(at('.claude/commands'))) {
323 const m = /^speckit\.(.+)\.md$/.exec(e.name)
324 if (m) names.add(m[1]!.replace(/\./g, '-'))
325 }
326 } catch {
327 // No commands folder.
328 }
329 return [...names].sort()
330}
331
332/** The feature's checklists with open items. */
333async function checklists($: $, featureDir: string): Promise<{ file: string; open: number }[]> {
334 const files: Record<string, string> = {}
335 try {
336 for (const e of await $.fs.list(at(`${featureDir}/checklists`))) {
337 if (e.kind !== 'file' || !e.name.endsWith('.md')) continue
338 const text = await readText($, `${featureDir}/checklists/${e.name}`)
339 if (text !== null) files[`checklists/${e.name}`] = text
340 }
341 } catch {
342 return []
343 }
344 return openChecklistItems(files)
345}
346
347/** The test files tied to the feature (anchored, linked or touched) that hold real tests, not only test.todo. */
348async function findRealTests($: $, ledger: Ledger): Promise<string[]> {
349 const candidates = new Set<string>()
350 for (const a of ledger.anchors) candidates.add(a.file)
351 for (const r of Object.values(ledger.requirements)) r.files.forEach(f => candidates.add(f))
352 for (const t of Object.values(ledger.tasks)) [...t.touched, ...t.linked].forEach(f => candidates.add(f))
353 const real: string[] = []
354 for (const file of [...candidates].filter(isTestFile).slice(0, 200)) {
355 const text = await readText($, file)
356 if (text !== null && hasRealTests(text)) real.push(file)
357 }
358 return real.sort()
359}
360
361/**
362 * Spec Kit's Claude Code integration: its core commands are there, as skills or as commands. integration.json
363 * alone is not enough, and neither is a stray speckit-* skill: the autopilot has to be able to run the next step.
364 */
365async function hasClaudeIntegration($: $): Promise<boolean> {
366 for (const name of ['plan', 'implement']) {
367 if ((await exists($, `.claude/skills/speckit-${name}/SKILL.md`)) || (await exists($, `.claude/commands/speckit.${name}.md`))) return true
368 }
369 return false
370}
371
372function speckitVersionOf(initOptions: string | null): string | null {
373 try {
374 const value = JSON.parse(initOptions ?? '') as { speckit_version?: unknown }
375 return typeof value.speckit_version === 'string' ? value.speckit_version : null
376 } catch {
377 return null
378 }
379}
380
381async function refresh($: $): Promise<void> {
382 const snap = await read($, snapshotA)
383 const featureDir = await resolveFeature($, undefined, snap?.branch)
384 if (featureDir !== (snap?.featureDir ?? null) || (await currentSignature($, featureDir)) !== signature) await scan($)
385}
386
387/** Every `@spec <id>` comment in the repository, by ripgrep or else git grep; none when neither runs. */
388async function scanAnchors($: $): Promise<Anchor[] | null> {
389 const runs = [
390 ['rg', '-n', '--no-heading', '-o', '-e', ANCHOR_RG, '--glob', '!specs/**', '--glob', '!.specify/**', '.'],
391 // --untracked: a new file the person has not committed yet holds anchors too.
392 ['git', 'grep', '--untracked', '-n', '-o', '-E', ANCHOR_GIT, '--', '.', ':!specs', ':!.specify'],
393 ]
394 for (const argv of runs) {
395 try {
396 const ran = await $.process.run(argv, { cwd: root, timeoutMs: 5000 })
397 if (ran.exitCode > 1) continue
398 const anchors: Anchor[] = []
399 for (const line of ran.stdout.split('\n')) {
400 const m = /^(.+?):(\d+):(.*)$/.exec(line)
401 if (!m) continue
402 for (const a of anchorsIn(m[3] ?? '', relPath(m[1] ?? '', root))) anchors.push({ ...a, line: Number(m[2]) })
403 }
404 return anchors
405 } catch {
406 continue
407 }
408 }
409 return null
410}
411
412/**
413 * Writes both halves of the ledger: the committed `xref.json` (deterministic, reviewable) and the local file
414 * under `.specify/xref/local/`, which a `.gitignore` of its own keeps out of commits.
415 */
416async function persist($: $): Promise<void> {
417 const snap = await read($, snapshotA)
418 if (!snap?.featureDir) return
419 const ledger = await read($, ledgerA)
420 const parts = ledgerToParts(ledger, snap.featureDir)
421 try {
422 if ((await readText($, `${snap.featureDir}/xref.json`)) !== parts.committed) await $.fs.write(at(`${snap.featureDir}/xref.json`), parts.committed)
423 if (!(await exists($, LOCAL_IGNORE.path))) await $.fs.write(at(LOCAL_IGNORE.path), LOCAL_IGNORE.text)
424 await $.fs.write(at(localPath(snap.featureDir)), parts.local)
425 } catch {
426 // A read-only checkout keeps the ledger for the session alone.
427 }
428}
429
430/** Tells the person once the drift gets worse, never on every write. */
431async function notifyLevel($: $): Promise<void> {
432 const snap = await read($, snapshotA)
433 if (!snap) return
434 const report = evaluate(snap, await read($, ledgerA))
435 if (RANK[report.level] > RANK[lastLevel] && (report.level === 'yellow' || report.level === 'red')) {
436 $.ui.toast(`Spec drift ${report.level}: ${report.findings[0]?.text ?? ''}`)
437 }
438 lastLevel = report.level
439}
440
441async function afterWrite($: $, rel: string, c: Classified, newText: string): Promise<void> {
442 if (c.verdict === 'spec') {
443 await scan($)
444 return
445 }
446 // Lockfiles, build output, caches: no evidence for a task, no drift, no part of the turn's changes.
447 if (c.verdict === 'exempt') return
448 const when = await stamp($)
449 // A file's anchors are read again when the write may have added or removed one.
450 const hadAnchors = (await read($, ledgerA)).anchors.some(a => a.file === rel)
451 const text = hadAnchors || newText.includes('@spec') ? await readText($, rel) : null
452 const fresh = text === null ? null : anchorsIn(text, rel)
453 await update($, ledgerA, l => {
454 const next = recordTouch(l, rel, c, when)
455 return fresh ? { ...next, anchors: [...next.anchors.filter(a => a.file !== rel), ...fresh] } : next
456 })
457 // A file of a finished task is rework: it does not pull the focus back to that task.
458 const task = (await read($, snapshotA))?.tasks.find(t => t.id === c.task)
459 if (task && !task.done && (c.verdict === 'in-scope' || c.verdict === 'other-task')) await update($, activeA, () => task.id)
460 await update($, turnFilesA, files => (files.includes(rel) ? files : [...files, rel]))
461 // A test file written is read once: whether it holds real tests decides the ladder's `tested`.
462 if (isTestFile(rel)) {
463 const body = text ?? (await readText($, rel))
464 await update($, snapshotA, s => (s ? { ...s, realTests: body && hasRealTests(body) ? [...new Set([...s.realTests, rel])].sort() : s.realTests.filter(f => f !== rel) } : s))
465 }
466 await notifyLevel($)
467}
468
469/**
470 * The files the working tree changed, for a check the person asked for between turns: relative to the project
471 * and only inside it, since the project may be one folder of a larger repository.
472 */
473async function changedFiles($: $): Promise<string[]> {
474 const files = new Set<string>()
475 for (const argv of [
476 ['git', 'diff', '--name-only', '--relative', 'HEAD'],
477 ['git', 'ls-files', '--others', '--exclude-standard'],
478 ]) {
479 try {
480 const ran = await $.process.run(argv, { cwd: root, timeoutMs: 5000 })
481 if (ran.exitCode === 0) for (const line of ran.stdout.split('\n')) if (line.trim()) files.add(line.trim())
482 } catch {
483 // No git, or no commit yet: the files of this session's turns are all there is.
484 }
485 }
486 return [...files].filter(f => !isSpecArtifact(f)).slice(0, 40)
487}
488
489async function diffOf($: $, files: string[]): Promise<string> {
490 let diff = ''
491 try {
492 diff = (await $.process.run(['git', 'diff', '--no-color', '--relative', '-U2', 'HEAD', '--', ...files], { cwd: root, timeoutMs: 5000 })).stdout
493 } catch {
494 diff = ''
495 }
496 // A new file is no part of `git diff`; its head stands in for it.
497 for (const file of files) {
498 if (diff.length > 6000 || diff.includes(`b/${file}`)) continue
499 const text = await readText($, file)
500 if (text !== null) diff += `\n--- new or untracked: ${file}\n${text.split('\n').slice(0, 60).join('\n')}\n`
501 }
502 return diff
503}
504
505/** The semantic check: one tool-less question over the session's own transcript. */
506async function runCheck($: $, files: string[], model: string): Promise<string> {
507 const snap = await read($, snapshotA)
508 if (!snap?.featureDir) return 'No Spec Kit feature to check against.'
509 if (await read($, checkingA)) return 'A drift check is already running.'
510 await update($, checkingA, () => true)
511 try {
512 const changed = files.length ? files : await changedFiles($)
513 const ledger = await read($, ledgerA)
514 const active = await read($, activeA)
515 const diff = await diffOf($, changed)
516 let reply = await $.model.fork({ prompt: forkPrompt(snap, ledger, active, changed, diff) })
517 // Before the session's first turn there is no transcript to fork; the intent log carries the user's words instead.
518 if (!reply.isAnswered && reply.reason === 'nothing-to-fork') {
519 reply = await $.model.complete({ model, prompt: forkPrompt(snap, ledger, active, changed, diff, true), maxTokens: 1000 })
520 }
521 if (!reply.isAnswered) return `The drift check got no answer (${reply.reason}).`
522 const semantic = parseSemantic(reply.text, await stamp($))
523 if (!semantic) return 'The drift check answered in a shape the mod could not read.'
524 await update($, ledgerA, l => applySemantic(l, semantic))
525 await persist($)
526 await notifyLevel($)
527 const conflict = (await read($, ledgerA)).semantic?.changes.find(c => c.kind === 'contradicts')
528 if (conflict && (await read($, autopilotA)).on) await pauseAutopilot($, `your request "${conflict.text}" contradicts the spec; decide in the pane`)
529 return `Intent check ${semantic.score}/100 (${semantic.verdict})${semantic.reasons[0] ? `: ${semantic.reasons[0]}` : ''}`
530 } finally {
531 await update($, checkingA, () => false)
532 }
533}
534
535async function runMap($: $, model: string): Promise<string> {
536 const snap = await read($, snapshotA)
537 if (!snap?.spec || snap.tasks.length === 0) return 'Mapping needs a spec.md with requirements and a tasks.md.'
538 const reply = await $.model.complete({ model, prompt: mapPrompt(snap.spec, snap.tasks), maxTokens: 2000 })
539 if (!reply.isAnswered) return `The mapping got no answer (${reply.reason}).`
540 let mapped = 0
541 await update($, ledgerA, l => {
542 const result = applyMapping(l, reply.text, snap)
543 mapped = result.mapped
544 return result.ledger
545 })
546 await persist($)
547 const report = evaluate(snap, await read($, ledgerA))
548 return `Mapped ${mapped} requirements to tasks; ${report.covered}/${report.total} FRs covered.`
549}
550
551async function toSpec($: $, intent: string): Promise<void> {
552 const snap = await read($, snapshotA)
553 if (!snap) return
554 await update($, ledgerA, l => resolveIntent(l, intent))
555 await persist($)
556 const args = `The user asked during implementation: "${intent}". Fold this into the spec, or record why it stays out of scope.`
557 // The extension's revise keeps ids stable (new, SUPERSEDED, RETIRED) and logs revisions.md; clarify is the fallback.
558 const revise = snap.extensions.includes('xref') && snap.commands.includes('xref-revise')
559 // A plugin runs a slash command through $.command.run; a prompt may not start with one.
560 const command = speckitCommand(snap, revise ? 'xref.revise' : 'clarify').slice(1)
561 try {
562 await $.command.run({ command, args })
563 } catch {
564 void $.prompt.submit({ text: `Spec Kit: ${args} Use ${speckitCommand(snap, 'clarify')} for it.` })
565 }
566}
567
568async function asTask($: $, intent: string): Promise<string> {
569 const snap = await read($, snapshotA)
570 if (!snap?.featureDir) return 'No active feature.'
571 const path = `${snap.featureDir}/tasks.md`
572 const markdown = (await readText($, path)) ?? '# Tasks\n'
573 const added = appendRemediation(markdown, snap.tasks, intent)
574 await $.fs.write(at(path), added.markdown)
575 await update($, ledgerA, l => resolveIntent(l, intent))
576 await scan($)
577 await persist($)
578 return `Added ${added.id} to ${path}.`
579}
580
581async function focus($: $, id: string): Promise<string> {
582 const snap = await read($, snapshotA)
583 const task = snap?.tasks.find(t => t.id === id.trim().toUpperCase())
584 if (!snap || !task) return `No task ${id} in tasks.md.`
585 await update($, activeA, () => task.id)
586 const reqs = reqsOf(task, await read($, ledgerA))
587 return [
588 `Current task: ${task.id}${task.story ? ` [${task.story}]` : ''} ${task.text}`,
589 `Planned files: ${task.paths.join(', ') || '(none named)'}`,
590 `Serves: ${reqs.join(', ') || '(no requirement mapped yet)'}`,
591 ].join('\n')
592}
593
594async function where($: $, file: string): Promise<string> {
595 const snap = await read($, snapshotA)
596 if (!snap?.featureDir) return 'No active Spec Kit feature.'
597 const rel = relPath(file, root)
598 const ledger = await read($, ledgerA)
599 const c = classify(rel, snap, ledger, await read($, activeA))
600 const planned = snap.tasks.filter(t => t.paths.some(p => rel === p || rel.endsWith('/' + p) || rel.startsWith(p.replace(/\/?$/, '/'))))
601 const touched = Object.entries(ledger.tasks).filter(([, e]) => e.touched.includes(rel) || e.linked.includes(rel)).map(([id]) => id)
602 const reqs = Object.entries(ledger.requirements).filter(([, r]) => r.files.includes(rel)).map(([id]) => id)
603 const anchors = ledger.anchors.filter(a => a.file === rel).map(a => `${a.id} (line ${a.line})`)
604 return [
605 `${rel} in ${snap.featureDir}: ${c.verdict}${c.task ? ` (${c.task})` : ''}`,
606 `Planned by: ${planned.map(t => t.id).join(', ') || '-'}`,
607 `Touched or linked by: ${touched.join(', ') || '-'}`,
608 `Requirements: ${[...reqs, ...anchors].join(', ') || '-'}`,
609 ].join('\n')
610}
611
612async function linkFile($: $, file: string, id: string): Promise<string> {
613 const snap = await read($, snapshotA)
614 if (!snap?.featureDir) return 'No active Spec Kit feature.'
615 const rel = relPath(file, root)
616 let error: string | undefined
617 await update($, ledgerA, l => {
618 const result = link(l, rel, id, snap)
619 error = result.error
620 return result.ledger
621 })
622 if (error) return error
623 await persist($)
624 await notifyLevel($)
625 return `Linked ${rel} to ${id}.`
626}
627
628/** What the workflow reads beside the snapshot: the review option and the last test run. */
629const flowOf = (ap: Autopilot): Flow => ({ review, lastTest: ap.lastTest, repairs: ap.repairs })
630
631/** Starts a fresh autopilot run; its first step follows as soon as the session is free. */
632async function startAutopilot($: $, max?: number, night = false): Promise<void> {
633 const cost = await sessionCost($)
634 const now = await $.clock.now()
635 await update($, autopilotA, a => ({ ...a, on: true, paused: null, steps: 0, stalls: 0, last: null, lastPhase: null, idea: null, max: max ?? defaultMax, repairs: 0, lastTest: null, startedAt: now, costAtStart: cost, night, scope: null }))
636 await update($, seenAtA, () => now)
637 await retitle($)
638 $.clock.after(0, () => void advance($).catch(() => undefined))
639}
640
641/** Goes on after a pause; a run whose budget is spent gets a new one. */
642async function resumeAutopilot($: $): Promise<void> {
643 await update($, autopilotA, a => ({ ...a, paused: null, stalls: 0, repairs: a.repairs > MAX_REPAIRS ? 0 : a.repairs, steps: a.steps >= a.max ? 0 : a.steps }))
644 await retitle($)
645 $.clock.after(0, () => void advance($).catch(() => undefined))
646}
647
648/** Stop: the autopilot goes off, and a turn it is running ends now rather than at its end. */
649async function turnOffAutopilot($: $): Promise<void> {
650 const wasOn = (await read($, autopilotA)).on
651 await update($, autopilotA, a => ({ ...a, on: false, paused: null, idea: null }))
652 await retitle($)
653 const turn = await read($, turnA)
654 if (wasOn && turn) await $.turn.abort({ turnId: turn.id }).catch(() => undefined)
655 if (wasOn) await writeBriefing($, 'stopped by you')
656}
657
658/** The session's cost so far in US dollars, or 0 where the host keeps none. */
659async function sessionCost($: $): Promise<number> {
660 try {
661 return (await $.session.usage()).cost?.usd ?? 0
662 } catch {
663 return 0
664 }
665}
666
667/** The pane's tab says when the autopilot waits, so it shows behind the changes pane too. */
668async function retitle($: $): Promise<void> {
669 try {
670 const pane = (await $.ui.panes()).find(p => p.id === PANE)
671 if (!pane) return
672 const ap = await read($, autopilotA)
673 await $.ui.open({ id: PANE, title: ap.on && ap.paused ? `${TITLE} ⏸` : TITLE })
674 } catch {
675 // A host without panes has no tab to name.
676 }
677}
678
679/** The requests the pane's setup actions hand to Claude, as the person's own words: a press is their consent. */
680const SETUP_ASKS = {
681 idea: 'I want to start something new in this empty folder (an app, a project, or a problem to solve). Ask me what it is, then set Spec Kit up for it with the speckit-xref:speckit skill and write the spec in my own words.',
682 setup: 'Set Spec Kit up in this existing project with the speckit-xref:speckit skill: run specify init, draft the constitution from the code and the README and mark the assumptions, then ask me which change to specify first.',
683 integration: "Install Spec Kit's Claude Code integration in this project (specify integration install claude), so its /speckit-* commands run here.",
684} as const
685
686async function askClaude($: $, kind: keyof typeof SETUP_ASKS): Promise<void> {
687 // A press answers what the autopilot waits on: it goes on once Claude has done it.
688 if ((await read($, autopilotA)).paused) await update($, autopilotA, a => ({ ...a, paused: null, stalls: 0 }))
689 void $.prompt.submit({ text: SETUP_ASKS[kind], asUser: true })
690}
691
692/** Approvals only the person gives: the spec, the plan, open checklists. The autopilot records checkpoints, never these. */
693const APPROVALS = new Set(['spec', 'plan', 'checklists'])
694const CHECKPOINTS = new Set(['analyze', 'converge'])
695
696/** The person approves what the autopilot waits on (spec, plan, open checklists); the run goes on. */
697async function approve($: $): Promise<string> {
698 const snap = await read($, snapshotA)
699 if (!snap) return 'Nothing to approve.'
700 const ap = await read($, autopilotA)
701 const step = nextStep(snap, await read($, ledgerA), ap.idea, flowOf(ap))
702 if (!step.approve || !APPROVALS.has(step.approve.key)) return 'Nothing waits for an approval.'
703 const { key, value } = step.approve
704 await update($, ledgerA, l => ({ ...l, approvals: { ...l.approvals, [key]: value } }))
705 await persist($)
706 if (ap.on && ap.paused) await resumeAutopilot($)
707 return key === 'checklists' ? 'Going on despite the open checklist items.' : `Approved the ${key} of ${snap.featureDir}.`
708}
709
710/** Stops the autopilot until the person speaks, says why, and asks again after five minutes. */
711async function pauseAutopilot($: $, reason: string, step?: Step): Promise<void> {
712 await update($, autopilotA, a => ({ ...a, paused: reason }))
713 $.ui.toast(`Autopilot waits for you: ${reason}`)
714 await retitle($)
715 $.clock.after(5 * 60_000, () => void remind($, reason).catch(() => undefined))
716 // At a review the native dialog answers it in place, where a person is there to answer.
717 if (step?.approve && APPROVALS.has(step.approve.key) && interactive) $.clock.after(0, () => void askApproval($, step).catch(() => undefined))
718}
719
720/** A native notification on top of the toast; switched off or without a channel, the band and the toast still say it. */
721async function notify($: $, text: string): Promise<void> {
722 try {
723 await $.ui.notify(text, { title: TITLE })
724 } catch {
725 // Notifications switched off, or no channel: the band and the toast still say it.
726 }
727}
728
729async function remind($: $, reason: string): Promise<void> {
730 const ap = await read($, autopilotA)
731 if (ap.on && ap.paused === reason) await notify($, `Autopilot waits for you: ${reason}`)
732}
733
734async function askApproval($: $, step: Step): Promise<void> {
735 const key = step.approve!.key
736 const yes = key === 'spec' ? 'Approve spec' : key === 'plan' ? 'Approve plan' : 'Proceed anyway'
737 const question = key === 'checklists' ? `${step.why} Go on implementing anyway?` : `${step.why} Approve the ${key} as the contract?`
738 let answer: string
739 try {
740 answer = await $.ui.ask(question, [yes, 'Not yet'])
741 } catch {
742 return
743 }
744 if (answer === yes) await approve($)
745 // Anything typed under "Other" is what to change: it goes to Claude as the person's words.
746 else if (answer !== 'Not yet' && answer.trim()) void $.prompt.submit({ text: answer, asUser: true })
747}
748
749async function stopAutopilot($: $, why: string): Promise<void> {
750 await update($, autopilotA, a => ({ ...a, on: false, paused: null }))
751 $.ui.toast(why)
752 await retitle($)
753 await writeBriefing($, why)
754 await notify($, why)
755}
756
757/** The git HEAD, short; empty outside git. */
758async function head($: $): Promise<string> {
759 try {
760 const ran = await $.process.run(['git', 'rev-parse', '--short', 'HEAD'], { cwd: root, timeoutMs: 3000 })
761 return ran.exitCode === 0 ? ran.stdout.trim() : ''
762 } catch {
763 return ''
764 }
765}
766
767/**
768 * Runs the project's tests and records what they prove: per requirement from JUnit when a report path is set,
769 * else every requirement with a real test passes when the whole suite does. `ran` is false when no runner started.
770 */
771async function runTests($: $): Promise<{ ran: boolean; ok: boolean; output: string }> {
772 const snap = await read($, snapshotA)
773 if (!snap?.testCommand || !snap.featureDir) return { ran: false, ok: true, output: '' }
774 let exitCode: number
775 let output: string
776 try {
777 const ran = await $.process.run(['sh', '-c', snap.testCommand], { cwd: root, timeoutMs: 600_000 })
778 exitCode = ran.exitCode
779 output = `${ran.stdout}\n${ran.stderr}`.trim().slice(-4000)
780 } catch (error) {
781 return { ran: false, ok: true, output: String(error) }
782 }
783 const at = await stamp($)
784 const commit = await head($)
785 const current = Object.fromEntries((snap.spec?.reqs ?? []).map(r => [r.id, fingerprint(r.text)]))
786 const ledger = await read($, ledgerA)
787 const proof = { featureDir: snap.featureDir, tasks: snap.tasks, reqs: snap.spec?.reqs ?? [], realTests: snap.realTests }
788 const xml = junitPath ? await readText($, junitPath) : null
789 const verified = xml
790 ? verificationFrom(parseJunit(xml), ledger.anchors, current, snap.featureDir, at, commit)
791 : verificationFromExit(exitCode, (snap.spec?.reqs ?? []).filter(r => LADDER.indexOf(levelOf(r.id, proof, ledger)) >= LADDER.indexOf('tested')).map(r => r.id), current, at, commit)
792 if (Object.keys(verified).length) {
793 await update($, ledgerA, l => ({
794 ...l,
795 verification: { ...l.verification, ...verified },
796 fingerprints: { ...l.fingerprints, ...Object.fromEntries(Object.keys(verified).filter(id => current[id]).map(id => [id, current[id]!])) },
797 }))
798 await persist($)
799 }
800 return { ran: true, ok: exitCode === 0, output }
801}
802
803/** Appends one finished step to the run log (local, never committed). */
804async function logRun($: $, entry: RunEntry): Promise<void> {
805 const snap = await read($, snapshotA)
806 if (!snap?.featureDir) return
807 const path = runLogPath(snap.featureDir)
808 const lines = ((await readText($, path)) ?? '').split('\n').filter(Boolean).slice(-499)
809 lines.push(JSON.stringify(entry))
810 await $.fs.write(at(path), lines.join('\n') + '\n').catch(() => undefined)
811}
812
813async function runLog($: $): Promise<RunEntry[]> {
814 const snap = await read($, snapshotA)
815 if (!snap?.featureDir) return []
816 return ((await readText($, runLogPath(snap.featureDir))) ?? '')
817 .split('\n')
818 .filter(Boolean)
819 .map(line => {
820 try {
821 return JSON.parse(line) as RunEntry
822 } catch {
823 return null
824 }
825 })
826 .filter((e): e is RunEntry => !!e)
827}
828
829/** What happened since the person last spoke: the "since you left" card and the night run's briefing. */
830async function sinceYouLeft($: $): Promise<string[]> {
831 const seen = await read($, seenAtA)
832 const entries = (await runLog($)).filter(e => Date.parse(e.at) >= seen)
833 if (!entries.length) return []
834 const snap = await read($, snapshotA)
835 const ap = await read($, autopilotA)
836 const files = new Set(entries.flatMap(e => e.files))
837 const minutes = Math.round(entries.reduce((n, e) => n + e.durationMs, 0) / 60_000)
838 const lines = [`${entries.length} steps in ${minutes} min: ${[...new Set(entries.map(e => e.phase))].join(', ')}; ${files.size} files changed.`]
839 if (snap) lines.push(`Tasks ${snap.tasks.filter(t => t.done).length}/${snap.tasks.length} checked${ap.lastTest ? `; tests ${ap.lastTest.ok ? 'pass' : 'fail'}` : ''}.`)
840 const ledger = await read($, ledgerA)
841 const open = ledger.decisions.filter(d => d.answer === null)
842 if (open.length) lines.push(`Waiting for you: ${open.map(d => d.question).join(' | ')}`)
843 if (ap.paused) lines.push(`Paused: ${ap.paused}`)
844 return lines
845}
846
847/** A night run leaves a briefing for the morning: `.specify/xref/local/briefing.md`. */
848async function writeBriefing($: $, why: string): Promise<void> {
849 const ap = await read($, autopilotA)
850 if (!ap.night) return
851 const lines = await sinceYouLeft($)
852 await $.fs.write(at('.specify/xref/local/briefing.md'), [`# Autopilot briefing (${await stamp($)})`, '', `Ended: ${why}`, '', ...lines.map(l => `- ${l}`), ''].join('\n')).catch(() => undefined)
853}
854
855/**
856 * One autopilot step: the next Spec Kit step is handed to the model as a prompt of its own, unless only the
857 * person can take it, the run stopped moving, or the step budget is spent.
858 */
859async function advance($: $): Promise<void> {
860 const ap = await read($, autopilotA)
861 if (!ap.on || ap.paused) return
862 await scan($)
863 const snap = await read($, snapshotA)
864 if (!snap) return
865 const ledger = await read($, ledgerA)
866 const step = nextStep(snap, ledger, ap.idea, flowOf(ap))
867 if (step.needsUser) return pauseAutopilot($, step.needsUser, step)
868 // A repair attempt is progress of its own: the stall check must not end the three repairs early.
869 const key = `${progressKey(snap, ledger)}|${ap.repairs}`
870 const stalls = key === ap.last ? ap.stalls + 1 : 0
871 if (stalls >= 3) return pauseAutopilot($, 'three steps without progress; look at what blocks it, then Resume')
872 if (ap.steps >= ap.max) return pauseAutopilot($, `the step budget (${ap.max}) is used up; /xref auto on starts a new one`)
873 if (step.phase === 'verify' && ap.lastPhase === 'verify') {
874 const report = evaluate(snap, ledger)
875 if (report.level === 'red') return pauseAutopilot($, 'the feature is built, but the drift is red; decide in the pane')
876 // Done means proven: with a test command, the suite has to pass once more before the run ends.
877 if (snap.testCommand) {
878 const tests = await runTests($)
879 if (tests.ran && !tests.ok) {
880 await update($, autopilotA, a => ({ ...a, lastTest: { ok: false, at: '', output: tests.output }, repairs: a.repairs + 1, lastPhase: 'repair' }))
881 return advance($)
882 }
883 return stopAutopilot($, `Autopilot done: every task of ${snap.featureDir} is checked, ${tests.ran ? 'the tests pass' : `the tests could not run (${snap.testCommand})`} and the drift is ${report.level}.`)
884 }
885 return stopAutopilot($, `Autopilot done: every task of ${snap.featureDir} is checked and the drift is ${report.level}. No test command is known, so "done" rests on the checkboxes.`)
886 }
887 await update($, autopilotA, a => ({ ...a, steps: a.steps + 1, last: key, stalls, lastPhase: step.phase, idea: step.phase === 'specify' ? null : a.idea, scope: step.phase === 'implement' ? snap.phase : a.scope }))
888 // Handing a step over is its checkpoint (analyze, converge); a revise takes the request it folds in off the list.
889 if (step.approve) {
890 const { key: point, value } = step.approve
891 await update($, ledgerA, l => ({ ...l, checkpoints: { ...l.checkpoints, [point]: value } }))
892 }
893 if (step.resolves) await update($, ledgerA, l => resolveIntent(l, step.resolves!))
894 await persist($)
895 // The mod maps requirements itself where no extension command does it, then moves on.
896 if (step.phase === 'map' && !snap.extensions.includes('xref')) {
897 await runMap($, mapModel)
898 return advance($)
899 }
900 const active = await read($, activeA)
901 const text = autopilotPrompt(step, ap.steps + 1, ap.max, turnContext(snap, ledger, active)) + parallelNote(snap, step)
902 await update($, autopilotA, a => ({ ...a, doneAtStep: snap.tasks.filter(t => t.done).map(t => t.id) }))
903 stepStartedAt = await $.clock.now()
904 // Nobody is at the prompt in a -p run: the Stop hook hands the step over as its re-prompt.
905 if (!interactive) {
906 pendingPrompt = text
907 return
908 }
909 void $.prompt.submit({ text })
910}
911
912const RUNNER = 'task-runner'
913
914/** With the parallel option, a phase's open [P] tasks go to subagents of their own, one per task. */
915function parallelNote(snap: Snapshot, step: Step): string {
916 if (!parallel || step.phase !== 'implement' || !snap.phase) return ''
917 const ready = snap.tasks.filter(t => !t.done && t.parallel && t.phase === snap.phase)
918 if (ready.length < 2) return ''
919 return [
920 '',
921 '',
922 `Parallel: ${ready.map(t => t.id).join(', ')} are marked [P] and touch different files. Spawn one ${PLUGIN}:${RUNNER} agent per task, all in one message (prompt: the task id and its line from tasks.md).`,
923 'When they report back, check those tasks off in tasks.md yourself, then do the rest of the phase.',
924 ].join('\n')
925}
926
927/** What a task-runner subagent is told: one task, its planned files, the xref rules, and no tasks.md. */
928const RUNNER_PROMPT = [
929 'You implement exactly one task of a GitHub Spec Kit feature; other agents implement its sibling tasks at the same time.',
930 'Read the task line you were given, the spec.md and plan.md of the active feature (specs/<feature>/), and the files the task names.',
931 `First call mcp__${PLUGIN}__focus with the task id. Change only the files the task names; if another file is unavoidable, call mcp__${PLUGIN}__link for it.`,
932 'Mark code that implements a requirement with a comment `@spec <feature>/FR-###`; tests carry the anchor of what they prove.',
933 'Do not edit tasks.md and do not commit: the main agent checks the task off. Run the tests that cover your files if you can.',
934 'Report in three lines: what you changed (files), what the tests said, and anything left open.',
935].join('\n')
936
937/**
938 * Commit per task (an option, off by default): after a step whose tests passed, the tasks it checked off are
939 * committed with their files, on a feature branch only. The mod checks the staged diff itself, since git hooks
940 * do not run under $.process.run.
941 */
942async function commitTasks($: $): Promise<string | null> {
943 if (!commitPerTask) return null
944 const snap = await read($, snapshotA)
945 const ap = await read($, autopilotA)
946 if (!snap?.featureDir || !snap.branch || /^(main|master|trunk|develop)$/.test(snap.branch)) return null
947 const fresh = snap.tasks.filter(t => t.done && !(ap.doneAtStep ?? []).includes(t.id))
948 if (!fresh.length) return null
949 const ledger = await read($, ledgerA)
950 const git = (argv: string[]) => $.process.run(['git', ...argv], { cwd: root, timeoutMs: 15_000 })
951 try {
952 // Never mix in what the person staged themselves.
953 if ((await git(['diff', '--cached', '--name-only'])).stdout.trim()) return 'commit skipped: the index already holds staged changes'
954 const files = [...new Set([...fresh.flatMap(t => ledger.tasks[t.id]?.touched ?? []), `${snap.featureDir}/tasks.md`, `${snap.featureDir}/xref.json`])]
955 const present = []
956 for (const f of files) if (await exists($, f)) present.push(f)
957 const secret = present.find(f => /(^|\/)\.env(?!\.example$)/.test(f))
958 if (secret) return `commit skipped: ${secret} looks like a secret`
959 const drift = present.filter(f => ledger.unplanned.some(u => u.file === f && !u.acknowledged))
960 if (drift.length) return `commit skipped: ${drift.join(', ')} ${drift.length === 1 ? 'is' : 'are'} outside the plan`
961 await git(['add', '--', ...present])
962 const subject = fresh.length === 1 ? `${fresh[0]!.id}: ${short(fresh[0]!.text, 60)}` : `${fresh.map(t => t.id).join(', ')} (${snap.featureDir.slice(snap.featureDir.lastIndexOf('/') + 1)})`
963 const ran = await git(['commit', '-m', subject, '-m', `Tasks checked by the speckit-xref autopilot; tests passed.`])
964 return ran.exitCode === 0 ? `committed ${fresh.map(t => t.id).join(', ')}` : `commit failed: ${ran.stderr.trim().slice(0, 200)}`
965 } catch (error) {
966 return `commit failed: ${String(error).slice(0, 200)}`
967 }
968}
969
970type TurnEnd = { answer: string; reason: string; isAborted: boolean; durationMs: number; usage?: { input_tokens?: number; output_tokens?: number } }
971
972/** After a turn: the autopilot pauses where the model or the person needs the person, and moves on otherwise. */
973async function afterTurn($: $, e: TurnEnd, files: string[]): Promise<void> {
974 const asked = await read($, askedA)
975 await update($, askedA, () => null)
976 const ap = await read($, autopilotA)
977 if (!ap.on) return
978 if (ap.lastPhase && ap.steps > 0) {
979 const snap = await read($, snapshotA)
980 const outcome = e.isAborted ? 'interrupted' : e.reason !== 'answer' ? e.reason : asked ? 'asked the person' : 'answered'
981 // The tasks this step checked off, else the one it was on.
982 const checked = snap ? snap.tasks.filter(t => t.done && !(ap.doneAtStep ?? []).includes(t.id)).map(t => t.id) : []
983 await logRun($, {
984 at: await stamp($),
985 step: ap.steps,
986 phase: ap.lastPhase,
987 task: checked.length ? checked.join(', ') : snap ? (currentTask(snap, await read($, activeA))?.id ?? null) : null,
988 files,
989 durationMs: e.durationMs || (stepStartedAt ? (await $.clock.now()) - stepStartedAt : 0),
990 tokens: (e.usage?.input_tokens ?? 0) + (e.usage?.output_tokens ?? 0),
991 outcome,
992 })
993 }
994 if (ap.paused) return
995 if (e.isAborted) return pauseAutopilot($, 'you interrupted the turn; Resume, or a message from you, goes on')
996 // An API error or a refusal is no answer: going on would hand the next step to a turn that never did this one.
997 if (e.reason !== 'answer') return pauseAutopilot($, `the turn ended with ${e.reason === 'refusal' ? 'a refusal' : 'an error'}; Resume tries again`)
998 if (asked) return pauseAutopilot($, asked)
999 if (endsWithQuestion(e.answer)) {
1000 // A question at the end is a stop only when it is a real decision; "shall I go on?" is not one.
1001 const label = await $.model.classify(e.answer.slice(-1500), [...QUESTION_LABELS]).catch(() => QUESTION_LABELS[0])
1002 if (label !== QUESTION_LABELS[1]) return pauseAutopilot($, 'the last answer asks you something')
1003 }
1004 // Code changed: the tests say whether the step holds. A failure becomes a repair step, at most three in a row.
1005 if (ap.lastPhase === 'implement' || ap.lastPhase === 'repair' || ap.lastPhase === 'converge') {
1006 const tests = await runTests($)
1007 if (tests.ran) {
1008 const when = await stamp($)
1009 await update($, autopilotA, a => ({ ...a, lastTest: { ok: tests.ok, at: when, output: tests.output }, repairs: tests.ok ? 0 : a.repairs + 1 }))
1010 if (tests.ok) {
1011 const committed = await commitTasks($)
1012 if (committed) $.ui.toast(`Autopilot: ${committed}`)
1013 }
1014 }
1015 }
1016 const snap = await read($, snapshotA)
1017 if (snap && evaluate(snap, await read($, ledgerA)).level === 'red') return pauseAutopilot($, 'the drift is red; look at the findings in the pane, then Resume')
1018 await advance($)
1019}
1020/** The chip a booked write gets in the transcript: the task and what it serves, or why it is drift. */
1021function chipFor(c: Classified, snap: Snapshot, ledger: Ledger): string | null {
1022 if (c.verdict === 'unplanned') return '▲ unplanned'
1023 if (c.verdict === 'unclear') return '· unclear'
1024 if (c.verdict === 'spec' || c.verdict === 'exempt' || c.verdict === 'untracked') return null
1025 const task = snap.tasks.find(t => t.id === c.task)
1026 if (!task) return c.verdict === 'linked' ? '● linked' : null
1027 const reqs = reqsOf(task, ledger)
1028 return `● ${task.id}${reqs.length ? ` · ${reqs.join(' ')}` : ''}${task.done ? ' · rework' : ''}`
1029}
1030
1031/** An autopilot prompt as one line in the transcript: `▶ auto 3/25 · implement · 4 of 8 tasks open; next T005.` */
1032export function compactPrompt(text: string): string | null {
1033 const m = /^\[speckit-xref autopilot · step (\d+\/\d+)\] (\w+): (.*)$/m.exec(text)
1034 return m ? `▶ auto ${m[1]} · ${m[2]} · ${m[3]}` : null
1035}
1036
1037/** What this run cost so far and how long it took: ` · $1.80 · 23m`. */
1038async function runCost($: $, ap: Autopilot): Promise<string> {
1039 if (!ap.startedAt) return ''
1040 const minutes = Math.max(0, Math.round(((await $.clock.now()) - ap.startedAt) / 60_000))
1041 const usd = Math.max(0, (await sessionCost($)) - ap.costAtStart)
1042 return ` · $${usd.toFixed(2)} · ${minutes}m`
1043}
1044
1045/** Accepts one edit outside the plan as intended. */
1046async function acceptOne($: $, file: string): Promise<void> {
1047 await update($, ledgerA, l => ({ ...l, unplanned: l.unplanned.map(u => (u.file === file ? { ...u, acknowledged: true } : u)) }))
1048 await persist($)
1049 lastLevel = 'none'
1050}
1051
1052/** Accept all asks first: a blind accept is the one way drift tracking quietly stops meaning anything. */
1053async function confirmAcceptAll($: $): Promise<void> {
1054 const open = (await read($, ledgerA)).unplanned.filter(u => !u.acknowledged).length
1055 let answer = ''
1056 try {
1057 answer = await $.ui.ask(`Accept all ${open} edits outside the plan as intended?`, ['Accept all', 'Cancel'])
1058 } catch {
1059 return
1060 }
1061 if (answer !== 'Accept all') return
1062 await update($, ledgerA, acknowledgeAll)
1063 await persist($)
1064 lastLevel = 'none'
1065}
1066
1067/** The git blob hash of each file, to tell a file Bash changed during a turn from one that was dirty before. */
1068async function hashFiles($: $, files: string[]): Promise<Record<string, string>> {
1069 if (!files.length) return {}
1070 try {
1071 const ran = await $.process.run(['git', 'hash-object', '--', ...files], { cwd: root, timeoutMs: 5000 })
1072 if (ran.exitCode !== 0) return {}
1073 const hashes = ran.stdout.split('\n').filter(Boolean)
1074 return Object.fromEntries(files.map((f, i) => [f, hashes[i] ?? '']))
1075 } catch {
1076 return {}
1077 }
1078}
1079
1080/**
1081 * Books what Bash wrote (sed, a code generator, npm): files the working tree changed during the turn that no
1082 * Write or Edit booked. Without this the live view and the CI check would disagree.
1083 */
1084async function bookShellWrites($: $): Promise<void> {
1085 const turn = await read($, turnA)
1086 const snap = await read($, snapshotA)
1087 if (!turn || !snap?.featureDir) return
1088 const booked = new Set(await read($, turnFilesA))
1089 const before = new Map(turn.dirty.map(entry => [entry.slice(0, entry.lastIndexOf(':')), entry.slice(entry.lastIndexOf(':') + 1)]))
1090 const now = (await changedFiles($)).filter(f => !booked.has(f) && !matchesAny(f, snap.rules.exempt))
1091 const hashes = await hashFiles($, now.filter(f => before.has(f)))
1092 // A file dirty before the turn counts only when its content provably changed.
1093 const fresh = now.filter(f => !before.has(f) || (hashes[f] && before.get(f) && hashes[f] !== before.get(f))).slice(0, 40)
1094 for (const rel of fresh) {
1095 const text = (await readText($, rel)) ?? ''
1096 const c = classify(rel, await read($, snapshotA) ?? snap, await read($, ledgerA), currentTask(snap, await read($, activeA))?.id ?? null, text)
1097 await afterWrite($, rel, c, text)
1098 }
1099}
1100
1101/**
1102 * The ask tool: where a person is there, the native dialog answers in the same turn and the run goes on. A question
1103 * that blocks only some stories becomes a decision in the queue, and the rest of the work goes on; otherwise the
1104 * autopilot waits once the turn ends.
1105 */
1106async function ask($: $, input: Record<string, unknown>): Promise<string> {
1107 const question = str(input.question) || 'a decision only you can make'
1108 const options = Array.isArray(input.options) ? input.options.filter((o): o is string => typeof o === 'string').slice(0, 4) : []
1109 const blocks = Array.isArray(input.blocks) ? input.blocks.filter((b): b is string => typeof b === 'string') : []
1110 const ap = await read($, autopilotA)
1111 if (interactive) {
1112 try {
1113 const answer = await $.ui.ask(question.endsWith('?') ? question : `${question}?`, options.length >= 2 ? options : ['Decide it yourself, record it as an assumption', 'Stop and wait for me'])
1114 if (answer && answer !== 'Stop and wait for me') {
1115 await recordDecision($, question, options, blocks, answer)
1116 return `The person answered: ${answer}. Go on with it${ap.on ? '; the autopilot keeps running' : ''}.`
1117 }
1118 } catch {
1119 // Dismissed, or nobody there to ask: the question waits for the end of the turn.
1120 }
1121 }
1122 if (ap.on && blocks.length) {
1123 await recordDecision($, question, options, blocks, null)
1124 return `Recorded for the person: "${question}". Leave ${blocks.join(', ')} alone and go on with the work that does not depend on it.`
1125 }
1126 await update($, askedA, () => question)
1127 return ap.on ? 'The autopilot will wait for the person. Put the question in your answer and end your turn.' : 'Noted. Put the question in your answer.'
1128}
1129
1130async function recordDecision($: $, question: string, options: string[], blocks: string[], answer: string | null): Promise<void> {
1131 const when = await stamp($)
1132 await update($, ledgerA, l => ({ ...l, decisions: [...l.decisions, { id: `D${l.decisions.length + 1}`, question, options, blocks, at: when, answer }].slice(-50) }))
1133 await persist($)
1134}
1135
1136/** The person answers a queued decision in the pane: Claude reads it on the next step. */
1137async function answerDecision($: $, id: string, answer: string): Promise<void> {
1138 await update($, ledgerA, l => ({ ...l, decisions: l.decisions.map(d => (d.id === id ? { ...d, answer } : d)) }))
1139 await persist($)
1140 const d = (await read($, ledgerA)).decisions.find(x => x.id === id)
1141 if (d) void $.prompt.submit({ text: `Decision ${d.id} (${d.question}): ${answer}. Apply it to ${d.blocks.join(', ') || 'the work it blocked'}.`, asUser: true })
1142}
1143
1144async function statusText($: $): Promise<string> {
1145 const snap = await read($, snapshotA)
1146 if (!snap) return 'Spec X-Ref has not read this project yet.'
1147 const ledger = await read($, ledgerA)
1148 const next = nextStep(snap, ledger)
1149 if (!snap.featureDir) return [next.why, next.needsUser ?? (next.command ? `Next: ${next.command}` : '')].filter(Boolean).join(' ')
1150 // The host already names the plugin in front of a command's answer.
1151 return (turnContext(snap, ledger, await read($, activeA), stepLine(next)) ?? '').replace(/^speckit-xref · /, '')
1152}
1153
1154/** The workflow state for the model: where the project stands in Spec Kit and what comes next. */
1155async function workflowStatus($: $): Promise<string> {
1156 await scan($)
1157 const snap = await read($, snapshotA)
1158 if (!snap) return 'Spec X-Ref has not read this project yet.'
1159 const ledger = await read($, ledgerA)
1160 const next = nextStep(snap, ledger)
1161 const report = evaluate(snap, ledger)
1162 const ap = await read($, autopilotA)
1163 const lines = [
1164 `Project root: ${root} (every path below is relative to it)`,
1165 `Spec Kit CLI: ${snap.tools.specify ? 'specify is installed' : snap.tools.uvx ? 'not installed; runs through uvx' : 'missing, and no uvx either'}`,
1166 `Spec Kit: ${snap.initialized ? `set up${snap.speckitVersion ? ` (${snap.speckitVersion})` : ''}` : 'not set up'}${snap.initialized ? ` · Claude Code integration ${snap.claudeIntegration ? `yes, commands as ${speckitCommand(snap, 'plan')}` : 'missing'}` : ''} · extensions: ${snap.extensions.join(', ') || 'none'}`,
1167 `Autopilot: ${ap.on ? (ap.paused ? `on, waiting: ${ap.paused}` : `on, step ${ap.steps}/${ap.max}`) : 'off (/xref auto on)'}`,
1168 ...(snap.initialized ? [] : [`Folder: ${snap.folder === 'empty' ? 'empty (only dotfiles, README, license)' : 'existing code'}`]),
1169 `Constitution: ${snap.constitution?.principles.length ? `${snap.constitution.principles.length} principles, ${snap.constitution.musts.length} MUST rules` : 'missing or still the template'}`,
1170 `Active feature: ${snap.featureDir ? `${snap.featureDir}${snap.spec ? ` (${snap.spec.title})` : ''}` : 'none'}`,
1171 ]
1172 if (snap.features.length > 1) lines.push(`All features: ${snap.features.join(', ')} (switch by writing {"feature_directory": "<dir>"} to .specify/feature.json)`)
1173 if (snap.featureDir) {
1174 lines.push(`Artifacts: spec.md ${snap.spec ? 'yes' : 'no'} · plan.md ${snap.hasPlan ? 'yes' : 'no'} · tasks.md ${snap.tasks.length ? `${report.done}/${report.tasks} done` : 'no'} · FR covered ${report.covered}/${report.total} · drift ${report.level}`)
1175 }
1176 lines.push(`Phase: ${next.phase}. ${next.why}`, `Next: ${next.command ?? next.needsUser ?? ''}`)
1177 if (next.needsUser) lines.push(`Needs the person: ${next.needsUser}`)
1178 if (next.notes.length) lines.push('Notes:', ...next.notes.map(n => `- ${n}`))
1179 return lines.join('\n')
1180}
1181
1182const str = (value: unknown) => (typeof value === 'string' ? value : '')
1183
1184/** The answer of one of the mod's own tools when its hook failed. */
1185function failed() {
1186 return { result: 'speckit-xref could not answer this call; claude --debug has the reason.' }
1187}
1188
1189/** A failed write guard: in strict mode it refuses rather than letting an unchecked edit through. */
1190function guardFailed<E, R>($: $, e: E, next: ((e: E) => R) & { readonly called: boolean }): R | { deny: string } {
1191 if (strict && !next.called) return { deny: 'speckit-xref (strict) could not check this edit against the plan; try again, or switch strict mode off.' }
1192 return next(e)
1193}
1194
1195/** What the autopilot never does unattended: these wait for the person (docs: D6, tighten-only). */
1196const DESTRUCTIVE: [RegExp, string][] = [
1197 [/\bgit\s+push\b/, 'pushing'],
1198 [/\bgit\s+reset\s+--hard\b/, 'git reset --hard'],
1199 [/\bgit\s+clean\s+-[a-z]*f/, 'git clean -f'],
1200 [/\bgit\s+(?:checkout|restore)\s+(?:--\s+)?\.(?:\s|$)/, 'discarding every change'],hooks/ledger.ts 185 lines1// The ledger's two files (docs/contract-0.4.md §1): what is committed and reviewed, what stays on this machine.
2
3import type { Ledger, Semantic, Unplanned } from '../types'
4
5export const emptyLedger = (): Ledger => ({
6 tasks: {},
7 requirements: {},
8 unplanned: [],
9 intents: [],
10 anchors: [],
11 semantic: null,
12 semanticHistory: [],
13 fingerprints: {},
14 verification: {},
15 approvals: {},
16 decisions: [],
17 checkpoints: {},
18 extra: { committed: {}, local: {} },
19})
20
21const COMMITTED_KEYS = new Set(['schema_version', 'feature', 'map', 'links', 'accepted', 'fingerprints'])
22const LOCAL_KEYS = new Set(['schema_version', 'feature', 'touched', 'unplanned', 'intents', 'semantic', 'semanticHistory', 'verification', 'approvals', 'decisions', 'anchors', 'checkpoints'])
23const MAX_HISTORY = 10
24
25/** Where the local half lives: `.specify/xref/local/<feature-slug>.json`. */
26export const localPath = (featureDir: string) => `.specify/xref/local/${featureDir.slice(featureDir.lastIndexOf('/') + 1)}.json`
27export const runLogPath = (featureDir: string) => `.specify/xref/local/${featureDir.slice(featureDir.lastIndexOf('/') + 1)}.run.jsonl`
28export const LOCAL_IGNORE = { path: '.specify/xref/.gitignore', text: 'local/\n' }
29
30type Raw = Record<string, unknown>
31const obj = (v: unknown): Raw => (v && typeof v === 'object' && !Array.isArray(v) ? (v as Raw) : {})
32const arr = <T>(v: unknown): T[] => (Array.isArray(v) ? (v as T[]) : [])
33const strs = (v: unknown) => arr<unknown>(v).filter((x): x is string => typeof x === 'string')
34const parse = (text: string | null): Raw | null => {
35 if (!text) return null
36 try {
37 return obj(JSON.parse(text))
38 } catch {
39 return null
40 }
41}
42const others = (raw: Raw, known: Set<string>) => Object.fromEntries(Object.entries(raw).filter(([k]) => !known.has(k)))
43
44/** A v1 ledger (one file, `schema_version` 1 or none) as the two v2 halves. */
45export function migrateV1(v1: Raw): { committed: Raw; local: Raw } {
46 const map: Record<string, string[]> = {}
47 const links: Record<string, string[]> = {}
48 const add = (file: string, id: string) => (links[file] = [...new Set([...(links[file] ?? []), id])].sort())
49 for (const [id, r] of Object.entries(obj(v1.requirements))) {
50 const entry = obj(r)
51 const tasks = strs(entry.tasks)
52 if (tasks.length) map[id] = [...new Set(tasks)].sort()
53 for (const file of strs(entry.files)) add(file, id)
54 }
55 const touched: Record<string, string[]> = {}
56 for (const [id, t] of Object.entries(obj(v1.tasks))) {
57 const entry = obj(t)
58 for (const file of strs(entry.linked)) add(file, id)
59 if (strs(entry.touched).length) touched[id] = strs(entry.touched)
60 }
61 const unplanned = arr<Raw>(v1.unplanned)
62 return {
63 committed: {
64 schema_version: 2,
65 feature: v1.feature ?? null,
66 map,
67 links,
68 accepted: [...new Set(unplanned.filter(u => u.acknowledged === true).map(u => String(u.file)))].sort(),
69 fingerprints: {},
70 },
71 local: {
72 schema_version: 2,
73 feature: v1.feature ?? null,
74 touched,
75 unplanned: unplanned.filter(u => u.acknowledged !== true).map(u => ({ file: u.file, at: u.at, task: u.task ?? null })),
76 intents: arr(v1.intents),
77 semantic: v1.semantic ?? null,
78 semanticHistory: [],
79 verification: {},
80 approvals: {},
81 decisions: [],
82 anchors: arr(v1.anchors),
83 },
84 }
85}
86
87/** The in-memory ledger from both files; a v1 file is migrated on the way in, the next write makes it v2. */
88export function ledgerFromParts(committedText: string | null, localText: string | null): Ledger {
89 let committed = parse(committedText) ?? {}
90 let local = parse(localText) ?? {}
91 if (committedText && committed.schema_version !== 2) {
92 const migrated = migrateV1(committed)
93 committed = migrated.committed
94 // A local file next to a v1 ledger is newer than the migration; it wins where it says something.
95 local = { ...migrated.local, ...local }
96 }
97 const ledger = emptyLedger()
98 for (const [id, tasks] of Object.entries(obj(committed.map))) {
99 if (strs(tasks).length) ledger.requirements[id] = { tasks: strs(tasks), files: [], sources: ['map'] }
100 }
101 for (const [file, ids] of Object.entries(obj(committed.links))) {
102 for (const id of strs(ids)) {
103 if (/^T-?\d{3,}$/.test(id)) {
104 const entry = (ledger.tasks[id] ??= { touched: [], linked: [] })
105 if (!entry.linked.includes(file)) entry.linked.push(file)
106 } else {
107 const entry = (ledger.requirements[id] ??= { tasks: [], files: [], sources: [] })
108 if (!entry.files.includes(file)) entry.files.push(file)
109 if (!entry.sources.includes('link')) entry.sources.push('link')
110 }
111 }
112 }
113 for (const [id, files] of Object.entries(obj(local.touched))) {
114 const entry = (ledger.tasks[id] ??= { touched: [], linked: [] })
115 entry.touched = strs(files)
116 }
117 ledger.unplanned = [
118 ...arr<Raw>(local.unplanned).map((u): Unplanned => ({ file: String(u.file), at: String(u.at ?? ''), task: typeof u.task === 'string' ? u.task : null, acknowledged: false })),
119 ...strs(committed.accepted).map((file): Unplanned => ({ file, at: '', task: null, acknowledged: true })),
120 ]
121 ledger.fingerprints = Object.fromEntries(Object.entries(obj(committed.fingerprints)).filter(([, v]) => typeof v === 'string')) as Record<string, string>
122 ledger.intents = arr(local.intents)
123 ledger.anchors = arr(local.anchors)
124 ledger.semantic = (local.semantic as Semantic | null) ?? null
125 ledger.semanticHistory = arr(local.semanticHistory)
126 ledger.verification = obj(local.verification) as Ledger['verification']
127 ledger.approvals = obj(local.approvals) as Ledger['approvals']
128 ledger.decisions = arr(local.decisions)
129 ledger.checkpoints = obj(local.checkpoints) as Ledger['checkpoints']
130 ledger.extra = { committed: others(committed, COMMITTED_KEYS), local: others(local, LOCAL_KEYS) }
131 return ledger
132}
133
134/** JSON with every object's keys sorted, so the committed file is the same bytes for the same state. */
135function sortedJson(value: unknown): unknown {
136 if (Array.isArray(value)) return value.map(sortedJson)
137 if (value && typeof value === 'object') return Object.fromEntries(Object.keys(value as Raw).sort().map(k => [k, sortedJson((value as Raw)[k])]))
138 return value
139}
140const uniqSorted = (xs: string[]) => [...new Set(xs)].sort()
141
142/** Both files' text for a ledger: the committed half deterministic and without timestamps. */
143export function ledgerToParts(ledger: Ledger, featureDir: string): { committed: string; local: string } {
144 const map: Record<string, string[]> = {}
145 const links: Record<string, string[]> = {}
146 const add = (file: string, id: string) => (links[file] = [...(links[file] ?? []), id])
147 for (const [id, r] of Object.entries(ledger.requirements)) {
148 if (r.tasks.length) map[id] = uniqSorted(r.tasks)
149 for (const file of r.files) add(file, id)
150 }
151 const touched: Record<string, string[]> = {}
152 for (const [id, t] of Object.entries(ledger.tasks)) {
153 for (const file of t.linked) add(file, id)
154 if (t.touched.length) touched[id] = t.touched
155 }
156 for (const file of Object.keys(links)) links[file] = uniqSorted(links[file]!)
157 // An accepted file that was linked since is the link's, not an acceptance.
158 const accepted = uniqSorted(ledger.unplanned.filter(u => u.acknowledged && !links[u.file]).map(u => u.file))
159 const committed = sortedJson({
160 ...ledger.extra.committed,
161 schema_version: 2,
162 feature: featureDir,
163 map,
164 links,
165 accepted,
166 fingerprints: ledger.fingerprints,
167 })
168 const local = {
169 ...ledger.extra.local,
170 schema_version: 2,
171 feature: featureDir,
172 touched,
173 unplanned: ledger.unplanned.filter(u => !u.acknowledged).map(u => ({ file: u.file, at: u.at, task: u.task })),
174 intents: ledger.intents,
175 semantic: ledger.semantic,
176 semanticHistory: ledger.semanticHistory.slice(-MAX_HISTORY),
177 verification: ledger.verification,
178 approvals: ledger.approvals,
179 decisions: ledger.decisions,
180 anchors: ledger.anchors,
181 checkpoints: ledger.checkpoints,
182 }
183 return { committed: JSON.stringify(committed, null, 2) + '\n', local: JSON.stringify(local, null, 2) + '\n' }
184}
185hooks/proof.ts 112 lines1// Proof instead of claims (docs/contract-0.4.md §4, §9): how far each requirement got, and what the tests said.
2
3import type { Anchor, Ledger, Req, Task, Verification } from '../types'
4import { fingerprint, isTestFile, localId, reqStatus } from './rules'
5
6export const LADDER = ['specified', 'planned', 'implemented', 'tested', 'passing'] as const
7export type Rung = (typeof LADDER)[number]
8
9/** What the ladder reads off the snapshot; `realTests` are the test files with real tests in them. */
10export type ProofSnap = { featureDir: string | null; tasks: Pick<Task, 'id' | 'done' | 'reqs'>[]; reqs: Pick<Req, 'id' | 'text'>[]; realTests: string[] }
11
12export const isActive = (req: Pick<Req, 'text'> & { status?: Req['status'] }) => (req.status ?? reqStatus(req.text).status) === 'active'
13
14/** The tasks that own a requirement: those naming it, and those mapped to it. */
15export function owners(id: string, snap: ProofSnap, ledger: Ledger): Pick<Task, 'id' | 'done' | 'reqs'>[] {
16 const mapped = new Set(ledger.requirements[id]?.tasks ?? [])
17 return snap.tasks.filter(t => t.reqs.includes(id) || mapped.has(t.id))
18}
19
20/** Whether a stored fingerprint says the requirement's text changed since it was last proven. */
21export function isStale(id: string, snap: ProofSnap, ledger: Ledger): boolean {
22 const stored = ledger.fingerprints[id]
23 const req = snap.reqs.find(r => r.id === id)
24 return !!stored && !!req && stored !== fingerprint(req.text)
25}
26
27export function levelOf(id: string, snap: ProofSnap, ledger: Ledger): Rung {
28 const level = rawLevel(id, snap, ledger)
29 // A changed requirement needs new proof: whatever it had, it is planned at most.
30 return isStale(id, snap, ledger) && LADDER.indexOf(level) > 1 ? 'planned' : level
31}
32
33function rawLevel(id: string, snap: ProofSnap, ledger: Ledger): Rung {
34 const own = owners(id, snap, ledger)
35 if (!own.length) return 'specified'
36 const anchored = (a: Anchor) => localId(a.id, snap.featureDir) === id
37 const linked = ledger.requirements[id]?.files ?? []
38 const touched = own.flatMap(t => ledger.tasks[t.id]?.touched ?? [])
39 const code = [...ledger.anchors.filter(anchored).map(a => a.file), ...linked, ...touched].some(f => !isTestFile(f))
40 if (!own.some(t => t.done) || !code) return 'planned'
41 const real = new Set(snap.realTests)
42 const tested = [...ledger.anchors.filter(anchored).map(a => a.file), ...linked].some(f => isTestFile(f) && real.has(f))
43 if (!tested) return 'implemented'
44 const req = snap.reqs.find(r => r.id === id)
45 const v = ledger.verification[id]
46 return v?.status === 'passing' && req && v.fingerprint === fingerprint(req.text) ? 'passing' : 'tested'
47}
48
49/** How many active functional requirements stand on each rung. */
50export function levels(snap: ProofSnap & { reqs: (Pick<Req, 'id' | 'text'> & { kind?: string; status?: Req['status'] })[] }, ledger: Ledger): Record<Rung, number> {
51 const counts = Object.fromEntries(LADDER.map(r => [r, 0])) as Record<Rung, number>
52 for (const req of snap.reqs) if ((req.kind ?? 'FR') === 'FR' && isActive(req)) counts[levelOf(req.id, snap, ledger)] += 1
53 return counts
54}
55
56// ---- Test results (§9) ----
57
58export type JunitCase = { name: string; classname: string; file: string; status: 'passed' | 'failed' | 'skipped' }
59
60const unescape = (s: string) => s.replace(/</g, '<').replace(/>/g, '>').replace(/"/g, '"').replace(/'/g, "'").replace(/&/g, '&')
61const attr = (tag: string, name: string) => {
62 const m = new RegExp(`\\s${name}\\s*=\\s*("([^"]*)"|'([^']*)')`).exec(tag)
63 return m ? unescape(m[2] ?? m[3] ?? '') : ''
64}
65
66export function parseJunit(xml: string): JunitCase[] {
67 const cases: JunitCase[] = []
68 const re = /<testcase\b([^>]*?)(\/>|>([\s\S]*?)<\/testcase>)/g
69 for (const m of xml.matchAll(re)) {
70 const tag = m[1] ?? ''
71 const body = m[3] ?? ''
72 const status = /<(failure|error)\b/.test(body) ? 'failed' : /<skipped\b/.test(body) ? 'skipped' : 'passed'
73 cases.push({ name: attr(tag, 'name'), classname: attr(tag, 'classname'), file: attr(tag, 'file'), status })
74 }
75 return cases
76}
77
78const IDS = /\b(?:(?:FR|SC)-\d{3,}|US\d+-AS\d+)\b/g
79
80export function verificationFrom(cases: JunitCase[], anchors: Anchor[], fingerprints: Record<string, string>, featureDir: string | null, at: string, commit: string): Record<string, Verification> {
81 const byId = new Map<string, { failed: boolean; passed: boolean; tests: Set<string> }>()
82 for (const c of cases) {
83 if (c.status === 'skipped') continue
84 const ids = new Set([...`${c.name} ${c.classname}`.matchAll(IDS)].map(m => m[0]))
85 for (const a of anchors) {
86 const id = a.file === c.file && c.file ? localId(a.id, featureDir) : null
87 if (id) ids.add(id)
88 }
89 const label = c.classname ? `${c.classname}.${c.name}` : c.name
90 for (const id of ids) {
91 const entry = byId.get(id) ?? { failed: false, passed: false, tests: new Set<string>() }
92 if (c.status === 'failed') entry.failed = true
93 else entry.passed = true
94 entry.tests.add(label)
95 byId.set(id, entry)
96 }
97 }
98 const out: Record<string, Verification> = {}
99 for (const [id, e] of byId) {
100 out[id] = { status: e.failed ? 'failing' : 'passing', tests: [...e.tests].sort(byCodePoint).slice(0, 10), at, commit, fingerprint: fingerprints[id] ?? '' }
101 }
102 return out
103}
104
105const byCodePoint = (a: string, b: string) => (a < b ? -1 : a > b ? 1 : 0)
106
107/** Without JUnit: a suite that exits 0 proves every requirement that has a real test; a failing one proves nothing. */
108export function verificationFromExit(exitCode: number, ids: string[], fingerprints: Record<string, string>, at: string, commit: string): Record<string, Verification> {
109 if (exitCode !== 0) return {}
110 return Object.fromEntries(ids.map(id => [id, { status: 'passing' as const, tests: ['<suite>'], at, commit, fingerprint: fingerprints[id] ?? '' }]))
111}
112hooks/rules.ts 167 lines1// Shared rules of docs/contract-0.4.md that both the mod and the extension implement: fingerprints, path globs,
2// what never counts as drift, what a test file is, which feature a branch names, a requirement's status.
3
4// ---- Fingerprints (§6) ----
5
6/** The bytes of a string in UTF-8, without TextEncoder (the mod's sandbox has none). */
7function utf8(text: string): number[] {
8 const out: number[] = []
9 for (const ch of text) {
10 const c = ch.codePointAt(0) ?? 0
11 if (c < 0x80) out.push(c)
12 else if (c < 0x800) out.push(0xc0 | (c >> 6), 0x80 | (c & 63))
13 else if (c < 0x10000) out.push(0xe0 | (c >> 12), 0x80 | ((c >> 6) & 63), 0x80 | (c & 63))
14 else out.push(0xf0 | (c >> 18), 0x80 | ((c >> 12) & 63), 0x80 | ((c >> 6) & 63), 0x80 | (c & 63))
15 }
16 return out
17}
18
19const K = [
20 0x428a2f98, 0x71374491, 0xb5c0fbcf, 0xe9b5dba5, 0x3956c25b, 0x59f111f1, 0x923f82a4, 0xab1c5ed5, 0xd807aa98, 0x12835b01, 0x243185be, 0x550c7dc3, 0x72be5d74, 0x80deb1fe, 0x9bdc06a7, 0xc19bf174,
21 0xe49b69c1, 0xefbe4786, 0x0fc19dc6, 0x240ca1cc, 0x2de92c6f, 0x4a7484aa, 0x5cb0a9dc, 0x76f988da, 0x983e5152, 0xa831c66d, 0xb00327c8, 0xbf597fc7, 0xc6e00bf3, 0xd5a79147, 0x06ca6351, 0x14292967,
22 0x27b70a85, 0x2e1b2138, 0x4d2c6dfc, 0x53380d13, 0x650a7354, 0x766a0abb, 0x81c2c92e, 0x92722c85, 0xa2bfe8a1, 0xa81a664b, 0xc24b8b70, 0xc76c51a3, 0xd192e819, 0xd6990624, 0xf40e3585, 0x106aa070,
23 0x19a4c116, 0x1e376c08, 0x2748774c, 0x34b0bcb5, 0x391c0cb3, 0x4ed8aa4a, 0x5b9cca4f, 0x682e6ff3, 0x748f82ee, 0x78a5636f, 0x84c87814, 0x8cc70208, 0x90befffa, 0xa4506ceb, 0xbef9a3f7, 0xc67178f2,
24]
25
26/** SHA-256 of a string's UTF-8, as lower-case hex. Synchronous, so pure code can fingerprint. */
27export function sha256(text: string): string {
28 const bytes = utf8(text)
29 const bitLength = bytes.length * 8
30 bytes.push(0x80)
31 while (bytes.length % 64 !== 56) bytes.push(0)
32 for (let i = 7; i >= 0; i--) bytes.push(i >= 4 ? 0 : (bitLength >>> (i * 8)) & 0xff)
33 const h = [0x6a09e667, 0xbb67ae85, 0x3c6ef372, 0xa54ff53a, 0x510e527f, 0x9b05688c, 0x1f83d9ab, 0x5be0cd19]
34 const w = new Array<number>(64)
35 const rotr = (x: number, n: number) => (x >>> n) | (x << (32 - n))
36 for (let off = 0; off < bytes.length; off += 64) {
37 for (let i = 0; i < 16; i++) w[i] = (bytes[off + i * 4]! << 24) | (bytes[off + i * 4 + 1]! << 16) | (bytes[off + i * 4 + 2]! << 8) | bytes[off + i * 4 + 3]!
38 for (let i = 16; i < 64; i++) {
39 const s0 = rotr(w[i - 15]!, 7) ^ rotr(w[i - 15]!, 18) ^ (w[i - 15]! >>> 3)
40 const s1 = rotr(w[i - 2]!, 17) ^ rotr(w[i - 2]!, 19) ^ (w[i - 2]! >>> 10)
41 w[i] = (w[i - 16]! + s0 + w[i - 7]! + s1) | 0
42 }
43 let [a, b, c, d, e, f, g, hh] = h as [number, number, number, number, number, number, number, number]
44 for (let i = 0; i < 64; i++) {
45 const t1 = (hh + (rotr(e, 6) ^ rotr(e, 11) ^ rotr(e, 25)) + ((e & f) ^ (~e & g)) + K[i]! + w[i]!) | 0
46 const t2 = ((rotr(a, 2) ^ rotr(a, 13) ^ rotr(a, 22)) + ((a & b) ^ (a & c) ^ (b & c))) | 0
47 hh = g
48 g = f
49 f = e
50 e = (d + t1) | 0
51 d = c
52 c = b
53 b = a
54 a = (t1 + t2) | 0
55 }
56 ;[a, b, c, d, e, f, g, hh].forEach((v, i) => (h[i] = (h[i]! + v) | 0))
57 }
58 return h.map(v => (v >>> 0).toString(16).padStart(8, '0')).join('')
59}
60
61export function normalizeText(text: string): string {
62 const nfkc = typeof text.normalize === 'function' ? text.normalize('NFKC') : text
63 return nfkc.toLowerCase().replace(/\s+/g, ' ').trim().replace(/[.;:!]+$/, '').trim()
64}
65
66export const fingerprint = (text: string) => sha256(normalizeText(text))
67
68/** What the spec gate keys on: the person's words, every requirement and every scenario. */
69export function specFingerprint(spec: { input: string | null; reqs: { text: string }[]; stories: { scenarios?: { text: string }[] }[] }): string {
70 const parts = [spec.input ?? '', ...spec.reqs.map(r => r.text), ...spec.stories.flatMap(s => (s.scenarios ?? []).map(x => x.text))]
71 return fingerprint(parts.join('\n'))
72}
73
74// ---- Path rules (§2) ----
75
76export const DEFAULT_EXEMPT = [
77 'package-lock.json', 'yarn.lock', 'pnpm-lock.yaml', 'bun.lock', 'bun.lockb', 'Cargo.lock', 'poetry.lock', 'uv.lock', 'Gemfile.lock', 'go.sum', 'composer.lock', 'Podfile.lock', 'pubspec.lock',
78 'node_modules/', '/dist/', '/build/', '/coverage/', '/.next/', '/out/', '/target/', '__snapshots__/', '*.min.js', '*.map', '*.generated.*', '*.g.dart',
79]
80export const DEFAULT_UNCLEAR = [
81 'package.json', 'pyproject.toml', 'Cargo.toml', 'go.mod', 'pubspec.yaml', 'Gemfile', 'requirements*.txt', '*.config.*', 'tsconfig*.json', '.eslintrc*', '.prettierrc*', 'Dockerfile', 'docker-compose*.yml',
82 '/.github/workflows/', '.env.example', '*.md',
83]
84const TEST_GLOBS = ['*.test.*', '*.spec.*', 'test_*.py', '*_test.py', '*_test.go', '*_test.dart', '/tests/', '/test/', '__tests__/']
85
86const SPECIAL = /[\\^$.|+(){}[\]]/
87
88export function globToRegExp(pattern: string): RegExp {
89 let p = pattern.trim().replace(/^\.\//, '')
90 const folder = p.endsWith('/')
91 if (folder) p = p.slice(0, -1)
92 if (p.startsWith('/')) p = p.slice(1)
93 else if (!p.includes('/')) p = `**/${p}`
94 if (folder) p += '/**'
95 let re = ''
96 for (let i = 0; i < p.length; ) {
97 if (p.startsWith('**/', i)) {
98 re += '(?:.*/)?'
99 i += 3
100 } else if (p.startsWith('/**', i) && i + 3 === p.length) {
101 re += '(?:/.*)?'
102 i += 3
103 } else if (p.startsWith('**', i)) {
104 re += '.*'
105 i += 2
106 } else {
107 const ch = p[i]!
108 re += ch === '*' ? '[^/]*' : ch === '?' ? '[^/]' : SPECIAL.test(ch) ? `\\${ch}` : ch
109 i += 1
110 }
111 }
112 return new RegExp(`^${re}$`)
113}
114
115export type Rules = { exempt: string[]; unclear: string[] }
116
117export function parseIgnore(text: string): Rules {
118 const rules: Rules = { exempt: [], unclear: [] }
119 for (const raw of text.split('\n')) {
120 const line = raw.trim()
121 if (!line || line.startsWith('#')) continue
122 const unclear = /^unclear:\s*(.+)$/.exec(line)
123 if (unclear) rules.unclear.push(unclear[1]!.trim())
124 else rules.exempt.push(line)
125 }
126 return rules
127}
128
129export function rulesFrom(ignoreText: string | null): Rules {
130 const own = ignoreText ? parseIgnore(ignoreText) : { exempt: [], unclear: [] }
131 return { exempt: [...DEFAULT_EXEMPT, ...own.exempt], unclear: [...DEFAULT_UNCLEAR, ...own.unclear] }
132}
133
134const compiled = new Map<string, RegExp>()
135export function matchesAny(rel: string, patterns: readonly string[]): boolean {
136 return patterns.some(p => {
137 let re = compiled.get(p)
138 if (!re) compiled.set(p, (re = globToRegExp(p)))
139 return re.test(rel)
140 })
141}
142
143export const isTestFile = (rel: string) => matchesAny(rel, TEST_GLOBS)
144export const hasRealTests = (text: string) => /\b(?:it|test|describe)\s*\(|\bdef test_|\bfunc Test[A-Z_]|\btestWidgets\s*\(/.test(text)
145
146// ---- Feature from the branch (§7) ----
147
148export function featureFromBranch(branch: string, dirs: readonly string[]): string | null {
149 const number = /(?:^|\/)(\d{3,})-/.exec(branch)?.[1]
150 if (!number) return null
151 const hits = dirs.filter(d => d.slice(d.lastIndexOf('/') + 1).startsWith(`${number}-`))
152 return hits.length === 1 ? hits[0]! : null
153}
154
155// ---- Requirement status (§11): in speckit.ts, which the parity script loads without the rest ----
156
157export { reqStatus } from './speckit'
158export type { ReqStatus } from './speckit'
159
160/** The bare id of an anchor or link (`001-auth/FR-003` → `FR-003`) when it belongs to `featureDir`, else null. */
161export function localId(id: string, featureDir: string | null): string | null {
162 const slash = id.lastIndexOf('/')
163 if (slash === -1) return id
164 const feature = id.slice(0, slash)
165 return featureDir && featureDir.slice(featureDir.lastIndexOf('/') + 1) === feature ? id.slice(slash + 1) : null
166}
167hooks/speckit.ts 170 lines1// Pure readers for GitHub Spec Kit artifacts. No IO: register.tsx reads the files and hands the text in.
2
3import type { Constitution, Scenario, Spec, Story, Req, Task } from '../types'
4
5/** A requirement's status (docs/contract-0.4.md §11): `SUPERSEDED by FR-006`, `RETIRED`, else active. */
6export type ReqStatus = { status: 'active' | 'superseded' | 'retired'; supersededBy: string | null }
7
8export function reqStatus(text: string): ReqStatus {
9 const by = /SUPERSEDED by ((?:FR|SC)-\d{3,})/.exec(text)
10 if (by) return { status: 'superseded', supersededBy: by[1]! }
11 if (/RETIRED/.test(text)) return { status: 'retired', supersededBy: null }
12 return { status: 'active', supersededBy: null }
13}
14
15const clip = (text: string, max: number) => (text.length > max ? text.slice(0, max - 1) + '…' : text)
16const unbold = (text: string) => text.replace(/\*\*/g, '').replace(/`/g, '').trim()
17
18/** The bullet lines under the first `##`/`###` heading whose text matches `heading`, up to the next heading. */
19function bulletsUnder(markdown: string, heading: RegExp, max: number): string[] {
20 const out: string[] = []
21 let inside = false
22 for (const line of markdown.split('\n')) {
23 const h = /^#{2,4}\s+(.*)$/.exec(line)
24 if (h) {
25 if (inside) break
26 inside = heading.test(h[1] ?? '')
27 continue
28 }
29 if (!inside) continue
30 const bullet = /^\s*[-*]\s+(.*)$/.exec(line)
31 if (bullet && bullet[1] && !/^\[.*\]$/.test(bullet[1].trim())) out.push(clip(unbold(bullet[1]), 200))
32 if (out.length >= max) break
33 }
34 return out
35}
36
37export function parseSpec(markdown: string): Spec {
38 const titleLine = /^#\s+(?:Feature Specification:\s*)?(.+)$/m.exec(markdown)
39 const inputLine = /^\*\*Input\*\*:\s*(?:User description:\s*)?(.+)$/m.exec(markdown)
40 const input = inputLine?.[1] ? inputLine[1].trim().replace(/^"(.*)"$/, '$1').trim() : null
41
42 const stories: Story[] = []
43 let story: Story | null = null
44 for (const line of markdown.split('\n')) {
45 const m = /^###\s+User Story\s+(\d+)\s*[-–—:]\s*(.+?)\s*(?:\(Priority:\s*(P\d+)\))?\s*(?:🎯.*)?$/.exec(line)
46 if (m) {
47 story = { id: `US${m[1]}`, title: unbold(m[2] ?? ''), priority: m[3] ?? null, scenarios: [] }
48 stories.push(story)
49 continue
50 }
51 if (/^#{1,3}\s/.test(line)) {
52 story = null
53 continue
54 }
55 // A numbered Given/When/Then line is scenario USn-AS<k>: the id an anchor or a test can name.
56 const scenario = story ? /^\s*(\d+)\.\s+(\*\*Given\*\*.*)$/.exec(line) : null
57 if (story && scenario) story.scenarios.push({ id: `${story.id}-AS${Number(scenario[1])}`, text: clip(unbold(scenario[2] ?? ''), 300) } satisfies Scenario)
58 }
59
60 const reqs: Req[] = []
61 const seen = new Set<string>()
62 for (const m of markdown.matchAll(/^\s*[-*]\s*\*\*((FR|SC)-\d{3,})\*\*\s*:?\s*(.*)$/gm)) {
63 const id = m[1] ?? ''
64 if (seen.has(id)) continue
65 seen.add(id)
66 const text = m[3] ?? ''
67 const clipped = clip(unbold(text), 300)
68 reqs.push({ id, kind: m[2] as 'FR' | 'SC', text: clipped, needsClarification: /NEEDS CLARIFICATION/i.test(text), ...reqStatus(unbold(text)) })
69 }
70
71 return {
72 title: titleLine?.[1]?.trim() ?? 'Untitled feature',
73 input: input && !/^\$ARGUMENTS$/.test(input) ? clip(input, 600) : null,
74 stories,
75 reqs,
76 outOfScope: bulletsUnder(markdown, /out of scope|non-goals?|not in scope/i, 8),
77 assumptions: bulletsUnder(markdown, /^assumptions/i, 6),
78 }
79}
80
81const PATH_EXT = /\.[a-z0-9]{1,6}$/i
82const SPEC_DOCS = /^(spec|plan|tasks|research|data-model|quickstart|constitution|checklist)\.md$/i
83
84/** File and directory paths named in a task's description: backticked, or bare tokens that look like paths. */
85export function extractPaths(text: string): string[] {
86 const found = new Set<string>()
87 const keep = (raw: string) => {
88 const path = raw.trim().replace(/^\.\//, '').replace(/[.,;:)]+$/, '')
89 if (!path || /^https?:/i.test(path) || path.includes(' ') || path.length > 200) return
90 // "as plan.md says" names Spec Kit's own documents, not a file the task writes.
91 if (SPEC_DOCS.test(path)) return
92 if (path.includes('/') || PATH_EXT.test(path)) found.add(path)
93 }
94 for (const m of text.matchAll(/`([^`]+)`/g)) keep(m[1] ?? '')
95 const bare = text.replace(/`[^`]*`/g, ' ')
96 for (const m of bare.matchAll(/(?:^|[\s("'])((?:\.\/)?(?:[\w@.-]+\/)+[\w@.\[\]-]*|[\w-]+\.[a-z][a-z0-9]{0,5})(?=[\s,;:)"']|\.(?:\s|$)|$)/gi)) {
97 const token = m[1] ?? ''
98 // "e.g." or a version like "v1.2" is no file; a path names a folder or a known-looking extension.
99 if (/^(e\.g|i\.e|etc|vs)\.?$/i.test(token) || /^v?\d+(\.\d+)+$/.test(token)) continue
100 keep(token)
101 }
102 return [...found]
103}
104
105export function parseTasks(markdown: string): Task[] {
106 const tasks: Task[] = []
107 let phase = ''
108 for (const line of markdown.split('\n')) {
109 const heading = /^##\s+(.+)$/.exec(line)
110 if (heading) {
111 phase = unbold(heading[1] ?? '')
112 continue
113 }
114 const m = /^\s*[-*]\s*\[( |x|X)\]\s*(?:\*\*)?(T\d{3,})(?:\*\*)?\s*(.*)$/.exec(line)
115 if (!m) continue
116 let rest = m[3] ?? ''
117 const parallel = /\[P\]/.test(rest)
118 const story = /\[(US\d+)\]/.exec(rest)?.[1] ?? null
119 rest = rest.replace(/\[(P|US\d+)\]\s*/g, '').trim()
120 tasks.push({
121 id: m[2] ?? '',
122 done: m[1] !== ' ',
123 parallel,
124 story,
125 text: clip(rest, 240),
126 paths: extractPaths(rest),
127 reqs: [...new Set(rest.match(/\b(?:FR|SC)-\d{3,}\b/g) ?? [])],
128 phase,
129 })
130 }
131 return tasks
132}
133
134export function parseConstitution(markdown: string): Constitution {
135 const principles: string[] = []
136 const musts: string[] = []
137 for (const line of markdown.split('\n')) {
138 const h = /^###\s+(.+)$/.exec(line)
139 // A template never filled in still reads [PRINCIPLE_1_NAME]; it says nothing yet.
140 if (h && h[1] && !/\[[A-Z0-9_]+\]/.test(h[1])) principles.push(clip(unbold(h[1]), 80))
141 if (/\bMUST\b/.test(line) && !/\[[A-Z0-9_]+\]/.test(line) && !/^#/.test(line)) {
142 const text = unbold(line.replace(/^\s*[-*]\s*/, ''))
143 if (text) musts.push(clip(text, 200))
144 }
145 }
146 return { principles: principles.slice(0, 12), musts: musts.slice(0, 12) }
147}
148
149/** `.specify/feature.json` names the active feature's directory. */
150export function parseFeatureJson(text: string): string | null {
151 try {
152 const value = JSON.parse(text) as { feature_directory?: unknown }
153 return typeof value.feature_directory === 'string' && value.feature_directory ? value.feature_directory : null
154 } catch {
155 return null
156 }
157}
158
159/** The separator `.specify/integration.json` records for the active integration: `-` for skills (`/speckit-plan`), `.` for commands. */
160export function invokeSeparator(text: string): string | null {
161 try {
162 const value = JSON.parse(text) as { integration?: unknown; integration_settings?: Record<string, { invoke_separator?: unknown }> }
163 const name = typeof value.integration === 'string' ? value.integration : null
164 const separator = name ? value.integration_settings?.[name]?.invoke_separator : undefined
165 return typeof separator === 'string' ? separator : null
166 } catch {
167 return null
168 }
169}
170hooks/workflow.ts 302 lines1// Where a project stands in Spec Kit's workflow, the step that comes next, and the autopilot's rules. Pure: register.tsx hands in the snapshot.
2
3import type { Autopilot, Ledger, Snapshot, Task } from '../types'
4import { fingerprint, specFingerprint } from './rules'
5import { evaluate, speckitCommand } from './xref'
6
7export type Phase =
8 | 'setup'
9 | 'integration'
10 | 'constitution'
11 | 'specify'
12 | 'clarify'
13 | 'review'
14 | 'plan'
15 | 'tasks'
16 | 'revise'
17 | 'map'
18 | 'analyze'
19 | 'checklists'
20 | 'implement'
21 | 'repair'
22 | 'converge'
23 | 'verify'
24export type Step = {
25 phase: Phase
26 command: string | null
27 why: string
28 notes: string[]
29 /** Set when only the person can take this step: what they have to do. */
30 needsUser: string | null
31 /** What the step hands Claude instead of "run <command>", for steps that are no single command. */
32 prompt?: string
33 /**
34 * What the step settles: an approval only the person gives (`spec`, `plan`, `checklists`), or a checkpoint the
35 * autopilot records when it hands the step over (`analyze`, `converge`); the value is the fingerprint it holds for.
36 */
37 approve?: { key: string; value: string }
38 /** The request a revise step folds into the spec: handing the step over takes it off the list. */
39 resolves?: string
40}
41
42/** What the workflow reads beyond the snapshot: the review option, the persistence model, the last test run. */
43export type Flow = {
44 review: 'spec' | 'spec+plan' | 'none'
45 lastTest?: Autopilot['lastTest']
46 repairs?: number
47}
48
49const INIT = 'init --here --force --non-interactive --integration claude'
50export const MAX_REPAIRS = 3
51
52/** How Spec Kit's CLI runs here: installed, or through uvx without installing. */
53export const specifyCli = (snap: Snapshot) => (snap.tools.specify ? 'specify' : `uvx --from 'specify-cli>=1.1,<2' specify`)
54
55/** What a spec change sets off (docs/research.de.md §5), as the constitution names it: `Persistence model: living`. */
56export type Persistence = 'flow-back' | 'flow-forward' | 'living'
57export function persistenceModel(constitutionMd: string | null): Persistence {
58 const m = /persistence(?:\s+model)?\s*[:=-]\s*\**\s*(flow[- ]back|flow[- ]forward|living(?:\s+spec)?)/i.exec(constitutionMd ?? '')
59 const value = m?.[1]?.toLowerCase().replace(' ', '-') ?? 'flow-back'
60 return value.startsWith('living') ? 'living' : value === 'flow-forward' ? 'flow-forward' : 'flow-back'
61}
62
63/** The test command plan.md's `**Testing**:` line implies: a backticked command as written, else the framework's usual one. */
64export function testCommandFrom(planMd: string | null): string | null {
65 const line = /^\*\*Testing\*\*:\s*(.+)$/m.exec(planMd ?? '')?.[1]?.trim()
66 if (!line || /NEEDS CLARIFICATION|^\[/.test(line)) return null
67 const ticked = /`([^`]+)`/.exec(line)?.[1]
68 if (ticked) return ticked
69 const known: [RegExp, string][] = [
70 [/vitest/i, 'npx vitest run'],
71 [/jest/i, 'npx jest'],
72 [/playwright/i, 'npx playwright test'],
73 [/pytest/i, 'pytest'],
74 [/cargo/i, 'cargo test'],
75 [/\bgo test|\bgo\b/i, 'go test ./...'],
76 [/flutter/i, 'flutter test'],
77 [/dart/i, 'dart test'],
78 [/rspec/i, 'bundle exec rspec'],
79 [/npm test|node:test|mocha/i, 'npm test'],
80 ]
81 return known.find(([re]) => re.test(line))?.[1] ?? null
82}
83
84/** The first `##` phase of tasks.md that still has an open task: what one implement step covers. */
85export const openPhase = (tasks: Task[]) => tasks.find(t => !t.done)?.phase || null
86
87/** A fingerprint of the task list that ignores the checkboxes: converge changes it, implement does not. */
88export const tasksFingerprint = (tasks: Task[]) => fingerprint(tasks.map(t => `${t.id} ${t.text}`).join('\n'))
89
90/** The open checklist items of a feature's `checklists/*.md`, by file. */
91export function openChecklistItems(files: Record<string, string>): { file: string; open: number }[] {
92 return Object.entries(files)
93 .map(([file, text]) => ({ file, open: (text.match(/^\s*[-*]\s*\[ \]/gm) ?? []).length }))
94 .filter(f => f.open > 0)
95}
96
97export function nextStep(snap: Snapshot, ledger: Ledger, idea: string | null = null, flow: Flow = { review: 'spec' }): Step {
98 const notes: string[] = []
99 const xref = snap.extensions.includes('xref')
100 const has = (name: string) => (snap.commands ?? []).includes(name)
101 const step = (phase: Phase, command: string | null, why: string, needsUser: string | null = null, extra: Partial<Step> = {}): Step => ({ phase, command, why, notes, needsUser, ...extra })
102
103 // Setting Spec Kit up writes into the repository: that is always the person's call, never the autopilot's.
104 if (!snap.initialized) {
105 if (!snap.tools.specify && !snap.tools.uvx) {
106 return step('setup', null, 'Spec Kit is not set up, and neither the specify CLI nor uvx is on this machine.', 'Install uv (https://docs.astral.sh/uv/) or the specify CLI (pipx install specify-cli).')
107 }
108 return snap.folder === 'empty'
109 ? step('setup', `${specifyCli(snap)} ${INIT}`, 'This folder is empty: a new app, project or problem can start here with Spec Kit.', 'Tell me what you want to build (an app, a project, a problem): I set Spec Kit up for it and write the spec in your words.')
110 : step('setup', `${specifyCli(snap)} ${INIT}`, 'This folder has code, but no Spec Kit yet (no .specify/).', 'Say "set up Spec Kit here" to bring this project under Spec Kit, or press Set up Spec Kit in the pane.')
111 }
112 if (!snap.claudeIntegration) {
113 return step('integration', `${specifyCli(snap)} integration install claude`, "Spec Kit is set up, but not for Claude Code: its /speckit-* skills are missing.", 'Say "add the Claude integration", or press Add Claude integration in the pane.')
114 }
115 if (!snap.constitution || snap.constitution.principles.length === 0) {
116 return step('constitution', speckitCommand(snap, 'constitution'), 'The constitution is missing or still the template.')
117 }
118 if (!snap.featureDir || !snap.spec) {
119 return idea
120 ? step('specify', `${speckitCommand(snap, 'specify')} ${idea}`, 'No feature yet; the person has said what to build.')
121 : step('specify', speckitCommand(snap, 'specify'), 'No feature yet: describe what to build.', 'Describe the feature you want built; your own words become the spec.')
122 }
123
124 const spec = snap.spec
125 const unclear = spec.reqs.filter(r => r.needsClarification).map(r => r.id)
126 const specFp = specFingerprint(spec)
127 // The one review: where the idea becomes the contract. A project planned before 0.4 counts as reviewed until its spec changes.
128 const specApproved = ledger.approvals.spec ? ledger.approvals.spec === specFp : snap.hasPlan
129 if (!snap.hasPlan && unclear.length) return step('clarify', speckitCommand(snap, 'clarify'), `${unclear.join(', ')} still marked NEEDS CLARIFICATION; settle them before planning.`)
130 if (flow.review !== 'none' && !specApproved) {
131 return step('review', null, `The spec of ${snap.featureDir} is written: ${spec.reqs.filter(r => r.kind === 'FR').length} requirements, ${spec.outOfScope.length} out of scope, ${spec.assumptions.length} assumptions.`,
132 'Review the spec against your words, then press Approve spec in the pane (or /xref approve); tell me what to change otherwise.', { approve: { key: 'spec', value: specFp } })
133 }
134 if (!snap.hasPlan) return step('plan', speckitCommand(snap, 'plan'), 'The spec is written; plan.md is missing.')
135 if (flow.review === 'spec+plan' && snap.planFingerprint) {
136 const planApproved = ledger.approvals.plan ? ledger.approvals.plan === snap.planFingerprint : snap.tasks.length > 0
137 if (!planApproved) {
138 return step('review', null, `The plan of ${snap.featureDir} is written.`, 'Review plan.md, then press Approve plan in the pane (or /xref approve); tell me what to change otherwise.', { approve: { key: 'plan', value: snap.planFingerprint } })
139 }
140 }
141 if (unclear.length) notes.push(`${unclear.join(', ')} still marked NEEDS CLARIFICATION (${speckitCommand(snap, 'clarify')}).`)
142 if (snap.tasks.length === 0) return step('tasks', speckitCommand(snap, 'tasks'), 'The plan is written; tasks.md is missing.')
143
144 // A request beyond the spec reaches the spec before more code does; a contradiction is always the person's.
145 const changes = ledger.semantic?.changes ?? []
146 const contradiction = changes.find(c => c.kind === 'contradicts')
147 if (contradiction) return step('revise', null, `"${contradiction.text}" contradicts the spec.`, `Your request "${contradiction.text}" contradicts the spec: keep the spec, or change it (To spec in the pane).`)
148 const extension = changes.find(c => c.kind === 'extends')
149 if (extension) {
150 if (snap.persistence === 'flow-forward') return step('revise', null, `"${extension.text}" goes beyond the spec.`, `You asked for "${extension.text}", which the spec does not cover: fold it into the spec (To spec), or make it a task (As task).`)
151 const revise = xref && has('xref-revise') ? speckitCommand(snap, 'xref.revise') : speckitCommand(snap, 'clarify')
152 return step('revise', `${revise} The user asked during implementation: "${extension.text}". Fold it into the spec, new requirements under new ids.`, `"${extension.text}" goes beyond the spec; the spec follows the request (${snap.persistence ?? 'flow-back'}).`, null, { resolves: extension.text })
153 }
154
155 const report = evaluate(snap, ledger)
156 if (report.level === 'red' || report.level === 'yellow') notes.push(`Drift ${report.level}: ${report.findings[0]?.text ?? ''}`)
157 const mapped = Object.values(ledger.requirements).some(r => r.tasks.length > 0)
158 const uncovered = report.uncovered.filter(id => !unclear.includes(id))
159 if (uncovered.length && !mapped) {
160 return step('map', xref ? speckitCommand(snap, 'xref.map') : '/xref map', `${uncovered.join(', ')} name no task yet; map requirements to tasks once.`)
161 }
162 const open = snap.tasks.filter(t => !t.done)
163 const started = snap.tasks.some(t => t.done)
164 // analyze before the first implement, and again after every change to the spec.
165 if (open.length && has('analyze') && (ledger.checkpoints.analyze ? ledger.checkpoints.analyze !== specFp : !started)) {
166 return step('analyze', speckitCommand(snap, 'analyze'), started ? 'The spec changed since the last analysis: check spec, plan and tasks against each other.' : 'Before the first task: check spec, plan and tasks against each other.', null, { approve: { key: 'analyze', value: specFp } })
167 }
168 if (flow.lastTest && !flow.lastTest.ok) {
169 if ((flow.repairs ?? 0) > MAX_REPAIRS) {
170 return step('repair', null, `The tests still fail after ${MAX_REPAIRS} repairs.`, `The tests still fail after ${MAX_REPAIRS} attempts: look at the failure in the pane and decide how to go on.`)
171 }
172 return step('repair', snap.testCommand, `The tests fail (repair ${Math.max(1, flow.repairs ?? 0)} of ${MAX_REPAIRS}).`, null, {
173 prompt: `The tests fail. Run \`${snap.testCommand}\`, find the cause and fix the code (not the tests, unless a test contradicts the spec). Do not check off new tasks in this step.\nLast output:\n${flow.lastTest.output.slice(-1500)}`,
174 })
175 }
176 if (open.length) {
177 const checklists = snap.checklists ?? []
178 const listFp = fingerprint(checklists.map(c => `${c.file}:${c.open}`).join('\n'))
179 if (checklists.length && !started && ledger.approvals.checklists !== listFp) {
180 return step('checklists', null, `${checklists.reduce((n, c) => n + c.open, 0)} checklist items are open (${checklists.map(c => c.file).join(', ')}).`, 'Checklist items are open: complete them, or press Proceed anyway in the pane (/xref approve).', { approve: { key: 'checklists', value: listFp } })
181 }
182 // One phase per step (Spec Kit's own advice for larger features): the budget and the stall check then measure real work.
183 const phase = snap.phase
184 const command = speckitCommand(snap, 'implement')
185 const ids = (phase ? open.filter(t => t.phase === phase) : open).map(t => t.id)
186 return step('implement', command, `${open.length} of ${snap.tasks.length} tasks open; next ${open[0]!.id}.`, null, {
187 prompt: phase
188 ? `Run ${command} now through the Skill tool, scoped by its argument to one phase: "Only the phase '${phase}' (${ids.join(', ')}); stop when it is done." Carry that phase through.`
189 : undefined,
190 })
191 }
192 const tasksFp = snap.tasksFingerprint || tasksFingerprint(snap.tasks)
193 if (has('converge') && ledger.checkpoints.converge !== tasksFp) {
194 return step('converge', speckitCommand(snap, 'converge'), 'Every task is checked: find the work the tasks missed.', null, { approve: { key: 'converge', value: tasksFp } })
195 }
196 return step('verify', xref ? speckitCommand(snap, 'xref.check') : '/xref check', 'Every task is checked: check the code against the spec.')
197}
198
199/** The step as one line, as the pane, /xref and the note on each prompt show it. */
200export const stepLine = (s: Step) => (s.command ? `${s.command} · ${s.why}` : s.why)
201
202export const idleAutopilot = (): Autopilot => ({
203 on: false,
204 paused: null,
205 steps: 0,
206 max: 25,
207 last: null,
208 stalls: 0,
209 lastPhase: null,
210 idea: null,
211 repairs: 0,
212 lastTest: null,
213 startedAt: null,
214 costAtStart: 0,
215 night: false,
216 scope: null,
217})
218
219/** Whether the person's next words are the idea to specify: the autopilot waits on them for exactly that. */
220export const takesIdea = (step: Step, snap: Snapshot) => step.phase === 'specify' || (step.phase === 'setup' && snap.folder === 'empty' && step.command !== null)
221
222/** What the model reads beside the person's prompt while the autopilot waits before Spec Kit is there. */
223export function setupNote(step: Step, snap: Snapshot): string {
224 const lines = [`speckit-xref · autopilot waiting · ${step.why}`]
225 if (step.phase === 'setup' && step.command && snap.folder === 'empty') {
226 lines.push(`If this message says what to build, set Spec Kit up now with the speckit-xref:speckit skill (\`${step.command}\`); the autopilot then goes on with the constitution and the spec in the person's words.`)
227 } else if (step.phase === 'setup' && step.command) {
228 lines.push(`Set Spec Kit up (speckit-xref:speckit skill, \`${step.command}\`) only if this message asks for it; otherwise leave the repository as it is.`)
229 } else if (step.phase === 'integration') {
230 lines.push(`Run \`${step.command}\` only if this message asks for it.`)
231 }
232 return lines.join('\n')
233}
234
235/** What has to change between two autopilot steps for the run to count as moving. */
236export function progressKey(snap: Snapshot, ledger: Ledger): string {
237 const report = evaluate(snap, ledger)
238 const mapped = Object.values(ledger.requirements).filter(r => r.tasks.length > 0).length
239 return [
240 snap.initialized,
241 snap.claudeIntegration,
242 snap.constitution?.principles.length ?? 0,
243 snap.featureDir ?? '-',
244 snap.spec?.reqs.length ?? 0,
245 snap.spec?.reqs.filter(r => r.needsClarification).length ?? 0,
246 snap.hasPlan,
247 `${report.done}/${report.tasks}`,
248 snap.tasksFingerprint ?? '-',
249 mapped,
250 report.level,
251 Object.values(report.levels).join('/'),
252 ledger.semantic?.at ?? '-',
253 JSON.stringify(ledger.checkpoints),
254 ].join('|')
255}
256
257/** The rules the autopilot gives the model with every step it hands over. */
258export const AUTONOMY_RULES = [
259 'Autopilot is on. Carry the step through without asking whether to continue: when you end your turn, the autopilot moves on by itself.',
260 'Decide what the person left open, from the spec, the constitution and the repository, and record those choices as assumptions in the artifact you write.',
261 'Stop only for what only the person can decide: the product idea, a [NEEDS CLARIFICATION] question, a conflict between their request and the spec, anything destructive or irreversible, credentials or payments. Then call mcp__speckit-xref__ask with the question (and the user stories it blocks, if any); if it answers with the person\'s choice, go on with it, otherwise put the question in your answer and end your turn.',
262]
263
264/** The prompt the autopilot submits for one step. */
265export function autopilotPrompt(step: Step, n: number, max: number, context: string | null): string {
266 return [
267 `[speckit-xref autopilot · step ${n}/${max}] ${step.phase}: ${step.why}`,
268 step.prompt ?? `Run ${step.command} now: invoke it through the Skill tool and carry it through.`,
269 ...step.notes.map(note => `Note: ${note}`),
270 '',
271 ...AUTONOMY_RULES,
272 ...(context ? ['', context] : []),
273 ].join('\n')
274}
275
276/** Whether a model's answer ends on a question, the one case the autopilot asks a small model about. */
277export const endsWithQuestion = (answer: string) => /\?["')\]*_\s]*$/.test(answer.trim().slice(-400))
278
279/** The labels for that question; the first means stop. A "shall I continue?" is the second. */
280export const QUESTION_LABELS = ['needs a decision only the person can make', 'asks only whether to continue, or reports progress'] as const
281
282/** The phase strip: `✓setup ✓const ✓spec ✓plan ✓tasks ▶impl 7/8 ·verify`. */
283export function phaseStrip(snap: Snapshot, step: Step): string {
284 const order: [string, Phase[]][] = [
285 ['setup', ['setup', 'integration']],
286 ['const', ['constitution']],
287 ['spec', ['specify', 'clarify', 'review', 'revise']],
288 ['plan', ['plan']],
289 ['tasks', ['tasks', 'map', 'analyze', 'checklists']],
290 ['impl', ['implement', 'repair', 'converge']],
291 ['verify', ['verify']],
292 ]
293 const at = order.findIndex(([, phases]) => phases.includes(step.phase))
294 const done = snap.tasks.filter(t => t.done).length
295 return order
296 .map(([label], i) => {
297 const extra = label === 'impl' && snap.tasks.length ? ` ${done}/${snap.tasks.length}` : ''
298 return i < at ? `✓${label}` : i === at ? `▶${label}${extra}` : `·${label}`
299 })
300 .join(' ')
301}
302hooks/xref.ts 462 lines1// Pure X-Ref and drift logic: which task a file serves, what counts as drift, what the model is told.
2
3import type { Anchor, Ledger, Semantic, Snapshot, Spec, Task } from '../types'
4import { emptyLedger } from './ledger'
5import { LADDER, isActive, isStale, levelOf } from './proof'
6import type { Rung } from './proof'
7import { fingerprint, isTestFile, localId, matchesAny } from './rules'
8import type { Rules } from './rules'
9
10export { emptyLedger, localId }
11
12/** The verdicts of docs/contract-0.4.md §3, in the order they are tried. */
13export type CoreVerdict = 'spec' | 'exempt' | 'task-linked' | 'planned' | 'linked' | 'unclear' | 'untracked' | 'unplanned'
14export type Verdict = 'in-scope' | 'other-task' | 'linked' | 'unplanned' | 'untracked' | 'spec' | 'exempt' | 'unclear'
15export type Classified = { verdict: Verdict; task: string | null }
16
17export type Level = 'none' | 'green' | 'yellow' | 'red'
18export type Finding = {
19 kind: 'unplanned' | 'dangling' | 'orphan-test' | 're-verify' | 'semantic' | 'intent'
20 level: 'yellow' | 'red'
21 text: string
22 file?: string
23 intent?: string
24 id?: string
25}
26export type Report = {
27 level: Level
28 findings: Finding[]
29 /** Active FRs at `planned` or above. */
30 covered: number
31 total: number
32 uncovered: string[]
33 done: number
34 tasks: number
35 levels: Record<Rung, number>
36 /** Spec gaps that do not raise the drift level. */
37 notes: string[]
38}
39
40/** A text cut at a word boundary, an ellipsis where it was cut. */
41export function short(text: string, max: number): string {
42 if (text.length <= max) return text
43 const cut = text.slice(0, max)
44 const space = cut.lastIndexOf(' ')
45 return (space > max * 0.5 ? cut.slice(0, space) : cut).replace(/[\s,;:(]+$/, '') + '…'
46}
47
48const MAX_UNPLANNED = 100
49const MAX_INTENTS = 50
50
51export function relPath(file: string, root: string): string {
52 let path = file.replace(/\\/g, '/')
53 const base = root.replace(/\\/g, '/').replace(/\/$/, '')
54 if (base && path.startsWith(base + '/')) path = path.slice(base.length + 1)
55 return path.replace(/^\.\//, '')
56}
57
58const basename = (path: string) => path.slice(path.lastIndexOf('/') + 1)
59
60/** Spec Kit's own files and the agent's configuration are never drift; CI workflows are code. */
61export const isSpecArtifact = (rel: string) =>
62 (/^(specs|\.specify|\.claude|\.github)\//.test(rel) && !/^\.github\/workflows\//.test(rel)) || /^(CLAUDE|AGENTS)\.md$/.test(rel)
63
64export function pathMatches(rel: string, planned: string): boolean {
65 const p = planned.replace(/^\.\//, '')
66 if (rel === p) return true
67 if (p.endsWith('/')) return rel.startsWith(p)
68 if (rel.endsWith('/' + p)) return true
69 if (!p.includes('/') && !p.includes('.')) return false
70 if (rel.startsWith(p + '/')) return true
71 return !p.includes('/') && basename(rel) === p
72}
73
74/** Whether a task plans a file: a file it names always; a folder it names only while the task is open. */
75export function plans(task: Task, rel: string): boolean {
76 return task.paths.some(p => {
77 const isFolder = p.endsWith('/') || rel.startsWith(p.replace(/^\.\//, '') + '/')
78 return pathMatches(rel, p) && (!isFolder || !task.done)
79 })
80}
81
82export const ANCHOR = /@spec\s+((?:[\w.-]+\/)?(?:(?:FR|SC)-\d{3,}|T-?\d{3,}|US\d+-AS\d+))/g
83
84export function anchorsIn(text: string, file: string): Anchor[] {
85 const out: Anchor[] = []
86 text.split('\n').forEach((line, i) => {
87 for (const m of line.matchAll(ANCHOR)) out.push({ id: m[1] ?? '', file, line: i + 1 })
88 })
89 return out
90}
91
92const DEFAULT_RULES: Rules = { exempt: [], unclear: [] }
93
94/** The shared classification (§3): first match wins. The extension's batch check reports these verdicts as they are. */
95export function classifyCore(
96 rel: string,
97 snap: Pick<Snapshot, 'tasks'>,
98 ledger: Ledger,
99 rules: Rules = DEFAULT_RULES,
100 newText = '',
101): { verdict: CoreVerdict; task: string | null } {
102 if (isSpecArtifact(rel)) return { verdict: 'spec', task: null }
103 if (matchesAny(rel, rules.exempt)) return { verdict: 'exempt', task: null }
104 for (const [id, entry] of Object.entries(ledger.tasks)) if (entry.linked.includes(rel)) return { verdict: 'task-linked', task: id }
105 // The task that plans a file comes before its anchors: it says which task the work is on.
106 const matches = snap.tasks.filter(t => plans(t, rel))
107 const match = matches.find(t => !t.done) ?? matches[0]
108 if (match) return { verdict: 'planned', task: match.id }
109 if (newText && anchorsIn(newText, rel).length > 0) return { verdict: 'linked', task: null }
110 if (Object.values(ledger.requirements).some(r => r.files.includes(rel))) return { verdict: 'linked', task: null }
111 if (ledger.anchors.some(a => a.file === rel)) return { verdict: 'linked', task: null }
112 if (matchesAny(rel, rules.unclear)) return { verdict: 'unclear', task: null }
113 if (snap.tasks.length === 0) return { verdict: 'untracked', task: null }
114 return { verdict: 'unplanned', task: null }
115}
116
117/** The live verdict: a planned file is in scope for the task in focus, or moves the focus to its own task. */
118export function classify(rel: string, snap: Snapshot, ledger: Ledger, active: string | null, newText = ''): Classified {
119 const core = classifyCore(rel, snap, ledger, snap.rules, newText)
120 switch (core.verdict) {
121 case 'task-linked':
122 return { verdict: 'linked', task: core.task }
123 case 'planned': {
124 const own = active ? snap.tasks.find(t => t.id === active && plans(t, rel)) : undefined
125 if (own) return { verdict: 'in-scope', task: own.id }
126 return { verdict: active ? 'other-task' : 'in-scope', task: core.task }
127 }
128 case 'linked':
129 case 'unplanned':
130 case 'unclear':
131 return { verdict: core.verdict, task: active }
132 default:
133 return { verdict: core.verdict, task: null }
134 }
135}
136
137/** Books one finished write into the ledger; returns the ledger it became. */
138export function recordTouch(ledger: Ledger, rel: string, c: Classified, at: string): Ledger {
139 const next: Ledger = { ...ledger, tasks: { ...ledger.tasks }, unplanned: ledger.unplanned }
140 // Exempt and unclear files are no evidence for a task and no drift either.
141 if (c.task && c.verdict !== 'unplanned' && c.verdict !== 'exempt' && c.verdict !== 'unclear') {
142 const entry = next.tasks[c.task] ?? { touched: [], linked: [] }
143 if (!entry.touched.includes(rel)) next.tasks[c.task] = { ...entry, touched: [...entry.touched, rel] }
144 }
145 if (c.verdict === 'unplanned' && !ledger.unplanned.some(u => u.file === rel && !u.acknowledged)) {
146 next.unplanned = [...ledger.unplanned, { file: rel, at, task: c.task, acknowledged: false }].slice(-MAX_UNPLANNED)
147 }
148 return next
149}
150
151/** Ties a file to a task or requirement; a file once unplanned stops counting. */
152export function link(ledger: Ledger, rel: string, id: string, snap: Snapshot, source = 'link'): { ledger: Ledger; error?: string } {
153 const bare = localId(id, snap.featureDir)
154 if (!bare) return { ledger, error: `${id} names another feature than ${snap.featureDir ?? 'the active one'}` }
155 const next: Ledger = { ...ledger, tasks: { ...ledger.tasks }, requirements: { ...ledger.requirements } }
156 if (/^T\d{3,}$/.test(bare)) {
157 if (!snap.tasks.some(t => t.id === bare)) return { ledger, error: `${bare} is no task in tasks.md` }
158 const entry = next.tasks[bare] ?? { touched: [], linked: [] }
159 if (!entry.linked.includes(rel)) next.tasks[bare] = { ...entry, linked: [...entry.linked, rel] }
160 } else if (/^(FR|SC)-\d{3,}$/.test(bare) || /^US\d+-AS\d+$/.test(bare)) {
161 const known = snap.spec?.reqs.some(r => r.id === bare) || snap.spec?.stories.some(s => s.scenarios.some(x => x.id === bare))
162 if (!known) return { ledger, error: `${bare} is no requirement or acceptance scenario in spec.md` }
163 const entry = next.requirements[bare] ?? { tasks: [], files: [], sources: [] }
164 next.requirements[bare] = {
165 ...entry,
166 files: entry.files.includes(rel) ? entry.files : [...entry.files, rel],
167 sources: entry.sources.includes(source) ? entry.sources : [...entry.sources, source],
168 }
169 } else {
170 return { ledger, error: `${id} is neither a task (T###), a requirement (FR-### / SC-###) nor a scenario (US1-AS1)` }
171 }
172 next.unplanned = ledger.unplanned.map(u => (u.file === rel ? { ...u, acknowledged: true } : u))
173 // The text a link was made against: when it changes, the link needs new proof.
174 const req = snap.spec?.reqs.find(r => r.id === bare)
175 if (req) next.fingerprints = { ...ledger.fingerprints, [bare]: fingerprint(req.text) }
176 return { ledger: next }
177}
178
179export function logIntent(ledger: Ledger, text: string, task: string | null, at: string): Ledger {
180 return { ...ledger, intents: [...ledger.intents, { at, text: text.slice(0, 500), task, status: 'logged' as const }].slice(-MAX_INTENTS) }
181}
182
183export function acknowledgeAll(ledger: Ledger): Ledger {
184 return { ...ledger, unplanned: ledger.unplanned.map(u => ({ ...u, acknowledged: true })) }
185}
186
187/** Whether a logged prompt and a change the check quoted from it are the same request. */
188const quotes = (logged: string, quoted: string) => logged.includes(quoted) || quoted.includes(logged.slice(0, 40))
189
190export function resolveIntent(ledger: Ledger, text: string): Ledger {
191 const semantic = ledger.semantic ? { ...ledger.semantic, changes: ledger.semantic.changes.filter(c => c.text !== text) } : null
192 return { ...ledger, semantic, intents: ledger.intents.map(i => (i.status !== 'resolved' && quotes(i.text, text) ? { ...i, status: 'resolved' as const } : i)) }
193}
194
195/** The task in focus while it is open; once it is checked off, the first open one. */
196export function currentTask(snap: Snapshot, active: string | null): Task | null {
197 const chosen = active ? snap.tasks.find(t => t.id === active && !t.done) : undefined
198 return chosen ?? snap.tasks.find(t => !t.done) ?? null
199}
200
201/** What the ladder reads from a snapshot. */
202const proofSnap = (snap: Snapshot) => ({ featureDir: snap.featureDir, tasks: snap.tasks, reqs: snap.spec?.reqs ?? [], realTests: snap.realTests ?? [] })
203
204/**
205 * The drift report (§8). Only deterministic findings decide unless `semantic` is on: the live mod lets the
206 * intent check count, the extension's exit code never does.
207 */
208export function evaluate(snap: Snapshot, ledger: Ledger, opts: { semantic?: boolean } = {}): Report {
209 const semanticOn = opts.semantic ?? true
210 const findings: Finding[] = []
211 const open = ledger.unplanned.filter(u => !u.acknowledged)
212 for (const u of open) findings.push({ kind: 'unplanned', level: 'yellow', file: u.file, text: `${u.file} is not planned for any task` })
213 if (open.length >= 3) findings.push({ kind: 'unplanned', level: 'red', text: `${open.length} edits outside the planned files` })
214
215 const reqs = snap.spec?.reqs ?? []
216 const byId = new Map(reqs.map(r => [r.id, r]))
217 const known = new Set([...reqs.filter(r => isActive(r)).map(r => r.id), ...snap.tasks.map(t => t.id), ...(snap.spec?.stories.flatMap(s => s.scenarios.map(x => x.id)) ?? [])])
218 for (const a of ledger.anchors) {
219 const bare = localId(a.id, snap.featureDir)
220 if (!bare || known.has(bare)) continue
221 const old = byId.get(bare)
222 const why = old ? (old.supersededBy ? `superseded by ${old.supersededBy}` : 'retired') : 'which the spec no longer has'
223 const test = isTestFile(a.file)
224 findings.push({
225 kind: test ? 'orphan-test' : 'dangling',
226 level: 'yellow',
227 file: a.file,
228 id: bare,
229 text: test ? `${a.file}:${a.line} tests ${a.id}, ${old ? `which is ${why}` : why}` : `${a.file}:${a.line} anchors ${a.id}, ${old ? `which is ${why}` : why}`,
230 })
231 }
232 const proof = proofSnap(snap)
233 for (const r of reqs) {
234 if (isActive(r) && isStale(r.id, proof, ledger)) findings.push({ kind: 're-verify', level: 'yellow', id: r.id, text: `${r.id} changed since it was last proven; it needs new proof` })
235 }
236
237 const s = ledger.semantic
238 if (s && semanticOn) {
239 if (s.verdict === 'drift' || s.score < 50) findings.push({ kind: 'semantic', level: 'red', text: `Intent check ${s.score}/100: ${s.reasons[0] ?? 'the change drifts from the spec'}` })
240 else if (s.verdict === 'minor' || s.score < 80) findings.push({ kind: 'semantic', level: 'yellow', text: `Intent check ${s.score}/100: ${s.reasons[0] ?? 'minor drift'}` })
241 for (const c of s.changes) {
242 findings.push({ kind: 'intent', level: c.kind === 'contradicts' ? 'red' : 'yellow', intent: c.text, text: `User asked: "${c.text}", ${c.kind === 'contradicts' ? 'which contradicts the spec' : 'which the spec does not cover'}` })
243 }
244 }
245
246 const frs = reqs.filter(r => r.kind === 'FR' && isActive(r))
247 const levels = Object.fromEntries(LADDER.map(r => [r, 0])) as Record<Rung, number>
248 const uncovered: string[] = []
249 for (const r of frs) {
250 const rung = levelOf(r.id, proof, ledger)
251 levels[rung] += 1
252 if (rung === 'specified') uncovered.push(r.id)
253 }
254 const notes = (snap.spec?.stories ?? []).filter(st => st.scenarios.length === 0).map(st => `${st.id} has no acceptance scenarios`)
255 const level: Level = !snap.featureDir ? 'none' : findings.some(f => f.level === 'red') ? 'red' : findings.length ? 'yellow' : 'green'
256 return {
257 level,
258 findings,
259 covered: frs.length - uncovered.length,
260 total: frs.length,
261 uncovered,
262 done: snap.tasks.filter(t => t.done).length,
263 tasks: snap.tasks.length,
264 levels,
265 notes,
266 }
267}
268
269/** The FRs a task serves: named in its text, or mapped to it in the ledger. */
270export function reqsOf(task: Task, ledger: Ledger): string[] {
271 const mapped = Object.entries(ledger.requirements).filter(([, r]) => r.tasks.includes(task.id)).map(([id]) => id)
272 return [...new Set([...task.reqs, ...mapped])]
273}
274
275/** How the project invokes a Spec Kit command: `xref.map` is `/speckit-xref-map` for skills, `/speckit.xref.map` for commands. */
276export const speckitCommand = (snap: Snapshot, name: string) =>
277 snap.commandStyle === 'skills' ? `/speckit-${name.replace(/\./g, '-')}` : `/speckit.${name}`
278
279/**
280 * The system prompt section: the active spec and the rules that keep the work on it.
281 * It changes only when the spec does, so the conversation's prompt cache survives task switches;
282 * the current task and the drift status ride each prompt instead (turnContext).
283 */
284export function composeSection(snap: Snapshot, mode: string, plugin: string): string | null {
285 if (!snap.featureDir || !snap.spec) return null
286 const spec = snap.spec
287 const lines: string[] = [
288 '# Spec Kit cross-reference (speckit-xref)',
289 'This repository follows GitHub Spec Kit. Keep the work inside the active spec and say so when it would leave it.',
290 '',
291 `Active feature: ${snap.featureDir} — ${spec.title}`,
292 ]
293 if (spec.input) lines.push(`The user's original words: "${spec.input}"`)
294 if (spec.stories.length) lines.push(`User stories: ${spec.stories.map(s => `${s.id}${s.priority ? ` (${s.priority})` : ''} ${s.title}`).join(' | ')}`)
295 if (snap.constitution?.musts.length) lines.push('Constitution, MUST rules:', ...snap.constitution.musts.slice(0, 8).map(m => `- ${m}`))
296 if (spec.outOfScope.length) lines.push('Out of scope:', ...spec.outOfScope.map(o => `- ${o}`))
297 const clar = spec.reqs.filter(r => r.needsClarification).map(r => r.id)
298 if (clar.length) lines.push(`Still marked NEEDS CLARIFICATION: ${clar.join(', ')}. Ask the user before implementing them.`)
299 lines.push(
300 '',
301 'Rules:',
302 `- Before editing a file the current task does not plan, name the task or requirement it serves; call mcp__${plugin}__focus to switch tasks or mcp__${plugin}__link to tie the file to one.`,
303 `- When the user asks for something the spec does not cover or contradicts, say so plainly and offer to update the spec (${speckitCommand(snap, 'clarify')}) before building it.`,
304 '- Check a task off in tasks.md ([X]) only once the requirements it serves are met.',
305 `- Mark code that implements a requirement with a comment at file or function level: @spec ${basename(snap.featureDir)}/FR-###.`,
306 '- Each user prompt carries a "speckit-xref" note with the current task and the drift status; follow it.',
307 )
308 if (mode === 'strict') lines.push("- Strict mode: edits outside the current task's planned or linked files are refused until you focus another task or link the file.")
309 return lines.join('\n')
310}
311
312/** The note beside a user's prompt: the current task, what it serves, and the drift status right now. */
313export function turnContext(snap: Snapshot, ledger: Ledger, active: string | null, next?: string): string | null {
314 if (!snap.featureDir || !snap.spec) return null
315 const report = evaluate(snap, ledger)
316 const task = currentTask(snap, active)
317 const lines = [`speckit-xref · ${snap.featureDir} · tasks ${report.done}/${report.tasks} · FR covered ${report.covered}/${report.total} · drift ${report.level}`]
318 if (task) {
319 lines.push(`Current task: ${task.id}${task.story ? ` [${task.story}]` : ''} ${task.text}`)
320 if (task.paths.length) lines.push(`Planned files: ${task.paths.join(', ')}`)
321 const reqs = reqsOf(task, ledger)
322 if (reqs.length) {
323 const byId = new Map(snap.spec.reqs.map(r => [r.id, r.text]))
324 lines.push('Serves:', ...reqs.map(id => `- ${id}: ${byId.get(id) ?? ''}`))
325 }
326 const next = snap.tasks.filter(t => !t.done && t.id !== task.id).slice(0, 3)
327 if (next.length) lines.push(`Next: ${next.map(t => `${t.id} ${short(t.text, 80)}`).join(' | ')}`)
328 } else if (snap.tasks.length === 0) {
329 lines.push(`No tasks.md yet; ${speckitCommand(snap, 'tasks')} derives them from the plan.`)
330 } else {
331 lines.push('Every task in tasks.md is checked.')
332 }
333 if (report.findings.length) lines.push('Drift:', ...report.findings.slice(0, 5).map(f => `- ${f.text}`))
334 if (next) lines.push(`Next Spec Kit step: ${next}`)
335 return lines.join('\n')
336}
337
338/** What the model reads beside a write's result when the write left the plan, or moved to another task. */
339export function editNote(rel: string, c: Classified, snap: Snapshot, previous: string | null, plugin: string): string | null {
340 if (c.verdict === 'unplanned') {
341 const task = currentTask(snap, previous)
342 const planned = task?.paths.length ? ` (it plans ${task.paths.join(', ')})` : ''
343 return `speckit-xref: ${rel} is not planned for ${task ? `the current task ${task.id}${planned}` : 'any task'}. If it serves the spec, call mcp__${plugin}__link with the task or requirement; if the user asked for it beyond the spec, tell them and offer ${speckitCommand(snap, 'clarify')}.`
344 }
345 if (c.verdict === 'other-task' && c.task && c.task !== previous) {
346 const task = snap.tasks.find(t => t.id === c.task)
347 if (task?.done) return `speckit-xref: ${rel} belongs to ${c.task}, which is checked off: this is rework on a finished task.`
348 return `speckit-xref: ${rel} belongs to ${c.task}${task ? ` (${task.text.slice(0, 80)})` : ''}; that is now the current task.`
349 }
350 return null
351}
352
353/** The question the drift check asks over the session's own transcript. */
354export function forkPrompt(snap: Snapshot, ledger: Ledger, active: string | null, changed: string[], diff: string, withLog = false): string {
355 const spec = snap.spec
356 // Without the transcript (a check run before the session's first turn) the logged prompts stand in for the user's words.
357 const said = withLog ? ledger.intents.filter(i => i.status !== 'resolved').slice(-10).map(i => `- ${i.text}`) : []
358 const task = currentTask(snap, active)
359 const reqs = spec?.reqs.map(r => `${r.id}: ${r.text}`).join('\n') ?? '(no spec)'
360 return [
361 'You are auditing this session for drift. Do not call any tool. Answer with ONE JSON object and nothing else.',
362 'Compare (1) what the user asked for in this conversation, in their own words, (2) the active Spec Kit spec below, and (3) the files changed in the last turn.',
363 '',
364 `Spec: ${spec?.title ?? '-'} (${snap.featureDir ?? '-'})`,
365 `Original request: ${spec?.input ?? '-'}`,
366 `Requirements:\n${reqs}`,
367 `Out of scope: ${spec?.outOfScope.join('; ') || '-'}`,
368 `Current task: ${task ? `${task.id} ${task.text}` : '-'}`,
369 ...(said.length ? ['What the user said in this session, oldest first:', ...said] : []),
370 `Changed files: ${changed.join(', ') || '-'}`,
371 `Diff (may be cut):\n${diff.slice(0, 6000) || '(no diff available)'}`,
372 '',
373 'JSON shape: {"score": 0-100 (100 = the change matches both the user\'s intent and the spec), "verdict": "aligned" | "minor" | "drift", "reasons": ["at most 3 short reasons"], "intent_changes": [{"text": "a short quote of something the user asked for during the session that the spec does not cover or contradicts", "kind": "extends" | "contradicts"}]}',
374 'intent_changes holds only requests the user made, never the spec\'s own text; code that contradicts the spec without the user asking for it belongs in reasons and the score. An empty list is the normal answer.',
375 ].join('\n')
376}
377
378/** The first JSON object in a model reply. */
379function firstJson(text: string): unknown {
380 const start = text.indexOf('{')
381 const end = text.lastIndexOf('}')
382 if (start === -1 || end <= start) return null
383 try {
384 return JSON.parse(text.slice(start, end + 1))
385 } catch {
386 return null
387 }
388}
389
390export function parseSemantic(text: string, at: string): Semantic | null {
391 const raw = firstJson(text) as Record<string, unknown> | null
392 if (!raw || typeof raw.score !== 'number') return null
393 const verdict = raw.verdict === 'aligned' || raw.verdict === 'minor' || raw.verdict === 'drift' ? raw.verdict : raw.score >= 80 ? 'aligned' : raw.score >= 50 ? 'minor' : 'drift'
394 const reasons = Array.isArray(raw.reasons) ? raw.reasons.filter((r): r is string => typeof r === 'string').slice(0, 3) : []
395 const changes = Array.isArray(raw.intent_changes)
396 ? raw.intent_changes
397 .filter((c): c is { text: string; kind?: string } => !!c && typeof (c as { text?: unknown }).text === 'string')
398 .map(c => ({ text: c.text.slice(0, 200), kind: c.kind === 'contradicts' ? ('contradicts' as const) : ('extends' as const) }))
399 .slice(0, 5)
400 : []
401 return { score: Math.max(0, Math.min(100, Math.round(raw.score))), verdict, reasons, changes, at }
402}
403
404export function applySemantic(ledger: Ledger, semantic: Semantic): Ledger {
405 // A change the user already resolved stays resolved when a later check quotes it again.
406 const resolved = ledger.intents.filter(i => i.status === 'resolved')
407 const changes = semantic.changes.filter(c => !resolved.some(i => quotes(i.text, c.text)))
408 const intents = ledger.intents.map(i => {
409 const hit = changes.find(c => i.status === 'logged' && quotes(i.text, c.text))
410 return hit ? { ...i, status: hit.kind } : i
411 })
412 return { ...ledger, intents, semantic: { ...semantic, changes } }
413}
414
415export function mapPrompt(spec: Spec, tasks: Task[]): string {
416 return [
417 'Map each functional requirement of a Spec Kit feature to the tasks that implement it. Answer with ONE JSON object and nothing else:',
418 '{"FR-001": ["T004", "T005"], ...}. Use only ids listed below; a requirement no task serves maps to [].',
419 '',
420 'Requirements:',
421 ...spec.reqs.filter(r => r.kind === 'FR').map(r => `${r.id}: ${r.text}`),
422 '',
423 'Tasks:',
424 ...tasks.map(t => `${t.id}${t.story ? ` [${t.story}]` : ''}: ${t.text}`),
425 ].join('\n')
426}
427
428export function applyMapping(ledger: Ledger, text: string, snap: Snapshot): { ledger: Ledger; mapped: number } {
429 const raw = firstJson(text) as Record<string, unknown> | null
430 if (!raw) return { ledger, mapped: 0 }
431 const reqIds = new Set(snap.spec?.reqs.map(r => r.id) ?? [])
432 const taskIds = new Set(snap.tasks.map(t => t.id))
433 const requirements = { ...ledger.requirements }
434 const fingerprints = { ...ledger.fingerprints }
435 let mapped = 0
436 for (const [id, value] of Object.entries(raw)) {
437 if (!reqIds.has(id) || !Array.isArray(value)) continue
438 const tasks = value.filter((t): t is string => typeof t === 'string' && taskIds.has(t))
439 const entry = requirements[id] ?? { tasks: [], files: [], sources: [] }
440 requirements[id] = { ...entry, tasks: [...new Set([...entry.tasks, ...tasks])], sources: entry.sources.includes('llm') ? entry.sources : [...entry.sources, 'llm'] }
441 if (tasks.length) {
442 mapped += 1
443 const req = snap.spec?.reqs.find(r => r.id === id)
444 if (req) fingerprints[id] = fingerprint(req.text)
445 }
446 }
447 return { ledger: { ...ledger, requirements, fingerprints }, mapped }
448}
449
450export const REMEDIATION_HEADING = '## Drift Remediation (speckit-xref)'
451
452/** tasks.md with one more open task under the remediation heading, numbered after the highest T###. */
453export function appendRemediation(tasksMd: string, tasks: Task[], text: string): { markdown: string; id: string } {
454 const highest = tasks.reduce((max, t) => Math.max(max, Number(t.id.slice(1)) || 0), 0)
455 const width = Math.max(3, tasks[0]?.id.length ? tasks[0].id.length - 1 : 3)
456 const id = 'T' + String(highest + 1).padStart(width, '0')
457 const line = `- [ ] ${id} [Drift] ${text.replace(/\s+/g, ' ').trim()}`
458 const base = tasksMd.replace(/\s*$/, '')
459 const markdown = base.includes(REMEDIATION_HEADING) ? `${base}\n${line}\n` : `${base}\n\n${REMEDIATION_HEADING}\n\n${line}\n`
460 return { markdown, id }
461}
462types/index.d.ts 183 lines1// The speckit-xref contract: what the mod reads out of Spec Kit, what it keeps per feature, and the session values it draws from.
2
3export type Scenario = { id: string; text: string }
4export type Story = { id: string; title: string; priority: string | null; scenarios: Scenario[] }
5export type Req = {
6 id: string
7 kind: 'FR' | 'SC'
8 text: string
9 needsClarification: boolean
10 /** Retired and superseded requirements stay in spec.md for history; they leave the ladder. */
11 status: 'active' | 'superseded' | 'retired'
12 supersededBy: string | null
13}
14export type Task = {
15 id: string
16 done: boolean
17 parallel: boolean
18 story: string | null
19 text: string
20 paths: string[]
21 reqs: string[]
22 phase: string
23}
24export type Spec = {
25 title: string
26 input: string | null
27 stories: Story[]
28 reqs: Req[]
29 outOfScope: string[]
30 assumptions: string[]
31}
32export type Constitution = { principles: string[]; musts: string[] }
33
34export type Snapshot = {
35 /** Whether Spec Kit is set up here (`.specify/` exists). */
36 initialized: boolean
37 featureDir: string | null
38 /** Every feature directory under specs/ (and .specify/specs/). */
39 features: string[]
40 hasPlan: boolean
41 /** Installed Spec Kit extensions, by id (`.specify/extensions/<id>/`). */
42 extensions: string[]
43 /** Whether Spec Kit's Claude Code integration is installed (its skills or commands are there). */
44 claudeIntegration: boolean
45 /** The Spec Kit release the project was set up with (`.specify/init-options.json`). */
46 speckitVersion: string | null
47 /** What can run Spec Kit's CLI on this machine. */
48 tools: { specify: boolean; uvx: boolean }
49 /** The project folder at a glance: nothing but dotfiles and a README (empty), or code already there (existing). */
50 folder: 'empty' | 'existing'
51 spec: Spec | null
52 tasks: Task[]
53 constitution: Constitution | null
54 /** How the project invokes Spec Kit's commands: `/speckit-clarify` (skills) or `/speckit.clarify` (commands). */
55 commandStyle: 'skills' | 'commands'
56 /** What never counts as drift (exempt) and what is listed but not drift (unclear): defaults plus `.xrefignore`. */
57 rules: { exempt: string[]; unclear: string[] }
58 /** Test files tied to the feature whose content holds real tests (not only test.todo). */
59 realTests: string[]
60 /** The fingerprint of plan.md, for the spec+plan review gate. */
61 planFingerprint: string | null
62 /** The command that runs the project's tests: the plugin option, else plan.md's `**Testing**:` line. */
63 testCommand: string | null
64 /** The git branch checked out, when there is one. */
65 branch: string | null
66 /** The open `##` phase of tasks.md the autopilot implements next, if any. */
67 phase: string | null
68 /** A fingerprint of tasks.md's task lines, to see whether converge changed them. */
69 tasksFingerprint: string
70 /** The Spec Kit commands installed for Claude Code, without prefix: `plan`, `analyze`, `xref-check`. */
71 commands: string[]
72 /** Checklists of the feature with open items. */
73 checklists: { file: string; open: number }[]
74 /** What a spec change sets off, from the constitution (flow-back by default). */
75 persistence: 'flow-back' | 'flow-forward' | 'living'
76}
77
78export type Unplanned = { file: string; at: string; task: string | null; acknowledged: boolean }
79export type IntentStatus = 'logged' | 'extends' | 'contradicts' | 'resolved'
80export type Intent = { at: string; text: string; task: string | null; status: IntentStatus }
81export type Anchor = { id: string; file: string; line: number }
82export type Semantic = {
83 score: number
84 verdict: 'aligned' | 'minor' | 'drift'
85 reasons: string[]
86 changes: { text: string; kind: 'extends' | 'contradicts' }[]
87 at: string
88 /** git HEAD and a hash of the diff the verdict looked at. */
89 head?: string
90 diffHash?: string
91}
92export type Verification = { status: 'passing' | 'failing'; tests: string[]; at: string; commit: string; fingerprint: string }
93export type Decision = { id: string; question: string; options: string[]; blocks: string[]; at: string; answer: string | null }
94
95/**
96 * What the mod keeps per feature. Persisted in two files (docs/contract-0.4.md §1): the committed
97 * `specs/<feature>/xref.json` (map, links, accepted, fingerprints) and the local `.specify/xref/local/<feature>.json`.
98 */
99export type Ledger = {
100 tasks: Record<string, { touched: string[]; linked: string[] }>
101 /** By requirement id (FR, SC) and by scenario id (USn-ASm). */
102 requirements: Record<string, { tasks: string[]; files: string[]; sources: string[] }>
103 unplanned: Unplanned[]
104 intents: Intent[]
105 anchors: Anchor[]
106 semantic: Semantic | null
107 semanticHistory: Semantic[]
108 /** A requirement's text fingerprint when it was last mapped, linked or verified. */
109 fingerprints: Record<string, string>
110 verification: Record<string, Verification>
111 /** `spec`/`plan`: the fingerprint the person approved. */
112 approvals: Record<string, string>
113 decisions: Decision[]
114 /** Workflow checkpoints: `analyze` (spec fingerprint analyzed), `converge` (tasks fingerprint converged). */
115 checkpoints: Record<string, string>
116 /** Keys of either file the mod does not know, written back as they were. */
117 extra: { committed: Record<string, unknown>; local: Record<string, unknown> }
118}
119
120/** The autopilot: it moves through Spec Kit's workflow on its own and stops only for the person. */
121export type Autopilot = {
122 on: boolean
123 /** Why it waits for the person; null while it runs. */
124 paused: string | null
125 steps: number
126 max: number
127 /** The progress key after the last step, the steps since it last changed, and the last phase. */
128 last: string | null
129 stalls: number
130 lastPhase: string | null
131 /** What the person asked for before a feature existed: the words /speckit-specify gets. */
132 idea: string | null
133 /** Failed test runs in a row; the fourth makes it the person's decision. */
134 repairs: number
135 /** The last test run: whether it passed and the tail of its output. */
136 lastTest: { ok: boolean; at: string; output: string } | null
137 /** When this run started, and the session's cost then, for the band's `$1.80 · 23m`. */
138 startedAt: number | null
139 costAtStart: number
140 /** A night run: a larger budget, a briefing when it stops. */
141 night: boolean
142 /** The phase the last implement step was scoped to. */
143 scope: string | null
144 /** The tasks checked when the last step was handed over: commit per task commits what came after. */
145 doneAtStep?: string[]
146}
147
148/** One autopilot step in the run log (`.specify/xref/local/<feature>.run.jsonl`). */
149export type RunEntry = {
150 at: string
151 step: number
152 phase: string
153 task: string | null
154 files: string[]
155 durationMs: number
156 tokens: number
157 outcome: string
158}
159
160declare module 'claude-code' {
161 interface PluginState {
162 'speckit-xref': {
163 snapshot: Snapshot | null
164 ledger: Ledger
165 active: string | null
166 checking: boolean
167 autopilot: Autopilot
168 /** The question raised through the ask tool during the running turn. */
169 asked: string | null
170 /** The files written during the running turn, by tools or found changed at its end. */
171 turnFiles: string[]
172 /** The running turn's id and start, for Stop and for finding Bash writes. */
173 turn: { id: string; startedAt: number; dirty: string[] } | null
174 /** The pane's detail view on narrow widths. */
175 details: boolean
176 /** When the person last spoke: the "since you left" card counts from here. */
177 seenAt: number
178 /** The chip under each booked write in the transcript, by tool_use_id: `T006 · FR-004` or `▲ unplanned`. */
179 chips: Record<string, string>
180 }
181 }
182}
183