SLOPSHOPPER

Spec Kit X-Ref

Keeps a GitHub Spec Kit feature and its code tied together inside Claude Code: the spec in the system prompt, every edit booked against a task, a proof ladder…

newpanebandrowsguardcommand
v0.4.0MITupdated 2026-10-09moinsen-dev/speckit-xref/mod
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · speckit-xref
│ ┃ speckit-xref ✕ › fix the failing auth test and add an audit log call │ ┃ This folder has code, but no Spec Kit yet │ ┃ (no .specify/). ⏺ Read(src/auth.ts) │ ┃ Folder: /work/app ⎿ Read 6 lines │ ┃ Say "set up Spec Kit here" to bring this ⏺ Update(src/auth.ts) │ ┃ project under Spec Kit, or press Set up Spec ⎿ Added 2 lines, removed 1 line │ ┃ Kit in the pane. ⏺ Bash(bun test) │ ┃ ⎿ 3 pass, 1 fail │ ┃ [ Set up Spec Kit here ] │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /xref │ ⎿ speckit-xref: This folder has code, but no Spec Kit yet (no .spe │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · speckit-xref
This folder has code, but no Spec Kit yet (no .specify/). Folder: /work/app Say "set up Spec Kit here" to bring this project under Spec Kit, or press Set up Spec Kit in the pane. [ Set up Spec Kit here ]
README

Spec Kit X-Ref for Claude Code

A Claude Code mod that keeps a GitHub Spec Kit feature and its code tied together while the agent works: the spec rides every request, every edit is booked against a task, and after each turn a drift check asks whether the change still matches what the user said and what the spec says.

It complements read-only viewers such as SpecKit Companion: they show the run; this one keeps the run on the spec.

What it does

How
Spec in the system promptprompt.compose adds one section: active feature, the user's original words, user stories, constitution MUST rules, out of scope, open NEEDS CLARIFICATION, and the rules of the game. It changes only when the spec does, so the prompt cache survives.
Current task on every promptprompt.submit attaches a note: current task, its planned files, the requirements it serves, the next tasks, the drift status.
Intent logEvery prompt the person types is logged (intents in xref.json): their own words, the ground truth of the idea.
Edit bookingtool.call on Edit/Write/NotebookEdit books the file against the current task (paths come from tasks.md). A file planned for another task makes that task current. A file planned for none is drift: the model reads a note beside the tool result right away. Files Bash changed during a turn (sed, a generator, npm) are found at its end and booked the same way. Lockfiles and build output never count; manifests and configs are unclear, not drift; .xrefignore adds your own patterns.
Strict modeOptional: refuses an unplanned edit and names the way forward (focus another task, link the file, or extend the spec). If the guard itself fails, it refuses too.
Proof ladderEach requirement is specified, planned, implemented, tested or passing. Tested means a test file with real tests anchors it; passing means the test run proved it for its current text. A changed requirement drops back and asks to be re-verified.
Anchors// @spec 001-feature/FR-003 comments count as coverage; one that names a requirement the spec no longer has is drift.
Intent checkAfter a turn that wrote files, one tool-less $.model.fork over the session compares the user's words, the spec and the diff. It returns a score, reasons and requests the spec lacks. With no transcript to fork yet, one completion fed by the intent log stands in.
Fix it from the paneTo spec runs /speckit-clarify with the request; As task appends a [Drift] task to tasks.md; Map FR→tasks lets a small model map requirements to tasks once and keeps the map.
Band and paneA line above the prompt (▲ xref T004 · tasks 3/8 · FR 4/5 · watch (1)) that stacks with other mods' bands; while the autopilot waits, it shows the whole question with Resume, Stop and Pane. The pane: a phase strip, Goal, Vision, Now, Todo, Status, Proof, Next, a review card at the spec gate, "Away" after an autopilot run, open decisions, Auto, Tests and Drift. Each unplanned file is a card with Link to (a task picker) and Accept this; Accept all asks first. Under 70 columns the pane is compact; d shows the rest.
TranscriptEach booked write carries a chip (● T006 · FR-003, ▲ unplanned), and the autopilot's long prompts fold to one line (▶ auto 3/25 · implement · …); ctrl+o shows them whole.
Spec Kit's own commands/speckit-implement, -plan, -tasks, -clarify, -analyze and -converge get the current task, its files, coverage and drift appended. A compaction keeps the task in focus, your open requests and the decisions.

The model gets five tools: mcp__speckit-xref__status (CLI, setup, integration, workflow phase, autopilot and the next command), mcp__speckit-xref__ask (a question for the person, answered in place where they are there; blocks names the stories it holds up), mcp__speckit-xref__focus (make a task current), __where (which tasks and requirements a file belongs to) and __link (tie a file to a task, requirement or scenario). In a repository without Spec Kit they wait behind ToolSearch, and nothing polls.

What the mod keeps per feature lives in two files:

  • specs/<feature>/xref.json, committed: which tasks serve which requirement, which files are linked to what, which unplanned files were accepted, and each requirement's fingerprint. Sorted, without timestamps, so it merges and reviews cleanly.
  • .specify/xref/local/<feature>.json, git-ignored by a .gitignore of its own: your logged prompts, the intent verdicts, test results, approvals, decisions, and the autopilot's run log (<feature>.run.jsonl). Your words stay on your machine.

A 0.3 xref.json is migrated on first load.

Install

At the prompt of a terminal session:

/plugin install speckit-xref --marketplace moinsen-dev/speckit-xref

Autopilot

/xref auto on [steps] lets the work run: after every turn the mod hands the next Spec Kit step to Claude by itself, so nobody has to answer "next is X, shall I go on?". It follows Spec Kit's whole loop: constitution, specify, clarify, your review of the spec, plan, tasks, map, analyze, implement one phase per step, converge until the tasks stop changing, verify.

It stops only for what is yours:

  • the product idea;
  • the spec gate: once the spec is written, the pane shows your words beside the requirements, the out-of-scope list and the assumptions Claude added. Approve spec (or /xref approve, or the dialog) lets it go on; the plugin option review sets spec+plan or none. An approval holds for that text: a changed spec asks again;
  • a [NEEDS CLARIFICATION] question, or open checklist items;
  • a request that contradicts the spec (one that only extends it becomes a revise step);
  • destructive or irreversible actions: while it runs, git push, reset --hard, rm -rf and writes to .env files are refused and turned into a question.

Claude asks through mcp__speckit-xref__ask. Where you are there, the question opens as a dialog and the run goes on with your answer in the same turn. A question that blocks only some stories waits in the pane's Decide row while the rest goes on. When an answer ends on a question anyway, a small model decides whether it is a real decision or just "shall I go on?".

Tests gate the progress. After each implement step the mod runs the test command (the testCommand option, else derived from plan.md's Testing: line). A failure becomes a repair step, at most three in a row; then it is your call. A passing suite proves every requirement with a real test (with junitPath, each one by the tests that name or anchor it). The run ends only when the suite passes once more.

It also stops on Esc or /xref-stop (both end the running turn), on an API error or a refusal, when the drift turns red, after three steps without progress, at the step budget (25 by default) and when the feature is done. After five minutes of waiting it sends a notification. The band shows auto ▶ 3/25 · $1.80 · 23m while it runs and the whole question when it waits; the pane tab reads Spec X-Ref ⏸. Back at the keyboard, the pane's Away card says what happened.

More, all off by default:

  • /xref auto night: a night run of up to 100 steps; it writes .specify/xref/local/briefing.md when it stops.
  • commitPerTask: on: a step whose tests pass commits the tasks it checked off, on a feature branch only; nothing you staged is mixed in, and files outside the plan or .env files stop it.
  • parallel: on: a phase's open [P] tasks go to speckit-xref:task-runner subagents, one per task.
  • Headless: SPECKIT_XREF_AUTOPILOT=on|night|<steps> claude -p "…" runs the autopilot in a -p session; each next step rides the Stop hook.

The pane's Auto row switches it:

  • off: Start autopilot (p);
  • running: Stop (p);
  • waiting: Resume (r) and Stop. A spent budget restarts on Resume.

Click the buttons, or give the pane the keyboard with ctrl+x tab (once more if the band takes it first). /xref pane looks at the folder first, opens the pane without taking the keyboard (so the next key never presses a button), and says where things stand.

The autopilot never sets Spec Kit up by itself: specify init and specify integration install write into the repository, so it waits there for you. To start every session with the autopilot on, set the plugin options autopilot: on and autopilotMaxSteps. It is off by default, because every step is a model turn.

Before Spec Kit is there

/xref pane and the status tool tell the folder apart, and the pane offers the one fitting start:

FolderPane offersWhat Claude then does
empty (only dotfiles, a README, a license)Start from an ideaasks what you want to build, sets Spec Kit up, writes the spec in your words
existing code, no .specify/Set up Spec Kit hereruns specify init, drafts the constitution from the code, asks which change comes first
Spec Kit without its Claude Code commandsAdd Claude integrationruns specify integration install claude
inside a Spec Kit projectthe project itselfthe mod looks upward for .specify/, as Spec Kit does

Pressing one of these counts as your go. Without a press or your request nothing is written. A repository without Spec Kit stays quiet otherwise, with no band and no unasked pane.

The speckit skill

/speckit-xref:speckit, or any request to set up or run Spec Kit, loads a skill that:

  • sets Spec Kit up: finds the specify CLI or runs it through uvx, says what specify init writes before it runs, then moves on to the constitution;
  • walks the workflow: constitution → specify → clarify → plan → tasks → map → implement → verify, steered by the status tool rather than by guesswork;
  • handles the day-to-day: switching features, requests beyond the spec, unplanned edits, strict mode, and the Spec Kit extension with its CI check.

Use

/xref                  status of the active feature
/xref check            run the intent check now
/xref map              map requirements to tasks with a small model
/xref ack              accept every edit outside the plan
/xref approve          approve what the autopilot waits on (spec, plan, open checklists)
/xref focus T004       make a task current
/xref auto on [n]      autopilot: work through Spec Kit by itself (night: a night run, off: stop)
/xref pane             open the pane
/xref-stop             stop the autopilot now, mid-turn

Options (/config or pluginConfigs in settings): mode (advisory | strict), driftCheck (fork | off), mapModel (default haiku), autopilot (off | on), autopilotMaxSteps (default 25), review (spec | spec+plan | none), testCommand, junitPath, commitPerTask (off | on), parallel (off | on).

The active feature is found as Spec Kit finds it, then by the git branch (001-…), else the spec written last.

Develop

claude --plugin-dir ./mod                 # load this checkout, reloading on save
claude plugin validate --strict ./mod
claude plugin test ./mod
node scripts/build-fixture.mjs            # after changing examples/demo or scripts/fixtures/cases.json

examples/demo is a small Spec Kit project (magic-link login) to try it on. In the desktop app's Code tab, name the folder in CLAUDE_CODE_PLUGIN_DIRS under env in ~/.claude/settings.json.

Tested on Claude Code 2.1.295 and Spec Kit 1.1.2. The mods API is early access and can change between releases.

Source 8 files
hooks/register.tsx 1839 lines
1import { atom, read, update } from 'claude-code'
2import type { Color, EngineInterface, Register } from 'claude-code'
3
4import type { Autopilot, Ledger, RunEntry, Snapshot } from '../types'
5import { LOCAL_IGNORE, ledgerFromParts, ledgerToParts, localPath, runLogPath } from './ledger'
6import { LADDER, levelOf, parseJunit, verificationFrom, verificationFromExit } from './proof'
7import { featureFromBranch, fingerprint, hasRealTests, isTestFile, matchesAny, rulesFrom } from './rules'
8import { invokeSeparator, parseConstitution, parseFeatureJson, parseSpec, parseTasks } from './speckit'
9import type { Flow, Step } from './workflow'
10import {
11  AUTONOMY_RULES,
12  QUESTION_LABELS,
13  autopilotPrompt,
14  endsWithQuestion,
15  idleAutopilot,
16  nextStep,
17  openChecklistItems,
18  openPhase,
19  persistenceModel,
20  progressKey,
21  MAX_REPAIRS,
22  phaseStrip,
23  setupNote,
24  stepLine,
25  takesIdea,
26  tasksFingerprint,
27  testCommandFrom,
28} from './workflow'
29import {
30  acknowledgeAll,
31  anchorsIn,
32  appendRemediation,
33  applyMapping,
34  applySemantic,
35  classify,
36  composeSection,
37  currentTask,
38  editNote,
39  emptyLedger,
40  evaluate,
41  isSpecArtifact,
42  forkPrompt,
43  link,
44  logIntent,
45  mapPrompt,
46  parseSemantic,
47  recordTouch,
48  relPath,
49  reqsOf,
50  resolveIntent,
51  short,
52  speckitCommand,
53  turnContext,
54} from './xref'
55import type { Anchor } from '../types'
56import type { Classified, Level, Report } from './xref'
57
58type $ = EngineInterface
59type Options = {
60  mode: string
61  driftCheck: string
62  mapModel: string
63  autopilot: string
64  autopilotMaxSteps: number
65  testCommand: string
66  junitPath: string
67  review: string
68  commitPerTask: string
69  parallel: string
70}
71
72const PLUGIN = 'speckit-xref'
73const PANE = 'speckit-xref'
74const TITLE = 'Spec X-Ref'
75const COMMAND = 'xref'
76const STOP_COMMAND = 'xref-stop'
77const REFRESH_MS = 4000
78const WRITE_TOOLS = new Set(['Edit', 'Write', 'NotebookEdit'])
79// Prompts the person typed, wherever they typed them; a plugin's or a peer's are not their intent.
80const PERSON = new Set(['composer', 'bridge', 'sdk'])
81const LEVEL_COLOR = { none: 'subtle', green: 'success', yellow: 'warning', red: 'error' } as const
82const RANK: Record<Level, number> = { none: 0, green: 1, yellow: 2, red: 3 }
83// What the person reads: a word and a glyph, the color only on top.
84const STATE: Record<Level, { glyph: string; word: string }> = {
85  none: { glyph: '·', word: 'no feature' },
86  green: { glyph: '●', word: 'ok' },
87  yellow: { glyph: '▲', word: 'watch' },
88  red: { glyph: '✖', word: 'off-spec' },
89}
90const ANCHOR_RG = '@spec\\s+(?:[\\w.-]+/)?(?:(?:FR|SC)-\\d{3,}|T-?\\d{3,}|US\\d+-AS\\d+)'
91const ANCHOR_GIT = '@spec[[:space:]]+([[:alnum:]_.-]+/)?((FR|SC)-[0-9]{3,}|T-?[0-9]{3,}|US[0-9]+-AS[0-9]+)'
92
93const snapshotA = atom({ plugin: 'speckit-xref', key: 'snapshot' } as const, null)
94const ledgerA = atom({ plugin: 'speckit-xref', key: 'ledger' } as const, emptyLedger())
95const activeA = atom({ plugin: 'speckit-xref', key: 'active' } as const, null)
96const checkingA = atom({ plugin: 'speckit-xref', key: 'checking' } as const, false)
97const autopilotA = atom({ plugin: 'speckit-xref', key: 'autopilot' } as const, idleAutopilot())
98// Turn state lives in atoms, so a reload or a /config change in the middle of a turn keeps it.
99const askedA = atom({ plugin: 'speckit-xref', key: 'asked' } as const, null)
100const turnFilesA = atom({ plugin: 'speckit-xref', key: 'turnFiles' } as const, [] as string[])
101const turnA = atom({ plugin: 'speckit-xref', key: 'turn' } as const, null)
102const detailsA = atom({ plugin: 'speckit-xref', key: 'details' } as const, false)
103const seenAtA = atom({ plugin: 'speckit-xref', key: 'seenAt' } as const, 0)
104const chipsA = atom({ plugin: 'speckit-xref', key: 'chips' } as const, {} as Record<string, string>)
105
106// The module's own: they start over on a reload, and session.start or register fills them again.
107let root = ''
108let signature = ''
109let offered = false
110let interactive = true
111let lastLevel: Level = 'none'
112let timer: { cancel: () => void } | null = null
113// What can run Spec Kit's CLI here, looked up once a session.
114let cli = { specify: false, uvx: false }
115let mapModel = 'haiku'
116let defaultMax = 25
117let testCommandOption = ''
118let junitPath = ''
119let review: Flow['review'] = 'spec'
120// Read by the strict guard's fallback, which has to be a top-level function.
121let strict = false
122let parallel = false
123let commitPerTask = false
124// In a `-p` run the next step rides the Stop hook's re-prompt instead of a prompt of its own.
125let pendingPrompt: string | null = null
126// When the step running now was handed over: its length in the run log, where a -p run reports none.
127let stepStartedAt = 0
128
129const at = (rel: string) => (rel.startsWith('/') ? rel : `${root}/${rel}`)
130const stamp = async ($: $) => new Date(await $.clock.now()).toISOString()
131
132async function readText($: $, rel: string): Promise<string | null> {
133  try {
134    return await $.fs.read(at(rel))
135  } catch {
136    return null
137  }
138}
139
140async function mtime($: $, rel: string): Promise<number> {
141  try {
142    return (await $.fs.stat(at(rel))).mtimeMs
143  } catch {
144    return 0
145  }
146}
147
148async function exists($: $, rel: string): Promise<boolean> {
149  try {
150    return await $.fs.exists(at(rel))
151  } catch {
152    return false
153  }
154}
155
156/**
157 * The project root: the nearest folder at or above the working directory that holds `.specify/`, as Spec Kit's
158 * own scripts find it. Claude started in a subfolder of a Spec Kit project must not be offered a nested setup.
159 */
160async function projectRoot($: $, cwd: string): Promise<string> {
161  let dir = cwd.replace(/\/+$/, '') || '/'
162  for (let depth = 0; depth < 16; depth++) {
163    if (await $.fs.exists(`${dir === '/' ? '' : dir}/.specify`).catch(() => false)) return dir
164    const parent = dir.slice(0, dir.lastIndexOf('/')) || '/'
165    if (parent === dir) break
166    dir = parent
167  }
168  return cwd
169}
170
171const IGNORABLE = /^(\.|README|LICENSE|CHANGELOG)/i
172
173/** Empty: nothing at the top but dotfiles, a README, a license or a changelog. */
174async function folderKind($: $): Promise<'empty' | 'existing'> {
175  try {
176    return (await $.fs.list(root)).some(entry => !IGNORABLE.test(entry.name)) ? 'existing' : 'empty'
177  } catch {
178    return 'existing'
179  }
180}
181
182/** Whether a program is on the PATH. */
183async function onPath($: $, program: string): Promise<boolean> {
184  try {
185    return (await $.process.run(['which', program], { cwd: root, timeoutMs: 3000 })).exitCode === 0
186  } catch {
187    return false
188  }
189}
190
191/** Every feature directory with a spec.md, and when its spec or tasks last changed. */
192async function listFeatures($: $): Promise<{ dir: string; time: number }[]> {
193  const found: { dir: string; time: number }[] = []
194  for (const base of ['specs', '.specify/specs']) {
195    let entries
196    try {
197      entries = await $.fs.list(at(base))
198    } catch {
199      continue
200    }
201    for (const entry of entries) {
202      if (entry.kind !== 'dir') continue
203      const dir = `${base}/${entry.name}`
204      const spec = await mtime($, `${dir}/spec.md`)
205      if (spec) found.push({ dir, time: Math.max(spec, await mtime($, `${dir}/tasks.md`)) })
206    }
207  }
208  return found.sort((a, b) => a.dir.localeCompare(b.dir))
209}
210
211/** The git branch checked out; null outside git or on a detached HEAD. */
212async function currentBranch($: $): Promise<string | null> {
213  try {
214    const ran = await $.process.run(['git', 'rev-parse', '--abbrev-ref', 'HEAD'], { cwd: root, timeoutMs: 3000 })
215    const branch = ran.exitCode === 0 ? ran.stdout.trim() : ''
216    return branch && branch !== 'HEAD' ? branch : null
217  } catch {
218    return null
219  }
220}
221
222/**
223 * The active feature, as docs/contract-0.4.md §7 orders it: SPECIFY_FEATURE_DIRECTORY, .specify/feature.json,
224 * SPECIFY_FEATURE, the feature the git branch names (`001-…`), else the spec written last.
225 */
226async function resolveFeature($: $, features?: { dir: string; time: number }[], branch?: string | null): Promise<string | null> {
227  const named = await $.env.get('SPECIFY_FEATURE')
228  const pointers = [
229    await $.env.get('SPECIFY_FEATURE_DIRECTORY'),
230    parseFeatureJson((await readText($, '.specify/feature.json')) ?? ''),
231    named ? `specs/${named}` : null,
232  ]
233  for (const pointer of pointers) {
234    if (!pointer) continue
235    const rel = relPath(pointer, root).replace(/\/$/, '')
236    if (await exists($, `${rel}/spec.md`)) return rel
237  }
238  const all = features ?? (await listFeatures($))
239  const fromBranch = featureFromBranch(branch === undefined ? ((await currentBranch($)) ?? '') : (branch ?? ''), all.map(f => f.dir))
240  if (fromBranch) return fromBranch
241  const latest = [...all].sort((a, b) => b.time - a.time)[0]
242  return latest?.dir ?? null
243}
244
245async function currentSignature($: $, featureDir: string | null): Promise<string> {
246  const files = ['.specify/feature.json', '.specify/memory/constitution.md', '.specify/integration.json', '.specify/extensions.yml']
247  if (featureDir) files.push(`${featureDir}/spec.md`, `${featureDir}/tasks.md`, `${featureDir}/plan.md`, '.xrefignore')
248  const times = await Promise.all(files.map(f => mtime($, f)))
249  return [featureDir ?? '-', ...times].join('|')
250}
251
252/** Reads Spec Kit's files into the snapshot; a new feature brings its own ledger. */
253async function scan($: $): Promise<void> {
254  const features = await listFeatures($)
255  const branch = await currentBranch($)
256  const featureDir = await resolveFeature($, features, branch)
257  const [specMd, tasksMd, planMd, constitutionMd, integrationJson, initOptions, ignoreText] = await Promise.all([
258    featureDir ? readText($, `${featureDir}/spec.md`) : null,
259    featureDir ? readText($, `${featureDir}/tasks.md`) : null,
260    featureDir ? readText($, `${featureDir}/plan.md`) : null,
261    readText($, '.specify/memory/constitution.md'),
262    readText($, '.specify/integration.json'),
263    readText($, '.specify/init-options.json'),
264    readText($, '.xrefignore'),
265  ])
266  let extensions: string[] = []
267  try {
268    extensions = (await $.fs.list(at('.specify/extensions'))).filter(e => e.kind === 'dir' && !e.name.startsWith('.')).map(e => e.name)
269  } catch {
270    extensions = []
271  }
272  // integration.json says how commands are invoked; without it, a commands-only layout is the old one.
273  const separator = integrationJson ? invokeSeparator(integrationJson) : null
274  const commandsOnly = separator ? separator === '.' : (await exists($, '.claude/commands/speckit.clarify.md')) && !(await exists($, '.claude/skills/speckit-clarify/SKILL.md'))
275  const tasks = tasksMd ? parseTasks(tasksMd) : []
276  const previous = await read($, snapshotA)
277  const snapshot: Snapshot = {
278    initialized: await exists($, '.specify'),
279    featureDir,
280    features: features.map(f => f.dir),
281    hasPlan: planMd !== null,
282    extensions,
283    claudeIntegration: await hasClaudeIntegration($),
284    speckitVersion: speckitVersionOf(initOptions),
285    tools: cli,
286    folder: await folderKind($),
287    spec: specMd ? parseSpec(specMd) : null,
288    tasks,
289    constitution: constitutionMd ? parseConstitution(constitutionMd) : null,
290    commandStyle: commandsOnly ? 'commands' : 'skills',
291    rules: rulesFrom(ignoreText),
292    realTests: previous?.featureDir === featureDir ? previous.realTests : [],
293    planFingerprint: planMd ? fingerprint(planMd) : null,
294    testCommand: testCommandOption || testCommandFrom(planMd),
295    branch,
296    phase: openPhase(tasks),
297    tasksFingerprint: tasksFingerprint(tasks),
298    commands: await installedCommands($),
299    checklists: featureDir ? await checklists($, featureDir) : [],
300    persistence: persistenceModel(constitutionMd),
301  }
302  if (!previous || previous.featureDir !== featureDir) {
303    const ledger = featureDir ? ledgerFromParts(await readText($, `${featureDir}/xref.json`), await readText($, localPath(featureDir))) : emptyLedger()
304    const anchors = featureDir ? await scanAnchors($) : null
305    await update($, ledgerA, () => (anchors ? { ...ledger, anchors } : ledger))
306    await update($, activeA, () => null)
307    snapshot.realTests = await findRealTests($, anchors ? { ...ledger, anchors } : ledger)
308  }
309  await update($, snapshotA, () => snapshot)
310  signature = await currentSignature($, featureDir)
311}
312
313/** Spec Kit's commands installed for Claude Code, by name without prefix: `plan`, `analyze`, `xref-check`. */
314async function installedCommands($: $): Promise<string[]> {
315  const names = new Set<string>()
316  try {
317    for (const e of await $.fs.list(at('.claude/skills'))) if (e.kind === 'dir' && e.name.startsWith('speckit-')) names.add(e.name.slice('speckit-'.length))
318  } catch {
319    // No skills folder: the commands layout, or no integration.
320  }
321  try {
322    for (const e of await $.fs.list(at('.claude/commands'))) {
323      const m = /^speckit\.(.+)\.md$/.exec(e.name)
324      if (m) names.add(m[1]!.replace(/\./g, '-'))
325    }
326  } catch {
327    // No commands folder.
328  }
329  return [...names].sort()
330}
331
332/** The feature's checklists with open items. */
333async function checklists($: $, featureDir: string): Promise<{ file: string; open: number }[]> {
334  const files: Record<string, string> = {}
335  try {
336    for (const e of await $.fs.list(at(`${featureDir}/checklists`))) {
337      if (e.kind !== 'file' || !e.name.endsWith('.md')) continue
338      const text = await readText($, `${featureDir}/checklists/${e.name}`)
339      if (text !== null) files[`checklists/${e.name}`] = text
340    }
341  } catch {
342    return []
343  }
344  return openChecklistItems(files)
345}
346
347/** The test files tied to the feature (anchored, linked or touched) that hold real tests, not only test.todo. */
348async function findRealTests($: $, ledger: Ledger): Promise<string[]> {
349  const candidates = new Set<string>()
350  for (const a of ledger.anchors) candidates.add(a.file)
351  for (const r of Object.values(ledger.requirements)) r.files.forEach(f => candidates.add(f))
352  for (const t of Object.values(ledger.tasks)) [...t.touched, ...t.linked].forEach(f => candidates.add(f))
353  const real: string[] = []
354  for (const file of [...candidates].filter(isTestFile).slice(0, 200)) {
355    const text = await readText($, file)
356    if (text !== null && hasRealTests(text)) real.push(file)
357  }
358  return real.sort()
359}
360
361/**
362 * Spec Kit's Claude Code integration: its core commands are there, as skills or as commands. integration.json
363 * alone is not enough, and neither is a stray speckit-* skill: the autopilot has to be able to run the next step.
364 */
365async function hasClaudeIntegration($: $): Promise<boolean> {
366  for (const name of ['plan', 'implement']) {
367    if ((await exists($, `.claude/skills/speckit-${name}/SKILL.md`)) || (await exists($, `.claude/commands/speckit.${name}.md`))) return true
368  }
369  return false
370}
371
372function speckitVersionOf(initOptions: string | null): string | null {
373  try {
374    const value = JSON.parse(initOptions ?? '') as { speckit_version?: unknown }
375    return typeof value.speckit_version === 'string' ? value.speckit_version : null
376  } catch {
377    return null
378  }
379}
380
381async function refresh($: $): Promise<void> {
382  const snap = await read($, snapshotA)
383  const featureDir = await resolveFeature($, undefined, snap?.branch)
384  if (featureDir !== (snap?.featureDir ?? null) || (await currentSignature($, featureDir)) !== signature) await scan($)
385}
386
387/** Every `@spec <id>` comment in the repository, by ripgrep or else git grep; none when neither runs. */
388async function scanAnchors($: $): Promise<Anchor[] | null> {
389  const runs = [
390    ['rg', '-n', '--no-heading', '-o', '-e', ANCHOR_RG, '--glob', '!specs/**', '--glob', '!.specify/**', '.'],
391    // --untracked: a new file the person has not committed yet holds anchors too.
392    ['git', 'grep', '--untracked', '-n', '-o', '-E', ANCHOR_GIT, '--', '.', ':!specs', ':!.specify'],
393  ]
394  for (const argv of runs) {
395    try {
396      const ran = await $.process.run(argv, { cwd: root, timeoutMs: 5000 })
397      if (ran.exitCode > 1) continue
398      const anchors: Anchor[] = []
399      for (const line of ran.stdout.split('\n')) {
400        const m = /^(.+?):(\d+):(.*)$/.exec(line)
401        if (!m) continue
402        for (const a of anchorsIn(m[3] ?? '', relPath(m[1] ?? '', root))) anchors.push({ ...a, line: Number(m[2]) })
403      }
404      return anchors
405    } catch {
406      continue
407    }
408  }
409  return null
410}
411
412/**
413 * Writes both halves of the ledger: the committed `xref.json` (deterministic, reviewable) and the local file
414 * under `.specify/xref/local/`, which a `.gitignore` of its own keeps out of commits.
415 */
416async function persist($: $): Promise<void> {
417  const snap = await read($, snapshotA)
418  if (!snap?.featureDir) return
419  const ledger = await read($, ledgerA)
420  const parts = ledgerToParts(ledger, snap.featureDir)
421  try {
422    if ((await readText($, `${snap.featureDir}/xref.json`)) !== parts.committed) await $.fs.write(at(`${snap.featureDir}/xref.json`), parts.committed)
423    if (!(await exists($, LOCAL_IGNORE.path))) await $.fs.write(at(LOCAL_IGNORE.path), LOCAL_IGNORE.text)
424    await $.fs.write(at(localPath(snap.featureDir)), parts.local)
425  } catch {
426    // A read-only checkout keeps the ledger for the session alone.
427  }
428}
429
430/** Tells the person once the drift gets worse, never on every write. */
431async function notifyLevel($: $): Promise<void> {
432  const snap = await read($, snapshotA)
433  if (!snap) return
434  const report = evaluate(snap, await read($, ledgerA))
435  if (RANK[report.level] > RANK[lastLevel] && (report.level === 'yellow' || report.level === 'red')) {
436    $.ui.toast(`Spec drift ${report.level}: ${report.findings[0]?.text ?? ''}`)
437  }
438  lastLevel = report.level
439}
440
441async function afterWrite($: $, rel: string, c: Classified, newText: string): Promise<void> {
442  if (c.verdict === 'spec') {
443    await scan($)
444    return
445  }
446  // Lockfiles, build output, caches: no evidence for a task, no drift, no part of the turn's changes.
447  if (c.verdict === 'exempt') return
448  const when = await stamp($)
449  // A file's anchors are read again when the write may have added or removed one.
450  const hadAnchors = (await read($, ledgerA)).anchors.some(a => a.file === rel)
451  const text = hadAnchors || newText.includes('@spec') ? await readText($, rel) : null
452  const fresh = text === null ? null : anchorsIn(text, rel)
453  await update($, ledgerA, l => {
454    const next = recordTouch(l, rel, c, when)
455    return fresh ? { ...next, anchors: [...next.anchors.filter(a => a.file !== rel), ...fresh] } : next
456  })
457  // A file of a finished task is rework: it does not pull the focus back to that task.
458  const task = (await read($, snapshotA))?.tasks.find(t => t.id === c.task)
459  if (task && !task.done && (c.verdict === 'in-scope' || c.verdict === 'other-task')) await update($, activeA, () => task.id)
460  await update($, turnFilesA, files => (files.includes(rel) ? files : [...files, rel]))
461  // A test file written is read once: whether it holds real tests decides the ladder's `tested`.
462  if (isTestFile(rel)) {
463    const body = text ?? (await readText($, rel))
464    await update($, snapshotA, s => (s ? { ...s, realTests: body && hasRealTests(body) ? [...new Set([...s.realTests, rel])].sort() : s.realTests.filter(f => f !== rel) } : s))
465  }
466  await notifyLevel($)
467}
468
469/**
470 * The files the working tree changed, for a check the person asked for between turns: relative to the project
471 * and only inside it, since the project may be one folder of a larger repository.
472 */
473async function changedFiles($: $): Promise<string[]> {
474  const files = new Set<string>()
475  for (const argv of [
476    ['git', 'diff', '--name-only', '--relative', 'HEAD'],
477    ['git', 'ls-files', '--others', '--exclude-standard'],
478  ]) {
479    try {
480      const ran = await $.process.run(argv, { cwd: root, timeoutMs: 5000 })
481      if (ran.exitCode === 0) for (const line of ran.stdout.split('\n')) if (line.trim()) files.add(line.trim())
482    } catch {
483      // No git, or no commit yet: the files of this session's turns are all there is.
484    }
485  }
486  return [...files].filter(f => !isSpecArtifact(f)).slice(0, 40)
487}
488
489async function diffOf($: $, files: string[]): Promise<string> {
490  let diff = ''
491  try {
492    diff = (await $.process.run(['git', 'diff', '--no-color', '--relative', '-U2', 'HEAD', '--', ...files], { cwd: root, timeoutMs: 5000 })).stdout
493  } catch {
494    diff = ''
495  }
496  // A new file is no part of `git diff`; its head stands in for it.
497  for (const file of files) {
498    if (diff.length > 6000 || diff.includes(`b/${file}`)) continue
499    const text = await readText($, file)
500    if (text !== null) diff += `\n--- new or untracked: ${file}\n${text.split('\n').slice(0, 60).join('\n')}\n`
501  }
502  return diff
503}
504
505/** The semantic check: one tool-less question over the session's own transcript. */
506async function runCheck($: $, files: string[], model: string): Promise<string> {
507  const snap = await read($, snapshotA)
508  if (!snap?.featureDir) return 'No Spec Kit feature to check against.'
509  if (await read($, checkingA)) return 'A drift check is already running.'
510  await update($, checkingA, () => true)
511  try {
512    const changed = files.length ? files : await changedFiles($)
513    const ledger = await read($, ledgerA)
514    const active = await read($, activeA)
515    const diff = await diffOf($, changed)
516    let reply = await $.model.fork({ prompt: forkPrompt(snap, ledger, active, changed, diff) })
517    // Before the session's first turn there is no transcript to fork; the intent log carries the user's words instead.
518    if (!reply.isAnswered && reply.reason === 'nothing-to-fork') {
519      reply = await $.model.complete({ model, prompt: forkPrompt(snap, ledger, active, changed, diff, true), maxTokens: 1000 })
520    }
521    if (!reply.isAnswered) return `The drift check got no answer (${reply.reason}).`
522    const semantic = parseSemantic(reply.text, await stamp($))
523    if (!semantic) return 'The drift check answered in a shape the mod could not read.'
524    await update($, ledgerA, l => applySemantic(l, semantic))
525    await persist($)
526    await notifyLevel($)
527    const conflict = (await read($, ledgerA)).semantic?.changes.find(c => c.kind === 'contradicts')
528    if (conflict && (await read($, autopilotA)).on) await pauseAutopilot($, `your request "${conflict.text}" contradicts the spec; decide in the pane`)
529    return `Intent check ${semantic.score}/100 (${semantic.verdict})${semantic.reasons[0] ? `: ${semantic.reasons[0]}` : ''}`
530  } finally {
531    await update($, checkingA, () => false)
532  }
533}
534
535async function runMap($: $, model: string): Promise<string> {
536  const snap = await read($, snapshotA)
537  if (!snap?.spec || snap.tasks.length === 0) return 'Mapping needs a spec.md with requirements and a tasks.md.'
538  const reply = await $.model.complete({ model, prompt: mapPrompt(snap.spec, snap.tasks), maxTokens: 2000 })
539  if (!reply.isAnswered) return `The mapping got no answer (${reply.reason}).`
540  let mapped = 0
541  await update($, ledgerA, l => {
542    const result = applyMapping(l, reply.text, snap)
543    mapped = result.mapped
544    return result.ledger
545  })
546  await persist($)
547  const report = evaluate(snap, await read($, ledgerA))
548  return `Mapped ${mapped} requirements to tasks; ${report.covered}/${report.total} FRs covered.`
549}
550
551async function toSpec($: $, intent: string): Promise<void> {
552  const snap = await read($, snapshotA)
553  if (!snap) return
554  await update($, ledgerA, l => resolveIntent(l, intent))
555  await persist($)
556  const args = `The user asked during implementation: "${intent}". Fold this into the spec, or record why it stays out of scope.`
557  // The extension's revise keeps ids stable (new, SUPERSEDED, RETIRED) and logs revisions.md; clarify is the fallback.
558  const revise = snap.extensions.includes('xref') && snap.commands.includes('xref-revise')
559  // A plugin runs a slash command through $.command.run; a prompt may not start with one.
560  const command = speckitCommand(snap, revise ? 'xref.revise' : 'clarify').slice(1)
561  try {
562    await $.command.run({ command, args })
563  } catch {
564    void $.prompt.submit({ text: `Spec Kit: ${args} Use ${speckitCommand(snap, 'clarify')} for it.` })
565  }
566}
567
568async function asTask($: $, intent: string): Promise<string> {
569  const snap = await read($, snapshotA)
570  if (!snap?.featureDir) return 'No active feature.'
571  const path = `${snap.featureDir}/tasks.md`
572  const markdown = (await readText($, path)) ?? '# Tasks\n'
573  const added = appendRemediation(markdown, snap.tasks, intent)
574  await $.fs.write(at(path), added.markdown)
575  await update($, ledgerA, l => resolveIntent(l, intent))
576  await scan($)
577  await persist($)
578  return `Added ${added.id} to ${path}.`
579}
580
581async function focus($: $, id: string): Promise<string> {
582  const snap = await read($, snapshotA)
583  const task = snap?.tasks.find(t => t.id === id.trim().toUpperCase())
584  if (!snap || !task) return `No task ${id} in tasks.md.`
585  await update($, activeA, () => task.id)
586  const reqs = reqsOf(task, await read($, ledgerA))
587  return [
588    `Current task: ${task.id}${task.story ? ` [${task.story}]` : ''} ${task.text}`,
589    `Planned files: ${task.paths.join(', ') || '(none named)'}`,
590    `Serves: ${reqs.join(', ') || '(no requirement mapped yet)'}`,
591  ].join('\n')
592}
593
594async function where($: $, file: string): Promise<string> {
595  const snap = await read($, snapshotA)
596  if (!snap?.featureDir) return 'No active Spec Kit feature.'
597  const rel = relPath(file, root)
598  const ledger = await read($, ledgerA)
599  const c = classify(rel, snap, ledger, await read($, activeA))
600  const planned = snap.tasks.filter(t => t.paths.some(p => rel === p || rel.endsWith('/' + p) || rel.startsWith(p.replace(/\/?$/, '/'))))
601  const touched = Object.entries(ledger.tasks).filter(([, e]) => e.touched.includes(rel) || e.linked.includes(rel)).map(([id]) => id)
602  const reqs = Object.entries(ledger.requirements).filter(([, r]) => r.files.includes(rel)).map(([id]) => id)
603  const anchors = ledger.anchors.filter(a => a.file === rel).map(a => `${a.id} (line ${a.line})`)
604  return [
605    `${rel} in ${snap.featureDir}: ${c.verdict}${c.task ? ` (${c.task})` : ''}`,
606    `Planned by: ${planned.map(t => t.id).join(', ') || '-'}`,
607    `Touched or linked by: ${touched.join(', ') || '-'}`,
608    `Requirements: ${[...reqs, ...anchors].join(', ') || '-'}`,
609  ].join('\n')
610}
611
612async function linkFile($: $, file: string, id: string): Promise<string> {
613  const snap = await read($, snapshotA)
614  if (!snap?.featureDir) return 'No active Spec Kit feature.'
615  const rel = relPath(file, root)
616  let error: string | undefined
617  await update($, ledgerA, l => {
618    const result = link(l, rel, id, snap)
619    error = result.error
620    return result.ledger
621  })
622  if (error) return error
623  await persist($)
624  await notifyLevel($)
625  return `Linked ${rel} to ${id}.`
626}
627
628/** What the workflow reads beside the snapshot: the review option and the last test run. */
629const flowOf = (ap: Autopilot): Flow => ({ review, lastTest: ap.lastTest, repairs: ap.repairs })
630
631/** Starts a fresh autopilot run; its first step follows as soon as the session is free. */
632async function startAutopilot($: $, max?: number, night = false): Promise<void> {
633  const cost = await sessionCost($)
634  const now = await $.clock.now()
635  await update($, autopilotA, a => ({ ...a, on: true, paused: null, steps: 0, stalls: 0, last: null, lastPhase: null, idea: null, max: max ?? defaultMax, repairs: 0, lastTest: null, startedAt: now, costAtStart: cost, night, scope: null }))
636  await update($, seenAtA, () => now)
637  await retitle($)
638  $.clock.after(0, () => void advance($).catch(() => undefined))
639}
640
641/** Goes on after a pause; a run whose budget is spent gets a new one. */
642async function resumeAutopilot($: $): Promise<void> {
643  await update($, autopilotA, a => ({ ...a, paused: null, stalls: 0, repairs: a.repairs > MAX_REPAIRS ? 0 : a.repairs, steps: a.steps >= a.max ? 0 : a.steps }))
644  await retitle($)
645  $.clock.after(0, () => void advance($).catch(() => undefined))
646}
647
648/** Stop: the autopilot goes off, and a turn it is running ends now rather than at its end. */
649async function turnOffAutopilot($: $): Promise<void> {
650  const wasOn = (await read($, autopilotA)).on
651  await update($, autopilotA, a => ({ ...a, on: false, paused: null, idea: null }))
652  await retitle($)
653  const turn = await read($, turnA)
654  if (wasOn && turn) await $.turn.abort({ turnId: turn.id }).catch(() => undefined)
655  if (wasOn) await writeBriefing($, 'stopped by you')
656}
657
658/** The session's cost so far in US dollars, or 0 where the host keeps none. */
659async function sessionCost($: $): Promise<number> {
660  try {
661    return (await $.session.usage()).cost?.usd ?? 0
662  } catch {
663    return 0
664  }
665}
666
667/** The pane's tab says when the autopilot waits, so it shows behind the changes pane too. */
668async function retitle($: $): Promise<void> {
669  try {
670    const pane = (await $.ui.panes()).find(p => p.id === PANE)
671    if (!pane) return
672    const ap = await read($, autopilotA)
673    await $.ui.open({ id: PANE, title: ap.on && ap.paused ? `${TITLE} ⏸` : TITLE })
674  } catch {
675    // A host without panes has no tab to name.
676  }
677}
678
679/** The requests the pane's setup actions hand to Claude, as the person's own words: a press is their consent. */
680const SETUP_ASKS = {
681  idea: 'I want to start something new in this empty folder (an app, a project, or a problem to solve). Ask me what it is, then set Spec Kit up for it with the speckit-xref:speckit skill and write the spec in my own words.',
682  setup: 'Set Spec Kit up in this existing project with the speckit-xref:speckit skill: run specify init, draft the constitution from the code and the README and mark the assumptions, then ask me which change to specify first.',
683  integration: "Install Spec Kit's Claude Code integration in this project (specify integration install claude), so its /speckit-* commands run here.",
684} as const
685
686async function askClaude($: $, kind: keyof typeof SETUP_ASKS): Promise<void> {
687  // A press answers what the autopilot waits on: it goes on once Claude has done it.
688  if ((await read($, autopilotA)).paused) await update($, autopilotA, a => ({ ...a, paused: null, stalls: 0 }))
689  void $.prompt.submit({ text: SETUP_ASKS[kind], asUser: true })
690}
691
692/** Approvals only the person gives: the spec, the plan, open checklists. The autopilot records checkpoints, never these. */
693const APPROVALS = new Set(['spec', 'plan', 'checklists'])
694const CHECKPOINTS = new Set(['analyze', 'converge'])
695
696/** The person approves what the autopilot waits on (spec, plan, open checklists); the run goes on. */
697async function approve($: $): Promise<string> {
698  const snap = await read($, snapshotA)
699  if (!snap) return 'Nothing to approve.'
700  const ap = await read($, autopilotA)
701  const step = nextStep(snap, await read($, ledgerA), ap.idea, flowOf(ap))
702  if (!step.approve || !APPROVALS.has(step.approve.key)) return 'Nothing waits for an approval.'
703  const { key, value } = step.approve
704  await update($, ledgerA, l => ({ ...l, approvals: { ...l.approvals, [key]: value } }))
705  await persist($)
706  if (ap.on && ap.paused) await resumeAutopilot($)
707  return key === 'checklists' ? 'Going on despite the open checklist items.' : `Approved the ${key} of ${snap.featureDir}.`
708}
709
710/** Stops the autopilot until the person speaks, says why, and asks again after five minutes. */
711async function pauseAutopilot($: $, reason: string, step?: Step): Promise<void> {
712  await update($, autopilotA, a => ({ ...a, paused: reason }))
713  $.ui.toast(`Autopilot waits for you: ${reason}`)
714  await retitle($)
715  $.clock.after(5 * 60_000, () => void remind($, reason).catch(() => undefined))
716  // At a review the native dialog answers it in place, where a person is there to answer.
717  if (step?.approve && APPROVALS.has(step.approve.key) && interactive) $.clock.after(0, () => void askApproval($, step).catch(() => undefined))
718}
719
720/** A native notification on top of the toast; switched off or without a channel, the band and the toast still say it. */
721async function notify($: $, text: string): Promise<void> {
722  try {
723    await $.ui.notify(text, { title: TITLE })
724  } catch {
725    // Notifications switched off, or no channel: the band and the toast still say it.
726  }
727}
728
729async function remind($: $, reason: string): Promise<void> {
730  const ap = await read($, autopilotA)
731  if (ap.on && ap.paused === reason) await notify($, `Autopilot waits for you: ${reason}`)
732}
733
734async function askApproval($: $, step: Step): Promise<void> {
735  const key = step.approve!.key
736  const yes = key === 'spec' ? 'Approve spec' : key === 'plan' ? 'Approve plan' : 'Proceed anyway'
737  const question = key === 'checklists' ? `${step.why} Go on implementing anyway?` : `${step.why} Approve the ${key} as the contract?`
738  let answer: string
739  try {
740    answer = await $.ui.ask(question, [yes, 'Not yet'])
741  } catch {
742    return
743  }
744  if (answer === yes) await approve($)
745  // Anything typed under "Other" is what to change: it goes to Claude as the person's words.
746  else if (answer !== 'Not yet' && answer.trim()) void $.prompt.submit({ text: answer, asUser: true })
747}
748
749async function stopAutopilot($: $, why: string): Promise<void> {
750  await update($, autopilotA, a => ({ ...a, on: false, paused: null }))
751  $.ui.toast(why)
752  await retitle($)
753  await writeBriefing($, why)
754  await notify($, why)
755}
756
757/** The git HEAD, short; empty outside git. */
758async function head($: $): Promise<string> {
759  try {
760    const ran = await $.process.run(['git', 'rev-parse', '--short', 'HEAD'], { cwd: root, timeoutMs: 3000 })
761    return ran.exitCode === 0 ? ran.stdout.trim() : ''
762  } catch {
763    return ''
764  }
765}
766
767/**
768 * Runs the project's tests and records what they prove: per requirement from JUnit when a report path is set,
769 * else every requirement with a real test passes when the whole suite does. `ran` is false when no runner started.
770 */
771async function runTests($: $): Promise<{ ran: boolean; ok: boolean; output: string }> {
772  const snap = await read($, snapshotA)
773  if (!snap?.testCommand || !snap.featureDir) return { ran: false, ok: true, output: '' }
774  let exitCode: number
775  let output: string
776  try {
777    const ran = await $.process.run(['sh', '-c', snap.testCommand], { cwd: root, timeoutMs: 600_000 })
778    exitCode = ran.exitCode
779    output = `${ran.stdout}\n${ran.stderr}`.trim().slice(-4000)
780  } catch (error) {
781    return { ran: false, ok: true, output: String(error) }
782  }
783  const at = await stamp($)
784  const commit = await head($)
785  const current = Object.fromEntries((snap.spec?.reqs ?? []).map(r => [r.id, fingerprint(r.text)]))
786  const ledger = await read($, ledgerA)
787  const proof = { featureDir: snap.featureDir, tasks: snap.tasks, reqs: snap.spec?.reqs ?? [], realTests: snap.realTests }
788  const xml = junitPath ? await readText($, junitPath) : null
789  const verified = xml
790    ? verificationFrom(parseJunit(xml), ledger.anchors, current, snap.featureDir, at, commit)
791    : verificationFromExit(exitCode, (snap.spec?.reqs ?? []).filter(r => LADDER.indexOf(levelOf(r.id, proof, ledger)) >= LADDER.indexOf('tested')).map(r => r.id), current, at, commit)
792  if (Object.keys(verified).length) {
793    await update($, ledgerA, l => ({
794      ...l,
795      verification: { ...l.verification, ...verified },
796      fingerprints: { ...l.fingerprints, ...Object.fromEntries(Object.keys(verified).filter(id => current[id]).map(id => [id, current[id]!])) },
797    }))
798    await persist($)
799  }
800  return { ran: true, ok: exitCode === 0, output }
801}
802
803/** Appends one finished step to the run log (local, never committed). */
804async function logRun($: $, entry: RunEntry): Promise<void> {
805  const snap = await read($, snapshotA)
806  if (!snap?.featureDir) return
807  const path = runLogPath(snap.featureDir)
808  const lines = ((await readText($, path)) ?? '').split('\n').filter(Boolean).slice(-499)
809  lines.push(JSON.stringify(entry))
810  await $.fs.write(at(path), lines.join('\n') + '\n').catch(() => undefined)
811}
812
813async function runLog($: $): Promise<RunEntry[]> {
814  const snap = await read($, snapshotA)
815  if (!snap?.featureDir) return []
816  return ((await readText($, runLogPath(snap.featureDir))) ?? '')
817    .split('\n')
818    .filter(Boolean)
819    .map(line => {
820      try {
821        return JSON.parse(line) as RunEntry
822      } catch {
823        return null
824      }
825    })
826    .filter((e): e is RunEntry => !!e)
827}
828
829/** What happened since the person last spoke: the "since you left" card and the night run's briefing. */
830async function sinceYouLeft($: $): Promise<string[]> {
831  const seen = await read($, seenAtA)
832  const entries = (await runLog($)).filter(e => Date.parse(e.at) >= seen)
833  if (!entries.length) return []
834  const snap = await read($, snapshotA)
835  const ap = await read($, autopilotA)
836  const files = new Set(entries.flatMap(e => e.files))
837  const minutes = Math.round(entries.reduce((n, e) => n + e.durationMs, 0) / 60_000)
838  const lines = [`${entries.length} steps in ${minutes} min: ${[...new Set(entries.map(e => e.phase))].join(', ')}; ${files.size} files changed.`]
839  if (snap) lines.push(`Tasks ${snap.tasks.filter(t => t.done).length}/${snap.tasks.length} checked${ap.lastTest ? `; tests ${ap.lastTest.ok ? 'pass' : 'fail'}` : ''}.`)
840  const ledger = await read($, ledgerA)
841  const open = ledger.decisions.filter(d => d.answer === null)
842  if (open.length) lines.push(`Waiting for you: ${open.map(d => d.question).join(' | ')}`)
843  if (ap.paused) lines.push(`Paused: ${ap.paused}`)
844  return lines
845}
846
847/** A night run leaves a briefing for the morning: `.specify/xref/local/briefing.md`. */
848async function writeBriefing($: $, why: string): Promise<void> {
849  const ap = await read($, autopilotA)
850  if (!ap.night) return
851  const lines = await sinceYouLeft($)
852  await $.fs.write(at('.specify/xref/local/briefing.md'), [`# Autopilot briefing (${await stamp($)})`, '', `Ended: ${why}`, '', ...lines.map(l => `- ${l}`), ''].join('\n')).catch(() => undefined)
853}
854
855/**
856 * One autopilot step: the next Spec Kit step is handed to the model as a prompt of its own, unless only the
857 * person can take it, the run stopped moving, or the step budget is spent.
858 */
859async function advance($: $): Promise<void> {
860  const ap = await read($, autopilotA)
861  if (!ap.on || ap.paused) return
862  await scan($)
863  const snap = await read($, snapshotA)
864  if (!snap) return
865  const ledger = await read($, ledgerA)
866  const step = nextStep(snap, ledger, ap.idea, flowOf(ap))
867  if (step.needsUser) return pauseAutopilot($, step.needsUser, step)
868  // A repair attempt is progress of its own: the stall check must not end the three repairs early.
869  const key = `${progressKey(snap, ledger)}|${ap.repairs}`
870  const stalls = key === ap.last ? ap.stalls + 1 : 0
871  if (stalls >= 3) return pauseAutopilot($, 'three steps without progress; look at what blocks it, then Resume')
872  if (ap.steps >= ap.max) return pauseAutopilot($, `the step budget (${ap.max}) is used up; /xref auto on starts a new one`)
873  if (step.phase === 'verify' && ap.lastPhase === 'verify') {
874    const report = evaluate(snap, ledger)
875    if (report.level === 'red') return pauseAutopilot($, 'the feature is built, but the drift is red; decide in the pane')
876    // Done means proven: with a test command, the suite has to pass once more before the run ends.
877    if (snap.testCommand) {
878      const tests = await runTests($)
879      if (tests.ran && !tests.ok) {
880        await update($, autopilotA, a => ({ ...a, lastTest: { ok: false, at: '', output: tests.output }, repairs: a.repairs + 1, lastPhase: 'repair' }))
881        return advance($)
882      }
883      return stopAutopilot($, `Autopilot done: every task of ${snap.featureDir} is checked, ${tests.ran ? 'the tests pass' : `the tests could not run (${snap.testCommand})`} and the drift is ${report.level}.`)
884    }
885    return stopAutopilot($, `Autopilot done: every task of ${snap.featureDir} is checked and the drift is ${report.level}. No test command is known, so "done" rests on the checkboxes.`)
886  }
887  await update($, autopilotA, a => ({ ...a, steps: a.steps + 1, last: key, stalls, lastPhase: step.phase, idea: step.phase === 'specify' ? null : a.idea, scope: step.phase === 'implement' ? snap.phase : a.scope }))
888  // Handing a step over is its checkpoint (analyze, converge); a revise takes the request it folds in off the list.
889  if (step.approve) {
890    const { key: point, value } = step.approve
891    await update($, ledgerA, l => ({ ...l, checkpoints: { ...l.checkpoints, [point]: value } }))
892  }
893  if (step.resolves) await update($, ledgerA, l => resolveIntent(l, step.resolves!))
894  await persist($)
895  // The mod maps requirements itself where no extension command does it, then moves on.
896  if (step.phase === 'map' && !snap.extensions.includes('xref')) {
897    await runMap($, mapModel)
898    return advance($)
899  }
900  const active = await read($, activeA)
901  const text = autopilotPrompt(step, ap.steps + 1, ap.max, turnContext(snap, ledger, active)) + parallelNote(snap, step)
902  await update($, autopilotA, a => ({ ...a, doneAtStep: snap.tasks.filter(t => t.done).map(t => t.id) }))
903  stepStartedAt = await $.clock.now()
904  // Nobody is at the prompt in a -p run: the Stop hook hands the step over as its re-prompt.
905  if (!interactive) {
906    pendingPrompt = text
907    return
908  }
909  void $.prompt.submit({ text })
910}
911
912const RUNNER = 'task-runner'
913
914/** With the parallel option, a phase's open [P] tasks go to subagents of their own, one per task. */
915function parallelNote(snap: Snapshot, step: Step): string {
916  if (!parallel || step.phase !== 'implement' || !snap.phase) return ''
917  const ready = snap.tasks.filter(t => !t.done && t.parallel && t.phase === snap.phase)
918  if (ready.length < 2) return ''
919  return [
920    '',
921    '',
922    `Parallel: ${ready.map(t => t.id).join(', ')} are marked [P] and touch different files. Spawn one ${PLUGIN}:${RUNNER} agent per task, all in one message (prompt: the task id and its line from tasks.md).`,
923    'When they report back, check those tasks off in tasks.md yourself, then do the rest of the phase.',
924  ].join('\n')
925}
926
927/** What a task-runner subagent is told: one task, its planned files, the xref rules, and no tasks.md. */
928const RUNNER_PROMPT = [
929  'You implement exactly one task of a GitHub Spec Kit feature; other agents implement its sibling tasks at the same time.',
930  'Read the task line you were given, the spec.md and plan.md of the active feature (specs/<feature>/), and the files the task names.',
931  `First call mcp__${PLUGIN}__focus with the task id. Change only the files the task names; if another file is unavoidable, call mcp__${PLUGIN}__link for it.`,
932  'Mark code that implements a requirement with a comment `@spec <feature>/FR-###`; tests carry the anchor of what they prove.',
933  'Do not edit tasks.md and do not commit: the main agent checks the task off. Run the tests that cover your files if you can.',
934  'Report in three lines: what you changed (files), what the tests said, and anything left open.',
935].join('\n')
936
937/**
938 * Commit per task (an option, off by default): after a step whose tests passed, the tasks it checked off are
939 * committed with their files, on a feature branch only. The mod checks the staged diff itself, since git hooks
940 * do not run under $.process.run.
941 */
942async function commitTasks($: $): Promise<string | null> {
943  if (!commitPerTask) return null
944  const snap = await read($, snapshotA)
945  const ap = await read($, autopilotA)
946  if (!snap?.featureDir || !snap.branch || /^(main|master|trunk|develop)$/.test(snap.branch)) return null
947  const fresh = snap.tasks.filter(t => t.done && !(ap.doneAtStep ?? []).includes(t.id))
948  if (!fresh.length) return null
949  const ledger = await read($, ledgerA)
950  const git = (argv: string[]) => $.process.run(['git', ...argv], { cwd: root, timeoutMs: 15_000 })
951  try {
952    // Never mix in what the person staged themselves.
953    if ((await git(['diff', '--cached', '--name-only'])).stdout.trim()) return 'commit skipped: the index already holds staged changes'
954    const files = [...new Set([...fresh.flatMap(t => ledger.tasks[t.id]?.touched ?? []), `${snap.featureDir}/tasks.md`, `${snap.featureDir}/xref.json`])]
955    const present = []
956    for (const f of files) if (await exists($, f)) present.push(f)
957    const secret = present.find(f => /(^|\/)\.env(?!\.example$)/.test(f))
958    if (secret) return `commit skipped: ${secret} looks like a secret`
959    const drift = present.filter(f => ledger.unplanned.some(u => u.file === f && !u.acknowledged))
960    if (drift.length) return `commit skipped: ${drift.join(', ')} ${drift.length === 1 ? 'is' : 'are'} outside the plan`
961    await git(['add', '--', ...present])
962    const subject = fresh.length === 1 ? `${fresh[0]!.id}: ${short(fresh[0]!.text, 60)}` : `${fresh.map(t => t.id).join(', ')} (${snap.featureDir.slice(snap.featureDir.lastIndexOf('/') + 1)})`
963    const ran = await git(['commit', '-m', subject, '-m', `Tasks checked by the speckit-xref autopilot; tests passed.`])
964    return ran.exitCode === 0 ? `committed ${fresh.map(t => t.id).join(', ')}` : `commit failed: ${ran.stderr.trim().slice(0, 200)}`
965  } catch (error) {
966    return `commit failed: ${String(error).slice(0, 200)}`
967  }
968}
969
970type TurnEnd = { answer: string; reason: string; isAborted: boolean; durationMs: number; usage?: { input_tokens?: number; output_tokens?: number } }
971
972/** After a turn: the autopilot pauses where the model or the person needs the person, and moves on otherwise. */
973async function afterTurn($: $, e: TurnEnd, files: string[]): Promise<void> {
974  const asked = await read($, askedA)
975  await update($, askedA, () => null)
976  const ap = await read($, autopilotA)
977  if (!ap.on) return
978  if (ap.lastPhase && ap.steps > 0) {
979    const snap = await read($, snapshotA)
980    const outcome = e.isAborted ? 'interrupted' : e.reason !== 'answer' ? e.reason : asked ? 'asked the person' : 'answered'
981    // The tasks this step checked off, else the one it was on.
982    const checked = snap ? snap.tasks.filter(t => t.done && !(ap.doneAtStep ?? []).includes(t.id)).map(t => t.id) : []
983    await logRun($, {
984      at: await stamp($),
985      step: ap.steps,
986      phase: ap.lastPhase,
987      task: checked.length ? checked.join(', ') : snap ? (currentTask(snap, await read($, activeA))?.id ?? null) : null,
988      files,
989      durationMs: e.durationMs || (stepStartedAt ? (await $.clock.now()) - stepStartedAt : 0),
990      tokens: (e.usage?.input_tokens ?? 0) + (e.usage?.output_tokens ?? 0),
991      outcome,
992    })
993  }
994  if (ap.paused) return
995  if (e.isAborted) return pauseAutopilot($, 'you interrupted the turn; Resume, or a message from you, goes on')
996  // An API error or a refusal is no answer: going on would hand the next step to a turn that never did this one.
997  if (e.reason !== 'answer') return pauseAutopilot($, `the turn ended with ${e.reason === 'refusal' ? 'a refusal' : 'an error'}; Resume tries again`)
998  if (asked) return pauseAutopilot($, asked)
999  if (endsWithQuestion(e.answer)) {
1000    // A question at the end is a stop only when it is a real decision; "shall I go on?" is not one.
1001    const label = await $.model.classify(e.answer.slice(-1500), [...QUESTION_LABELS]).catch(() => QUESTION_LABELS[0])
1002    if (label !== QUESTION_LABELS[1]) return pauseAutopilot($, 'the last answer asks you something')
1003  }
1004  // Code changed: the tests say whether the step holds. A failure becomes a repair step, at most three in a row.
1005  if (ap.lastPhase === 'implement' || ap.lastPhase === 'repair' || ap.lastPhase === 'converge') {
1006    const tests = await runTests($)
1007    if (tests.ran) {
1008      const when = await stamp($)
1009      await update($, autopilotA, a => ({ ...a, lastTest: { ok: tests.ok, at: when, output: tests.output }, repairs: tests.ok ? 0 : a.repairs + 1 }))
1010      if (tests.ok) {
1011        const committed = await commitTasks($)
1012        if (committed) $.ui.toast(`Autopilot: ${committed}`)
1013      }
1014    }
1015  }
1016  const snap = await read($, snapshotA)
1017  if (snap && evaluate(snap, await read($, ledgerA)).level === 'red') return pauseAutopilot($, 'the drift is red; look at the findings in the pane, then Resume')
1018  await advance($)
1019}
1020/** The chip a booked write gets in the transcript: the task and what it serves, or why it is drift. */
1021function chipFor(c: Classified, snap: Snapshot, ledger: Ledger): string | null {
1022  if (c.verdict === 'unplanned') return '▲ unplanned'
1023  if (c.verdict === 'unclear') return '· unclear'
1024  if (c.verdict === 'spec' || c.verdict === 'exempt' || c.verdict === 'untracked') return null
1025  const task = snap.tasks.find(t => t.id === c.task)
1026  if (!task) return c.verdict === 'linked' ? '● linked' : null
1027  const reqs = reqsOf(task, ledger)
1028  return `● ${task.id}${reqs.length ? ` · ${reqs.join(' ')}` : ''}${task.done ? ' · rework' : ''}`
1029}
1030
1031/** An autopilot prompt as one line in the transcript: `▶ auto 3/25 · implement · 4 of 8 tasks open; next T005.` */
1032export function compactPrompt(text: string): string | null {
1033  const m = /^\[speckit-xref autopilot · step (\d+\/\d+)\] (\w+): (.*)$/m.exec(text)
1034  return m ? `▶ auto ${m[1]} · ${m[2]} · ${m[3]}` : null
1035}
1036
1037/** What this run cost so far and how long it took: ` · $1.80 · 23m`. */
1038async function runCost($: $, ap: Autopilot): Promise<string> {
1039  if (!ap.startedAt) return ''
1040  const minutes = Math.max(0, Math.round(((await $.clock.now()) - ap.startedAt) / 60_000))
1041  const usd = Math.max(0, (await sessionCost($)) - ap.costAtStart)
1042  return ` · $${usd.toFixed(2)} · ${minutes}m`
1043}
1044
1045/** Accepts one edit outside the plan as intended. */
1046async function acceptOne($: $, file: string): Promise<void> {
1047  await update($, ledgerA, l => ({ ...l, unplanned: l.unplanned.map(u => (u.file === file ? { ...u, acknowledged: true } : u)) }))
1048  await persist($)
1049  lastLevel = 'none'
1050}
1051
1052/** Accept all asks first: a blind accept is the one way drift tracking quietly stops meaning anything. */
1053async function confirmAcceptAll($: $): Promise<void> {
1054  const open = (await read($, ledgerA)).unplanned.filter(u => !u.acknowledged).length
1055  let answer = ''
1056  try {
1057    answer = await $.ui.ask(`Accept all ${open} edits outside the plan as intended?`, ['Accept all', 'Cancel'])
1058  } catch {
1059    return
1060  }
1061  if (answer !== 'Accept all') return
1062  await update($, ledgerA, acknowledgeAll)
1063  await persist($)
1064  lastLevel = 'none'
1065}
1066
1067/** The git blob hash of each file, to tell a file Bash changed during a turn from one that was dirty before. */
1068async function hashFiles($: $, files: string[]): Promise<Record<string, string>> {
1069  if (!files.length) return {}
1070  try {
1071    const ran = await $.process.run(['git', 'hash-object', '--', ...files], { cwd: root, timeoutMs: 5000 })
1072    if (ran.exitCode !== 0) return {}
1073    const hashes = ran.stdout.split('\n').filter(Boolean)
1074    return Object.fromEntries(files.map((f, i) => [f, hashes[i] ?? '']))
1075  } catch {
1076    return {}
1077  }
1078}
1079
1080/**
1081 * Books what Bash wrote (sed, a code generator, npm): files the working tree changed during the turn that no
1082 * Write or Edit booked. Without this the live view and the CI check would disagree.
1083 */
1084async function bookShellWrites($: $): Promise<void> {
1085  const turn = await read($, turnA)
1086  const snap = await read($, snapshotA)
1087  if (!turn || !snap?.featureDir) return
1088  const booked = new Set(await read($, turnFilesA))
1089  const before = new Map(turn.dirty.map(entry => [entry.slice(0, entry.lastIndexOf(':')), entry.slice(entry.lastIndexOf(':') + 1)]))
1090  const now = (await changedFiles($)).filter(f => !booked.has(f) && !matchesAny(f, snap.rules.exempt))
1091  const hashes = await hashFiles($, now.filter(f => before.has(f)))
1092  // A file dirty before the turn counts only when its content provably changed.
1093  const fresh = now.filter(f => !before.has(f) || (hashes[f] && before.get(f) && hashes[f] !== before.get(f))).slice(0, 40)
1094  for (const rel of fresh) {
1095    const text = (await readText($, rel)) ?? ''
1096    const c = classify(rel, await read($, snapshotA) ?? snap, await read($, ledgerA), currentTask(snap, await read($, activeA))?.id ?? null, text)
1097    await afterWrite($, rel, c, text)
1098  }
1099}
1100
1101/**
1102 * The ask tool: where a person is there, the native dialog answers in the same turn and the run goes on. A question
1103 * that blocks only some stories becomes a decision in the queue, and the rest of the work goes on; otherwise the
1104 * autopilot waits once the turn ends.
1105 */
1106async function ask($: $, input: Record<string, unknown>): Promise<string> {
1107  const question = str(input.question) || 'a decision only you can make'
1108  const options = Array.isArray(input.options) ? input.options.filter((o): o is string => typeof o === 'string').slice(0, 4) : []
1109  const blocks = Array.isArray(input.blocks) ? input.blocks.filter((b): b is string => typeof b === 'string') : []
1110  const ap = await read($, autopilotA)
1111  if (interactive) {
1112    try {
1113      const answer = await $.ui.ask(question.endsWith('?') ? question : `${question}?`, options.length >= 2 ? options : ['Decide it yourself, record it as an assumption', 'Stop and wait for me'])
1114      if (answer && answer !== 'Stop and wait for me') {
1115        await recordDecision($, question, options, blocks, answer)
1116        return `The person answered: ${answer}. Go on with it${ap.on ? '; the autopilot keeps running' : ''}.`
1117      }
1118    } catch {
1119      // Dismissed, or nobody there to ask: the question waits for the end of the turn.
1120    }
1121  }
1122  if (ap.on && blocks.length) {
1123    await recordDecision($, question, options, blocks, null)
1124    return `Recorded for the person: "${question}". Leave ${blocks.join(', ')} alone and go on with the work that does not depend on it.`
1125  }
1126  await update($, askedA, () => question)
1127  return ap.on ? 'The autopilot will wait for the person. Put the question in your answer and end your turn.' : 'Noted. Put the question in your answer.'
1128}
1129
1130async function recordDecision($: $, question: string, options: string[], blocks: string[], answer: string | null): Promise<void> {
1131  const when = await stamp($)
1132  await update($, ledgerA, l => ({ ...l, decisions: [...l.decisions, { id: `D${l.decisions.length + 1}`, question, options, blocks, at: when, answer }].slice(-50) }))
1133  await persist($)
1134}
1135
1136/** The person answers a queued decision in the pane: Claude reads it on the next step. */
1137async function answerDecision($: $, id: string, answer: string): Promise<void> {
1138  await update($, ledgerA, l => ({ ...l, decisions: l.decisions.map(d => (d.id === id ? { ...d, answer } : d)) }))
1139  await persist($)
1140  const d = (await read($, ledgerA)).decisions.find(x => x.id === id)
1141  if (d) void $.prompt.submit({ text: `Decision ${d.id} (${d.question}): ${answer}. Apply it to ${d.blocks.join(', ') || 'the work it blocked'}.`, asUser: true })
1142}
1143
1144async function statusText($: $): Promise<string> {
1145  const snap = await read($, snapshotA)
1146  if (!snap) return 'Spec X-Ref has not read this project yet.'
1147  const ledger = await read($, ledgerA)
1148  const next = nextStep(snap, ledger)
1149  if (!snap.featureDir) return [next.why, next.needsUser ?? (next.command ? `Next: ${next.command}` : '')].filter(Boolean).join(' ')
1150  // The host already names the plugin in front of a command's answer.
1151  return (turnContext(snap, ledger, await read($, activeA), stepLine(next)) ?? '').replace(/^speckit-xref · /, '')
1152}
1153
1154/** The workflow state for the model: where the project stands in Spec Kit and what comes next. */
1155async function workflowStatus($: $): Promise<string> {
1156  await scan($)
1157  const snap = await read($, snapshotA)
1158  if (!snap) return 'Spec X-Ref has not read this project yet.'
1159  const ledger = await read($, ledgerA)
1160  const next = nextStep(snap, ledger)
1161  const report = evaluate(snap, ledger)
1162  const ap = await read($, autopilotA)
1163  const lines = [
1164    `Project root: ${root} (every path below is relative to it)`,
1165    `Spec Kit CLI: ${snap.tools.specify ? 'specify is installed' : snap.tools.uvx ? 'not installed; runs through uvx' : 'missing, and no uvx either'}`,
1166    `Spec Kit: ${snap.initialized ? `set up${snap.speckitVersion ? ` (${snap.speckitVersion})` : ''}` : 'not set up'}${snap.initialized ? ` · Claude Code integration ${snap.claudeIntegration ? `yes, commands as ${speckitCommand(snap, 'plan')}` : 'missing'}` : ''} · extensions: ${snap.extensions.join(', ') || 'none'}`,
1167    `Autopilot: ${ap.on ? (ap.paused ? `on, waiting: ${ap.paused}` : `on, step ${ap.steps}/${ap.max}`) : 'off (/xref auto on)'}`,
1168    ...(snap.initialized ? [] : [`Folder: ${snap.folder === 'empty' ? 'empty (only dotfiles, README, license)' : 'existing code'}`]),
1169    `Constitution: ${snap.constitution?.principles.length ? `${snap.constitution.principles.length} principles, ${snap.constitution.musts.length} MUST rules` : 'missing or still the template'}`,
1170    `Active feature: ${snap.featureDir ? `${snap.featureDir}${snap.spec ? ` (${snap.spec.title})` : ''}` : 'none'}`,
1171  ]
1172  if (snap.features.length > 1) lines.push(`All features: ${snap.features.join(', ')} (switch by writing {"feature_directory": "<dir>"} to .specify/feature.json)`)
1173  if (snap.featureDir) {
1174    lines.push(`Artifacts: spec.md ${snap.spec ? 'yes' : 'no'} · plan.md ${snap.hasPlan ? 'yes' : 'no'} · tasks.md ${snap.tasks.length ? `${report.done}/${report.tasks} done` : 'no'} · FR covered ${report.covered}/${report.total} · drift ${report.level}`)
1175  }
1176  lines.push(`Phase: ${next.phase}. ${next.why}`, `Next: ${next.command ?? next.needsUser ?? ''}`)
1177  if (next.needsUser) lines.push(`Needs the person: ${next.needsUser}`)
1178  if (next.notes.length) lines.push('Notes:', ...next.notes.map(n => `- ${n}`))
1179  return lines.join('\n')
1180}
1181
1182const str = (value: unknown) => (typeof value === 'string' ? value : '')
1183
1184/** The answer of one of the mod's own tools when its hook failed. */
1185function failed() {
1186  return { result: 'speckit-xref could not answer this call; claude --debug has the reason.' }
1187}
1188
1189/** A failed write guard: in strict mode it refuses rather than letting an unchecked edit through. */
1190function guardFailed<E, R>($: $, e: E, next: ((e: E) => R) & { readonly called: boolean }): R | { deny: string } {
1191  if (strict && !next.called) return { deny: 'speckit-xref (strict) could not check this edit against the plan; try again, or switch strict mode off.' }
1192  return next(e)
1193}
1194
1195/** What the autopilot never does unattended: these wait for the person (docs: D6, tighten-only). */
1196const DESTRUCTIVE: [RegExp, string][] = [
1197  [/\bgit\s+push\b/, 'pushing'],
1198  [/\bgit\s+reset\s+--hard\b/, 'git reset --hard'],
1199  [/\bgit\s+clean\s+-[a-z]*f/, 'git clean -f'],
1200  [/\bgit\s+(?:checkout|restore)\s+(?:--\s+)?\.(?:\s|$)/, 'discarding every change'],
hooks/ledger.ts 185 lines
1// The ledger's two files (docs/contract-0.4.md §1): what is committed and reviewed, what stays on this machine.
2
3import type { Ledger, Semantic, Unplanned } from '../types'
4
5export const emptyLedger = (): Ledger => ({
6  tasks: {},
7  requirements: {},
8  unplanned: [],
9  intents: [],
10  anchors: [],
11  semantic: null,
12  semanticHistory: [],
13  fingerprints: {},
14  verification: {},
15  approvals: {},
16  decisions: [],
17  checkpoints: {},
18  extra: { committed: {}, local: {} },
19})
20
21const COMMITTED_KEYS = new Set(['schema_version', 'feature', 'map', 'links', 'accepted', 'fingerprints'])
22const LOCAL_KEYS = new Set(['schema_version', 'feature', 'touched', 'unplanned', 'intents', 'semantic', 'semanticHistory', 'verification', 'approvals', 'decisions', 'anchors', 'checkpoints'])
23const MAX_HISTORY = 10
24
25/** Where the local half lives: `.specify/xref/local/<feature-slug>.json`. */
26export const localPath = (featureDir: string) => `.specify/xref/local/${featureDir.slice(featureDir.lastIndexOf('/') + 1)}.json`
27export const runLogPath = (featureDir: string) => `.specify/xref/local/${featureDir.slice(featureDir.lastIndexOf('/') + 1)}.run.jsonl`
28export const LOCAL_IGNORE = { path: '.specify/xref/.gitignore', text: 'local/\n' }
29
30type Raw = Record<string, unknown>
31const obj = (v: unknown): Raw => (v && typeof v === 'object' && !Array.isArray(v) ? (v as Raw) : {})
32const arr = <T>(v: unknown): T[] => (Array.isArray(v) ? (v as T[]) : [])
33const strs = (v: unknown) => arr<unknown>(v).filter((x): x is string => typeof x === 'string')
34const parse = (text: string | null): Raw | null => {
35  if (!text) return null
36  try {
37    return obj(JSON.parse(text))
38  } catch {
39    return null
40  }
41}
42const others = (raw: Raw, known: Set<string>) => Object.fromEntries(Object.entries(raw).filter(([k]) => !known.has(k)))
43
44/** A v1 ledger (one file, `schema_version` 1 or none) as the two v2 halves. */
45export function migrateV1(v1: Raw): { committed: Raw; local: Raw } {
46  const map: Record<string, string[]> = {}
47  const links: Record<string, string[]> = {}
48  const add = (file: string, id: string) => (links[file] = [...new Set([...(links[file] ?? []), id])].sort())
49  for (const [id, r] of Object.entries(obj(v1.requirements))) {
50    const entry = obj(r)
51    const tasks = strs(entry.tasks)
52    if (tasks.length) map[id] = [...new Set(tasks)].sort()
53    for (const file of strs(entry.files)) add(file, id)
54  }
55  const touched: Record<string, string[]> = {}
56  for (const [id, t] of Object.entries(obj(v1.tasks))) {
57    const entry = obj(t)
58    for (const file of strs(entry.linked)) add(file, id)
59    if (strs(entry.touched).length) touched[id] = strs(entry.touched)
60  }
61  const unplanned = arr<Raw>(v1.unplanned)
62  return {
63    committed: {
64      schema_version: 2,
65      feature: v1.feature ?? null,
66      map,
67      links,
68      accepted: [...new Set(unplanned.filter(u => u.acknowledged === true).map(u => String(u.file)))].sort(),
69      fingerprints: {},
70    },
71    local: {
72      schema_version: 2,
73      feature: v1.feature ?? null,
74      touched,
75      unplanned: unplanned.filter(u => u.acknowledged !== true).map(u => ({ file: u.file, at: u.at, task: u.task ?? null })),
76      intents: arr(v1.intents),
77      semantic: v1.semantic ?? null,
78      semanticHistory: [],
79      verification: {},
80      approvals: {},
81      decisions: [],
82      anchors: arr(v1.anchors),
83    },
84  }
85}
86
87/** The in-memory ledger from both files; a v1 file is migrated on the way in, the next write makes it v2. */
88export function ledgerFromParts(committedText: string | null, localText: string | null): Ledger {
89  let committed = parse(committedText) ?? {}
90  let local = parse(localText) ?? {}
91  if (committedText && committed.schema_version !== 2) {
92    const migrated = migrateV1(committed)
93    committed = migrated.committed
94    // A local file next to a v1 ledger is newer than the migration; it wins where it says something.
95    local = { ...migrated.local, ...local }
96  }
97  const ledger = emptyLedger()
98  for (const [id, tasks] of Object.entries(obj(committed.map))) {
99    if (strs(tasks).length) ledger.requirements[id] = { tasks: strs(tasks), files: [], sources: ['map'] }
100  }
101  for (const [file, ids] of Object.entries(obj(committed.links))) {
102    for (const id of strs(ids)) {
103      if (/^T-?\d{3,}$/.test(id)) {
104        const entry = (ledger.tasks[id] ??= { touched: [], linked: [] })
105        if (!entry.linked.includes(file)) entry.linked.push(file)
106      } else {
107        const entry = (ledger.requirements[id] ??= { tasks: [], files: [], sources: [] })
108        if (!entry.files.includes(file)) entry.files.push(file)
109        if (!entry.sources.includes('link')) entry.sources.push('link')
110      }
111    }
112  }
113  for (const [id, files] of Object.entries(obj(local.touched))) {
114    const entry = (ledger.tasks[id] ??= { touched: [], linked: [] })
115    entry.touched = strs(files)
116  }
117  ledger.unplanned = [
118    ...arr<Raw>(local.unplanned).map((u): Unplanned => ({ file: String(u.file), at: String(u.at ?? ''), task: typeof u.task === 'string' ? u.task : null, acknowledged: false })),
119    ...strs(committed.accepted).map((file): Unplanned => ({ file, at: '', task: null, acknowledged: true })),
120  ]
121  ledger.fingerprints = Object.fromEntries(Object.entries(obj(committed.fingerprints)).filter(([, v]) => typeof v === 'string')) as Record<string, string>
122  ledger.intents = arr(local.intents)
123  ledger.anchors = arr(local.anchors)
124  ledger.semantic = (local.semantic as Semantic | null) ?? null
125  ledger.semanticHistory = arr(local.semanticHistory)
126  ledger.verification = obj(local.verification) as Ledger['verification']
127  ledger.approvals = obj(local.approvals) as Ledger['approvals']
128  ledger.decisions = arr(local.decisions)
129  ledger.checkpoints = obj(local.checkpoints) as Ledger['checkpoints']
130  ledger.extra = { committed: others(committed, COMMITTED_KEYS), local: others(local, LOCAL_KEYS) }
131  return ledger
132}
133
134/** JSON with every object's keys sorted, so the committed file is the same bytes for the same state. */
135function sortedJson(value: unknown): unknown {
136  if (Array.isArray(value)) return value.map(sortedJson)
137  if (value && typeof value === 'object') return Object.fromEntries(Object.keys(value as Raw).sort().map(k => [k, sortedJson((value as Raw)[k])]))
138  return value
139}
140const uniqSorted = (xs: string[]) => [...new Set(xs)].sort()
141
142/** Both files' text for a ledger: the committed half deterministic and without timestamps. */
143export function ledgerToParts(ledger: Ledger, featureDir: string): { committed: string; local: string } {
144  const map: Record<string, string[]> = {}
145  const links: Record<string, string[]> = {}
146  const add = (file: string, id: string) => (links[file] = [...(links[file] ?? []), id])
147  for (const [id, r] of Object.entries(ledger.requirements)) {
148    if (r.tasks.length) map[id] = uniqSorted(r.tasks)
149    for (const file of r.files) add(file, id)
150  }
151  const touched: Record<string, string[]> = {}
152  for (const [id, t] of Object.entries(ledger.tasks)) {
153    for (const file of t.linked) add(file, id)
154    if (t.touched.length) touched[id] = t.touched
155  }
156  for (const file of Object.keys(links)) links[file] = uniqSorted(links[file]!)
157  // An accepted file that was linked since is the link's, not an acceptance.
158  const accepted = uniqSorted(ledger.unplanned.filter(u => u.acknowledged && !links[u.file]).map(u => u.file))
159  const committed = sortedJson({
160    ...ledger.extra.committed,
161    schema_version: 2,
162    feature: featureDir,
163    map,
164    links,
165    accepted,
166    fingerprints: ledger.fingerprints,
167  })
168  const local = {
169    ...ledger.extra.local,
170    schema_version: 2,
171    feature: featureDir,
172    touched,
173    unplanned: ledger.unplanned.filter(u => !u.acknowledged).map(u => ({ file: u.file, at: u.at, task: u.task })),
174    intents: ledger.intents,
175    semantic: ledger.semantic,
176    semanticHistory: ledger.semanticHistory.slice(-MAX_HISTORY),
177    verification: ledger.verification,
178    approvals: ledger.approvals,
179    decisions: ledger.decisions,
180    anchors: ledger.anchors,
181    checkpoints: ledger.checkpoints,
182  }
183  return { committed: JSON.stringify(committed, null, 2) + '\n', local: JSON.stringify(local, null, 2) + '\n' }
184}
185
hooks/proof.ts 112 lines
1// Proof instead of claims (docs/contract-0.4.md §4, §9): how far each requirement got, and what the tests said.
2
3import type { Anchor, Ledger, Req, Task, Verification } from '../types'
4import { fingerprint, isTestFile, localId, reqStatus } from './rules'
5
6export const LADDER = ['specified', 'planned', 'implemented', 'tested', 'passing'] as const
7export type Rung = (typeof LADDER)[number]
8
9/** What the ladder reads off the snapshot; `realTests` are the test files with real tests in them. */
10export type ProofSnap = { featureDir: string | null; tasks: Pick<Task, 'id' | 'done' | 'reqs'>[]; reqs: Pick<Req, 'id' | 'text'>[]; realTests: string[] }
11
12export const isActive = (req: Pick<Req, 'text'> & { status?: Req['status'] }) => (req.status ?? reqStatus(req.text).status) === 'active'
13
14/** The tasks that own a requirement: those naming it, and those mapped to it. */
15export function owners(id: string, snap: ProofSnap, ledger: Ledger): Pick<Task, 'id' | 'done' | 'reqs'>[] {
16  const mapped = new Set(ledger.requirements[id]?.tasks ?? [])
17  return snap.tasks.filter(t => t.reqs.includes(id) || mapped.has(t.id))
18}
19
20/** Whether a stored fingerprint says the requirement's text changed since it was last proven. */
21export function isStale(id: string, snap: ProofSnap, ledger: Ledger): boolean {
22  const stored = ledger.fingerprints[id]
23  const req = snap.reqs.find(r => r.id === id)
24  return !!stored && !!req && stored !== fingerprint(req.text)
25}
26
27export function levelOf(id: string, snap: ProofSnap, ledger: Ledger): Rung {
28  const level = rawLevel(id, snap, ledger)
29  // A changed requirement needs new proof: whatever it had, it is planned at most.
30  return isStale(id, snap, ledger) && LADDER.indexOf(level) > 1 ? 'planned' : level
31}
32
33function rawLevel(id: string, snap: ProofSnap, ledger: Ledger): Rung {
34  const own = owners(id, snap, ledger)
35  if (!own.length) return 'specified'
36  const anchored = (a: Anchor) => localId(a.id, snap.featureDir) === id
37  const linked = ledger.requirements[id]?.files ?? []
38  const touched = own.flatMap(t => ledger.tasks[t.id]?.touched ?? [])
39  const code = [...ledger.anchors.filter(anchored).map(a => a.file), ...linked, ...touched].some(f => !isTestFile(f))
40  if (!own.some(t => t.done) || !code) return 'planned'
41  const real = new Set(snap.realTests)
42  const tested = [...ledger.anchors.filter(anchored).map(a => a.file), ...linked].some(f => isTestFile(f) && real.has(f))
43  if (!tested) return 'implemented'
44  const req = snap.reqs.find(r => r.id === id)
45  const v = ledger.verification[id]
46  return v?.status === 'passing' && req && v.fingerprint === fingerprint(req.text) ? 'passing' : 'tested'
47}
48
49/** How many active functional requirements stand on each rung. */
50export function levels(snap: ProofSnap & { reqs: (Pick<Req, 'id' | 'text'> & { kind?: string; status?: Req['status'] })[] }, ledger: Ledger): Record<Rung, number> {
51  const counts = Object.fromEntries(LADDER.map(r => [r, 0])) as Record<Rung, number>
52  for (const req of snap.reqs) if ((req.kind ?? 'FR') === 'FR' && isActive(req)) counts[levelOf(req.id, snap, ledger)] += 1
53  return counts
54}
55
56// ---- Test results (§9) ----
57
58export type JunitCase = { name: string; classname: string; file: string; status: 'passed' | 'failed' | 'skipped' }
59
60const unescape = (s: string) => s.replace(/&lt;/g, '<').replace(/&gt;/g, '>').replace(/&quot;/g, '"').replace(/&apos;/g, "'").replace(/&amp;/g, '&')
61const attr = (tag: string, name: string) => {
62  const m = new RegExp(`\\s${name}\\s*=\\s*("([^"]*)"|'([^']*)')`).exec(tag)
63  return m ? unescape(m[2] ?? m[3] ?? '') : ''
64}
65
66export function parseJunit(xml: string): JunitCase[] {
67  const cases: JunitCase[] = []
68  const re = /<testcase\b([^>]*?)(\/>|>([\s\S]*?)<\/testcase>)/g
69  for (const m of xml.matchAll(re)) {
70    const tag = m[1] ?? ''
71    const body = m[3] ?? ''
72    const status = /<(failure|error)\b/.test(body) ? 'failed' : /<skipped\b/.test(body) ? 'skipped' : 'passed'
73    cases.push({ name: attr(tag, 'name'), classname: attr(tag, 'classname'), file: attr(tag, 'file'), status })
74  }
75  return cases
76}
77
78const IDS = /\b(?:(?:FR|SC)-\d{3,}|US\d+-AS\d+)\b/g
79
80export function verificationFrom(cases: JunitCase[], anchors: Anchor[], fingerprints: Record<string, string>, featureDir: string | null, at: string, commit: string): Record<string, Verification> {
81  const byId = new Map<string, { failed: boolean; passed: boolean; tests: Set<string> }>()
82  for (const c of cases) {
83    if (c.status === 'skipped') continue
84    const ids = new Set([...`${c.name} ${c.classname}`.matchAll(IDS)].map(m => m[0]))
85    for (const a of anchors) {
86      const id = a.file === c.file && c.file ? localId(a.id, featureDir) : null
87      if (id) ids.add(id)
88    }
89    const label = c.classname ? `${c.classname}.${c.name}` : c.name
90    for (const id of ids) {
91      const entry = byId.get(id) ?? { failed: false, passed: false, tests: new Set<string>() }
92      if (c.status === 'failed') entry.failed = true
93      else entry.passed = true
94      entry.tests.add(label)
95      byId.set(id, entry)
96    }
97  }
98  const out: Record<string, Verification> = {}
99  for (const [id, e] of byId) {
100    out[id] = { status: e.failed ? 'failing' : 'passing', tests: [...e.tests].sort(byCodePoint).slice(0, 10), at, commit, fingerprint: fingerprints[id] ?? '' }
101  }
102  return out
103}
104
105const byCodePoint = (a: string, b: string) => (a < b ? -1 : a > b ? 1 : 0)
106
107/** Without JUnit: a suite that exits 0 proves every requirement that has a real test; a failing one proves nothing. */
108export function verificationFromExit(exitCode: number, ids: string[], fingerprints: Record<string, string>, at: string, commit: string): Record<string, Verification> {
109  if (exitCode !== 0) return {}
110  return Object.fromEntries(ids.map(id => [id, { status: 'passing' as const, tests: ['<suite>'], at, commit, fingerprint: fingerprints[id] ?? '' }]))
111}
112
hooks/rules.ts 167 lines
1// Shared rules of docs/contract-0.4.md that both the mod and the extension implement: fingerprints, path globs,
2// what never counts as drift, what a test file is, which feature a branch names, a requirement's status.
3
4// ---- Fingerprints (§6) ----
5
6/** The bytes of a string in UTF-8, without TextEncoder (the mod's sandbox has none). */
7function utf8(text: string): number[] {
8  const out: number[] = []
9  for (const ch of text) {
10    const c = ch.codePointAt(0) ?? 0
11    if (c < 0x80) out.push(c)
12    else if (c < 0x800) out.push(0xc0 | (c >> 6), 0x80 | (c & 63))
13    else if (c < 0x10000) out.push(0xe0 | (c >> 12), 0x80 | ((c >> 6) & 63), 0x80 | (c & 63))
14    else out.push(0xf0 | (c >> 18), 0x80 | ((c >> 12) & 63), 0x80 | ((c >> 6) & 63), 0x80 | (c & 63))
15  }
16  return out
17}
18
19const K = [
20  0x428a2f98, 0x71374491, 0xb5c0fbcf, 0xe9b5dba5, 0x3956c25b, 0x59f111f1, 0x923f82a4, 0xab1c5ed5, 0xd807aa98, 0x12835b01, 0x243185be, 0x550c7dc3, 0x72be5d74, 0x80deb1fe, 0x9bdc06a7, 0xc19bf174,
21  0xe49b69c1, 0xefbe4786, 0x0fc19dc6, 0x240ca1cc, 0x2de92c6f, 0x4a7484aa, 0x5cb0a9dc, 0x76f988da, 0x983e5152, 0xa831c66d, 0xb00327c8, 0xbf597fc7, 0xc6e00bf3, 0xd5a79147, 0x06ca6351, 0x14292967,
22  0x27b70a85, 0x2e1b2138, 0x4d2c6dfc, 0x53380d13, 0x650a7354, 0x766a0abb, 0x81c2c92e, 0x92722c85, 0xa2bfe8a1, 0xa81a664b, 0xc24b8b70, 0xc76c51a3, 0xd192e819, 0xd6990624, 0xf40e3585, 0x106aa070,
23  0x19a4c116, 0x1e376c08, 0x2748774c, 0x34b0bcb5, 0x391c0cb3, 0x4ed8aa4a, 0x5b9cca4f, 0x682e6ff3, 0x748f82ee, 0x78a5636f, 0x84c87814, 0x8cc70208, 0x90befffa, 0xa4506ceb, 0xbef9a3f7, 0xc67178f2,
24]
25
26/** SHA-256 of a string's UTF-8, as lower-case hex. Synchronous, so pure code can fingerprint. */
27export function sha256(text: string): string {
28  const bytes = utf8(text)
29  const bitLength = bytes.length * 8
30  bytes.push(0x80)
31  while (bytes.length % 64 !== 56) bytes.push(0)
32  for (let i = 7; i >= 0; i--) bytes.push(i >= 4 ? 0 : (bitLength >>> (i * 8)) & 0xff)
33  const h = [0x6a09e667, 0xbb67ae85, 0x3c6ef372, 0xa54ff53a, 0x510e527f, 0x9b05688c, 0x1f83d9ab, 0x5be0cd19]
34  const w = new Array<number>(64)
35  const rotr = (x: number, n: number) => (x >>> n) | (x << (32 - n))
36  for (let off = 0; off < bytes.length; off += 64) {
37    for (let i = 0; i < 16; i++) w[i] = (bytes[off + i * 4]! << 24) | (bytes[off + i * 4 + 1]! << 16) | (bytes[off + i * 4 + 2]! << 8) | bytes[off + i * 4 + 3]!
38    for (let i = 16; i < 64; i++) {
39      const s0 = rotr(w[i - 15]!, 7) ^ rotr(w[i - 15]!, 18) ^ (w[i - 15]! >>> 3)
40      const s1 = rotr(w[i - 2]!, 17) ^ rotr(w[i - 2]!, 19) ^ (w[i - 2]! >>> 10)
41      w[i] = (w[i - 16]! + s0 + w[i - 7]! + s1) | 0
42    }
43    let [a, b, c, d, e, f, g, hh] = h as [number, number, number, number, number, number, number, number]
44    for (let i = 0; i < 64; i++) {
45      const t1 = (hh + (rotr(e, 6) ^ rotr(e, 11) ^ rotr(e, 25)) + ((e & f) ^ (~e & g)) + K[i]! + w[i]!) | 0
46      const t2 = ((rotr(a, 2) ^ rotr(a, 13) ^ rotr(a, 22)) + ((a & b) ^ (a & c) ^ (b & c))) | 0
47      hh = g
48      g = f
49      f = e
50      e = (d + t1) | 0
51      d = c
52      c = b
53      b = a
54      a = (t1 + t2) | 0
55    }
56    ;[a, b, c, d, e, f, g, hh].forEach((v, i) => (h[i] = (h[i]! + v) | 0))
57  }
58  return h.map(v => (v >>> 0).toString(16).padStart(8, '0')).join('')
59}
60
61export function normalizeText(text: string): string {
62  const nfkc = typeof text.normalize === 'function' ? text.normalize('NFKC') : text
63  return nfkc.toLowerCase().replace(/\s+/g, ' ').trim().replace(/[.;:!]+$/, '').trim()
64}
65
66export const fingerprint = (text: string) => sha256(normalizeText(text))
67
68/** What the spec gate keys on: the person's words, every requirement and every scenario. */
69export function specFingerprint(spec: { input: string | null; reqs: { text: string }[]; stories: { scenarios?: { text: string }[] }[] }): string {
70  const parts = [spec.input ?? '', ...spec.reqs.map(r => r.text), ...spec.stories.flatMap(s => (s.scenarios ?? []).map(x => x.text))]
71  return fingerprint(parts.join('\n'))
72}
73
74// ---- Path rules (§2) ----
75
76export const DEFAULT_EXEMPT = [
77  'package-lock.json', 'yarn.lock', 'pnpm-lock.yaml', 'bun.lock', 'bun.lockb', 'Cargo.lock', 'poetry.lock', 'uv.lock', 'Gemfile.lock', 'go.sum', 'composer.lock', 'Podfile.lock', 'pubspec.lock',
78  'node_modules/', '/dist/', '/build/', '/coverage/', '/.next/', '/out/', '/target/', '__snapshots__/', '*.min.js', '*.map', '*.generated.*', '*.g.dart',
79]
80export const DEFAULT_UNCLEAR = [
81  'package.json', 'pyproject.toml', 'Cargo.toml', 'go.mod', 'pubspec.yaml', 'Gemfile', 'requirements*.txt', '*.config.*', 'tsconfig*.json', '.eslintrc*', '.prettierrc*', 'Dockerfile', 'docker-compose*.yml',
82  '/.github/workflows/', '.env.example', '*.md',
83]
84const TEST_GLOBS = ['*.test.*', '*.spec.*', 'test_*.py', '*_test.py', '*_test.go', '*_test.dart', '/tests/', '/test/', '__tests__/']
85
86const SPECIAL = /[\\^$.|+(){}[\]]/
87
88export function globToRegExp(pattern: string): RegExp {
89  let p = pattern.trim().replace(/^\.\//, '')
90  const folder = p.endsWith('/')
91  if (folder) p = p.slice(0, -1)
92  if (p.startsWith('/')) p = p.slice(1)
93  else if (!p.includes('/')) p = `**/${p}`
94  if (folder) p += '/**'
95  let re = ''
96  for (let i = 0; i < p.length; ) {
97    if (p.startsWith('**/', i)) {
98      re += '(?:.*/)?'
99      i += 3
100    } else if (p.startsWith('/**', i) && i + 3 === p.length) {
101      re += '(?:/.*)?'
102      i += 3
103    } else if (p.startsWith('**', i)) {
104      re += '.*'
105      i += 2
106    } else {
107      const ch = p[i]!
108      re += ch === '*' ? '[^/]*' : ch === '?' ? '[^/]' : SPECIAL.test(ch) ? `\\${ch}` : ch
109      i += 1
110    }
111  }
112  return new RegExp(`^${re}$`)
113}
114
115export type Rules = { exempt: string[]; unclear: string[] }
116
117export function parseIgnore(text: string): Rules {
118  const rules: Rules = { exempt: [], unclear: [] }
119  for (const raw of text.split('\n')) {
120    const line = raw.trim()
121    if (!line || line.startsWith('#')) continue
122    const unclear = /^unclear:\s*(.+)$/.exec(line)
123    if (unclear) rules.unclear.push(unclear[1]!.trim())
124    else rules.exempt.push(line)
125  }
126  return rules
127}
128
129export function rulesFrom(ignoreText: string | null): Rules {
130  const own = ignoreText ? parseIgnore(ignoreText) : { exempt: [], unclear: [] }
131  return { exempt: [...DEFAULT_EXEMPT, ...own.exempt], unclear: [...DEFAULT_UNCLEAR, ...own.unclear] }
132}
133
134const compiled = new Map<string, RegExp>()
135export function matchesAny(rel: string, patterns: readonly string[]): boolean {
136  return patterns.some(p => {
137    let re = compiled.get(p)
138    if (!re) compiled.set(p, (re = globToRegExp(p)))
139    return re.test(rel)
140  })
141}
142
143export const isTestFile = (rel: string) => matchesAny(rel, TEST_GLOBS)
144export const hasRealTests = (text: string) => /\b(?:it|test|describe)\s*\(|\bdef test_|\bfunc Test[A-Z_]|\btestWidgets\s*\(/.test(text)
145
146// ---- Feature from the branch (§7) ----
147
148export function featureFromBranch(branch: string, dirs: readonly string[]): string | null {
149  const number = /(?:^|\/)(\d{3,})-/.exec(branch)?.[1]
150  if (!number) return null
151  const hits = dirs.filter(d => d.slice(d.lastIndexOf('/') + 1).startsWith(`${number}-`))
152  return hits.length === 1 ? hits[0]! : null
153}
154
155// ---- Requirement status (§11): in speckit.ts, which the parity script loads without the rest ----
156
157export { reqStatus } from './speckit'
158export type { ReqStatus } from './speckit'
159
160/** The bare id of an anchor or link (`001-auth/FR-003` → `FR-003`) when it belongs to `featureDir`, else null. */
161export function localId(id: string, featureDir: string | null): string | null {
162  const slash = id.lastIndexOf('/')
163  if (slash === -1) return id
164  const feature = id.slice(0, slash)
165  return featureDir && featureDir.slice(featureDir.lastIndexOf('/') + 1) === feature ? id.slice(slash + 1) : null
166}
167
hooks/speckit.ts 170 lines
1// Pure readers for GitHub Spec Kit artifacts. No IO: register.tsx reads the files and hands the text in.
2
3import type { Constitution, Scenario, Spec, Story, Req, Task } from '../types'
4
5/** A requirement's status (docs/contract-0.4.md §11): `SUPERSEDED by FR-006`, `RETIRED`, else active. */
6export type ReqStatus = { status: 'active' | 'superseded' | 'retired'; supersededBy: string | null }
7
8export function reqStatus(text: string): ReqStatus {
9  const by = /SUPERSEDED by ((?:FR|SC)-\d{3,})/.exec(text)
10  if (by) return { status: 'superseded', supersededBy: by[1]! }
11  if (/RETIRED/.test(text)) return { status: 'retired', supersededBy: null }
12  return { status: 'active', supersededBy: null }
13}
14
15const clip = (text: string, max: number) => (text.length > max ? text.slice(0, max - 1) + '…' : text)
16const unbold = (text: string) => text.replace(/\*\*/g, '').replace(/`/g, '').trim()
17
18/** The bullet lines under the first `##`/`###` heading whose text matches `heading`, up to the next heading. */
19function bulletsUnder(markdown: string, heading: RegExp, max: number): string[] {
20  const out: string[] = []
21  let inside = false
22  for (const line of markdown.split('\n')) {
23    const h = /^#{2,4}\s+(.*)$/.exec(line)
24    if (h) {
25      if (inside) break
26      inside = heading.test(h[1] ?? '')
27      continue
28    }
29    if (!inside) continue
30    const bullet = /^\s*[-*]\s+(.*)$/.exec(line)
31    if (bullet && bullet[1] && !/^\[.*\]$/.test(bullet[1].trim())) out.push(clip(unbold(bullet[1]), 200))
32    if (out.length >= max) break
33  }
34  return out
35}
36
37export function parseSpec(markdown: string): Spec {
38  const titleLine = /^#\s+(?:Feature Specification:\s*)?(.+)$/m.exec(markdown)
39  const inputLine = /^\*\*Input\*\*:\s*(?:User description:\s*)?(.+)$/m.exec(markdown)
40  const input = inputLine?.[1] ? inputLine[1].trim().replace(/^"(.*)"$/, '$1').trim() : null
41
42  const stories: Story[] = []
43  let story: Story | null = null
44  for (const line of markdown.split('\n')) {
45    const m = /^###\s+User Story\s+(\d+)\s*[-–—:]\s*(.+?)\s*(?:\(Priority:\s*(P\d+)\))?\s*(?:🎯.*)?$/.exec(line)
46    if (m) {
47      story = { id: `US${m[1]}`, title: unbold(m[2] ?? ''), priority: m[3] ?? null, scenarios: [] }
48      stories.push(story)
49      continue
50    }
51    if (/^#{1,3}\s/.test(line)) {
52      story = null
53      continue
54    }
55    // A numbered Given/When/Then line is scenario USn-AS<k>: the id an anchor or a test can name.
56    const scenario = story ? /^\s*(\d+)\.\s+(\*\*Given\*\*.*)$/.exec(line) : null
57    if (story && scenario) story.scenarios.push({ id: `${story.id}-AS${Number(scenario[1])}`, text: clip(unbold(scenario[2] ?? ''), 300) } satisfies Scenario)
58  }
59
60  const reqs: Req[] = []
61  const seen = new Set<string>()
62  for (const m of markdown.matchAll(/^\s*[-*]\s*\*\*((FR|SC)-\d{3,})\*\*\s*:?\s*(.*)$/gm)) {
63    const id = m[1] ?? ''
64    if (seen.has(id)) continue
65    seen.add(id)
66    const text = m[3] ?? ''
67    const clipped = clip(unbold(text), 300)
68    reqs.push({ id, kind: m[2] as 'FR' | 'SC', text: clipped, needsClarification: /NEEDS CLARIFICATION/i.test(text), ...reqStatus(unbold(text)) })
69  }
70
71  return {
72    title: titleLine?.[1]?.trim() ?? 'Untitled feature',
73    input: input && !/^\$ARGUMENTS$/.test(input) ? clip(input, 600) : null,
74    stories,
75    reqs,
76    outOfScope: bulletsUnder(markdown, /out of scope|non-goals?|not in scope/i, 8),
77    assumptions: bulletsUnder(markdown, /^assumptions/i, 6),
78  }
79}
80
81const PATH_EXT = /\.[a-z0-9]{1,6}$/i
82const SPEC_DOCS = /^(spec|plan|tasks|research|data-model|quickstart|constitution|checklist)\.md$/i
83
84/** File and directory paths named in a task's description: backticked, or bare tokens that look like paths. */
85export function extractPaths(text: string): string[] {
86  const found = new Set<string>()
87  const keep = (raw: string) => {
88    const path = raw.trim().replace(/^\.\//, '').replace(/[.,;:)]+$/, '')
89    if (!path || /^https?:/i.test(path) || path.includes(' ') || path.length > 200) return
90    // "as plan.md says" names Spec Kit's own documents, not a file the task writes.
91    if (SPEC_DOCS.test(path)) return
92    if (path.includes('/') || PATH_EXT.test(path)) found.add(path)
93  }
94  for (const m of text.matchAll(/`([^`]+)`/g)) keep(m[1] ?? '')
95  const bare = text.replace(/`[^`]*`/g, ' ')
96  for (const m of bare.matchAll(/(?:^|[\s("'])((?:\.\/)?(?:[\w@.-]+\/)+[\w@.\[\]-]*|[\w-]+\.[a-z][a-z0-9]{0,5})(?=[\s,;:)"']|\.(?:\s|$)|$)/gi)) {
97    const token = m[1] ?? ''
98    // "e.g." or a version like "v1.2" is no file; a path names a folder or a known-looking extension.
99    if (/^(e\.g|i\.e|etc|vs)\.?$/i.test(token) || /^v?\d+(\.\d+)+$/.test(token)) continue
100    keep(token)
101  }
102  return [...found]
103}
104
105export function parseTasks(markdown: string): Task[] {
106  const tasks: Task[] = []
107  let phase = ''
108  for (const line of markdown.split('\n')) {
109    const heading = /^##\s+(.+)$/.exec(line)
110    if (heading) {
111      phase = unbold(heading[1] ?? '')
112      continue
113    }
114    const m = /^\s*[-*]\s*\[( |x|X)\]\s*(?:\*\*)?(T\d{3,})(?:\*\*)?\s*(.*)$/.exec(line)
115    if (!m) continue
116    let rest = m[3] ?? ''
117    const parallel = /\[P\]/.test(rest)
118    const story = /\[(US\d+)\]/.exec(rest)?.[1] ?? null
119    rest = rest.replace(/\[(P|US\d+)\]\s*/g, '').trim()
120    tasks.push({
121      id: m[2] ?? '',
122      done: m[1] !== ' ',
123      parallel,
124      story,
125      text: clip(rest, 240),
126      paths: extractPaths(rest),
127      reqs: [...new Set(rest.match(/\b(?:FR|SC)-\d{3,}\b/g) ?? [])],
128      phase,
129    })
130  }
131  return tasks
132}
133
134export function parseConstitution(markdown: string): Constitution {
135  const principles: string[] = []
136  const musts: string[] = []
137  for (const line of markdown.split('\n')) {
138    const h = /^###\s+(.+)$/.exec(line)
139    // A template never filled in still reads [PRINCIPLE_1_NAME]; it says nothing yet.
140    if (h && h[1] && !/\[[A-Z0-9_]+\]/.test(h[1])) principles.push(clip(unbold(h[1]), 80))
141    if (/\bMUST\b/.test(line) && !/\[[A-Z0-9_]+\]/.test(line) && !/^#/.test(line)) {
142      const text = unbold(line.replace(/^\s*[-*]\s*/, ''))
143      if (text) musts.push(clip(text, 200))
144    }
145  }
146  return { principles: principles.slice(0, 12), musts: musts.slice(0, 12) }
147}
148
149/** `.specify/feature.json` names the active feature's directory. */
150export function parseFeatureJson(text: string): string | null {
151  try {
152    const value = JSON.parse(text) as { feature_directory?: unknown }
153    return typeof value.feature_directory === 'string' && value.feature_directory ? value.feature_directory : null
154  } catch {
155    return null
156  }
157}
158
159/** The separator `.specify/integration.json` records for the active integration: `-` for skills (`/speckit-plan`), `.` for commands. */
160export function invokeSeparator(text: string): string | null {
161  try {
162    const value = JSON.parse(text) as { integration?: unknown; integration_settings?: Record<string, { invoke_separator?: unknown }> }
163    const name = typeof value.integration === 'string' ? value.integration : null
164    const separator = name ? value.integration_settings?.[name]?.invoke_separator : undefined
165    return typeof separator === 'string' ? separator : null
166  } catch {
167    return null
168  }
169}
170
hooks/workflow.ts 302 lines
1// Where a project stands in Spec Kit's workflow, the step that comes next, and the autopilot's rules. Pure: register.tsx hands in the snapshot.
2
3import type { Autopilot, Ledger, Snapshot, Task } from '../types'
4import { fingerprint, specFingerprint } from './rules'
5import { evaluate, speckitCommand } from './xref'
6
7export type Phase =
8  | 'setup'
9  | 'integration'
10  | 'constitution'
11  | 'specify'
12  | 'clarify'
13  | 'review'
14  | 'plan'
15  | 'tasks'
16  | 'revise'
17  | 'map'
18  | 'analyze'
19  | 'checklists'
20  | 'implement'
21  | 'repair'
22  | 'converge'
23  | 'verify'
24export type Step = {
25  phase: Phase
26  command: string | null
27  why: string
28  notes: string[]
29  /** Set when only the person can take this step: what they have to do. */
30  needsUser: string | null
31  /** What the step hands Claude instead of "run <command>", for steps that are no single command. */
32  prompt?: string
33  /**
34   * What the step settles: an approval only the person gives (`spec`, `plan`, `checklists`), or a checkpoint the
35   * autopilot records when it hands the step over (`analyze`, `converge`); the value is the fingerprint it holds for.
36   */
37  approve?: { key: string; value: string }
38  /** The request a revise step folds into the spec: handing the step over takes it off the list. */
39  resolves?: string
40}
41
42/** What the workflow reads beyond the snapshot: the review option, the persistence model, the last test run. */
43export type Flow = {
44  review: 'spec' | 'spec+plan' | 'none'
45  lastTest?: Autopilot['lastTest']
46  repairs?: number
47}
48
49const INIT = 'init --here --force --non-interactive --integration claude'
50export const MAX_REPAIRS = 3
51
52/** How Spec Kit's CLI runs here: installed, or through uvx without installing. */
53export const specifyCli = (snap: Snapshot) => (snap.tools.specify ? 'specify' : `uvx --from 'specify-cli>=1.1,<2' specify`)
54
55/** What a spec change sets off (docs/research.de.md §5), as the constitution names it: `Persistence model: living`. */
56export type Persistence = 'flow-back' | 'flow-forward' | 'living'
57export function persistenceModel(constitutionMd: string | null): Persistence {
58  const m = /persistence(?:\s+model)?\s*[:=-]\s*\**\s*(flow[- ]back|flow[- ]forward|living(?:\s+spec)?)/i.exec(constitutionMd ?? '')
59  const value = m?.[1]?.toLowerCase().replace(' ', '-') ?? 'flow-back'
60  return value.startsWith('living') ? 'living' : value === 'flow-forward' ? 'flow-forward' : 'flow-back'
61}
62
63/** The test command plan.md's `**Testing**:` line implies: a backticked command as written, else the framework's usual one. */
64export function testCommandFrom(planMd: string | null): string | null {
65  const line = /^\*\*Testing\*\*:\s*(.+)$/m.exec(planMd ?? '')?.[1]?.trim()
66  if (!line || /NEEDS CLARIFICATION|^\[/.test(line)) return null
67  const ticked = /`([^`]+)`/.exec(line)?.[1]
68  if (ticked) return ticked
69  const known: [RegExp, string][] = [
70    [/vitest/i, 'npx vitest run'],
71    [/jest/i, 'npx jest'],
72    [/playwright/i, 'npx playwright test'],
73    [/pytest/i, 'pytest'],
74    [/cargo/i, 'cargo test'],
75    [/\bgo test|\bgo\b/i, 'go test ./...'],
76    [/flutter/i, 'flutter test'],
77    [/dart/i, 'dart test'],
78    [/rspec/i, 'bundle exec rspec'],
79    [/npm test|node:test|mocha/i, 'npm test'],
80  ]
81  return known.find(([re]) => re.test(line))?.[1] ?? null
82}
83
84/** The first `##` phase of tasks.md that still has an open task: what one implement step covers. */
85export const openPhase = (tasks: Task[]) => tasks.find(t => !t.done)?.phase || null
86
87/** A fingerprint of the task list that ignores the checkboxes: converge changes it, implement does not. */
88export const tasksFingerprint = (tasks: Task[]) => fingerprint(tasks.map(t => `${t.id} ${t.text}`).join('\n'))
89
90/** The open checklist items of a feature's `checklists/*.md`, by file. */
91export function openChecklistItems(files: Record<string, string>): { file: string; open: number }[] {
92  return Object.entries(files)
93    .map(([file, text]) => ({ file, open: (text.match(/^\s*[-*]\s*\[ \]/gm) ?? []).length }))
94    .filter(f => f.open > 0)
95}
96
97export function nextStep(snap: Snapshot, ledger: Ledger, idea: string | null = null, flow: Flow = { review: 'spec' }): Step {
98  const notes: string[] = []
99  const xref = snap.extensions.includes('xref')
100  const has = (name: string) => (snap.commands ?? []).includes(name)
101  const step = (phase: Phase, command: string | null, why: string, needsUser: string | null = null, extra: Partial<Step> = {}): Step => ({ phase, command, why, notes, needsUser, ...extra })
102
103  // Setting Spec Kit up writes into the repository: that is always the person's call, never the autopilot's.
104  if (!snap.initialized) {
105    if (!snap.tools.specify && !snap.tools.uvx) {
106      return step('setup', null, 'Spec Kit is not set up, and neither the specify CLI nor uvx is on this machine.', 'Install uv (https://docs.astral.sh/uv/) or the specify CLI (pipx install specify-cli).')
107    }
108    return snap.folder === 'empty'
109      ? step('setup', `${specifyCli(snap)} ${INIT}`, 'This folder is empty: a new app, project or problem can start here with Spec Kit.', 'Tell me what you want to build (an app, a project, a problem): I set Spec Kit up for it and write the spec in your words.')
110      : step('setup', `${specifyCli(snap)} ${INIT}`, 'This folder has code, but no Spec Kit yet (no .specify/).', 'Say "set up Spec Kit here" to bring this project under Spec Kit, or press Set up Spec Kit in the pane.')
111  }
112  if (!snap.claudeIntegration) {
113    return step('integration', `${specifyCli(snap)} integration install claude`, "Spec Kit is set up, but not for Claude Code: its /speckit-* skills are missing.", 'Say "add the Claude integration", or press Add Claude integration in the pane.')
114  }
115  if (!snap.constitution || snap.constitution.principles.length === 0) {
116    return step('constitution', speckitCommand(snap, 'constitution'), 'The constitution is missing or still the template.')
117  }
118  if (!snap.featureDir || !snap.spec) {
119    return idea
120      ? step('specify', `${speckitCommand(snap, 'specify')} ${idea}`, 'No feature yet; the person has said what to build.')
121      : step('specify', speckitCommand(snap, 'specify'), 'No feature yet: describe what to build.', 'Describe the feature you want built; your own words become the spec.')
122  }
123
124  const spec = snap.spec
125  const unclear = spec.reqs.filter(r => r.needsClarification).map(r => r.id)
126  const specFp = specFingerprint(spec)
127  // The one review: where the idea becomes the contract. A project planned before 0.4 counts as reviewed until its spec changes.
128  const specApproved = ledger.approvals.spec ? ledger.approvals.spec === specFp : snap.hasPlan
129  if (!snap.hasPlan && unclear.length) return step('clarify', speckitCommand(snap, 'clarify'), `${unclear.join(', ')} still marked NEEDS CLARIFICATION; settle them before planning.`)
130  if (flow.review !== 'none' && !specApproved) {
131    return step('review', null, `The spec of ${snap.featureDir} is written: ${spec.reqs.filter(r => r.kind === 'FR').length} requirements, ${spec.outOfScope.length} out of scope, ${spec.assumptions.length} assumptions.`,
132      'Review the spec against your words, then press Approve spec in the pane (or /xref approve); tell me what to change otherwise.', { approve: { key: 'spec', value: specFp } })
133  }
134  if (!snap.hasPlan) return step('plan', speckitCommand(snap, 'plan'), 'The spec is written; plan.md is missing.')
135  if (flow.review === 'spec+plan' && snap.planFingerprint) {
136    const planApproved = ledger.approvals.plan ? ledger.approvals.plan === snap.planFingerprint : snap.tasks.length > 0
137    if (!planApproved) {
138      return step('review', null, `The plan of ${snap.featureDir} is written.`, 'Review plan.md, then press Approve plan in the pane (or /xref approve); tell me what to change otherwise.', { approve: { key: 'plan', value: snap.planFingerprint } })
139    }
140  }
141  if (unclear.length) notes.push(`${unclear.join(', ')} still marked NEEDS CLARIFICATION (${speckitCommand(snap, 'clarify')}).`)
142  if (snap.tasks.length === 0) return step('tasks', speckitCommand(snap, 'tasks'), 'The plan is written; tasks.md is missing.')
143
144  // A request beyond the spec reaches the spec before more code does; a contradiction is always the person's.
145  const changes = ledger.semantic?.changes ?? []
146  const contradiction = changes.find(c => c.kind === 'contradicts')
147  if (contradiction) return step('revise', null, `"${contradiction.text}" contradicts the spec.`, `Your request "${contradiction.text}" contradicts the spec: keep the spec, or change it (To spec in the pane).`)
148  const extension = changes.find(c => c.kind === 'extends')
149  if (extension) {
150    if (snap.persistence === 'flow-forward') return step('revise', null, `"${extension.text}" goes beyond the spec.`, `You asked for "${extension.text}", which the spec does not cover: fold it into the spec (To spec), or make it a task (As task).`)
151    const revise = xref && has('xref-revise') ? speckitCommand(snap, 'xref.revise') : speckitCommand(snap, 'clarify')
152    return step('revise', `${revise} The user asked during implementation: "${extension.text}". Fold it into the spec, new requirements under new ids.`, `"${extension.text}" goes beyond the spec; the spec follows the request (${snap.persistence ?? 'flow-back'}).`, null, { resolves: extension.text })
153  }
154
155  const report = evaluate(snap, ledger)
156  if (report.level === 'red' || report.level === 'yellow') notes.push(`Drift ${report.level}: ${report.findings[0]?.text ?? ''}`)
157  const mapped = Object.values(ledger.requirements).some(r => r.tasks.length > 0)
158  const uncovered = report.uncovered.filter(id => !unclear.includes(id))
159  if (uncovered.length && !mapped) {
160    return step('map', xref ? speckitCommand(snap, 'xref.map') : '/xref map', `${uncovered.join(', ')} name no task yet; map requirements to tasks once.`)
161  }
162  const open = snap.tasks.filter(t => !t.done)
163  const started = snap.tasks.some(t => t.done)
164  // analyze before the first implement, and again after every change to the spec.
165  if (open.length && has('analyze') && (ledger.checkpoints.analyze ? ledger.checkpoints.analyze !== specFp : !started)) {
166    return step('analyze', speckitCommand(snap, 'analyze'), started ? 'The spec changed since the last analysis: check spec, plan and tasks against each other.' : 'Before the first task: check spec, plan and tasks against each other.', null, { approve: { key: 'analyze', value: specFp } })
167  }
168  if (flow.lastTest && !flow.lastTest.ok) {
169    if ((flow.repairs ?? 0) > MAX_REPAIRS) {
170      return step('repair', null, `The tests still fail after ${MAX_REPAIRS} repairs.`, `The tests still fail after ${MAX_REPAIRS} attempts: look at the failure in the pane and decide how to go on.`)
171    }
172    return step('repair', snap.testCommand, `The tests fail (repair ${Math.max(1, flow.repairs ?? 0)} of ${MAX_REPAIRS}).`, null, {
173      prompt: `The tests fail. Run \`${snap.testCommand}\`, find the cause and fix the code (not the tests, unless a test contradicts the spec). Do not check off new tasks in this step.\nLast output:\n${flow.lastTest.output.slice(-1500)}`,
174    })
175  }
176  if (open.length) {
177    const checklists = snap.checklists ?? []
178    const listFp = fingerprint(checklists.map(c => `${c.file}:${c.open}`).join('\n'))
179    if (checklists.length && !started && ledger.approvals.checklists !== listFp) {
180      return step('checklists', null, `${checklists.reduce((n, c) => n + c.open, 0)} checklist items are open (${checklists.map(c => c.file).join(', ')}).`, 'Checklist items are open: complete them, or press Proceed anyway in the pane (/xref approve).', { approve: { key: 'checklists', value: listFp } })
181    }
182    // One phase per step (Spec Kit's own advice for larger features): the budget and the stall check then measure real work.
183    const phase = snap.phase
184    const command = speckitCommand(snap, 'implement')
185    const ids = (phase ? open.filter(t => t.phase === phase) : open).map(t => t.id)
186    return step('implement', command, `${open.length} of ${snap.tasks.length} tasks open; next ${open[0]!.id}.`, null, {
187      prompt: phase
188        ? `Run ${command} now through the Skill tool, scoped by its argument to one phase: "Only the phase '${phase}' (${ids.join(', ')}); stop when it is done." Carry that phase through.`
189        : undefined,
190    })
191  }
192  const tasksFp = snap.tasksFingerprint || tasksFingerprint(snap.tasks)
193  if (has('converge') && ledger.checkpoints.converge !== tasksFp) {
194    return step('converge', speckitCommand(snap, 'converge'), 'Every task is checked: find the work the tasks missed.', null, { approve: { key: 'converge', value: tasksFp } })
195  }
196  return step('verify', xref ? speckitCommand(snap, 'xref.check') : '/xref check', 'Every task is checked: check the code against the spec.')
197}
198
199/** The step as one line, as the pane, /xref and the note on each prompt show it. */
200export const stepLine = (s: Step) => (s.command ? `${s.command} · ${s.why}` : s.why)
201
202export const idleAutopilot = (): Autopilot => ({
203  on: false,
204  paused: null,
205  steps: 0,
206  max: 25,
207  last: null,
208  stalls: 0,
209  lastPhase: null,
210  idea: null,
211  repairs: 0,
212  lastTest: null,
213  startedAt: null,
214  costAtStart: 0,
215  night: false,
216  scope: null,
217})
218
219/** Whether the person's next words are the idea to specify: the autopilot waits on them for exactly that. */
220export const takesIdea = (step: Step, snap: Snapshot) => step.phase === 'specify' || (step.phase === 'setup' && snap.folder === 'empty' && step.command !== null)
221
222/** What the model reads beside the person's prompt while the autopilot waits before Spec Kit is there. */
223export function setupNote(step: Step, snap: Snapshot): string {
224  const lines = [`speckit-xref · autopilot waiting · ${step.why}`]
225  if (step.phase === 'setup' && step.command && snap.folder === 'empty') {
226    lines.push(`If this message says what to build, set Spec Kit up now with the speckit-xref:speckit skill (\`${step.command}\`); the autopilot then goes on with the constitution and the spec in the person's words.`)
227  } else if (step.phase === 'setup' && step.command) {
228    lines.push(`Set Spec Kit up (speckit-xref:speckit skill, \`${step.command}\`) only if this message asks for it; otherwise leave the repository as it is.`)
229  } else if (step.phase === 'integration') {
230    lines.push(`Run \`${step.command}\` only if this message asks for it.`)
231  }
232  return lines.join('\n')
233}
234
235/** What has to change between two autopilot steps for the run to count as moving. */
236export function progressKey(snap: Snapshot, ledger: Ledger): string {
237  const report = evaluate(snap, ledger)
238  const mapped = Object.values(ledger.requirements).filter(r => r.tasks.length > 0).length
239  return [
240    snap.initialized,
241    snap.claudeIntegration,
242    snap.constitution?.principles.length ?? 0,
243    snap.featureDir ?? '-',
244    snap.spec?.reqs.length ?? 0,
245    snap.spec?.reqs.filter(r => r.needsClarification).length ?? 0,
246    snap.hasPlan,
247    `${report.done}/${report.tasks}`,
248    snap.tasksFingerprint ?? '-',
249    mapped,
250    report.level,
251    Object.values(report.levels).join('/'),
252    ledger.semantic?.at ?? '-',
253    JSON.stringify(ledger.checkpoints),
254  ].join('|')
255}
256
257/** The rules the autopilot gives the model with every step it hands over. */
258export const AUTONOMY_RULES = [
259  'Autopilot is on. Carry the step through without asking whether to continue: when you end your turn, the autopilot moves on by itself.',
260  'Decide what the person left open, from the spec, the constitution and the repository, and record those choices as assumptions in the artifact you write.',
261  'Stop only for what only the person can decide: the product idea, a [NEEDS CLARIFICATION] question, a conflict between their request and the spec, anything destructive or irreversible, credentials or payments. Then call mcp__speckit-xref__ask with the question (and the user stories it blocks, if any); if it answers with the person\'s choice, go on with it, otherwise put the question in your answer and end your turn.',
262]
263
264/** The prompt the autopilot submits for one step. */
265export function autopilotPrompt(step: Step, n: number, max: number, context: string | null): string {
266  return [
267    `[speckit-xref autopilot · step ${n}/${max}] ${step.phase}: ${step.why}`,
268    step.prompt ?? `Run ${step.command} now: invoke it through the Skill tool and carry it through.`,
269    ...step.notes.map(note => `Note: ${note}`),
270    '',
271    ...AUTONOMY_RULES,
272    ...(context ? ['', context] : []),
273  ].join('\n')
274}
275
276/** Whether a model's answer ends on a question, the one case the autopilot asks a small model about. */
277export const endsWithQuestion = (answer: string) => /\?["')\]*_\s]*$/.test(answer.trim().slice(-400))
278
279/** The labels for that question; the first means stop. A "shall I continue?" is the second. */
280export const QUESTION_LABELS = ['needs a decision only the person can make', 'asks only whether to continue, or reports progress'] as const
281
282/** The phase strip: `✓setup ✓const ✓spec ✓plan ✓tasks ▶impl 7/8 ·verify`. */
283export function phaseStrip(snap: Snapshot, step: Step): string {
284  const order: [string, Phase[]][] = [
285    ['setup', ['setup', 'integration']],
286    ['const', ['constitution']],
287    ['spec', ['specify', 'clarify', 'review', 'revise']],
288    ['plan', ['plan']],
289    ['tasks', ['tasks', 'map', 'analyze', 'checklists']],
290    ['impl', ['implement', 'repair', 'converge']],
291    ['verify', ['verify']],
292  ]
293  const at = order.findIndex(([, phases]) => phases.includes(step.phase))
294  const done = snap.tasks.filter(t => t.done).length
295  return order
296    .map(([label], i) => {
297      const extra = label === 'impl' && snap.tasks.length ? ` ${done}/${snap.tasks.length}` : ''
298      return i < at ? `✓${label}` : i === at ? `▶${label}${extra}` : `·${label}`
299    })
300    .join(' ')
301}
302
hooks/xref.ts 462 lines
1// Pure X-Ref and drift logic: which task a file serves, what counts as drift, what the model is told.
2
3import type { Anchor, Ledger, Semantic, Snapshot, Spec, Task } from '../types'
4import { emptyLedger } from './ledger'
5import { LADDER, isActive, isStale, levelOf } from './proof'
6import type { Rung } from './proof'
7import { fingerprint, isTestFile, localId, matchesAny } from './rules'
8import type { Rules } from './rules'
9
10export { emptyLedger, localId }
11
12/** The verdicts of docs/contract-0.4.md §3, in the order they are tried. */
13export type CoreVerdict = 'spec' | 'exempt' | 'task-linked' | 'planned' | 'linked' | 'unclear' | 'untracked' | 'unplanned'
14export type Verdict = 'in-scope' | 'other-task' | 'linked' | 'unplanned' | 'untracked' | 'spec' | 'exempt' | 'unclear'
15export type Classified = { verdict: Verdict; task: string | null }
16
17export type Level = 'none' | 'green' | 'yellow' | 'red'
18export type Finding = {
19  kind: 'unplanned' | 'dangling' | 'orphan-test' | 're-verify' | 'semantic' | 'intent'
20  level: 'yellow' | 'red'
21  text: string
22  file?: string
23  intent?: string
24  id?: string
25}
26export type Report = {
27  level: Level
28  findings: Finding[]
29  /** Active FRs at `planned` or above. */
30  covered: number
31  total: number
32  uncovered: string[]
33  done: number
34  tasks: number
35  levels: Record<Rung, number>
36  /** Spec gaps that do not raise the drift level. */
37  notes: string[]
38}
39
40/** A text cut at a word boundary, an ellipsis where it was cut. */
41export function short(text: string, max: number): string {
42  if (text.length <= max) return text
43  const cut = text.slice(0, max)
44  const space = cut.lastIndexOf(' ')
45  return (space > max * 0.5 ? cut.slice(0, space) : cut).replace(/[\s,;:(]+$/, '') + '…'
46}
47
48const MAX_UNPLANNED = 100
49const MAX_INTENTS = 50
50
51export function relPath(file: string, root: string): string {
52  let path = file.replace(/\\/g, '/')
53  const base = root.replace(/\\/g, '/').replace(/\/$/, '')
54  if (base && path.startsWith(base + '/')) path = path.slice(base.length + 1)
55  return path.replace(/^\.\//, '')
56}
57
58const basename = (path: string) => path.slice(path.lastIndexOf('/') + 1)
59
60/** Spec Kit's own files and the agent's configuration are never drift; CI workflows are code. */
61export const isSpecArtifact = (rel: string) =>
62  (/^(specs|\.specify|\.claude|\.github)\//.test(rel) && !/^\.github\/workflows\//.test(rel)) || /^(CLAUDE|AGENTS)\.md$/.test(rel)
63
64export function pathMatches(rel: string, planned: string): boolean {
65  const p = planned.replace(/^\.\//, '')
66  if (rel === p) return true
67  if (p.endsWith('/')) return rel.startsWith(p)
68  if (rel.endsWith('/' + p)) return true
69  if (!p.includes('/') && !p.includes('.')) return false
70  if (rel.startsWith(p + '/')) return true
71  return !p.includes('/') && basename(rel) === p
72}
73
74/** Whether a task plans a file: a file it names always; a folder it names only while the task is open. */
75export function plans(task: Task, rel: string): boolean {
76  return task.paths.some(p => {
77    const isFolder = p.endsWith('/') || rel.startsWith(p.replace(/^\.\//, '') + '/')
78    return pathMatches(rel, p) && (!isFolder || !task.done)
79  })
80}
81
82export const ANCHOR = /@spec\s+((?:[\w.-]+\/)?(?:(?:FR|SC)-\d{3,}|T-?\d{3,}|US\d+-AS\d+))/g
83
84export function anchorsIn(text: string, file: string): Anchor[] {
85  const out: Anchor[] = []
86  text.split('\n').forEach((line, i) => {
87    for (const m of line.matchAll(ANCHOR)) out.push({ id: m[1] ?? '', file, line: i + 1 })
88  })
89  return out
90}
91
92const DEFAULT_RULES: Rules = { exempt: [], unclear: [] }
93
94/** The shared classification (§3): first match wins. The extension's batch check reports these verdicts as they are. */
95export function classifyCore(
96  rel: string,
97  snap: Pick<Snapshot, 'tasks'>,
98  ledger: Ledger,
99  rules: Rules = DEFAULT_RULES,
100  newText = '',
101): { verdict: CoreVerdict; task: string | null } {
102  if (isSpecArtifact(rel)) return { verdict: 'spec', task: null }
103  if (matchesAny(rel, rules.exempt)) return { verdict: 'exempt', task: null }
104  for (const [id, entry] of Object.entries(ledger.tasks)) if (entry.linked.includes(rel)) return { verdict: 'task-linked', task: id }
105  // The task that plans a file comes before its anchors: it says which task the work is on.
106  const matches = snap.tasks.filter(t => plans(t, rel))
107  const match = matches.find(t => !t.done) ?? matches[0]
108  if (match) return { verdict: 'planned', task: match.id }
109  if (newText && anchorsIn(newText, rel).length > 0) return { verdict: 'linked', task: null }
110  if (Object.values(ledger.requirements).some(r => r.files.includes(rel))) return { verdict: 'linked', task: null }
111  if (ledger.anchors.some(a => a.file === rel)) return { verdict: 'linked', task: null }
112  if (matchesAny(rel, rules.unclear)) return { verdict: 'unclear', task: null }
113  if (snap.tasks.length === 0) return { verdict: 'untracked', task: null }
114  return { verdict: 'unplanned', task: null }
115}
116
117/** The live verdict: a planned file is in scope for the task in focus, or moves the focus to its own task. */
118export function classify(rel: string, snap: Snapshot, ledger: Ledger, active: string | null, newText = ''): Classified {
119  const core = classifyCore(rel, snap, ledger, snap.rules, newText)
120  switch (core.verdict) {
121    case 'task-linked':
122      return { verdict: 'linked', task: core.task }
123    case 'planned': {
124      const own = active ? snap.tasks.find(t => t.id === active && plans(t, rel)) : undefined
125      if (own) return { verdict: 'in-scope', task: own.id }
126      return { verdict: active ? 'other-task' : 'in-scope', task: core.task }
127    }
128    case 'linked':
129    case 'unplanned':
130    case 'unclear':
131      return { verdict: core.verdict, task: active }
132    default:
133      return { verdict: core.verdict, task: null }
134  }
135}
136
137/** Books one finished write into the ledger; returns the ledger it became. */
138export function recordTouch(ledger: Ledger, rel: string, c: Classified, at: string): Ledger {
139  const next: Ledger = { ...ledger, tasks: { ...ledger.tasks }, unplanned: ledger.unplanned }
140  // Exempt and unclear files are no evidence for a task and no drift either.
141  if (c.task && c.verdict !== 'unplanned' && c.verdict !== 'exempt' && c.verdict !== 'unclear') {
142    const entry = next.tasks[c.task] ?? { touched: [], linked: [] }
143    if (!entry.touched.includes(rel)) next.tasks[c.task] = { ...entry, touched: [...entry.touched, rel] }
144  }
145  if (c.verdict === 'unplanned' && !ledger.unplanned.some(u => u.file === rel && !u.acknowledged)) {
146    next.unplanned = [...ledger.unplanned, { file: rel, at, task: c.task, acknowledged: false }].slice(-MAX_UNPLANNED)
147  }
148  return next
149}
150
151/** Ties a file to a task or requirement; a file once unplanned stops counting. */
152export function link(ledger: Ledger, rel: string, id: string, snap: Snapshot, source = 'link'): { ledger: Ledger; error?: string } {
153  const bare = localId(id, snap.featureDir)
154  if (!bare) return { ledger, error: `${id} names another feature than ${snap.featureDir ?? 'the active one'}` }
155  const next: Ledger = { ...ledger, tasks: { ...ledger.tasks }, requirements: { ...ledger.requirements } }
156  if (/^T\d{3,}$/.test(bare)) {
157    if (!snap.tasks.some(t => t.id === bare)) return { ledger, error: `${bare} is no task in tasks.md` }
158    const entry = next.tasks[bare] ?? { touched: [], linked: [] }
159    if (!entry.linked.includes(rel)) next.tasks[bare] = { ...entry, linked: [...entry.linked, rel] }
160  } else if (/^(FR|SC)-\d{3,}$/.test(bare) || /^US\d+-AS\d+$/.test(bare)) {
161    const known = snap.spec?.reqs.some(r => r.id === bare) || snap.spec?.stories.some(s => s.scenarios.some(x => x.id === bare))
162    if (!known) return { ledger, error: `${bare} is no requirement or acceptance scenario in spec.md` }
163    const entry = next.requirements[bare] ?? { tasks: [], files: [], sources: [] }
164    next.requirements[bare] = {
165      ...entry,
166      files: entry.files.includes(rel) ? entry.files : [...entry.files, rel],
167      sources: entry.sources.includes(source) ? entry.sources : [...entry.sources, source],
168    }
169  } else {
170    return { ledger, error: `${id} is neither a task (T###), a requirement (FR-### / SC-###) nor a scenario (US1-AS1)` }
171  }
172  next.unplanned = ledger.unplanned.map(u => (u.file === rel ? { ...u, acknowledged: true } : u))
173  // The text a link was made against: when it changes, the link needs new proof.
174  const req = snap.spec?.reqs.find(r => r.id === bare)
175  if (req) next.fingerprints = { ...ledger.fingerprints, [bare]: fingerprint(req.text) }
176  return { ledger: next }
177}
178
179export function logIntent(ledger: Ledger, text: string, task: string | null, at: string): Ledger {
180  return { ...ledger, intents: [...ledger.intents, { at, text: text.slice(0, 500), task, status: 'logged' as const }].slice(-MAX_INTENTS) }
181}
182
183export function acknowledgeAll(ledger: Ledger): Ledger {
184  return { ...ledger, unplanned: ledger.unplanned.map(u => ({ ...u, acknowledged: true })) }
185}
186
187/** Whether a logged prompt and a change the check quoted from it are the same request. */
188const quotes = (logged: string, quoted: string) => logged.includes(quoted) || quoted.includes(logged.slice(0, 40))
189
190export function resolveIntent(ledger: Ledger, text: string): Ledger {
191  const semantic = ledger.semantic ? { ...ledger.semantic, changes: ledger.semantic.changes.filter(c => c.text !== text) } : null
192  return { ...ledger, semantic, intents: ledger.intents.map(i => (i.status !== 'resolved' && quotes(i.text, text) ? { ...i, status: 'resolved' as const } : i)) }
193}
194
195/** The task in focus while it is open; once it is checked off, the first open one. */
196export function currentTask(snap: Snapshot, active: string | null): Task | null {
197  const chosen = active ? snap.tasks.find(t => t.id === active && !t.done) : undefined
198  return chosen ?? snap.tasks.find(t => !t.done) ?? null
199}
200
201/** What the ladder reads from a snapshot. */
202const proofSnap = (snap: Snapshot) => ({ featureDir: snap.featureDir, tasks: snap.tasks, reqs: snap.spec?.reqs ?? [], realTests: snap.realTests ?? [] })
203
204/**
205 * The drift report (§8). Only deterministic findings decide unless `semantic` is on: the live mod lets the
206 * intent check count, the extension's exit code never does.
207 */
208export function evaluate(snap: Snapshot, ledger: Ledger, opts: { semantic?: boolean } = {}): Report {
209  const semanticOn = opts.semantic ?? true
210  const findings: Finding[] = []
211  const open = ledger.unplanned.filter(u => !u.acknowledged)
212  for (const u of open) findings.push({ kind: 'unplanned', level: 'yellow', file: u.file, text: `${u.file} is not planned for any task` })
213  if (open.length >= 3) findings.push({ kind: 'unplanned', level: 'red', text: `${open.length} edits outside the planned files` })
214
215  const reqs = snap.spec?.reqs ?? []
216  const byId = new Map(reqs.map(r => [r.id, r]))
217  const known = new Set([...reqs.filter(r => isActive(r)).map(r => r.id), ...snap.tasks.map(t => t.id), ...(snap.spec?.stories.flatMap(s => s.scenarios.map(x => x.id)) ?? [])])
218  for (const a of ledger.anchors) {
219    const bare = localId(a.id, snap.featureDir)
220    if (!bare || known.has(bare)) continue
221    const old = byId.get(bare)
222    const why = old ? (old.supersededBy ? `superseded by ${old.supersededBy}` : 'retired') : 'which the spec no longer has'
223    const test = isTestFile(a.file)
224    findings.push({
225      kind: test ? 'orphan-test' : 'dangling',
226      level: 'yellow',
227      file: a.file,
228      id: bare,
229      text: test ? `${a.file}:${a.line} tests ${a.id}, ${old ? `which is ${why}` : why}` : `${a.file}:${a.line} anchors ${a.id}, ${old ? `which is ${why}` : why}`,
230    })
231  }
232  const proof = proofSnap(snap)
233  for (const r of reqs) {
234    if (isActive(r) && isStale(r.id, proof, ledger)) findings.push({ kind: 're-verify', level: 'yellow', id: r.id, text: `${r.id} changed since it was last proven; it needs new proof` })
235  }
236
237  const s = ledger.semantic
238  if (s && semanticOn) {
239    if (s.verdict === 'drift' || s.score < 50) findings.push({ kind: 'semantic', level: 'red', text: `Intent check ${s.score}/100: ${s.reasons[0] ?? 'the change drifts from the spec'}` })
240    else if (s.verdict === 'minor' || s.score < 80) findings.push({ kind: 'semantic', level: 'yellow', text: `Intent check ${s.score}/100: ${s.reasons[0] ?? 'minor drift'}` })
241    for (const c of s.changes) {
242      findings.push({ kind: 'intent', level: c.kind === 'contradicts' ? 'red' : 'yellow', intent: c.text, text: `User asked: "${c.text}", ${c.kind === 'contradicts' ? 'which contradicts the spec' : 'which the spec does not cover'}` })
243    }
244  }
245
246  const frs = reqs.filter(r => r.kind === 'FR' && isActive(r))
247  const levels = Object.fromEntries(LADDER.map(r => [r, 0])) as Record<Rung, number>
248  const uncovered: string[] = []
249  for (const r of frs) {
250    const rung = levelOf(r.id, proof, ledger)
251    levels[rung] += 1
252    if (rung === 'specified') uncovered.push(r.id)
253  }
254  const notes = (snap.spec?.stories ?? []).filter(st => st.scenarios.length === 0).map(st => `${st.id} has no acceptance scenarios`)
255  const level: Level = !snap.featureDir ? 'none' : findings.some(f => f.level === 'red') ? 'red' : findings.length ? 'yellow' : 'green'
256  return {
257    level,
258    findings,
259    covered: frs.length - uncovered.length,
260    total: frs.length,
261    uncovered,
262    done: snap.tasks.filter(t => t.done).length,
263    tasks: snap.tasks.length,
264    levels,
265    notes,
266  }
267}
268
269/** The FRs a task serves: named in its text, or mapped to it in the ledger. */
270export function reqsOf(task: Task, ledger: Ledger): string[] {
271  const mapped = Object.entries(ledger.requirements).filter(([, r]) => r.tasks.includes(task.id)).map(([id]) => id)
272  return [...new Set([...task.reqs, ...mapped])]
273}
274
275/** How the project invokes a Spec Kit command: `xref.map` is `/speckit-xref-map` for skills, `/speckit.xref.map` for commands. */
276export const speckitCommand = (snap: Snapshot, name: string) =>
277  snap.commandStyle === 'skills' ? `/speckit-${name.replace(/\./g, '-')}` : `/speckit.${name}`
278
279/**
280 * The system prompt section: the active spec and the rules that keep the work on it.
281 * It changes only when the spec does, so the conversation's prompt cache survives task switches;
282 * the current task and the drift status ride each prompt instead (turnContext).
283 */
284export function composeSection(snap: Snapshot, mode: string, plugin: string): string | null {
285  if (!snap.featureDir || !snap.spec) return null
286  const spec = snap.spec
287  const lines: string[] = [
288    '# Spec Kit cross-reference (speckit-xref)',
289    'This repository follows GitHub Spec Kit. Keep the work inside the active spec and say so when it would leave it.',
290    '',
291    `Active feature: ${snap.featureDir} — ${spec.title}`,
292  ]
293  if (spec.input) lines.push(`The user's original words: "${spec.input}"`)
294  if (spec.stories.length) lines.push(`User stories: ${spec.stories.map(s => `${s.id}${s.priority ? ` (${s.priority})` : ''} ${s.title}`).join(' | ')}`)
295  if (snap.constitution?.musts.length) lines.push('Constitution, MUST rules:', ...snap.constitution.musts.slice(0, 8).map(m => `- ${m}`))
296  if (spec.outOfScope.length) lines.push('Out of scope:', ...spec.outOfScope.map(o => `- ${o}`))
297  const clar = spec.reqs.filter(r => r.needsClarification).map(r => r.id)
298  if (clar.length) lines.push(`Still marked NEEDS CLARIFICATION: ${clar.join(', ')}. Ask the user before implementing them.`)
299  lines.push(
300    '',
301    'Rules:',
302    `- Before editing a file the current task does not plan, name the task or requirement it serves; call mcp__${plugin}__focus to switch tasks or mcp__${plugin}__link to tie the file to one.`,
303    `- When the user asks for something the spec does not cover or contradicts, say so plainly and offer to update the spec (${speckitCommand(snap, 'clarify')}) before building it.`,
304    '- Check a task off in tasks.md ([X]) only once the requirements it serves are met.',
305    `- Mark code that implements a requirement with a comment at file or function level: @spec ${basename(snap.featureDir)}/FR-###.`,
306    '- Each user prompt carries a "speckit-xref" note with the current task and the drift status; follow it.',
307  )
308  if (mode === 'strict') lines.push("- Strict mode: edits outside the current task's planned or linked files are refused until you focus another task or link the file.")
309  return lines.join('\n')
310}
311
312/** The note beside a user's prompt: the current task, what it serves, and the drift status right now. */
313export function turnContext(snap: Snapshot, ledger: Ledger, active: string | null, next?: string): string | null {
314  if (!snap.featureDir || !snap.spec) return null
315  const report = evaluate(snap, ledger)
316  const task = currentTask(snap, active)
317  const lines = [`speckit-xref · ${snap.featureDir} · tasks ${report.done}/${report.tasks} · FR covered ${report.covered}/${report.total} · drift ${report.level}`]
318  if (task) {
319    lines.push(`Current task: ${task.id}${task.story ? ` [${task.story}]` : ''} ${task.text}`)
320    if (task.paths.length) lines.push(`Planned files: ${task.paths.join(', ')}`)
321    const reqs = reqsOf(task, ledger)
322    if (reqs.length) {
323      const byId = new Map(snap.spec.reqs.map(r => [r.id, r.text]))
324      lines.push('Serves:', ...reqs.map(id => `- ${id}: ${byId.get(id) ?? ''}`))
325    }
326    const next = snap.tasks.filter(t => !t.done && t.id !== task.id).slice(0, 3)
327    if (next.length) lines.push(`Next: ${next.map(t => `${t.id} ${short(t.text, 80)}`).join(' | ')}`)
328  } else if (snap.tasks.length === 0) {
329    lines.push(`No tasks.md yet; ${speckitCommand(snap, 'tasks')} derives them from the plan.`)
330  } else {
331    lines.push('Every task in tasks.md is checked.')
332  }
333  if (report.findings.length) lines.push('Drift:', ...report.findings.slice(0, 5).map(f => `- ${f.text}`))
334  if (next) lines.push(`Next Spec Kit step: ${next}`)
335  return lines.join('\n')
336}
337
338/** What the model reads beside a write's result when the write left the plan, or moved to another task. */
339export function editNote(rel: string, c: Classified, snap: Snapshot, previous: string | null, plugin: string): string | null {
340  if (c.verdict === 'unplanned') {
341    const task = currentTask(snap, previous)
342    const planned = task?.paths.length ? ` (it plans ${task.paths.join(', ')})` : ''
343    return `speckit-xref: ${rel} is not planned for ${task ? `the current task ${task.id}${planned}` : 'any task'}. If it serves the spec, call mcp__${plugin}__link with the task or requirement; if the user asked for it beyond the spec, tell them and offer ${speckitCommand(snap, 'clarify')}.`
344  }
345  if (c.verdict === 'other-task' && c.task && c.task !== previous) {
346    const task = snap.tasks.find(t => t.id === c.task)
347    if (task?.done) return `speckit-xref: ${rel} belongs to ${c.task}, which is checked off: this is rework on a finished task.`
348    return `speckit-xref: ${rel} belongs to ${c.task}${task ? ` (${task.text.slice(0, 80)})` : ''}; that is now the current task.`
349  }
350  return null
351}
352
353/** The question the drift check asks over the session's own transcript. */
354export function forkPrompt(snap: Snapshot, ledger: Ledger, active: string | null, changed: string[], diff: string, withLog = false): string {
355  const spec = snap.spec
356  // Without the transcript (a check run before the session's first turn) the logged prompts stand in for the user's words.
357  const said = withLog ? ledger.intents.filter(i => i.status !== 'resolved').slice(-10).map(i => `- ${i.text}`) : []
358  const task = currentTask(snap, active)
359  const reqs = spec?.reqs.map(r => `${r.id}: ${r.text}`).join('\n') ?? '(no spec)'
360  return [
361    'You are auditing this session for drift. Do not call any tool. Answer with ONE JSON object and nothing else.',
362    'Compare (1) what the user asked for in this conversation, in their own words, (2) the active Spec Kit spec below, and (3) the files changed in the last turn.',
363    '',
364    `Spec: ${spec?.title ?? '-'} (${snap.featureDir ?? '-'})`,
365    `Original request: ${spec?.input ?? '-'}`,
366    `Requirements:\n${reqs}`,
367    `Out of scope: ${spec?.outOfScope.join('; ') || '-'}`,
368    `Current task: ${task ? `${task.id} ${task.text}` : '-'}`,
369    ...(said.length ? ['What the user said in this session, oldest first:', ...said] : []),
370    `Changed files: ${changed.join(', ') || '-'}`,
371    `Diff (may be cut):\n${diff.slice(0, 6000) || '(no diff available)'}`,
372    '',
373    'JSON shape: {"score": 0-100 (100 = the change matches both the user\'s intent and the spec), "verdict": "aligned" | "minor" | "drift", "reasons": ["at most 3 short reasons"], "intent_changes": [{"text": "a short quote of something the user asked for during the session that the spec does not cover or contradicts", "kind": "extends" | "contradicts"}]}',
374    'intent_changes holds only requests the user made, never the spec\'s own text; code that contradicts the spec without the user asking for it belongs in reasons and the score. An empty list is the normal answer.',
375  ].join('\n')
376}
377
378/** The first JSON object in a model reply. */
379function firstJson(text: string): unknown {
380  const start = text.indexOf('{')
381  const end = text.lastIndexOf('}')
382  if (start === -1 || end <= start) return null
383  try {
384    return JSON.parse(text.slice(start, end + 1))
385  } catch {
386    return null
387  }
388}
389
390export function parseSemantic(text: string, at: string): Semantic | null {
391  const raw = firstJson(text) as Record<string, unknown> | null
392  if (!raw || typeof raw.score !== 'number') return null
393  const verdict = raw.verdict === 'aligned' || raw.verdict === 'minor' || raw.verdict === 'drift' ? raw.verdict : raw.score >= 80 ? 'aligned' : raw.score >= 50 ? 'minor' : 'drift'
394  const reasons = Array.isArray(raw.reasons) ? raw.reasons.filter((r): r is string => typeof r === 'string').slice(0, 3) : []
395  const changes = Array.isArray(raw.intent_changes)
396    ? raw.intent_changes
397        .filter((c): c is { text: string; kind?: string } => !!c && typeof (c as { text?: unknown }).text === 'string')
398        .map(c => ({ text: c.text.slice(0, 200), kind: c.kind === 'contradicts' ? ('contradicts' as const) : ('extends' as const) }))
399        .slice(0, 5)
400    : []
401  return { score: Math.max(0, Math.min(100, Math.round(raw.score))), verdict, reasons, changes, at }
402}
403
404export function applySemantic(ledger: Ledger, semantic: Semantic): Ledger {
405  // A change the user already resolved stays resolved when a later check quotes it again.
406  const resolved = ledger.intents.filter(i => i.status === 'resolved')
407  const changes = semantic.changes.filter(c => !resolved.some(i => quotes(i.text, c.text)))
408  const intents = ledger.intents.map(i => {
409    const hit = changes.find(c => i.status === 'logged' && quotes(i.text, c.text))
410    return hit ? { ...i, status: hit.kind } : i
411  })
412  return { ...ledger, intents, semantic: { ...semantic, changes } }
413}
414
415export function mapPrompt(spec: Spec, tasks: Task[]): string {
416  return [
417    'Map each functional requirement of a Spec Kit feature to the tasks that implement it. Answer with ONE JSON object and nothing else:',
418    '{"FR-001": ["T004", "T005"], ...}. Use only ids listed below; a requirement no task serves maps to [].',
419    '',
420    'Requirements:',
421    ...spec.reqs.filter(r => r.kind === 'FR').map(r => `${r.id}: ${r.text}`),
422    '',
423    'Tasks:',
424    ...tasks.map(t => `${t.id}${t.story ? ` [${t.story}]` : ''}: ${t.text}`),
425  ].join('\n')
426}
427
428export function applyMapping(ledger: Ledger, text: string, snap: Snapshot): { ledger: Ledger; mapped: number } {
429  const raw = firstJson(text) as Record<string, unknown> | null
430  if (!raw) return { ledger, mapped: 0 }
431  const reqIds = new Set(snap.spec?.reqs.map(r => r.id) ?? [])
432  const taskIds = new Set(snap.tasks.map(t => t.id))
433  const requirements = { ...ledger.requirements }
434  const fingerprints = { ...ledger.fingerprints }
435  let mapped = 0
436  for (const [id, value] of Object.entries(raw)) {
437    if (!reqIds.has(id) || !Array.isArray(value)) continue
438    const tasks = value.filter((t): t is string => typeof t === 'string' && taskIds.has(t))
439    const entry = requirements[id] ?? { tasks: [], files: [], sources: [] }
440    requirements[id] = { ...entry, tasks: [...new Set([...entry.tasks, ...tasks])], sources: entry.sources.includes('llm') ? entry.sources : [...entry.sources, 'llm'] }
441    if (tasks.length) {
442      mapped += 1
443      const req = snap.spec?.reqs.find(r => r.id === id)
444      if (req) fingerprints[id] = fingerprint(req.text)
445    }
446  }
447  return { ledger: { ...ledger, requirements, fingerprints }, mapped }
448}
449
450export const REMEDIATION_HEADING = '## Drift Remediation (speckit-xref)'
451
452/** tasks.md with one more open task under the remediation heading, numbered after the highest T###. */
453export function appendRemediation(tasksMd: string, tasks: Task[], text: string): { markdown: string; id: string } {
454  const highest = tasks.reduce((max, t) => Math.max(max, Number(t.id.slice(1)) || 0), 0)
455  const width = Math.max(3, tasks[0]?.id.length ? tasks[0].id.length - 1 : 3)
456  const id = 'T' + String(highest + 1).padStart(width, '0')
457  const line = `- [ ] ${id} [Drift] ${text.replace(/\s+/g, ' ').trim()}`
458  const base = tasksMd.replace(/\s*$/, '')
459  const markdown = base.includes(REMEDIATION_HEADING) ? `${base}\n${line}\n` : `${base}\n\n${REMEDIATION_HEADING}\n\n${line}\n`
460  return { markdown, id }
461}
462
types/index.d.ts 183 lines
1// The speckit-xref contract: what the mod reads out of Spec Kit, what it keeps per feature, and the session values it draws from.
2
3export type Scenario = { id: string; text: string }
4export type Story = { id: string; title: string; priority: string | null; scenarios: Scenario[] }
5export type Req = {
6  id: string
7  kind: 'FR' | 'SC'
8  text: string
9  needsClarification: boolean
10  /** Retired and superseded requirements stay in spec.md for history; they leave the ladder. */
11  status: 'active' | 'superseded' | 'retired'
12  supersededBy: string | null
13}
14export type Task = {
15  id: string
16  done: boolean
17  parallel: boolean
18  story: string | null
19  text: string
20  paths: string[]
21  reqs: string[]
22  phase: string
23}
24export type Spec = {
25  title: string
26  input: string | null
27  stories: Story[]
28  reqs: Req[]
29  outOfScope: string[]
30  assumptions: string[]
31}
32export type Constitution = { principles: string[]; musts: string[] }
33
34export type Snapshot = {
35  /** Whether Spec Kit is set up here (`.specify/` exists). */
36  initialized: boolean
37  featureDir: string | null
38  /** Every feature directory under specs/ (and .specify/specs/). */
39  features: string[]
40  hasPlan: boolean
41  /** Installed Spec Kit extensions, by id (`.specify/extensions/<id>/`). */
42  extensions: string[]
43  /** Whether Spec Kit's Claude Code integration is installed (its skills or commands are there). */
44  claudeIntegration: boolean
45  /** The Spec Kit release the project was set up with (`.specify/init-options.json`). */
46  speckitVersion: string | null
47  /** What can run Spec Kit's CLI on this machine. */
48  tools: { specify: boolean; uvx: boolean }
49  /** The project folder at a glance: nothing but dotfiles and a README (empty), or code already there (existing). */
50  folder: 'empty' | 'existing'
51  spec: Spec | null
52  tasks: Task[]
53  constitution: Constitution | null
54  /** How the project invokes Spec Kit's commands: `/speckit-clarify` (skills) or `/speckit.clarify` (commands). */
55  commandStyle: 'skills' | 'commands'
56  /** What never counts as drift (exempt) and what is listed but not drift (unclear): defaults plus `.xrefignore`. */
57  rules: { exempt: string[]; unclear: string[] }
58  /** Test files tied to the feature whose content holds real tests (not only test.todo). */
59  realTests: string[]
60  /** The fingerprint of plan.md, for the spec+plan review gate. */
61  planFingerprint: string | null
62  /** The command that runs the project's tests: the plugin option, else plan.md's `**Testing**:` line. */
63  testCommand: string | null
64  /** The git branch checked out, when there is one. */
65  branch: string | null
66  /** The open `##` phase of tasks.md the autopilot implements next, if any. */
67  phase: string | null
68  /** A fingerprint of tasks.md's task lines, to see whether converge changed them. */
69  tasksFingerprint: string
70  /** The Spec Kit commands installed for Claude Code, without prefix: `plan`, `analyze`, `xref-check`. */
71  commands: string[]
72  /** Checklists of the feature with open items. */
73  checklists: { file: string; open: number }[]
74  /** What a spec change sets off, from the constitution (flow-back by default). */
75  persistence: 'flow-back' | 'flow-forward' | 'living'
76}
77
78export type Unplanned = { file: string; at: string; task: string | null; acknowledged: boolean }
79export type IntentStatus = 'logged' | 'extends' | 'contradicts' | 'resolved'
80export type Intent = { at: string; text: string; task: string | null; status: IntentStatus }
81export type Anchor = { id: string; file: string; line: number }
82export type Semantic = {
83  score: number
84  verdict: 'aligned' | 'minor' | 'drift'
85  reasons: string[]
86  changes: { text: string; kind: 'extends' | 'contradicts' }[]
87  at: string
88  /** git HEAD and a hash of the diff the verdict looked at. */
89  head?: string
90  diffHash?: string
91}
92export type Verification = { status: 'passing' | 'failing'; tests: string[]; at: string; commit: string; fingerprint: string }
93export type Decision = { id: string; question: string; options: string[]; blocks: string[]; at: string; answer: string | null }
94
95/**
96 * What the mod keeps per feature. Persisted in two files (docs/contract-0.4.md §1): the committed
97 * `specs/<feature>/xref.json` (map, links, accepted, fingerprints) and the local `.specify/xref/local/<feature>.json`.
98 */
99export type Ledger = {
100  tasks: Record<string, { touched: string[]; linked: string[] }>
101  /** By requirement id (FR, SC) and by scenario id (USn-ASm). */
102  requirements: Record<string, { tasks: string[]; files: string[]; sources: string[] }>
103  unplanned: Unplanned[]
104  intents: Intent[]
105  anchors: Anchor[]
106  semantic: Semantic | null
107  semanticHistory: Semantic[]
108  /** A requirement's text fingerprint when it was last mapped, linked or verified. */
109  fingerprints: Record<string, string>
110  verification: Record<string, Verification>
111  /** `spec`/`plan`: the fingerprint the person approved. */
112  approvals: Record<string, string>
113  decisions: Decision[]
114  /** Workflow checkpoints: `analyze` (spec fingerprint analyzed), `converge` (tasks fingerprint converged). */
115  checkpoints: Record<string, string>
116  /** Keys of either file the mod does not know, written back as they were. */
117  extra: { committed: Record<string, unknown>; local: Record<string, unknown> }
118}
119
120/** The autopilot: it moves through Spec Kit's workflow on its own and stops only for the person. */
121export type Autopilot = {
122  on: boolean
123  /** Why it waits for the person; null while it runs. */
124  paused: string | null
125  steps: number
126  max: number
127  /** The progress key after the last step, the steps since it last changed, and the last phase. */
128  last: string | null
129  stalls: number
130  lastPhase: string | null
131  /** What the person asked for before a feature existed: the words /speckit-specify gets. */
132  idea: string | null
133  /** Failed test runs in a row; the fourth makes it the person's decision. */
134  repairs: number
135  /** The last test run: whether it passed and the tail of its output. */
136  lastTest: { ok: boolean; at: string; output: string } | null
137  /** When this run started, and the session's cost then, for the band's `$1.80 · 23m`. */
138  startedAt: number | null
139  costAtStart: number
140  /** A night run: a larger budget, a briefing when it stops. */
141  night: boolean
142  /** The phase the last implement step was scoped to. */
143  scope: string | null
144  /** The tasks checked when the last step was handed over: commit per task commits what came after. */
145  doneAtStep?: string[]
146}
147
148/** One autopilot step in the run log (`.specify/xref/local/<feature>.run.jsonl`). */
149export type RunEntry = {
150  at: string
151  step: number
152  phase: string
153  task: string | null
154  files: string[]
155  durationMs: number
156  tokens: number
157  outcome: string
158}
159
160declare module 'claude-code' {
161  interface PluginState {
162    'speckit-xref': {
163      snapshot: Snapshot | null
164      ledger: Ledger
165      active: string | null
166      checking: boolean
167      autopilot: Autopilot
168      /** The question raised through the ask tool during the running turn. */
169      asked: string | null
170      /** The files written during the running turn, by tools or found changed at its end. */
171      turnFiles: string[]
172      /** The running turn's id and start, for Stop and for finding Bash writes. */
173      turn: { id: string; startedAt: number; dirty: string[] } | null
174      /** The pane's detail view on narrow widths. */
175      details: boolean
176      /** When the person last spoke: the "since you left" card counts from here. */
177      seenAt: number
178      /** The chip under each booked write in the transcript, by tool_use_id: `T006 · FR-004` or `▲ unplanned`. */
179      chips: Record<string, string>
180    }
181  }
182}
183